Multi-modal interview training method, device and equipment and storage medium

Through multimodal data fusion analysis of interviewers' visual, text and voice data, we generate emotional abnormality reports, solving the problem of lack of non-verbal performance analysis in the existing technology, and achieving comprehensive and personalized feedback of interview training.

CN120336814APending Publication Date: 2025-07-18SHANGHAI JUNXING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510445200.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing interview training methods are mainly text-based, and the lack of analysis of students' non-verbal performance has led to limited training results, making it difficult to comprehensively evaluate students' interview performance and provide personalized feedback.

Method used

By analyzing the interviewer's visual data, text data and voice data separately, the initial interview results are generated, and the timestamps of emotional abnormalities are fused, and a multimodal feedback report is generated to provide personalized feedback.

Benefits of technology

It improves the comprehensiveness and reliability of interview training, provides real-time and comprehensive emotional status assessment and feedback, and helps students improve their interview performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336814A_ABST
    Figure CN120336814A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal interview training method, device and equipment and a storage medium, and relates to the technical field of multi-modal information processing. The method comprises the steps of performing emotional state analysis according to current visual data, current text data and current voice data of a current interviewer to obtain an initial interview result of the current interviewer; wherein the initial interview result comprises a visual modal result, a text modal result and a voice modal result; fusing the visual modal result, the text modal result and the voice modal result to obtain a fused interview result of the current interviewer; if the fused interview result does not meet the interview passing condition, screening out at least one abnormal timestamp of emotion abnormality of the current interviewer from the fused interview result; and determining a multi-mode feedback report according to the visual mode, the text mode and the voice mode corresponding to the at least one abnormal timestamp. According to the technical scheme, through multi-modal data processing, the comprehensiveness of interview training can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technology, and particularly to the field of multimodal information processing technology. Specifically, the present application relates to a multimodal interview training method, device, equipment, and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, natural language processing and computer vision technologies based on large models have been widely applied in the field of education and training.

[0003] Existing interview training methods mainly focus on text, lacking the analysis of trainees' non-verbal expressions, resulting in limited training effects. The main challenges faced by current technologies include how to effectively integrate voice, vision, and text information to comprehensively evaluate trainees' interview performances, and how to provide personalized feedback to enhance trainees' confidence and expression abilities. Summary of the Invention

[0004] The present application provides a multimodal interview training method, device, equipment, and storage medium to improve the comprehensiveness of interview training.

[0005] According to one aspect of the present application, there is provided a multimodal interview training method, which includes:

[0006] Respectively perform emotional state analysis based on the current visual data, current text data, and current voice data of the current interviewee to obtain an initial interview result of the current interviewee; wherein, the initial interview result includes a visual modality result, a text modality result, and a voice modality result;

[0007] Fuse the visual modality result, the text modality result, and the voice modality result to obtain a fused interview result of the current interviewee; wherein, the fused interview result is used to characterize the emotional stability of the current interviewee;

[0008] If the fused interview result does not meet the interview passing condition, screen out at least one abnormal timestamp at which the current interviewee has emotional abnormalities from the fused interview result;

[0009] Determine a multimodal feedback report according to the visual modality, text modality, and voice modality corresponding to the at least one abnormal timestamp, and feedback the multimodal feedback report to the current interviewee in the form of text and / or voice.

[0010] According to another aspect of the present application, there is provided a multimodal interview training device, which includes:

[0011] A status analysis module, configured to perform emotional state analysis based on the current visual data, current text data, and current voice data of the current interviewee respectively, to obtain an initial interview result of the current interviewee; wherein, the initial interview result includes a visual modality result, a text modality result, and a voice modality result;

[0012] A result fusion module, configured to fuse the visual modality result, the text modality result, and the voice modality result to obtain a fused interview result of the current interviewee; wherein, the fused interview result is used to characterize the emotional stability of the current interviewee;

[0013] A timestamp screening module, configured to, if the fused interview result does not meet the interview passing condition, screen out at least one abnormal timestamp at which the current interviewee has emotional abnormalities from the fused interview result;

[0014] A report generation module, configured to determine a multimodal feedback report according to the visual modality, text modality, and voice modality corresponding to the at least one abnormal timestamp, and feedback the multimodal feedback report to the current interviewee in the form of text and / or voice.

[0015] According to another aspect of the present application, there is provided an electronic device, which includes:

[0016] One or more processors;

[0017] A memory, configured to store one or more programs;

[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement any one of the multimodal interview training methods provided by the embodiments of the present application.

[0019] According to another aspect of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any one of the multimodal interview training methods provided by the embodiments of the present application.

[0020] According to another aspect of the present application, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, it implements any one of the multimodal interview training methods provided by the embodiments of the present application.

[0021] This application analyzes the emotional state respectively based on the current visual data, current text data, and current voice data of the current interviewee to obtain the initial interview result of the current interviewee; wherein, the initial interview result includes a visual modality result, a text modality result, and a voice modality result; the visual modality result, the text modality result, and the voice modality result are fused to obtain the fused interview result of the current interviewee; wherein, the fused interview result is used to characterize the emotional stability of the current interviewee; if the fused interview result does not meet the interview passing condition, at least one abnormal timestamp when the current interviewee has emotional abnormalities is screened out from the fused interview result; according to the visual modality, text modality, and voice modality corresponding to at least one abnormal timestamp, a multimodal feedback report is determined and the multimodal feedback report is fed back to the current interviewee in the form of text and / or voice. The above technical solution comprehensively evaluates the emotional state of the current interviewee through multimodal data to generate a return report, which can improve the reliability and comprehensiveness of interview training, enabling the current interviewee to obtain real-time and comprehensive feedback. Description of the Drawings

[0022] Figure 1 is a flowchart of a multimodal interview training method provided in Embodiment 1 of the present application;

[0023] Figure 2 is a flowchart of a multimodal interview training method provided in Embodiment 2 of the present application;

[0024] Figure 3 is a schematic structural diagram of a multimodal interview training device provided in Embodiment 3 of the present application;

[0025] Figure 4 is a schematic structural diagram of an electronic device implementing the multimodal interview training method of Embodiment 4 of the present application. Detailed Embodiments

[0026] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0027] It should be noted that in the description of the present application, the claims and the above-mentioned drawings, the terms "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0028] In addition, it should also be noted that in the technical solution of the present application, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of relevant data such as the current visual data and the current text data complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0029] Embodiment 1

[0030] Figure 1 is a flowchart of a multi-modal interview training method provided according to Embodiment 1 of the present application. This embodiment is applicable to the situation of simulating an interview for trainees and can be executed by a multi-modal interview training device. The multi-modal interview training device can be implemented in the form of hardware and / or software, and the multi-modal interview training device can be configured in a computer device, such as a server. As Figure 1 shown, the method includes:

[0031] S110. Perform emotional state analysis respectively according to the current visual data, current text data, and current voice data of the current interviewee to obtain the initial interview result of the current interviewee; wherein, the initial interview result includes a visual modality result, a text modality result, and a voice modality result.

[0032] In this embodiment, the current interviewee refers to the object undergoing a mock interview. The current visual data refers to all visual information related to the interviewee captured by a camera during the interview; this data mainly involves the interviewee's facial expressions, eye contact, body language, posture changes, etc. The current text data refers to the text data generated during the interview through text or language input (such as answering questions, written tests, chat records, etc.); generally, this data includes the interviewee's speech content, grammar structure, sentence composition, expression methods, etc. The current voice data refers to the data recorded by voice during the interview, including the interviewee's pronunciation, intonation, speech rate, pauses, volume, etc. The initial interview result refers to the scores of the current interviewee in the visual modality, text modality, and voice modality, used to characterize the emotional stability of the current interviewee in the visual modality, text modality, and voice modality. The visual modality result refers to the result of the emotional state or performance obtained through the analysis of visual signals such as the interviewee's facial expressions, body language, and eye contact. The text modality result refers to the emotional state or emotional tendency obtained through the analysis of the interviewee's language expression (text or speech-to-text). The voice modality result refers to the emotional state or emotional information obtained through the analysis of the interviewee's voice signal; the voice modality analysis mainly considers features such as pitch, speech rate, pauses, and volume, and these factors can reveal the interviewee's emotional fluctuations.

[0033] Exemplarily, a deep learning framework can be used to synchronously analyze the emotional state of the current visual data, current text data, and current voice data of the current interviewee to obtain the initial interview result of the current interviewee.

[0034] Optionally, the current visual data, current text data, and current voice data of the current interviewee are respectively converted into multi-dimensional data to obtain the multi-dimensional visual data, multi-dimensional text data, and multi-dimensional voice data of the current interviewee; the emotional state of the multi-dimensional visual data, multi-dimensional text data, and multi-dimensional voice data is analyzed to obtain the initial interview result of the current interviewee.

[0035] In this embodiment, the multi-dimensional visual data refers to the data of a multi-dimensional data structure formed by organizing the current visual data in space and time. The multi-dimensional text data refers to the data of a multi-dimensional data structure formed by the current text data through methods such as word vectors or embedding vectors. The multi-dimensional voice data refers to the data of a multi-dimensional data structure formed by the current voice data through signal processing methods.

[0036] Furthermore, the local visual features, voice time series features, and text sequence features of the current interviewee are respectively extracted from the multi-dimensional visual data, multi-dimensional text data, and multi-dimensional voice data; the emotional state of the local visual features, voice time series features, and text sequence features is respectively analyzed to obtain the interview result of the current interviewee.

[0037] In this embodiment, the local visual features may include at least one of facial local features and limb local features, etc.; among them, the facial local features refer to the features extracted from the facial image related to facial expressions, facial surface details (such as eyes, mouth, nose, etc.), and facial features are usually used to judge the emotional state (for example, smiling, frowning, etc.); the limb local features refer to the features extracted from the whole body movements related to limb postures, gestures, etc., and this information helps to evaluate the behavioral patterns and emotional states of the interviewee (such as anxiety, relaxation, etc.). The speech time series features refer to the features extracted when performing time series analysis on the speech signal; these features may include the frequency, pitch, duration, speech rate, etc. of the audio, and these features can reflect the emotions and psychological states of the speaker. The text sequence features refer to the sequence data features extracted from the text spoken by the interviewee; these features usually include the emotional tendency, keywords, grammatical structure, etc. of the text, and are used to judge the information and emotional attitudes conveyed by the interviewee during the interview.

[0038] S120. Fuse the visual modality result, the text modality result, and the speech modality result to obtain the fusion interview result of the current interviewee; among them, the fusion interview result is used to characterize the emotional stability of the current interviewee.

[0039] In this embodiment, the fusion interview result refers to the final evaluation result after comprehensively analyzing the visual, text, and speech modality results; through multimodal fusion (combining and analyzing data from different sources), a more comprehensive and accurate evaluation of the emotional state is obtained. Emotional stability refers to whether the emotional performance of the interviewee during the interview is stable and whether they can control their emotions.

[0040] In an alternative manner, the visual modality result, the text modality result, and the speech modality result can be presented in the form of scores; correspondingly, perform weighted summation on the visual modality result, the text modality result, and the speech modality result to obtain the fusion interview result of the current interviewee.

[0041] Exemplarily, set the weights a, b, and c to correspond to the outputs of the visual, text, and audio modalities respectively, and the finally output fusion interview result can be expressed as: O = aA + bB + cC; where A is the visual modality result, that is, the confidence score of the visual modality; B is the text modality result, that is, the positive emotion score of the text modality; C is the speech modality result, that is, the tension score of the speech modality (here the tension can be converted into the positive degree, such as 1 - tension).

[0042] It should be noted that the score is only a part of the result; presenting the result in the form of a score is to facilitate the interviewee to more intuitively understand how they performed in the current mock interview.

[0043] S130. If the integrated interview result does not meet the interview passing criteria, at least one abnormal timestamp when the current interviewee has abnormal emotions is screened out from the integrated interview result.

[0044] In this embodiment, the interview passing criteria is a set standard used to determine whether an interviewee meets the requirements for passing the interview. This criteria may include factors such as emotional stability, communication ability, professional ability, etc. If the integrated interview result does not meet these criteria, it may be regarded as not passing the interview. An abnormal timestamp refers to the fact that during the emotional analysis process, the emotional data at certain time points may deviate from the normal, showing abnormal emotional fluctuations. An abnormal timestamp refers to these abnormal moments, that is, the moments when the interviewee's emotional performance is unstable or there are sudden emotional fluctuations.

[0045] In an alternative embodiment, there are real-time emotional scores for the visual modality, text modality, and speech modality corresponding to each timestamp in the integrated interview result. If there is at least one real-time emotional score of a modality exceeding the set score threshold, it is determined that the timestamp is abnormal.

[0046] In this embodiment, the real-time emotional score refers to the system's immediate scoring and evaluation of the interviewee's emotional state during the interview. Each timestamp (i.e., each moment or period) will have a real-time emotional score, reflecting the interviewee's emotional performance at that moment (such as happy, angry, nervous, etc.). The set score threshold refers to a critical point set by the system for each emotional state in emotional analysis. If the real-time emotional score of a modality exceeds this threshold, it means that the emotional performance at that timestamp is abnormal, which may reflect abnormal states such as the interviewee's emotional fluctuations and nervousness.

[0047] S140. According to the visual modality, text modality, and speech modality corresponding to at least one abnormal timestamp, a multimodal feedback report is determined and the multimodal feedback report is fed back to the current interviewee in the form of text and / or speech.

[0048] In this embodiment, the multimodal feedback report refers to a feedback report generated by combining the visual, text, and speech data of the interviewee. This report summarizes the interviewee's emotional state during the interview, especially during the time periods when abnormal emotional performances are found. The report is usually provided to the interviewee in the form of text or speech to help them understand their performance and emotional fluctuations during the interview.

[0049] Exemplarily, the system generates a feedback report according to the visual modality, text modality, and speech modality corresponding to at least one abnormal timestamp, including emotional analysis results, non-verbal performance evaluations, and improvement suggestions. For example, the feedback report may indicate that the interviewee visually shows a high level of confidence, but there is a certain degree of nervousness in speech, and it is recommended that the interviewee pay attention to controlling the speech rate and intonation in subsequent practice.

[0050] In an alternative embodiment, after determining the multi-modal feedback report, if a report generation requirement of the current interviewee is received, the multi-modal feedback report is adjusted according to the report generation requirement.

[0051] Exemplarily, a setting option is added to the user interface to allow the trainee to select the level of detail of the feedback. Several preset options can be provided, such as brief feedback, medium feedback, and detailed feedback; each option corresponds to different feedback content and depth; for example, when brief feedback is selected, the system only provides key improvement suggestions, while when detailed feedback is selected, it will include sentiment analysis, non-verbal performance assessment, and specific improvement strategies; and the trainee is allowed to customize the feedback content according to their own needs; for example, the trainee can choose to focus on a certain aspect of language expression, emotional state, or non-verbal performance; the system will adjust the generation logic of the feedback report according to the trainee's selection to ensure that the most relevant information is provided.

[0052] In the embodiment of the present application, the emotional state of the current interviewee is analyzed respectively according to the current visual data, current text data, and current voice data of the current interviewee to obtain the initial interview result of the current interviewee; wherein, the initial interview result includes a visual modality result, a text modality result, and a voice modality result; the visual modality result, the text modality result, and the voice modality result are fused to obtain the fused interview result of the current interviewee; wherein, the fused interview result is used to characterize the emotional stability of the current interviewee; if the fused interview result does not meet the interview passing condition, at least one abnormal timestamp when the current interviewee has emotional abnormalities is screened out from the fused interview result; according to the visual modality, text modality, and voice modality corresponding to the at least one abnormal timestamp, a multi-modal feedback report is determined, and the multi-modal feedback report is fed back to the current interviewee in the form of text and / or voice. The above technical solution comprehensively evaluates the emotional state of the current interviewee through multi-modal data to generate a feedback report, which can improve the reliability and comprehensiveness of the interview training, enabling the current interviewee to obtain real-time and comprehensive feedback.

[0053] Embodiment II

[0054] Figure 2It is a flowchart of a multi-modal interview training method provided in Embodiment 2 of the present application. Based on the technical solutions of the above embodiments, this embodiment refines "determining a multi-modal feedback report according to the visual modality, text modality, and speech modality corresponding to at least one abnormal timestamp" into "for each abnormal timestamp, based on the modality priority, sequentially perform abnormal emotion detection on the visual modality, text modality, and speech modality corresponding to the abnormal timestamp to obtain at least one candidate abnormal modality; perform abnormal persistence detection on at least one candidate abnormal modality to obtain at least one target abnormal modality; determine the multi-modal feedback report corresponding to the abnormal timestamp according to at least one target abnormal modality". It should be noted that for the parts not described in detail in the embodiments of the present application, reference may be made to the relevant descriptions of other embodiments. As Figure 2 shown, the method includes:

[0055] S210. Respectively perform emotion state analysis on the current visual data, current text data, and current speech data of the current interviewee to obtain the initial interview result of the current interviewee; wherein, the initial interview result includes a visual modality result, a text modality result, and a speech modality result.

[0056] In an alternative embodiment, a multi-modal model combined by constructing a visual model, a speech model, and a text model can be used; the multi-modal model respectively performs emotion state analysis on the current visual data, current text data, and current speech data of the current interviewee to obtain the initial interview result of the current interviewee.

[0057] Optionally, the visual model can use the application programming interface (API) of the deep learning framework to construct a convolutional layer to extract local features of the input image. For example, through the convolutional kernel, local features such as the eyebrows and eyes of the interviewee's face, as well as features such as the posture of the limbs, can be extracted; a pooling layer is added for downsampling to reduce the amount of calculation and control overfitting; for example, using max pooling to select the maximum value within the local area and retain the most important feature information; using the ReLU (Rectified Linear Unit) activation function to increase the non-linearity of the model, enabling the model to learn more complex feature representations, and integrating the features extracted by the convolutional layer through a fully connected layer, and the output layer outputs the prediction result according to the specific task; in the interview scenario, scores such as the confidence level and nervousness of the interviewee can be output.

[0058] Optionally, the text model can convert the integer IDs (Identifiers) of the text into low-dimensional word vector representations through an embedding layer, enabling the text data to be better processed by the model; add an RNN (Recurrent Neural Network) or LSTM (Long Short-Term Memory) layer to process sequential data and learn the context information of the text; for example, it can understand the logical relationships between sentences in the interviewee's answer. Multiple layers of RNN or LSTM can be added as needed to enhance the learning ability of the model; finally, add a fully connected layer and an output layer for prediction, and output scores such as the positive emotion and logic of the interviewee's answer.

[0059] Optionally, the speech model can use a convolutional layer to extract the acoustic features of the speech and capture local frequency and time information; for example, it can extract features such as the pitch and timbre of the speech; pass the output of the convolutional layer to an RNN or LSTM layer to capture the time series features of the speech signal, such as changes in speech rate; finally, add a fully connected layer and an output layer for prediction of speech-related tasks, and output scores such as the nervousness and clarity of expression of the interviewee.

[0060] S220. Fuse the visual modality results, text modality results, and speech modality results to obtain the integrated interview result of the current interviewee; among them, the integrated interview result is used to characterize the emotional stability of the current interviewee.

[0061] Optionally, the visual modality results, text modality results, and speech modality results can be sorted in chronological order to obtain the integrated interview result of the current interviewee.

[0062] Optionally, based on the modality priority, the visual modality results, text modality results, and speech modality results can be sorted to obtain the integrated interview result of the current interviewee.

[0063] S230. If the integrated interview result does not meet the interview passing criteria, at least one abnormal timestamp when the current interviewee has abnormal emotions is screened out from the integrated interview result.

[0064] S240. For each abnormal timestamp, based on the modality priority, perform abnormal emotion detection on the visual modality, text modality, and speech modality corresponding to the abnormal timestamp in sequence to obtain at least one candidate abnormal modality.

[0065] In this embodiment, modal priority refers to the concept in a multimodal system that determines the importance or priority of different modalities when processing information. A multimodal system involves multiple data sources or perception methods (such as vision, speech, text, etc.), and each modality can provide different types of information. In such a system, the data of different modalities may have different importance in certain tasks. Therefore, it is necessary to rank the modalities by priority to ensure that the system can reasonably combine the information of each modality, thereby improving the decision-making effect or the accuracy of task completion. Abnormal emotion detection is a technique for analyzing and identifying abnormal emotional states. It is usually based on multimodal inputs (vision, text, speech), and detects abnormal fluctuations of emotions through machine learning or deep learning models. It is often used in emotion analysis and emotion computing tasks. For example, detecting abnormal emotions (such as sudden changes in emotions like anxiety, anger, etc.) in a conversation. A candidate abnormal modality refers to a modality that may be abnormal selected through the multimodal emotion detection results at each abnormal timestamp. It can be any one of the vision, text, or speech modalities.

[0066] S250. Perform abnormal persistence detection on at least one candidate abnormal modality to obtain at least one target abnormal modality.

[0067] In this embodiment, abnormal persistence detection is used to evaluate whether an abnormal emotion persists within a given time period. This helps to determine whether it is just a short-term abnormal fluctuation or a persistent emotional abnormality, and the latter may require further attention and processing. A target abnormal modality refers to a modality that is finally determined to have a persistent abnormality after abnormal persistence detection. It usually refers to those modalities with obvious and persistent abnormal emotions.

[0068] Optionally, for each candidate abnormal modality, perform abnormal emotion detection on the modality performance of the candidate abnormal modality at the previous timestamp and the next timestamp. If the modality performance of the candidate abnormal modality at the previous timestamp and the next timestamp fails the abnormal emotion detection, then determine the candidate abnormal modality as the target abnormal modality.

[0069] In this embodiment, the modality performance at the previous timestamp and the next timestamp refers to the data performance at the moments before and after the candidate abnormal modality, that is, the data performance at the two consecutive observation moments of the system.

[0070] S260. Determine a multimodal feedback report corresponding to the abnormal timestamp according to at least one target abnormal modality, and feedback the multimodal feedback report to the current interviewee in the form of text and / or speech.

[0071] Optionally, extract features from at least one target abnormal modality to obtain at least one target abnormal feature; based on the correspondence between candidate abnormal features and candidate feedback suggestions, determine at least one target feedback suggestion according to the at least one target abnormal feature; fuse the at least one target feedback suggestion to obtain a multi-modal feedback report corresponding to the abnormal timestamp.

[0072] In this embodiment, the target abnormal feature refers to the feature extracted from the target abnormal modality and used to describe the abnormal performance of the modality; the target abnormal feature is important input information for the system to analyze and process abnormal modalities; among them, the target abnormal feature may include visual abnormal features, text abnormal features or voice abnormal features. The candidate abnormal feature refers to the feature that may cause abnormalities, usually pre-configured from different modalities or moments. The candidate feedback suggestion refers to some candidate feedback suggestions that may be generated based on the currently analyzed features; these suggestions are usually pre-configured by the system as possible countermeasures or adjustment strategies according to different situations. The target feedback suggestion refers to the suggestion proposed for the target abnormal feature.

[0073] Exemplarily, the system captures the facial expression of the trainee through a camera and identifies that the expression is nervous (equivalent to a visual abnormal feature); at the same time, the system analyzes the voice characteristics of the trainee and finds that the speech rate is 120 words per minute and the intonation is low (equivalent to a voice abnormal feature); based on this information, the system generates a feedback report indicating that the trainee shows positive emotions when expressing interest, but the nervous emotion affects his confidence level.

[0074] In the embodiment of the present application, by performing emotional state analysis based on the current visual data, current text data, and current voice data of the current interviewee respectively, an initial interview result of the current interviewee is obtained; wherein, the initial interview result includes a visual modality result, a text modality result, and a voice modality result; the visual modality result, the text modality result, and the voice modality result are fused to obtain a fused interview result of the current interviewee; wherein, the fused interview result is used to characterize the emotional stability of the current interviewee; if the fused interview result does not meet the interview passing condition, at least one abnormal timestamp when the current interviewee has emotional abnormalities is screened out from the fused interview result; for each abnormal timestamp, based on the modality priority, abnormal emotion detection is sequentially performed on the visual modality, text modality, and voice modality corresponding to the abnormal timestamp to obtain at least one candidate abnormal modality; abnormal persistence detection is performed on at least one candidate abnormal modality to obtain at least one target abnormal modality; according to at least one target abnormal modality, a multimodal feedback report corresponding to the abnormal timestamp is determined, and the multimodal feedback report is fed back to the current interviewee in the form of text and / or voice. Through the above technical solution, the emotional state of the current interviewee is comprehensively evaluated through multimodal data to generate a return report, which can improve the reliability and comprehensiveness of interview training, enabling the current interviewee to obtain real-time and comprehensive feedback.

[0075] Embodiment III

[0076] Figure 3 FIG. 7 is a schematic structural diagram of a multimodal interview training device provided according to Embodiment III of the present application, which is applicable to the situation of simulating an interview for trainees. The multimodal interview training device can be implemented in the form of hardware and / or software, and the multimodal interview training device can be configured in a computer device, such as a server. As Figure 3 shown, the device includes:

[0077] A state analysis module 310, configured to perform emotional state analysis based on the current visual data, current text data, and current voice data of the current interviewee respectively, to obtain an initial interview result of the current interviewee; wherein, the initial interview result includes a visual modality result, a text modality result, and a voice modality result;

[0078] A result fusion module 320, configured to fuse the visual modality result, the text modality result, and the voice modality result to obtain a fused interview result of the current interviewee; wherein, the fused interview result is used to characterize the emotional stability of the current interviewee;

[0079] A timestamp screening module 330, configured to, if the fused interview result does not meet the interview passing condition, screen out at least one abnormal timestamp when the current interviewee has emotional abnormalities from the fused interview result;

[0080] A report generation module 340 is configured to determine a multimodal feedback report based on the visual modality, text modality, and speech modality corresponding to at least one abnormal timestamp, and feedback the multimodal feedback report to the current interviewee in the form of text and / or speech.

[0081] In the embodiment of the present application, by respectively performing emotional state analysis on the current visual data, current text data, and current speech data of the current interviewee, an initial interview result of the current interviewee is obtained; wherein, the initial interview result includes a visual modality result, a text modality result, and a speech modality result; the visual modality result, the text modality result, and the speech modality result are fused to obtain a fused interview result of the current interviewee; wherein, the fused interview result is used to characterize the emotional stability of the current interviewee; if the fused interview result does not meet the interview passing condition, at least one abnormal timestamp when the current interviewee has abnormal emotions is screened out from the fused interview result; a multimodal feedback report is determined based on the visual modality, text modality, and speech modality corresponding to at least one abnormal timestamp, and the multimodal feedback report is feedback to the current interviewee in the form of text and / or speech. Through the above technical solution, the emotional state of the current interviewee is comprehensively evaluated through multimodal data to generate a feedback report, which can improve the reliability and comprehensiveness of interview training, enabling the current interviewee to obtain real-time and comprehensive feedback.

[0082] Optionally, the report generation module 340 includes:

[0083] An abnormal emotion detection unit is configured to perform abnormal emotion detection on the visual modality, text modality, and speech modality corresponding to each abnormal timestamp in sequence based on the modality priority to obtain at least one candidate abnormal modality;

[0084] An abnormal modality screening unit is configured to perform abnormal persistence detection on at least one candidate abnormal modality to obtain at least one target abnormal modality;

[0085] A report generation unit is configured to determine a multimodal feedback report corresponding to the abnormal timestamp based on at least one target abnormal modality.

[0086] Optionally, the abnormal modality screening unit is specifically configured to:

[0087] For each candidate abnormal modality, perform abnormal emotion detection on the modality performance of the candidate abnormal modality at the previous timestamp and the next timestamp;

[0088] If the modality performance of the candidate abnormal modality at the previous timestamp and the next timestamp fails the abnormal emotion detection, the candidate abnormal modality is determined as the target abnormal modality.

[0089] Optionally, the report generation unit is specifically configured to:

[0090] Extract features from at least one target abnormal modality to obtain at least one target abnormal feature; wherein, the target abnormal feature includes a visual abnormal feature, a text abnormal feature, or a voice abnormal feature;

[0091] Based on the correspondence between the candidate abnormal feature and the candidate feedback suggestion, determine at least one target feedback suggestion according to at least one target abnormal feature;

[0092] Fuse at least one target feedback suggestion to obtain a multimodal feedback report corresponding to the abnormal timestamp.

[0093] Optionally, the state analysis module 310 includes:

[0094] A data conversion unit for respectively converting the current visual data, current text data, and current voice data of the current interviewee into multi-dimensional data to obtain the multi-dimensional visual data, multi-dimensional text data, and multi-dimensional voice data of the current interviewee;

[0095] A state analysis unit for performing an emotional state analysis on the multi-dimensional visual data, multi-dimensional text data, and multi-dimensional voice data to obtain an initial interview result of the current interviewee.

[0096] Optionally, the state analysis unit is specifically used for:

[0097] Respectively extract features from the multi-dimensional visual data, multi-dimensional text data, and multi-dimensional voice data to obtain the local visual features, voice time series features, and text sequence features of the current interviewee; wherein, the local visual features include facial local features and limb local features;

[0098] Respectively perform an emotional state analysis on the local visual features, voice time series features, and text sequence features to obtain an interview result of the current interviewee.

[0099] The multimodal interview training device provided by the embodiments of the present application can execute the multimodal interview training method provided by any embodiment of the present application, and has corresponding functional modules and beneficial effects for executing each multimodal interview training method.

[0100] According to the embodiments of the present application, the present application also provides an electronic device, a readable storage medium, and a computer program product.

[0101] Embodiment 4

[0102] Figure 4It is a schematic structural diagram of an electronic device 410 for implementing the multi-modal interview training method of the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described herein and / or claimed.

[0103] As Figure 4 shown, the electronic device 410 includes at least one processor 411, and a memory communicatively connected to the at least one processor 411, such as a read-only memory (ROM) 412, a random access memory (RAM) 413, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 411 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 412 or the computer program loaded from the storage unit 418 into the random access memory (RAM) 413. In the RAM 413, various programs and data required for the operation of the electronic device 410 can also be stored. The processor 411, the ROM 412, and the RAM 413 are connected to each other through a bus 414. The input / output (I / O) interface 415 is also connected to the bus 414.

[0104] Multiple components in the electronic device 410 are connected to the I / O interface 415, including: an input unit 416, such as a keyboard, a mouse, etc.; an output unit 417, such as various types of displays, speakers, etc.; a storage unit 418, such as a disk, an optical disc, etc.; and a communication unit 419, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 419 allows the electronic device 410 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0105] The processor 411 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 411 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 411 executes the various methods and processes described above, such as the multi-modal interview training method.

[0106] In some embodiments, the multimodal interview training method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 418. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 410 via the ROM 412 and / or the communication unit 419. When the computer program is loaded into the RAM 413 and executed by the processor 411, one or more steps of the multimodal interview training method described above can be performed. Alternatively, in other embodiments, the processor 411 can be configured for the multimodal interview training method by any other suitable means (e.g., by means of firmware).

[0107] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0108] The computer programs for implementing the methods of this application can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable multimodal interview training device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0109] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0111] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0112] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is created by computer programs that run on respective computers and have a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0113] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution of this application can be achieved, and no limitation is imposed herein.

[0114] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. A multi-modal interview training method, characterized in that, Including: Respectively perform emotional state analysis based on the current visual data, current text data, and current voice data of the current interviewee to obtain the initial interview result of the current interviewee; wherein, the initial interview result includes a visual modality result, a text modality result, and a voice modality result; Fuse the visual modality result, the text modality result, and the voice modality result to obtain the fused interview result of the current interviewee; wherein, the fused interview result is used to characterize the emotional stability of the current interviewee; If the fused interview result does not meet the interview passing condition, then screen out at least one abnormal timestamp when the current interviewee has abnormal emotions from the fused interview result; Determine a multimodal feedback report according to the visual modality, text modality, and voice modality corresponding to the at least one abnormal timestamp, and feedback the multimodal feedback report to the current interviewee in the form of text and / or voice.

2. The method according to claim 1, wherein Determine a multimodal feedback report according to the visual modality, text modality, and voice modality corresponding to the at least one abnormal timestamp, including: For each abnormal timestamp, based on the modality priority, sequentially perform abnormal emotion detection on the visual modality, text modality, and voice modality corresponding to this abnormal timestamp to obtain at least one candidate abnormal modality; Perform abnormal persistence detection on the at least one candidate abnormal modality to obtain at least one target abnormal modality; Determine the multimodal feedback report corresponding to this abnormal timestamp according to the at least one target abnormal modality.

3. The method according to claim 2, characterized in that Perform abnormal persistence detection on the at least one candidate abnormal modality to obtain at least one target abnormal modality, including: For each candidate abnormal modality, perform abnormal emotion detection on the modality performance of this candidate abnormal modality at the previous timestamp and the next timestamp; If the modality performance of this candidate abnormal modality at the previous timestamp and the next timestamp fails the abnormal emotion detection, then determine this candidate abnormal modality as the target abnormal modality.

4. The method according to claim 2, wherein Determine the multimodal feedback report corresponding to this abnormal timestamp according to the at least one target abnormal modality, including: Extract features from the at least one target abnormal modality to obtain at least one target abnormal feature; wherein, the target abnormal feature includes a visual abnormal feature, a text abnormal feature, or a voice abnormal feature; Based on the correspondence between candidate abnormal features and candidate feedback suggestions, determine at least one target feedback suggestion according to the at least one target abnormal feature; Fuse the at least one target feedback suggestion to obtain the multimodal feedback report corresponding to this abnormal timestamp.

5. The method according to claim 1, wherein Respectively perform emotional state analysis based on the current visual data, current text data, and current voice data of the current interviewee to obtain the initial interview result of the current interviewee, including: Respectively convert the current visual data, current text data, and current voice data of the current interviewee into multi-dimensional data to obtain the multi-dimensional visual data, multi-dimensional text data, and multi-dimensional voice data of the current interviewee; Perform emotional state analysis on the multi-dimensional visual data, multi-dimensional text data, and multi-dimensional voice data to obtain the initial interview result of the current interviewee.

6. The method according to claim 5, characterized in that, Perform emotional state analysis on the multi-dimensional visual data, the multi-dimensional text data, and the multi-dimensional voice data to obtain the interview result of the current interviewee, including: Extract features from the multi-dimensional visual data, the multi-dimensional text data, and the multi-dimensional voice data respectively to obtain the local visual features, speech time series features, and text sequence features of the current interviewee; wherein, the local visual features include facial local features and limb local features; Perform emotional state analysis on the local visual features, the speech time series features, and the text sequence features respectively to obtain the interview result of the current interviewee.

7. A multimodal interview training device, characterized in that, Including: A state analysis module for performing emotional state analysis according to the current visual data, current text data, and current voice data of the current interviewee respectively to obtain the initial interview result of the current interviewee; wherein, the initial interview result includes a visual modality result, a text modality result, and a voice modality result; A result fusion module for fusing the visual modality result, the text modality result, and the voice modality result to obtain the fused interview result of the current interviewee; wherein, the fused interview result is used to characterize the emotional stability of the current interviewee; A timestamp screening module for screening out at least one abnormal timestamp when the current interviewee has emotional abnormalities from the fused interview result if the fused interview result does not meet the interview passing condition; A report generation module for determining a multi-modal feedback report according to the visual modality, text modality, and voice modality corresponding to the at least one abnormal timestamp, and feeding back the multi-modal feedback report to the current interviewee in the form of text and / or voice.

8. An electronic device, characterized in that, Including: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the multi-modal interview training method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the multi-modal interview training method according to any one of claims 1-6.

10. A computer program product, including a computer program, where the computer program implements the multi-modal interview training method according to any one of claims 1-6 when executed by a processor.