An AI interview evaluation method and system based on multi-modal adaptive fusion
Patent Information
- Application Number
- CN202611029867.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明的目的在于提供一种基于多模态自适应融合的AI面试评估方法及系统,以解决现有技术中AI面试评估模型对优秀表现激励不足、评估结果可解释性差的技术问题
[0022]1.通过引入协同增强机制,本发明能更好地区分表现平庸与表现卓越的候选人,对各方面能力均衡且出色的个体给予正向激励,使得高分段的区分更为明显和合理。
Smart Images

Figure CN122820156A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and talent assessment technology, and in particular to an AI interview assessment method and system based on multimodal adaptive fusion. Background Technology
[0002] In the corporate recruitment process, interviews are a crucial step in evaluating candidates. Traditional in-person interviews rely heavily on the interviewer's personal experience and subjective judgment, resulting in inefficiency, inconsistent standards, and susceptibility to bias. To address these issues, artificial intelligence (AI) interview systems have emerged.
[0003] Existing AI interview systems typically analyze candidate information such as video and audio. Some early systems could only perform single-dimensional analysis, such as evaluating answers solely through text keyword matching or analyzing candidate emotions based solely on tone of voice, resulting in rather one-sided assessments. To overcome this limitation, subsequent technological advancements introduced the concept of multimodal fusion, which comprehensively analyzes text, audio, and video information. For example, some technical solutions propose weighted fusion of extracted multimodal features based on a pre-defined competency model, and can identify and handle evaluation conflicts between different modalities.
[0004] While this approach improves the comprehensiveness of the assessment to some extent, it still has significant technical shortcomings: First, when dealing with intermodal information relationships, its logic focuses primarily on correcting negative biases or conflicts, lacking a mechanism to positively incentivize outstanding candidates who demonstrate high consistency across modalities. This makes it difficult to effectively distinguish between candidates with mediocre performance but no conflicts and those with consistent and excellent performance across all aspects. Second, the presentation of assessment results is rather rudimentary, typically providing only final scores, ratings, or simple charts, lacking in-depth, traceable analytical details and intuitive evidence. This makes the assessment process still like a "black box," making it difficult for human resource managers to fully trust the assessment results and to make targeted decisions based on the reports. Summary of the Invention
[0005] The purpose of this invention is to provide an AI interview evaluation method and system based on multimodal adaptive fusion, so as to solve the technical problems of insufficient incentive for excellent performance and poor interpretability of evaluation results in existing AI interview evaluation models.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] An AI interview assessment method based on multimodal adaptive fusion includes the following steps:
[0008] Calculate the candidate's unimodal scores for text, voice, and video during the interview process;
[0009] Based on a pre-defined job competency model library, multimodal fusion weights are configured for the target job.
[0010] Based on the multimodal fusion weights and the single-modal scores, the candidate's comprehensive score is calculated. When the standard deviation among multiple single-modal scores is less than a preset consistency threshold, a synergy factor is introduced to enhance the calculation of the comprehensive score. The synergy factor is equal to the product of a first synergy coefficient and the reciprocal of a target value. The target value is the sum of the standard deviation and a preset smoothing constant, and the preset smoothing constant is greater than zero.
[0011] An evaluation report is generated, which includes multimodal analysis details and key segments of the interview video that are automatically tagged based on the unimodal score, the composite score, and the multimodal analysis details.
[0012] Optionally, the comprehensive score is calculated using the following formula: Comprehensive score = (α × text score + β × voice score + γ × video score) × (1 + δ × collaboration factor), where α, β, and γ are the weight coefficients in the multimodal fusion weights corresponding to the text score, voice score, and video score, respectively, and δ is a preset collaboration enhancement adjustment coefficient.
[0013] Optionally, it also includes: triggering a conflict handling mechanism when the absolute value of the difference between any two scores among the multiple single-modal scores is greater than a preset difference threshold; the conflict handling mechanism specifically includes: identifying the target single modality with the lowest score among the multiple single-modal scores, increasing the weight coefficient corresponding to the target single modality by a preset first step length, and reducing the weight coefficients of the other single modalities by a preset proportion, so as to amplify the evaluation weight of the candidate in the weakness dimension.
[0014] Optionally, the key segments of the interview video include highlight moments and risk moments determined according to preset rules.
[0015] Optionally, the multimodal analysis details include at least one of the following: capability radar chart, speech rate change curve, speech emotion temporal distribution map, keyword cloud map, or gaze point heat map.
[0016] Optionally, the multimodal fusion weights can be manually adjusted or automatically optimized by performing regression analysis on historical interview data and hiring results.
[0017] Optionally, the calculation of the single-modal score is based on at least one video feature extracted from the video data stream, selected from: facial expression features, eye tracking features, head pose features, or body movement features.
[0018] Optionally, the calculation of the single-modal score is based on at least one speech feature extracted from the audio data stream, selected from: speech rate, volume, fundamental frequency, pauses, or speech emotion features.
[0019] Optionally, the calculation of the unimodal score is based on at least one text feature extracted from the text data stream, selected from: content keyword matching, semantic integrity, logical coherence, or terminology usage.
[0020] This invention also provides an AI interview assessment system based on multimodal adaptive fusion, characterized by comprising: a single-modal score calculation module for calculating the single-modal scores of a candidate's text, voice, and video during the interview process; a weight configuration module for configuring multimodal fusion weights for a target position based on a preset job competency model library; a comprehensive score calculation module for calculating the candidate's comprehensive score based on the multimodal fusion weights and the single-modal scores, wherein when the standard deviation among multiple single-modal scores is less than a preset consistency threshold, a synergy factor is introduced to enhance the calculation of the comprehensive score, the synergy factor being equal to the product of a first synergy coefficient and the reciprocal of a target value, the target value being the sum of the standard deviation and a preset smoothing constant, the preset smoothing constant being greater than zero; and a report generation module for generating an assessment report, the assessment report including multimodal analysis details and key segments of the interview video automatically marked based on the single-modal scores, the comprehensive score, and the multimodal analysis details.
[0021] Compared with the prior art, the present invention has the following beneficial effects:
[0022] 1. By introducing a synergistic enhancement mechanism, this invention can better distinguish between candidates with mediocre performance and those with outstanding performance, and provide positive incentives to individuals with balanced and excellent abilities in all aspects, making the distinction of high-scoring segments more obvious and reasonable.
[0023] 2. By generating reports containing multimodal analysis details and key video clip clips, this invention directly links abstract scores with specific data and behavioral performance. Recruiters can intuitively see changes in candidates' speech rate, the distribution of facial expressions, and even directly review video clips marked as risks or highlights, making the evaluation process no longer a "black box," thereby enhancing the trust and acceptance of AI evaluation results.
[0024] 3. By dynamically adjusting weights based on the job competency model, the core evaluation criteria are highly consistent with the core requirements of the job. Whether it is a research and development position that focuses on technical depth or a sales position that focuses on communication skills, a more targeted and scientific evaluation can be obtained.
[0025] 4. By generating structured, visual evaluation reports with intuitive evidence, recruiters can quickly grasp the key qualities of candidates, reducing the time spent watching complete interview videos from beginning to end and effectively improving the efficiency of screening and decision-making. Attached Figure Description
[0026] Figure 1 A flowchart illustrating an AI interview assessment method based on multimodal adaptive fusion, provided for an embodiment of the present invention;
[0027] Figure 2 A schematic diagram of the functional module structure of an AI interview assessment system based on multimodal adaptive fusion provided in an embodiment of the present invention;
[0028] Figure 3 This is a schematic diagram of the internal working logic of the fusion scoring engine provided in an embodiment of the present invention;
[0029] Figure 4 This is a schematic diagram of an interface for a deeply interpretable evaluation report provided in an embodiment of the present invention.
[0030] In the diagram: 10. Data Acquisition and Preprocessing Module; 20. Feature Extraction Engine; 30. Job Competency Model Library; 40. Fusion Scoring Engine; 41. Weight Acquisition Unit; 42. Score Receiving Unit; 43. Conflict Collaboration Judgment Unit; 44. Score Calculation Unit; 50. Report Generation Module; 61. Comprehensive Score Area; 62. Capability Radar Chart; 63. Multimodal Analysis Details Area; 64. Key Video Segment Area. Detailed Implementation
[0031] To better understand the technical solutions of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. It should be understood that the specific embodiments described herein are intended to explain the present invention, and not to limit the present invention.
[0032] Example 1
[0033] This embodiment provides a specific implementation of an AI interview evaluation method and system based on multimodal adaptive fusion. This embodiment focuses on demonstrating how the collaborative enhancement mechanism functions when a candidate performs exceptionally well in all aspects, thereby providing positive incentives for candidates with consistent and outstanding performance.
[0034] Please see Figure 1 This illustrates the complete process of an AI interview assessment method based on multimodal adaptive fusion provided by an embodiment of the present invention. This method can be... Figure 2 The AI interview assessment system based on multimodal adaptive fusion is shown in the figure. Figure 2The functional module structure of an AI interview assessment system based on multimodal adaptive fusion was demonstrated. It should be noted that the system can be deployed on cloud servers, local servers, or terminal devices with sufficient computing power.
[0035] In a specific application scenario, a company uses the AI interview assessment system provided by this invention to recruit a sales manager position. The AI interview assessment system pre-sets a set of multimodal fusion weights for the position in the job competency model library 30. These multimodal fusion weights reflect the emphasis placed on various abilities of the candidate for the sales position. For example, content depth (text), expression ability (voice), and professional image (video) are set to be equally important.
[0036] The evaluation process specifically includes the following steps:
[0037] Step S100: Data acquisition and preprocessing.
[0038] When a candidate (Candidate A) begins an AI interview, the data acquisition and preprocessing module 10 of the AI interview assessment system is activated. This module connects to the camera and microphone of the interview device (such as a computer or mobile terminal) and simultaneously acquires video and audio data streams during the candidate A's interview. Specifically, the video data stream is acquired at a rate of 30 frames per second and a resolution of 1920x1080 pixels; the audio data stream is acquired at a sampling rate of 44.1kHz and a sampling depth of 16 bits. The data acquisition and preprocessing module 10 first performs noise reduction processing on the acquired audio data stream, for example, using spectral subtraction or a deep learning noise reduction model to filter out background noise and current noise, thereby improving the accuracy of subsequent speech recognition. After noise reduction, the module performs endpoint detection to accurately identify the start and end points of speech to segment effective speech segments. Subsequently, through the integrated speech recognition engine, the processed audio data stream is converted into a text data stream in real time and pre-formatted, such as adding punctuation and distinguishing paragraphs. For the video data stream, the data acquisition and preprocessing module 10 performs face detection and alignment. As an optional implementation, algorithms such as multi-task cascaded convolutional networks can be used to accurately locate the face region in the video frame, and rotate and scale it based on key points such as the eyes, nose, and mouth to maintain the face's standard posture in subsequent analysis and eliminate interference caused by head movements. After preprocessing, the system obtains aligned and standardized video, audio, and text data streams, laying a solid foundation for subsequent feature extraction.
[0039] Step S200: Multimodal feature extraction.
[0040] The preprocessed data stream is fed into the feature extraction engine 20. This feature extraction engine 20 contains analysis sub-engines for different modalities, which extract rich quantized features from the three data streams in parallel.
[0041] Text Feature Extraction: For the text data stream generated by speech recognition, the text analysis sub-engine extracts multiple features. First, it performs content keyword matching analysis, comparing the candidate A's answer with relevant question answer points or keyword libraries (e.g., "understand the reasons," "provide alternatives," "maintain relationships," "build trust") pre-set in the job competency model library 30, calculating coverage and matching degree. Second, it analyzes semantic completeness and logical coherence using natural language processing techniques. For example, dependency parsing can be used to assess the completeness of sentence structure, and the frequency and correctness of discourse conjunctions (such as firstly, secondly, therefore) can be combined to assess logical consistency. In addition, the use of professional terminology is also statistically analyzed.
[0042] Speech Feature Extraction: The speech analysis sub-engine performs deep analysis of the audio data stream. The extracted acoustic features include: speech rate (words per minute), for example, candidate A's average speech rate is 180 words per minute, within the range of confident fluency; volume (decibels), analyzing its variations to determine emphasis and fluctuations in expression; fundamental frequency (Hertz), analyzing its mean and variance to reflect the smoothness or intonation of the tone; pauses, statistically analyzing the number and duration of effective pauses (used for thinking and organizing language) and ineffective pauses (such as stuttering or pausing). Additionally, a pre-trained speech emotion recognition model classifies speech segments by emotion, identifying emotional states such as positive, neutral, tense, and annoyed, along with their confidence levels. For candidate A, the emotion recognition results consistently show positive and confident.
[0043] Video Feature Extraction: The video analysis sub-engine analyzes the face-aligned video data stream. Extracted facial expression features include identifying the frequency and duration of expressions such as smiling, surprise, and seriousness by analyzing the movement patterns of key facial points. Extracted eye-tracking features quantify the duration and frequency of eye contact between the candidate and the camera; Candidate A demonstrated stable and continuous eye contact in this aspect. Extracted head posture features analyze head movements such as nodding, shaking, and tilting to assess focus and agreement. Extracted body movement features (in half-body or full-body videos) identify the amplitude and frequency of gestures; for example, Candidate A used appropriate, open gestures to aid expression. The system also detects involuntary nervous movements such as rubbing hands or scratching ears.
[0044] Step S300: Single-modal scoring.
[0045] After feature extraction, the system converts the multi-dimensional features extracted from each modality into a normalized single-modality score (e.g., a 0-100 score) according to preset scoring rules. For example, the text modality score is a weighted average of multiple features, including keyword matching (50% weight), logical coherence (30% weight), and semantic completeness (20% weight). Candidate A's answer is logically clear and comprehensive, earning a text score of 88. Correspondingly, the speech modality score is a weighted result of features such as moderate speaking speed (30% weight), confident tone (30% weight), positive emotional expression (20% weight), and fluency of pauses (20% weight). Candidate A's speech is engaging and fluent, earning a speech score of 92. The video modality score is a weighted result of features such as positive facial expressions (40% weight), quality of eye contact (40% weight), and appropriate posture (20% weight). Candidate A maintained a smile throughout and made ample eye contact; their video received a score of 90.
[0046] Step S400: Multimodal fusion weight configuration.
[0047] At this point, the system enters the fusion scoring phase, which is executed by the fusion scoring engine 40. Please refer to [link / reference]. Figure 3 The weight acquisition unit 41 within the fusion scoring engine 40 first queries and loads the multimodal fusion weights for the sales manager position from the job competency model library 30. In this embodiment, the multimodal fusion weights are configured as follows: text weight coefficient α = 0.35, voice weight coefficient β = 0.30, and video weight coefficient γ = 0.35. These weight coefficients satisfy the normalization condition α + β + γ = 1. It is understood that these weights can be manually set by the system administrator based on recruitment experience, or they can be automatically optimized by machine learning algorithms through regression analysis of the interview data of successfully interviewed sales managers in the company's history and their actual performance, to achieve more accurate job matching.
[0048] As a specific implementation method, the process of automatically optimizing weights based on historical interview data and hiring results through regression analysis specifically includes:
[0049] Construct a feature vector with the candidate's historical text score, voice score, and video score as input. The target variable is the candidate's actual performance appraisal score after joining the company. The supervised learning dataset is used. The model is trained using the Ridge Regression algorithm with L2 regularization, and its loss function is defined as:
[0050]
[0051] in, That is, the coefficient vector of the multimodal fusion weights to be optimized (corresponding to...) , , ), This is a regularization parameter to prevent overfitting. The system iteratively updates the parameters using gradient descent. Until the loss function converges, then the converged loss function is analyzed. Normalization is performed, and the processed values are saved as the automatically optimized multimodal fusion weights for the target job in the job competency model library.
[0052] Step S500: Conflict and Collaboration Handling.
[0053] The scoring receiving unit 42 receives three single-modal scores from step S300: a text score of 88 points, a voice score of 92 points, and a video score of 90 points.
[0054] Subsequently, the conflict coordination judgment unit 43 analyzes the scores of this group. The conflict coordination judgment unit 43 first calculates statistical indicators for these three scores, such as the standard deviation. For the scoring group (88, 92, 90), the standard deviation is approximately 1.63. The system has a preset consistency threshold, for example, 5.0. Since 1.63 is much smaller than 5.0, it indicates that candidate A's performance in the three modalities of content, expression, and image is highly consistent and at an excellent level.
[0055] Based on this judgment, the system triggers the cooperation enhancement mechanism. The conflict cooperation judgment unit 43 calculates the cooperation factor. The cooperation factor can be calculated using the following formula: ,in A preset minimal smoothing constant (e.g., with a value of 0.01) is used to prevent division by zero errors.
[0056] Assuming the first synergy coefficient is 0.003, then the synergy factor is... Specifically, when the scores of each single mode are completely identical (i.e., the standard deviation is 0), the synergy factor is calculated as follows: It can provide a preset penalty or maximum gain cap to prevent algorithm crashes. This positive co-factor will be used to provide a reward-based enhancement to the base weighted score. Afterwards, the scoring calculation unit 44 performs the calculation of the comprehensive score using the following formula: The overall score is calculated as follows: (α×text score + β×voice score + γ×video score)×(1+δ×synergy factor), where δ is a preset synergy enhancement adjustment coefficient used to control the intensity of synergy enhancement. In this embodiment, it is set to 0.5.
[0057] Substitute all values into the formula: Overall Score =(0.35×88+0.30×92+0.35×90)×(1+0.5×0.1)=(30.8+27.6+31.5)×(1+0.05)=89.9×1.05≈94.4 points.
[0058] As can be seen, due to the triggering of the synergy enhancement mechanism, candidate A's final score increased from the base of 89.9 to 94.4, which effectively distinguished him from candidates who only excelled in one aspect but had mediocre overall performance (whose synergy factor would approach 0).
[0059] Step S600: Evaluation report generated.
[0060] In the final stage of the evaluation process, the report generation module 50 generates a highly interpretable evaluation report based on the comprehensive score of 94.4 points output by the scoring calculation unit 44, and all intermediate analysis data generated in the preceding steps (such as scores for each single modality, extracted feature values, etc.). This evaluation report highly recommends Candidate A and automatically marks the video clips of their fluent answers and confident expressions as highlights for quick review by human resource managers.
[0061] Through this embodiment, the method and system of the present invention can identify candidates who perform well in all aspects and give them higher evaluations through a collaborative enhancement mechanism, thereby improving the accuracy and discriminativeness of the evaluation.
[0062] Example 2
[0063] This embodiment aims to illustrate the response mechanism of the AI interview evaluation method based on multimodal adaptive fusion provided by the present invention when there are significant conflicts in the performance of candidates across different modalities. This embodiment follows the scenario of recruiting a sales manager in Embodiment 1, but the evaluation subject is another candidate (Candidate B).
[0064] The initial steps S100, S200, and S300 of the method are similar to those described in Example 1. However, for candidate B, after analyzing their performance in answering the question "How do you view customer rejection?", the system yielded a drastically different unimodal scoring result:
[0065] Text score: 86 points. Candidate B's response was well-prepared, with a reasonable logical structure, and covered most of the key points, thus earning a high score in the text modality.
[0066] Voice score: 65 points. Voice analysis shows that Candidate B speaks too fast, reaching approximately 300 words per minute, far exceeding the normal communication range; at the same time, their volume is too low, their tone is flat, and they lack persuasiveness; the voice emotion recognition model detected a consistently high probability of their nervousness.
[0067] Video score: 70 points. Video analysis revealed that Candidate B frequently avoided eye contact during the answer, dared not look directly at the camera, and kept his gaze fixed on the corner of the screen for a long time; in addition, the system also detected several unconscious hand-rubbing movements, which are typical signs of nervousness and lack of confidence.
[0068] After entering step S500, the score receiving unit 42 of the fusion scoring engine 40 receives this set of scores with huge differences: text score 86 points, voice score 65 points, and video score 70 points.
[0069] The conflict resolution unit 43 then analyzes the scores of the group. This unit first calculates the difference between the scores. For example, the absolute value of the difference between the text score and the speech score is |86-65|=21. The system has a preset difference threshold, for example, 15. Since 21 is greater than 15, the system determines that there is a significant conflict between the evaluation results of different modalities.
[0070] According to a preferred embodiment of the present invention, when such a conflict is detected, the system will trigger a conflict resolution mechanism. The core of this mechanism lies in dynamically adjusting the multimodal fusion weights based on the core competency requirements of the target position. For the sales manager position, the competency model emphasizes on-the-spot communication skills and a confident professional image, primarily manifested through voice and video modalities. Therefore, although candidate B's textual content (representing their theoretical knowledge) is good, their nervousness and lack of confidence in actual expression are weaknesses that require more attention for this position.
[0071] The conflict resolution mechanism then makes temporary adjustments to the initial fusion weights accordingly.
[0072] For example, the original text weight α=0.35 can be reduced to α'=0.25, while the voice weight β=0.30 can be increased to β'=0.40. The video weight γ=0.35 can remain unchanged or be increased accordingly (in this example, γ'=0.35 to ensure the sum is 1). The purpose of this adjustment is to amplify the candidate's deficiencies in key competency dimensions to obtain evaluation results that are more in line with job requirements.
[0073] Meanwhile, due to the inconsistent scores of each modality, the standard deviation calculated by the conflict coordination judgment unit 43 is very large, far exceeding the consistency threshold. Therefore, the coordination factor is calculated as 0 or a negligible minimum value, and the coordination enhancement mechanism is not triggered.
[0074] Subsequently, the scoring calculation unit 44 uses the adjusted weighted synergy factor, which sums to 0, to calculate the overall score: Overall score = (α'×text score + β'×voice score + γ'×video score)×(1+δ×0) = (0.25×86+0.40×65+0.35×70)×1 = 21.5+26+24.5 = 72 points.
[0075] Ultimately, Candidate B's overall score was 72. This score was far below what his seemingly good written score should have indicated, but it accurately reflected a significant potential risk for him as a sales manager candidate: poor on-the-spot communication skills.
[0076] In step S600, the report generated by the report generation module 50 will not only present a comprehensive score of 72 points, but will also point out in the comprehensive evaluation summary that the candidate's content preparation was sufficient, but the candidate was nervous during the on-site performance, which posed a communication risk. Based on the analysis process, the report will automatically mark the segments in the video that show the candidate avoiding eye contact, speaking too fast, and rubbing their hands as risk moments, providing human resource managers with a strong and intuitive basis for judgment.
[0077] This embodiment demonstrates that, through the conflict resolution mechanism, the present invention can intelligently identify and handle information contradictions between modalities and make adaptive adjustments according to job characteristics. This avoids the fatal flaws in other key dimensions being masked by excellent performance in a single dimension (such as text content), and significantly improves the robustness and practical guiding significance of the evaluation.
[0078] Example 3
[0079] This embodiment demonstrates how the AI interview assessment method provided by the present invention can accurately assess the core competencies of different positions by configuring different multimodal fusion weights, thereby reflecting its high job adaptability.
[0080] In this scenario, the same company uses this system to recruit for a technical R&D position. Unlike sales positions, technical R&D positions place greater emphasis on candidates' technical depth, logical thinking, and problem-solving abilities, which are primarily reflected in the professionalism and rigor of their answers; while fluency in oral expression and on-camera performance are relatively less important.
[0081] Therefore, in step S400, the system administrator, or an algorithm optimized using historical data, configures a completely different set of fusion weights for the technical R&D position in the job competency model library 30: text weight α=0.50, voice weight β=0.25, and video weight γ=0.25. This configuration explicitly places the focus of the evaluation on the text modality.
[0082] A candidate (Candidate C) participated in the interview for this technical position and answered a technical question: Please explain your understanding of microservice architecture.
[0083] The system also executed steps S100 to S300 to analyze and score candidate C's performance using a single modality:
[0084] Text score: 95 points. After candidate C's answer was converted into text by speech recognition, the text analysis sub-engine found that its content was extremely in-depth, systematically covering multiple core knowledge points such as service decomposition principles, inter-service communication mechanisms, data consistency solutions, and service governance, and accurately using a large number of professional terms. Therefore, the text modality received a very high score.
[0085] Voice score: 75 points. Voice analysis shows that Candidate C's speaking speed is relatively slow, and there were several pauses exceeding 3 seconds during the answer. It should be noted that this is not due to nervousness or hesitation, but rather a typical manifestation of deep thinking. The tone is flat, lacking the expressive intonation of a salesperson. Therefore, from the perspective of pure fluency, the voice score is not high.
[0086] Video score: 80 points. Video analysis shows that Candidate C appeared serious and focused while answering questions, occasionally frowning in thought. His gaze was mostly not directed at the camera, but rather focused on one side of the screen, possibly indicating he was organizing his thoughts or reviewing prepared notes. From a professional camera presence perspective, his performance was adequate.
[0087] In step S500, the fusion scoring engine 40 receives the following scores: 95, 75, and 80.
[0088] After analysis, the conflict coordination judgment unit 43 found that although there were certain differences in the scores of each modality, there was no serious negative conflict similar to that in Example 2 (for example, no strong tension or avoidance behavior was detected).
[0089] Therefore, the conflict resolution mechanism was not triggered, and the weights remained at the initial settings of α=0.50, β=0.25, and γ=0.25. Meanwhile, due to the large score differences, the synergy factor was also calculated as 0.
[0090] Scoring calculation unit 44 calculates the overall score based on this:
[0091] Overall score = (0.50×95+0.25×75+0.25×80)×1=47.5+18.75+20=86.25 points.
[0092] Ultimately, Candidate C achieved a high score of 86.25. This result clearly demonstrates that although Candidate C's performance in fluency and on-screen presence was mediocre, the system, by giving high weight to the evaluation of text content, still accurately identified his core value as a technical professional—profound professional knowledge and rigorous logical thinking.
[0093] In step S600, the evaluation report will highlight the candidate's high score of 95 points for text content and emphasize in the overall comments that the candidate has a deep and systematic understanding of microservice architecture, rigorous logic, and may objectively describe their expression style as typical of a technical expert—composed and thoughtful. This report provides human resource managers with a judgment criterion that is perfectly aligned with the characteristics of this position, avoiding the mistake of missing out on an excellent technical expert by using the standards for evaluating sales personnel.
[0094] This embodiment fully demonstrates that by flexibly configuring multimodal fusion weights, the present invention can achieve adaptive and accurate assessment of the core competency requirements of different positions, thereby achieving a better person-job matching effect.
[0095] Example 4
[0096] This embodiment will elaborate on the specific implementation of step S600 (evaluation report generation), especially the composition and presentation of the deeply interpretable report. Using candidate B (final score 72 points) from Embodiment 2, which exhibited conflicting performance, this embodiment illustrates how the system transforms abstract scores into intuitive, evidence-supported decision-making support information.
[0097] After the fusion scoring engine 40 completes its evaluation of candidate B, it transmits the overall score, individual modality scores, all extracted raw feature data, and conflict resolution decision records to the report generation module 50. Based on this information, the report generation module 50 generates a structured evaluation report, which can be in electronic document format or an interactive webpage.
[0098] Please see Figure 4 This is a schematic diagram of the interface of a deeply interpretable evaluation report provided in an embodiment of the present invention. The report typically includes the following core components:
[0099] Overall Score Section 61: At the very top or most prominent position in the evaluation report is the overall score section 61. This area typically displays candidate B's basic information (such as photo and name), the "Sales Manager" position applied for, and the final overall score of "72 points." A brief evaluation conclusion may be attached, such as "Further investigation recommended" or "Significant communication risks exist." This provides decision-makers with the most direct first impression.
[0100] Competency Radar Chart 62: Following this, the evaluation report will present a multi-dimensional competency radar chart 62. This multi-dimensional competency radar chart 62 visually displays the score distribution of candidate B across multiple sub-competency items. These competency items are aggregated based on the previously extracted features and single-modal scores. For example, the five vertices of a pentagram radar chart can represent: content depth (mainly from text score, 86 points), logic (from text analysis, 85 points), language expression (mainly from voice score, 65 points), emotional state (from sentiment analysis of voice and video, 68 points), and professional image (mainly from video score, 70 points). Decision-makers can clearly see from this chart that candidate B performs reasonably well in the content and logic dimensions, but has significant weaknesses in language expression, emotional state, and professional image, resulting in a clear imbalance in the chart.
[0101] Multimodal Analysis Details Section 63: This is the most interpretable part of the assessment report, presenting the data evidence behind the scores in graphical form. This area is typically organized in tabs or collapsible panels, containing detailed analyses of each modality:
[0102] Voice Analysis Module: This module displays a speech rate variation graph, clearly showing that candidate B's speech rate rapidly increased from an initial 150 words / minute to a peak of 300 words / minute during the question-and-answer session, accompanied by a label indicating excessively fast speech. Additionally, it displays a speech emotion temporal distribution graph, showing that the probability of nervousness remained consistently above 70% for most of the interview.
[0103] Video Analysis Module: This module displays a gaze-focus heatmap. This heatmap shows the distribution of the candidate's gaze by overlaying a heat map over their face. The heatmap shows that over 90% of the gaze is focused on the lower right corner of the screen, rather than the camera area representing the person being interacted with, and is marked as eye avoidance. There may also be a bar chart or counter showing that the algorithm recognized and recorded the subtle hand-rubbing gesture 15 times during the interview.
[0104] Text Analysis Module: This module displays a keyword cloud, showing frequently occurring words from the candidate's answers at different sizes. Words like "correct," "but," and "possibly" are highlighted, potentially reflecting their thought process or common phrases. It also lists the match between their answers and the key points of the standard answer.
[0105] Key Video Segment Section 64: This section of the evaluation report provides the most direct evidence. The report generation module 50 automatically marks and edits the complete interview video timeline based on multimodal analysis details. For example, the system automatically marks the time period "01:15-01:30" as a "risk moment: speaking too fast" based on the rule of "speech rate exceeding 250 words / minute". Based on the rule of "eyes leaving the camera area for more than 3 seconds", "02:05-02:10" is marked as a "risk moment: avoiding eye contact". In Key Video Segment Section 64, the left side of the interface may contain an embedded video player, while the right side contains two lists: "Highlight Moments" and "Risk Moments". The "Risk Moments" list clearly lists the marked items. When the decision-maker clicks on the "01:15-01:30 speaking too fast" item, the video player on the left automatically jumps to 1 minute and 15 seconds to start playing that 15-second video. Decision-makers can hear the candidate's rapid speech and see his nervous expression, thus forming a deep and emotional connection to the assessment result of "72 points".
[0106] Through this in-depth, interpretable report that includes macro-level scoring, multi-dimensional diagnostics, micro-level data, and video evidence, HR managers can gain a clear, intuitive, and evidence-based understanding of candidate B's strengths and fatal weaknesses within minutes, without having to watch a full interview recording lasting tens of minutes. This allows them to quickly determine that "this candidate is not suitable for a sales position requiring strong communication skills." This embodiment fully demonstrates the significant benefits of this invention in improving the interpretability and credibility of evaluation results, thereby enhancing decision-making efficiency.
[0107] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. An AI interview assessment method based on multimodal adaptive fusion, characterized in that, Includes the following steps: Calculate the candidate's unimodal scores for text, voice, and video during the interview process; Based on a pre-defined job competency model library, multimodal fusion weights are configured for the target job. Based on the multimodal fusion weights and the single-modal scores, the candidate's comprehensive score is calculated. When the standard deviation among multiple single-modal scores is less than a preset consistency threshold, a synergy factor is introduced to enhance the calculation of the comprehensive score. The synergy factor is equal to the product of a first synergy coefficient and the reciprocal of a target value. The target value is the sum of the standard deviation and a preset smoothing constant, and the preset smoothing constant is greater than zero. An evaluation report is generated, which includes multimodal analysis details and key segments of the interview video that are automatically tagged based on the unimodal score, the composite score, and the multimodal analysis details.
2. The AI interview assessment method based on multimodal adaptive fusion according to claim 1, characterized in that, The overall score is calculated using the following formula: Overall score = (α × text score + β × voice score + γ × video score) × (1 + δ × co-factor) Wherein, α, β, and γ are the weight coefficients in the multimodal fusion weights corresponding to the text score, speech score, and video score, respectively, and δ is a preset collaborative enhancement adjustment coefficient.
3. The AI interview assessment method based on multimodal adaptive fusion according to claim 1, characterized in that, Also includes: When the absolute value of the difference between any two scores in the multiple single-modal scores is greater than a preset difference threshold, the conflict handling mechanism is triggered. The conflict resolution mechanism specifically includes: identifying the target single modality with the lowest score among multiple single modality scores, increasing the weight coefficient corresponding to the target single modality by a preset first step length, and reducing the weight coefficients of the remaining single modalities by a preset proportion, so as to amplify the evaluation weight of the candidate in the weakness dimension.
4. The AI interview assessment method based on multimodal adaptive fusion according to claim 1, characterized in that, The key segments of the interview video include highlight moments and risk moments determined according to preset rules.
5. The AI interview assessment method based on multimodal adaptive fusion according to claim 1, characterized in that, The details of the multimodal analysis include at least one selected from the following: Capability radar chart, speech rate change curve, speech emotion temporal distribution map, keyword cloud map, or gaze point heat map.
6. The AI interview assessment method based on multimodal adaptive fusion according to claim 1, characterized in that, The multimodal fusion weights can be manually adjusted or automatically optimized through regression analysis of historical interview data and hiring results.
7. The AI interview assessment method based on multimodal adaptive fusion according to claim 1, characterized in that, The single-modal score is calculated based on at least one video feature extracted from the video data stream, selected from the following: Facial expression features, eye tracking features, head posture features, or body movement features.
8. The AI interview assessment method based on multimodal adaptive fusion according to claim 1, characterized in that, The single-modal score is calculated based on at least one speech feature selected from the following, extracted from the audio data stream: Speech rate, volume, fundamental frequency, pauses, or emotional characteristics of speech.
9. The AI interview assessment method based on multimodal adaptive fusion according to claim 1, characterized in that, The unimodal score is calculated based on at least one text feature extracted from the text data stream, selected from the following: Content keyword matching accuracy, semantic completeness, logical coherence, and the use of technical terms.
10. An AI interview assessment system based on multimodal adaptive fusion, characterized in that, include: A single-modal score calculation module is used to calculate the single-modal scores of candidates for text, voice and video during the interview process; A weight configuration module is used to configure multimodal fusion weights for target positions based on a preset job competency model library; A comprehensive score calculation module is used to calculate the comprehensive score of a candidate based on the multimodal fusion weights and the single-modal scores. When the standard deviation among multiple single-modal scores is less than a preset consistency threshold, a synergy factor is introduced to enhance the calculation of the comprehensive score. The synergy factor is equal to the product of a first synergy coefficient and the reciprocal of a target value. The target value is the sum of the standard deviation and a preset smoothing constant. The preset smoothing constant is greater than zero. A report generation module is used to generate an evaluation report, which includes multimodal analysis details and key segments of the interview video automatically marked based on the unimodal score, the comprehensive score, and the multimodal analysis details.