Multi-modal fusion recruitment interview method and device based on AI and medium
By collecting and processing multimodal data, and combining reinforcement learning and debiasing algorithms, the problems of insufficient multimodal analysis and high interaction latency in AI interview systems have been solved, enabling dynamic scoring and rapid evaluation of interviewees.
Patent Information
- Application Number
- CN202510911245.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-21
AI Technical Summary
Existing AI interview systems suffer from insufficient multimodal data analysis capabilities, difficulty in eliminating algorithmic biases, and high real-time interaction latency, especially when dealing with interviewees' micro-expressions and complex scenarios.
By using multiple sensors to collect multimodal data from interviewees, extracting text, voice, and video features, and combining reinforcement learning-based scoring weight optimization and debiasing algorithms, the scoring weights are adjusted in real time. Edge computing is then used to integrate the dynamic scoring and debiasing results to generate an interview evaluation report.
It improves multimodal analysis capabilities, reduces algorithmic bias, lowers interaction latency, and enables dynamic scoring and rapid evaluation of interviewees.
Smart Images

Figure CN120822932A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and specifically relates to an AI-based multimodal fusion recruitment interview method, device and medium. Background Art
[0002] Global companies now spend over a trillion dollars annually on recruiting, but traditional recruitment processes suffer from inefficiencies, high costs, and subjectivity. To alleviate these pain points in HR recruitment, AI interview systems have been introduced.
[0003] Currently, the market size of AI interview systems is growing at an annual rate of 35%. Core requirements include automated assessment, multilingual support, and global deployment. In the early days of AI interview systems, most were resume screening tools based on keyword matching, capable only of processing structured data. Today, AI interview systems incorporate natural language processing (NLP) and speech recognition technologies, enabling semantic analysis of language expressions, reducing labor costs, and automating the scoring of interview questions and answers.
[0004] However, existing AI interview systems suffer from insufficient multimodal data analysis capabilities, difficulty eradicating algorithmic bias, and high real-time interaction latency. Specifically, they struggle with multimodal data analysis, often overlooking interviewees' micro-expressions. Algorithmic bias is difficult to eradicate, such as historical biases implicit in training data (such as gender and age discrimination), leading to unfair algorithmic decision-making. Furthermore, they lack dynamic optimization mechanisms. Furthermore, they perform poorly in real-time interaction, multilingual support, and complex scenarios (such as evaluating programming skills for technical positions), resulting in excessive interaction latency. Summary of the Invention
[0005] This application aims to provide an AI-based multimodal fusion recruitment interview solution. It aims to address the problems of insufficient multimodal data analysis capabilities, algorithmic bias, and excessive interaction delay in existing technologies.
[0006] According to the first aspect of the present application, the present application provides an AI-based multimodal fusion recruitment interview method, comprising:
[0007] Use multiple sensors to collect multimodal data from interviewees;
[0008] Extracting corresponding multimodal features from the multimodal data according to a multimodal feature extraction algorithm corresponding to the data type of the multimodal data;
[0009] According to the reinforcement learning-based scoring weight optimization algorithm, the scoring weights of multimodal features are adjusted in real time. The scoring weights are used to fuse multimodal features to obtain dynamic scores for interviewees.
[0010] Use a debiasing algorithm to debias the interviewee's personal information and obtain a debiasing result;
[0011] Edge computing is used to integrate dynamic scoring and de-biasing processing results to generate interview evaluation reports for interviewees.
[0012] Preferably, in the above-mentioned multimodal fusion recruitment interview method, the step of extracting corresponding multimodal features from the multimodal data according to a multimodal feature extraction algorithm corresponding to the data type of the multimodal data includes:
[0013] Use multiple sensors to collect multimodal data from interviewees, where the multimodal data includes text data, voice data, and video data;
[0014] Use semantic models to extract text features from text data;
[0015] Use visual recognition technology to extract frames and mark key points of video data to obtain video features;
[0016] Using speech processing technology to reduce noise and extract features from speech data to obtain speech features;
[0017] Use a multimodal fusion architecture to fuse text features, video features, and speech features.
[0018] Preferably, in the above-mentioned multimodal fusion recruitment interview method, the step of adjusting the scoring weights of the multimodal features in real time according to the scoring weight optimization algorithm based on reinforcement learning includes:
[0019] Using a dynamic evaluation model, we optimize the scoring weights using reinforcement learning:
[0020] Q(s,a)=Q(s,a)+α*[r+γ*max(Q(s',a'))-Q(s,a)]
[0021] Update the scoring weights of the multimodal features; where s = current scoring state, a = weight adjustment action, and r = manual feedback correction reward;
[0022] At predetermined intervals, the scoring weights of the multimodal features are re-optimized according to the scoring weight optimization algorithm.
[0023] Preferably, in the above-mentioned multimodal fusion recruitment interview method, the step of using the scoring weight to fuse the multimodal features to obtain the dynamic score of the interviewee includes:
[0024] The scoring weights are fused with the corresponding modal features respectively;
[0025] According to the similarity calculation model, calculate the feature similarity between the modal features after fusion of the scoring weights and the corresponding features of the job capabilities;
[0026] Generate a capability radar chart based on the relationship between job capabilities and modal features after integrating scoring weights;
[0027] Calculate the dynamic score of the interviewee according to the feature similarity and ability radar chart.
[0028] Preferably, in the above-mentioned multimodal fusion recruitment interview method, before the step of adjusting the scoring weights of the multimodal features in real time according to the scoring weight optimization algorithm based on reinforcement learning, and using the scoring weights to fuse the multimodal features to obtain the dynamic score of the interviewee, the method further includes:
[0029] Input the job requirements of the interview position and the interviewee's multimodal features into the AI large language model;
[0030] According to the job requirements, the AI language model is controlled to generate follow-up interview questions based on the interviewee's multimodal characteristics;
[0031] Receive new multimodal data generated by interviewees’ answers to follow-up questions;
[0032] Use the multimodal feature extraction algorithm to extract the multimodal features corresponding to the new multimodal data.
[0033] Preferably, in the above-mentioned multimodal fusion recruitment interview method, the step of using a debiasing algorithm to debias the interviewee's personal information to obtain a debiasing result includes:
[0034] Use the synthetic data enhancement mechanism to enhance the personal information and obtain the debiased processing results;
[0035] Alternatively, an adversarial training model can be used to jointly optimize personal information in combination with a bias discriminator.
[0036] Preferably, the multimodal fusion recruitment interview method uses edge computing to integrate dynamic scoring and debiasing processing results to generate an interview evaluation report for the interviewee, including the following steps:
[0037] Using edge computing nodes, extracting corresponding multimodal features from multimodal data and compressing and uploading the multimodal features to the cloud platform;
[0038] Using the cloud platform, the scoring weights are integrated with multimodal features to obtain dynamic scores for interviewees.
[0039] Using the cloud platform, the debiasing processing results are obtained according to the debiasing algorithm, and the dynamic scoring and debiasing processing results are integrated to obtain the interview evaluation report of the interviewee.
[0040] Preferably, the multimodal fusion recruitment interview method further comprises, after the step of using a debiasing algorithm to debias the interviewee's personal information to obtain a debiasing result:
[0041] Upload multimodal features, interviewees’ dynamic scores, and debiasing processing results to the real-time evaluation panel;
[0042] Use a real-time evaluation dashboard to display multimodal features, dynamic scoring, and debiasing results;
[0043] Obtain a scoring weight adjustment signal, and readjust the scoring weight of the multimodal feature in an edge computing manner according to the scoring weight adjustment signal.
[0044] According to the second aspect of the present application, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, it implements the AI-based multimodal fusion recruitment interview method provided by any of the above technical solutions.
[0045] According to the third aspect of the present application, the present application also provides a computer storage medium on which a computer program is stored. When the computer program is executed, the AI-based multimodal fusion recruitment interview method provided by any of the above technical solutions is implemented.
[0046] The technical solution of this application has at least the following technical effects:
[0047] The AI-based multimodal fusion recruitment interview solution provided in this application collects the multimodal data of the interviewee through the use of multiple sensors, and then extracts the multimodal features according to the multimodal feature extraction algorithm corresponding to the type of the multimodal data, thereby improving the multimodal analysis capability. After extracting the multimodal features, the weight score of the multimodal features is adjusted in real time according to the scoring weight optimization algorithm based on reinforcement learning, and the multimodal weight score is integrated with the multimodal feature to improve the multimodal analysis capability, obtain the dynamic score of the interviewee, and then use the debiasing algorithm to debias the personal information of the interviewee, thereby removing the bias against the interviewee, and finally using edge computing to integrate the dynamic score and debiasing processing results, so as to quickly obtain the interviewee's interview evaluation report, reduce the interaction delay rate, and improve the interaction efficiency. In summary, the technical solution of this application can solve the problems of insufficient multimodal analysis capability, algorithm bias and high real-time interaction delay in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0049] Figure 1 A schematic diagram of the structure of an intelligent interview system provided in an embodiment of the present application;
[0050] Figure 2 A flowchart of an AI-based multimodal fusion recruitment interview method provided in an embodiment of the present application;
[0051] Figure 3 A schematic diagram of a candidate-side process provided in an embodiment of the present application;
[0052] Figure 4 A schematic diagram of a server-side process provided in an embodiment of the present application;
[0053] Figure 5 A schematic diagram of an interviewer-side process provided in an embodiment of the present application;
[0054] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] In order to more clearly illustrate the overall concept of the present application, a detailed description is given below in an illustrative manner in conjunction with the accompanying drawings.
[0056] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application may also be implemented in other ways than those described herein, and therefore, the scope of protection of the present application is not limited by the specific embodiments disclosed below. It should be noted that the embodiments of the present application and the features of each embodiment may be combined with each other unless there is a conflict.
[0057] In this application, unless otherwise clearly specified and limited, a first feature "above" or "below" a second feature may be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in an appropriate manner in any one or more embodiments or examples.
[0058] The existing technology has the following defects:
[0059] Existing AI interview systems have the following problems:
[0060] (1) Subjectivity and inefficiency: Traditional interviews rely on manual evaluation, are susceptible to interviewer bias, and are inefficient in processing massive resumes and interview feedback.
[0061] (2) Single evaluation dimension: Existing AI interview systems mostly focus on single-dimensional analysis of text or voice, and lack in-depth mining of candidates' non-verbal information (such as facial expressions and body language).
[0062] (3) Algorithmic bias and lack of fairness: Historical biases implicit in training data (such as gender and age discrimination) lead to unfair algorithmic decisions and lack of dynamic optimization mechanisms.
[0063] This results in the existing AI interview system having insufficient multimodal data analysis capabilities, algorithmic bias and excessive interaction delays.
[0064] In order to overcome the above technical difficulties, the following embodiments of this application provide an AI-based multimodal fusion recruitment interview solution, which aims to obtain dynamic scores by fusing multimodal features and their scoring weights, use a debiasing algorithm to debias user personal information, and implement the above operations in an edge computing manner, thereby achieving the purpose of improving multimodal data analysis capabilities, reducing algorithm bias against interviewees and improving interaction efficiency.
[0065] To achieve the above purpose, see Figure 1 , Figure 1 An overall architecture diagram of an intelligent interview system provided for an embodiment of the present application. The intelligent interview system is used to implement the AI-based multimodal fusion recruitment interview solution of the following embodiment of the present application. The intelligent interview system includes a data acquisition layer, an algorithm model layer, and an interactive feedback layer; wherein, the data acquisition layer includes hardware devices including various sensors, a data preprocessing module, and an execution subject (here mainly a client device). The data preprocessing module includes processing of video streams and voice signals. The algorithm model layer includes a feature extraction module, a dynamic evaluation module, and a debiasing module. The interactive feedback layer includes a real-time feedback module, a report generation module, and an execution subject, and the execution subject here is the server of the cloud platform.
[0066] To achieve the above purpose, see Figure 1 , Figure 1 A flowchart of a multimodal fusion recruitment interview method based on AI provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, according to the first aspect of the present application, the present application provides an AI-based multimodal fusion recruitment interview method, comprising:
[0067] S110: Use multiple sensors to collect multimodal data of interviewees. Figure 1As shown in the intelligent interview system architecture diagram, the various sensors provided in the embodiment of the present application include cameras (to capture the candidate's facial expressions and body language), microphones (voice input) and client devices (PC / mobile). These sensors and client devices need to be set up at the recruitment interview site.
[0068] S120: extracting corresponding multimodal features from the multimodal data according to a multimodal feature extraction algorithm corresponding to the data type of the multimodal data.
[0069] Specifically, as a preferred embodiment, Figure 2 In the multimodal fusion recruitment interview method shown, step S120: extracting corresponding multimodal features from the multimodal data according to a multimodal feature extraction algorithm corresponding to the data type of the multimodal data includes:
[0070] S121: Use multiple sensors to collect multimodal data of the interviewee, where the multimodal data includes text data, voice data, and video data.
[0071] S122: Extract text features of the text data using a semantic model. The semantic model here can be a GPT or LSTM semantic matching model.
[0072] S123: Using visual recognition technology to extract frames and mark key points of the video data to obtain video features.
[0073] S124: Using speech processing technology, perform noise reduction and feature extraction on the speech data to obtain speech features.
[0074] S125: Use a multimodal fusion architecture to fuse text features, video features, and speech features.
[0075] The technical solution provided in the embodiment of the present application executes the processing of video streams and voice signals on the client device. Specifically, for the video stream: frame extraction (30 frames per second) and key point annotation (such as 68 facial feature points) are performed through OpenCV. Voice signal: MFCC (Mel-Frequency Cepstral Coefficient) is used for noise reduction and feature extraction. The execution subject here is the client device on the candidate's side, that is, the interviewer's side. In addition, the multimodal feature extraction algorithm used specifically uses the Transformer architecture to fuse text (BERT), voice (Wav2Vec 2.0) and video (3D-CNN) features.
[0076] Figure 2 The technical solution shown, after extracting corresponding multimodal features from multimodal data, further includes:
[0077] S130: According to the scoring weight optimization algorithm based on reinforcement learning, the scoring weight of the multimodal features is adjusted in real time, and the multimodal features are fused using the scoring weight to obtain a dynamic score of the interviewee.
[0078] Multimodal fusion can include the following methods:
[0079] Early fusion: Fusion at the raw data level (such as concatenating speech MFCC with video optical flow features), replacing the late fusion of Transformer.
[0080] Applicable scenarios: When computing resources are limited, a lightweight model can be used (such as MobileNet instead of 3D-CNN).
[0081] This embodiment of the application uses reinforcement learning-based scoring weight optimization (e.g., Q-Learning) to adjust the scoring weights of multimodal features, such as language logic, code quality, and emotional stability, in real time. Specifically, the scoring weights are adjusted for language logic (weight 0.4), code quality (weight 0.3), and emotional stability (weight 0.3).
[0082] After obtaining the scoring weight, the scoring weight is multiplied by the multimodal feature to obtain the dynamic score of the interviewee.
[0083] Specifically, as a preferred embodiment, in the above-mentioned multimodal fusion recruitment interview method, step S130: adjusting the scoring weights of the multimodal features in real time according to the reinforcement learning-based scoring weight optimization algorithm includes:
[0084] S131: Using a dynamic evaluation model, we optimize the scoring weights based on reinforcement learning:
[0085] Q(s,a)=Q(s,a)+α*[r+γ*max(Q(s',a'))-Q(s,a)]
[0086] Update the scoring weights of the multimodal features. Here, s = the current scoring state, a = the weight adjustment action, and r = the reward for manual feedback correction. Q(s,a) represents the value Q of selecting dynamic a (i.e., the multimodal feature) under the current scoring state s. α represents the learning rate, ranging from 0 to 1. γ represents the discount factor, which measures the current value of future drops. max(Q(s',a')) represents the maximum Q value among all possible dynamics preferred for the next scoring state.
[0087] S132: Re-optimize the score weights of the multimodal features according to the score weight optimization algorithm at predetermined intervals. Specifically, perform abnormal weight allocation every 30 seconds.
[0088] In the technical solution provided in the embodiments of the present application, Q-Learning is a classic model-free reinforcement learning algorithm that can be used to learn the optimal strategy for multimodal features, thereby obtaining the scoring weight corresponding to the multimodal features under the current scoring state.
[0089] In addition, as a preferred embodiment, in the above multimodal fusion recruitment interview method, step S130: using the scoring weight to fuse multimodal features to obtain the dynamic score of the interviewee includes:
[0090] S133: Fusing the scoring weights with the corresponding modality features respectively;
[0091] S134: Calculate the feature similarity between the modal feature after fusion of the scoring weights and the feature corresponding to the job capability according to the similarity calculation model;
[0092] S135: Generate a capability radar chart based on the relationship between job capabilities and modal features after integrating scoring weights;
[0093] S136: Calculate the dynamic score of the interviewee according to the feature similarity and ability radar chart.
[0094] In the embodiment of the present application, after the weight of the equal distribution of the prize is fused with the corresponding modal feature, the modal feature will have a corresponding value, so that the modal feature and the corresponding feature of the job ability can be used to perform similarity evaluation. For example, it is determined which modal features are included in the features required for the job, and then what the score of these features is. In this way, the feature similarity can be calculated using the above score according to the similarity calculation formula; and according to the relationship between the job ability and the modal feature, a radar chart of the ability is generated, including the score of the corresponding feature of the job ability. In this way, the dynamic score of the interviewee can be calculated by combining the feature similarity and the ability radar chart, so as to achieve a comprehensive evaluation of the interviewee.
[0095] In addition, as a preferred embodiment, in the above-mentioned multimodal fusion recruitment interview method, before the step S130 of adjusting the scoring weights of the multimodal features in real time according to the reinforcement learning-based scoring weight optimization algorithm, and using the scoring weights to fuse the multimodal features to obtain a dynamic score of the interviewee, the method further includes:
[0096] Input the job requirement information of the interview position and the multimodal features of the interviewee into the AI large language model.
[0097] According to the job requirements, the AI large language model is controlled to generate follow-up interview questions based on the interviewee's multimodal characteristics.
[0098] Receive new multimodal data generated by interviewees' answers to follow-up interview questions.
[0099] Use the multimodal feature extraction algorithm to extract the multimodal features corresponding to the new multimodal data.
[0100] The technical solution provided in the embodiment of this application can use the AI large language model to generate follow-up questions based on the scoring results (such as automatically triggering programming questions for technical positions, please explain the time complexity of the code just now, etc.). Through the "Flower Girl" method, the new multimodal data generated by the interviewee's answers to the questions can be obtained through sensors, and then the multimodal features corresponding to the new multimodal data can be extracted according to the above-mentioned multimodal feature extraction algorithm. The specific method is as above and will not be repeated here.
[0101] Figure 2 The technical solution provided by the illustrated embodiment, after using the scoring weights to fuse multimodal features to obtain the dynamic score of the interviewee, further includes:
[0102] S140: Debias the interviewee's personal information using a debiasing algorithm to obtain a debiasing result. The debiasing algorithm of the embodiment of the present application can use synthetic data augmentation technology (such as GAN to generate gender / age balanced training samples).
[0103] Specifically, as a preferred embodiment, in the above-mentioned multimodal fusion recruitment interview method, step S140: using a debiasing algorithm to debias the interviewee's personal information to obtain a debiasing result includes:
[0104] S141: Use the synthetic data enhancement mechanism to perform synthetic data enhancement on the personal information to obtain a debiasing processing result.
[0105] S142: Use adversarial training models and bias discriminators to jointly optimize personal information.
[0106] The debiasing mechanism of the embodiment of the present application can use a synthetic data augmentation mechanism or alternatively use adversarial training. Specifically, adversarial training is used instead of synthetic data augmentation, and a gender / age discriminator is introduced into the model training for joint optimization.
[0107] The loss function formula of the adversarial training model is described as follows:
[0108] L_total=L_task+λ*L_adv
[0109] Where L_task = scoring task loss, L_adv = adversarial loss (deceiving the discriminator, i.e. the biased discriminator mentioned above), and λ is the weight of the discriminator.
[0110] In addition, as a preferred embodiment, the multimodal fusion recruitment interview method further includes, after the step S140 of using a debiasing algorithm to debias the interviewee's personal information and obtaining the debiasing result, the following steps:
[0111] S143: Uploading the multimodal features, the interviewee's dynamic score, and the debiasing processing results to the real-time evaluation panel;
[0112] S144: Use a real-time evaluation panel to display multimodal features, dynamic scoring, and debiasing results;
[0113] S145: Obtain a scoring weight adjustment signal, and readjust the scoring weight of the multimodal feature in an edge computing manner according to the scoring weight adjustment signal.
[0114] The technical solution provided by the embodiments of the present application allows the interviewer to view the candidate in real time, namely, the interviewer's real-time evaluation panel. This panel also includes emotional heat maps and voice waveforms, which can display the aforementioned multimodal features, dynamic scoring, and debiasing processing results. The interviewer can send a scoring weight adjustment signal, which is then sent from the real-time evaluation panel to the cloud platform based on the scoring weight adjustment signal, i.e., the aforementioned edge computing method, to readjust the scoring weights of the multimodal features.
[0115] in addition, Figure 2 The technical solution provided by the illustrated embodiment further includes, after obtaining the debiasing processing result:
[0116] S150: Use edge computing to integrate dynamic scoring and debiasing processing results to generate interview evaluation reports for interviewees.
[0117] Among them, as a preferred embodiment, in the above-mentioned multimodal fusion recruitment interview method, step S150: using edge computing to integrate dynamic scoring and debiasing processing results to generate an interview evaluation report for the interviewee includes:
[0118] S151: Using the edge computing node, executing the step of extracting corresponding multimodal features from the multimodal data, and compressing the multimodal features and uploading them to the cloud platform;
[0119] S152: Using the cloud platform, the scoring weights are integrated with multimodal features to obtain dynamic scores for the interviewees.
[0120] S153: Using the cloud platform, obtain the debiasing processing result according to the debiasing algorithm, integrate the dynamic scoring and the debiasing processing result, and obtain the interview evaluation report of the interviewee.
[0121] The technical solution of this application adopts the edge computing model as the overall architecture, and its architecture is adjusted as follows:
[0122] Deploy the feature extraction module to edge nodes (such as the candidate's local device) and upload only the compressed feature vectors to the cloud. This approach has been proven to reduce latency (from 500ms to 200ms), making it suitable for areas with unstable networks.
[0123] In summary, the AI-based multimodal fusion recruitment interview method provided in the embodiment of the present application collects the multimodal data of the interviewee by using a variety of sensors, and then extracts the multimodal features according to the multimodal feature extraction algorithm corresponding to the type of the multimodal data, thereby improving the multimodal analysis capability. After extracting the multimodal features, the weight score of the multimodal features is adjusted in real time according to the scoring weight optimization algorithm based on reinforcement learning, and the multimodal weight score is integrated with the multimodal feature, so as to improve the multimodal analysis capability, obtain the dynamic score of the interviewee, and then use the debiasing algorithm to debias the personal information of the interviewee, thereby removing the bias against the interviewee, and finally using the edge computing method to integrate the dynamic score and debiasing processing results, so as to quickly obtain the interviewee's interview evaluation report, reduce the interaction delay rate, and improve the interaction efficiency. In summary, the technical solution of the present application can solve the problems of insufficient multimodal analysis capability, algorithm bias and high real-time interaction delay in the existing technology.
[0124] In addition, the AI-based multimodal fusion recruitment interview method of this application has three core process interactions: candidate-side interview process, server-side interview process, and interviewer-side process. Specifically, the following core process interactions include:
[0125] like Figure 3 As shown, the process on the candidate (i.e., interviewer) side is:
[0126] Step 201: Log in to the system and complete identity verification (liveness detection + document OCR).
[0127] Step 202: Enter the interview interface and authorize camera / microphone permissions.
[0128] Step 203: Answer preset questions (voice + video recording). Programming positions need to complete online coding (integrated with LeetCode API) simultaneously.
[0129] Step 204: Receive a real-time follow-up question (such as "Please explain the time complexity of the code just now") and respond.
[0130] Step 205: The collection result is forwarded to the server.
[0131] like Figure 4 As shown, the server-side process:
[0132] Step 301: Receive multimodal data streams and start parallel processing threads.
[0133] Step 302: Feature extraction and fusion (text semantic analysis → speech emotion recognition → micro-expression classification).
[0134] Step 303: Dynamic scoring (the comprehensive score is updated every 5 seconds).
[0135] Step 304: Trigger exception handling (such as automatically saving progress when the network is disconnected).
[0136] Step 305: Save the scoring results.
[0137] like Figure 5 As shown, the interviewer-side process:
[0138] Step 401: View the candidate's real-time evaluation panel (including emotion heat map and voice waveform graph).
[0139] Step 402: Manually intervene to adjust the scoring weight (e.g., a technical position needs to increase the code weight to 0.5).
[0140] Step 403: Download the final report and mark the interview conclusions.
[0141] In summary, the AI-based multimodal fusion recruitment interview method provided in the above embodiments of the present application can achieve multimodal fusion, dynamic debiasing mechanism and real-time interactive optimization.
[0142] Multimodal fusion: Integrate voice, text, video, and programming test data to build a comprehensive evaluation model;
[0143] Dynamic debiasing mechanism: Balances the training set through synthetic data augmentation technology and optimizes the algorithm with real-time feedback; Real-time interaction optimization: Uses edge computing to reduce latency and support cross-border and multilingual interview scenarios.
[0144] This application lies at the intersection of artificial intelligence technology and human resource management. This multimodal fusion recruitment interview method, based on the design of an intelligent interview system using natural language processing (NLP), speech recognition, computer vision, and multimodal data analysis, can be used to automate, standardize, and objectify candidate competency assessments in recruitment scenarios. Technically, it implements the application of artificial intelligence algorithms in the recruitment process; it can also be combined with 2D / 3D AI virtual humans; implements fusion analysis technology for multimodal data (speech, text, and video); and ultimately constructs a deep learning-based interview behavior modeling and evaluation method.
[0145] In addition, the beneficial effects of the product embodiments provided in the following embodiments of this application are the same as the beneficial effects of the AI-based multimodal fusion recruitment interview method provided in the above embodiments, and the other technical features in the product embodiments are the same as the features disclosed in the above embodiment methods, which will not be repeated here.
[0146] In addition, see Figure 6 The electronic device provided in this application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the AI-based multimodal fusion recruitment interview method provided in any of the above embodiments.
[0147] Reference below Figure 6 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application can include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.
[0148] like Figure 6As shown, the electronic device can include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the electronic device are also stored in RAM 1004. The processing device 1001, ROM 1002, and RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can enable the above-mentioned electronic devices to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a model building device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or provided instead.
[0149] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0150] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any appropriate manner in any one or more embodiments or examples.
[0151] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart can be implemented using a dedicated hardware-based system that performs the specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0153] The modules described in the embodiments of the present application can be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0154] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0155] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included in the protection scope of the present application.
Claims
1. A multimodal fusion recruitment interview method based on AI, characterized by: include: Use multiple sensors to collect multimodal data from interviewees; Extracting corresponding multimodal features from the multimodal data according to a multimodal feature extraction algorithm corresponding to the data type of the multimodal data; According to a scoring weight optimization algorithm based on reinforcement learning, the scoring weights of the multimodal features are adjusted in real time, and the multimodal features are integrated with the scoring weights to obtain a dynamic score of the interviewee; Using a debiasing algorithm to perform debiasing processing on the personal information of the interviewee to obtain a debiasing processing result; The dynamic scoring and de-biasing processing results are integrated using edge computing to generate an interview evaluation report for the interviewee.
2. The method according to claim 1, wherein The step of extracting corresponding multimodal features from the multimodal data according to a multimodal feature extraction algorithm corresponding to the data type of the multimodal data includes: Using multiple sensors to collect multimodal data of the interviewee, wherein the multimodal data includes text data, voice data, and video data; Using a semantic model, extracting text features of the text data; Using visual recognition technology to extract frames and mark key points of the video data to obtain video features; Using speech processing technology to perform noise reduction and feature extraction on the speech data to obtain speech features; A multimodal fusion architecture is used to fuse the text features, video features, and speech features.
3. The method according to claim 1, wherein The step of adjusting the scoring weights of the multimodal features in real time according to the scoring weight optimization algorithm based on reinforcement learning includes: Using a dynamic evaluation model, we optimize the scoring weights using reinforcement learning: Q(s,a)=Q(s,a)+α*[r+γ*max(Q(s',a'))-Q(s,a)] Update the scoring weight of the multimodal feature; where s = current scoring state, a = weight adjustment action, and r = manual feedback correction reward; At predetermined intervals, the score weights of the multimodal features are re-optimized according to the score weight optimization algorithm.
4. The method according to claim 1, wherein The step of using the scoring weight to fuse the multimodal features to obtain a dynamic score of the interviewee includes: fusing the scoring weights with corresponding modal features respectively; Calculate the feature similarity between the modal features after integrating the scoring weights and the features corresponding to the job capabilities according to the similarity calculation model; Generate a capability radar chart according to the relationship between the position capability and the modal characteristics after integrating the scoring weights; The dynamic score of the interviewee is calculated according to the feature similarity and the ability radar chart.
5. The method according to claim 1, wherein Before the step of adjusting the scoring weights of the multimodal features in real time according to the reinforcement learning-based scoring weight optimization algorithm, and fusing the multimodal features with the scoring weights to obtain a dynamic score of the interviewee, the method further includes: Inputting the job requirements of the interview position and the multimodal features of the interviewee into the AI large language model; According to the job requirement information, control the AI large language model to generate interview follow-up questions based on the multimodal characteristics of the interviewee; Receiving new multimodal data generated by the interviewee's answers to the follow-up interview questions; The multimodal features corresponding to the new multimodal data are extracted using the multimodal feature extraction algorithm.
6. The method according to claim 1, wherein The step of using a debiasing algorithm to debias the interviewee's personal information to obtain a debiasing result includes: Performing synthetic data enhancement on the personal information using a synthetic data enhancement mechanism to obtain a debiased processing result; Alternatively, an adversarial training model is used in combination with a bias discriminator to jointly optimize the personal information.
7. The method according to claim 1, wherein The step of integrating the dynamic scoring and debiasing processing results using edge computing to generate an interview evaluation report for the interviewee includes: Using an edge computing node, executing the step of extracting corresponding multimodal features from the multimodal data, and compressing the multimodal features and uploading them to a cloud platform; Using the cloud platform, the scoring weight is integrated with the multimodal features to obtain a dynamic score of the interviewee; The cloud platform is used to obtain the debiasing processing result according to the debiasing algorithm, and the dynamic score and the debiasing processing result are integrated to obtain an interview evaluation report of the interviewee.
8. The method according to claim 1, wherein After the step of performing debias processing on the interviewee's personal information using a debiasing algorithm to obtain a debiasing processing result, the method further includes: Uploading the multimodal features, the interviewee's dynamic score, and the debiasing processing results to a real-time evaluation panel respectively; Using the real-time evaluation panel to display the multimodal features, dynamic scoring and debiasing processing results; A scoring weight adjustment signal is obtained, and the scoring weight of the multimodal feature is readjusted in an edge computing manner according to the scoring weight adjustment signal.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the AI-based multimodal fusion recruitment interview method as described in any one of claims 1 to 8.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the AI-based multimodal fusion recruitment interview method as described in any one of claims 1 to 8 is implemented.