Interview auxiliary cooperation system based on intelligent glasses
By using multimodal data fusion analysis from smart glasses terminals to generate AR prompts, the problem of information isolation in interview systems is solved, interview efficiency and interactivity are improved, and the naturalness and immersion of the interview process are enhanced.
Patent Information
- Application Number
- CN202511480430.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-02
AI Technical Summary
In existing interview systems, candidates' static historical data is isolated from dynamic multidimensional information, leading to inefficiency, strong subjectivity, and easy omission of key information by interviewers. Furthermore, the interactive experience lacks naturalness and immersion.
An interview assistance and collaboration system based on smart glasses is adopted. Interview data is acquired through image acquisition and audio acquisition modules, and multimodal fusion analysis is performed using a real-time analysis engine to generate AR prompts for interviewers to refer to in real time.
It enables efficient comparison of candidate performance and static data during the interview process, reduces subjectivity, increases the naturalness and immersion of the interview, and enhances interactivity.
Smart Images

Figure CN121256707A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an interview assistance collaborative system. Background Technology
[0002] As enterprises accelerate their digital transformation, interviews, as a crucial part of talent selection, have become an indispensable part of the recruitment process. When evaluating candidates, interviewers need to consider both static background information (such as resumes and assessment reports) and dynamic behavioral data generated during the interview (such as verbal expression, emotional reactions, and nonverbal cues) to make a comprehensive and objective recruitment decision.
[0003] Currently, interviewers primarily obtain candidate information through two methods:
[0004] Firstly, online interviews are conducted using personal computers or tablets via video conferencing software. However, while conducting the interview on a computer, interviewers need to manually switch browser or software windows to view candidates' electronic resumes, assessment reports, and other materials. This method leads to distraction for interviewers, frequent interruptions in the interview process, and low efficiency.
[0005] Secondly, in offline interviews, interviewers need to talk to candidates while simultaneously looking at paper resumes and test scores (or using tablets to view paper resumes and test scores, etc.). Interviewers need to frequently switch between "communicating with candidates" and "viewing information," relying on visual attention to manually integrate information from different sources.
[0006] Therefore, all of the above methods have the following problems:
[0007] (1) Information fusion bottleneck: The candidate's static historical data (resume, assessment report) and the dynamic multidimensional information generated during the interview (language content, tone of voice, micro-expression, body language) are completely isolated. The interviewer has to rely on his own brainpower to recall, compare and verify in real time. This process is inefficient, highly subjective and very easy to miss key information.
[0008] (2) Interactive experience bottleneck: Existing auxiliary tools (such as online interview systems or tablet computers) rely on traditional screen and keyboard interaction, forcing interviewers to frequently switch between "making eye contact and communicating deeply with candidates" and "looking down to operate the device and view the information", which seriously disrupts the naturalness, immersion and interactivity of the interview process. Summary of the Invention
[0009] Based on this, and in response to the aforementioned technical problems, an interview assistance and collaboration system based on smart glasses is provided to solve the problems of low efficiency, strong subjectivity, easy omission of key information, and reduced immersion and interactivity of existing technologies.
[0010] A collaborative interview assistance system based on smart glasses, the system comprising: a smart glasses terminal and an interview management server;
[0011] The smart glasses terminal includes: an image acquisition module, an audio acquisition module, an AR display module, a first communication module, and a local processing module; the smart glasses terminal is for the interviewer to wear on their head;
[0012] The interview management server includes: a candidate information database, a real-time analysis engine, and a second communication module; the candidate information database stores historical static data of candidates; the historical static data includes: resume data and assessment test paper data; the interview management server is connected to the first communication module of the smart glasses terminal through the second communication module.
[0013] The real-time analysis engine is configured to include: an identity recognition module, a multimodal fusion analysis module, and a prompt information generation and output module;
[0014] The identity recognition module is used to receive video and audio information collected by the image acquisition module and audio acquisition module of the smart glasses terminal during the initial stage of the interview; extract features that can characterize the identity based on the video and audio information of the initial stage of the interview; and obtain historical static data of the corresponding current candidate from the candidate information database based on the features that can characterize the identity.
[0015] The multimodal analysis module is used to continuously receive video and audio information of the current candidate's interview process from the image acquisition module and audio acquisition module of the smart glasses terminal in real time during the interview process; and to extract keywords and sentiment features based on the real-time received video and audio information, and to perform multimodal fusion analysis based on the keywords, sentiment features and the current candidate's historical static data to generate a credibility assessment result.
[0016] The prompt information generation and output module is used to: generate contextualized prompt information related to the current dialogue context based on the credibility assessment results, and present the prompt information in the AR display module of the smart glasses terminal in an AR overlay manner.
[0017] Optionally, in the above scheme, the identity recognition module is specifically used for:
[0018] During the initial stage of the interview, the system receives video and audio information from the image acquisition module and audio acquisition module of the smart glasses terminal.
[0019] The facial image of the current candidate is collected from the video information at the initial stage of the interview. If the facial image is compared with the facial image in the current candidate information database, and the comparison is successful, the identity of the current candidate and the corresponding historical static data are determined.
[0020] If the comparison fails or there is no facial image in the current candidate information database, then name keywords are extracted from the audio information of the initial stage of the interview, and the name keywords are matched with the name field in the candidate information database to determine the current candidate's identity and corresponding historical static data.
[0021] Optionally, in the above solution, the smart glasses terminal includes: an audio output module; the audio output module is used to provide audio prompts to the interviewer.
[0022] In the above solution, optionally, the audio output module is a bone conduction headphone.
[0023] In the above scheme, optionally, the prompt information generation and output module further includes:
[0024] The prompt information is converted into a voice signal and output to the interviewer through the audio output module.
[0025] Optionally, the prompting information in the above scheme may include: conflict warning prompts and emotional state reminder prompts.
[0026] Optionally, in the above scheme, the real-time analysis engine further includes: a question generation module; the question generation module is used to: generate follow-up questions based on the keywords and prompt information, and present them in an AR overlay manner on the AR display module of the smart glasses terminal.
[0027] Optionally, in the above scheme, the multimodal fusion analysis module is specifically used for:
[0028] The real-time received video and audio information is timestamped and aligned.
[0029] Micro-expression recognition and body language analysis are performed on the real-time received video information to extract video emotional features;
[0030] The real-time received audio information is converted from speech to text, and keywords and audio emotional features are extracted.
[0031] The semantic consistency analysis results are obtained by performing semantic consistency analysis on the keywords and the historical static data of the corresponding current candidates. The credibility analysis is then performed by combining the semantic consistency analysis results of the keywords with the video sentiment features and audio sentiment features of the corresponding timestamps. If the credibility is less than a certain value, a prompt message is generated based on the keywords.
[0032] Optionally, in the above scheme, the real-time analysis engine further includes: a closed-loop feedback module: the closed-loop feedback module is used for:
[0033] After the prompt information is presented in an AR overlay manner on the AR display module of the smart glasses terminal, the system receives real-time video and audio information during the interview process.
[0034] Identify predefined interactive behaviors of the interviewer from the video and audio information; the interactive behaviors include: predefined voice command information and predefined gesture information;
[0035] Based on predefined voice command information and predefined gesture information, the interviewer determines the prompt information for the keyword and labels it. The labeling includes at least: confirmation, ignoring, and marking as important.
[0036] Based on the statistics of interactive behaviors marked by the tags, the output frequency and content of subsequent prompts are adjusted to optimize the prompt strategy.
[0037] This application has at least the following beneficial effects:
[0038] This application acquires video and audio information from the interview process through the image and audio acquisition modules of a smart glasses terminal. The identity recognition module of the real-time analysis engine extracts identity-representing features from the initial video and audio information, and retrieves the corresponding historical static data of the current candidate from the candidate information database. Subsequently, the multimodal analysis module extracts keywords and sentiment features from the video and audio information during the interview process. Based on the keywords, sentiment features, and the candidate's historical static data, a multimodal fusion analysis is performed to generate a credibility assessment result; and a prompt message is generated accordingly, presented in an AR overlay manner on the AR display module of the smart glasses terminal for the interviewer to view in real time. Thus, the real-time analysis engine can comprehensively determine the credibility of the candidate's conversation by combining keywords from the interview, the candidate's historical static data, and sentiment features. If the credibility is low, a prompt message is output for the interviewer to ask further questions. This application can automatically compare and verify the candidate's performance and static data in real time during the interview process; this process is highly efficient, has low subjectivity, and does not miss key information. Furthermore, the interviewer only needs to maintain eye contact and in-depth communication with the candidate without needing to switch gazes, thus increasing the naturalness, immersion, and interactivity of the interview process. Attached Figure Description
[0039] Figure 1 This is a structural diagram of an interview assistance collaborative system based on smart glasses, provided as an embodiment of this application. Detailed Implementation
[0040] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0041] In one embodiment, such as Figure 1 As shown, an interview assistance and collaboration system based on smart glasses is provided. The system includes: a smart glasses terminal and an interview management server.
[0042] The smart glasses terminal includes: an image acquisition module, an audio acquisition module, an AR display module, a first communication module, and a local processing module; the smart glasses terminal is for the interviewer to wear on their head;
[0043] Specifically, the image acquisition module can use a miniature camera to capture candidate images and environmental information without being noticed; the audio acquisition module can use a microphone array to capture interview dialogue audio with high fidelity; the AR display module uses an optical waveguide lens to display visual cues to the interviewer in an AR overlay manner; the first communication module is used to exchange data with the server; and the local processing module is used to perform necessary local calculations (such as preliminary face detection and audio noise reduction).
[0044] The interview management server includes: a candidate information database, a real-time analysis engine, and a second communication module; the candidate information database stores historical static data of candidates; the historical static data includes: resume data and assessment test paper data; the interview management server is connected to the first communication module of the smart glasses terminal through the second communication module.
[0045] Specifically, the candidate information database stores historical data such as resumes and assessment scores. The real-time analysis engine performs real-time analysis, fusion, and decision-making on multimodal data; the second communication module connects to the smart glasses terminal.
[0046] The real-time analysis engine is configured to include: an identity recognition module, a multimodal fusion analysis module, and a prompt information generation and output module;
[0047] The identity recognition module is used to receive video and audio information collected by the image acquisition module and audio acquisition module of the smart glasses terminal during the initial stage of the interview; extract features that can characterize the identity based on the video and audio information during the initial stage of the interview; and obtain historical static data of the corresponding current candidate from the candidate information database based on the features that can characterize the identity.
[0048] The identity recognition module initiates a seamless identity recognition process. Specifically, at the initial stage of the interview, it receives video and audio information from the image and audio acquisition modules of the smart glasses terminal.
[0049] The facial image of the current candidate is collected from the video information at the initial stage of the interview. If the facial image is compared with the facial image in the current candidate information database, and the comparison is successful, the identity of the current candidate and the corresponding historical static data are determined.
[0050] If the comparison fails or there is no facial image in the current candidate information database, then name keywords are extracted from the audio information of the initial stage of the interview, and the name keywords are matched with the name field in the candidate information database to determine the current candidate's identity and corresponding historical static data.
[0051] Specifically, this operation adaptively selects one of the following two paths or executes them in parallel based on data completeness:
[0052] S1.1 Visual Recognition: If the candidate's profile on the server contains a photo, the image acquisition module captures the candidate's facial image. After verifying the candidate's identity through facial recognition, the server is automatically triggered to preload the candidate's complete historical data into the cache.
[0053] S1.2 Audio Recognition: If there is no available photo in the file, the candidate's self-introduction is captured through the audio acquisition module. Information such as name is extracted through speech recognition (ASR) and natural language processing (NLP). After the identity is verified by matching with the database, data preloading is also triggered.
[0054] Multimodal analysis module: During the interview process, it continuously receives video and audio information of the current candidate's interview process synchronously uploaded by the image acquisition module and audio acquisition module of the smart glasses terminal; and performs keyword extraction and sentiment feature extraction based on the real-time received video and audio information. Based on the keywords, sentiment features, and the historical static data of the current candidate, it performs multimodal fusion analysis to generate a credibility assessment result.
[0055] Specifically, the multimodal analysis module performs real-time speech-to-text (ASR), keyword extraction, and sentiment analysis (NLP) on the audio stream; and real-time micro-expression recognition and body language analysis (CV) on the video stream. Following this, the real-time analysis results are cross-referenced, correlated, and fused with historical static data pre-loaded by the identity recognition module. For example, the consistency between the candidate's claimed project roles and their resume is verified; and credibility is assessed by combining micro-expression changes when the candidate claims a particular skill.
[0056] The prompt information generation and output module is used to: generate contextualized prompt information related to the current dialogue context based on the credibility assessment results, and present the prompt information in the AR display module of the smart glasses terminal in an AR overlay manner.
[0057] Specifically, based on credibility analysis, the system generates contextualized prompts closely related to the current moment of the conversation (e.g., warnings of logical contradictions, reminders of emotional states, etc.). Through the AR display module of the smart glasses, these prompts are visually overlaid and presented precisely and unobstructed in the appropriate location within the interviewer's field of vision (e.g., a prompt box floating next to the candidate's image). Simultaneously, brief, personalized voice prompts can be provided via bone conduction headphones.
[0058] In the aforementioned interview assistance and collaboration system based on smart glasses, video and audio information from the interview process is acquired through the image acquisition module and audio acquisition module of the smart glasses terminal. The identity recognition module of the real-time analysis engine extracts identity-representing features from the video and audio information in the initial stage of the interview, and retrieves the corresponding historical static data of the current candidate from the candidate information database. Subsequently, the multimodal analysis module extracts keywords and sentiment features from the video and audio information during the interview process. Based on the keywords, sentiment features, and the historical static data of the current candidate, a multimodal fusion analysis is performed to generate a credibility assessment result; and a prompt message is generated accordingly, presented in an AR overlay manner on the AR display module of the smart glasses terminal for the interviewer to view in real time. Thus, the real-time analysis engine can comprehensively determine the credibility of the candidate's conversation by combining keywords from the interview, the current candidate's historical static data, and sentiment features. If the credibility is low, a prompt message is output for the interviewer to ask further questions. This application can automatically compare and verify the candidate's performance and static data in real time during the interview process; this process is highly efficient, has low subjectivity, and does not miss key information. Furthermore, the interviewer only needs to maintain eye contact and in-depth communication with the candidate, without switching eyes, which increases the naturalness, immersion, and interactivity of the interview process.
[0059] In one embodiment, the smart glasses terminal includes: an audio output module; the audio output module is used to provide audio prompts to the interviewer. The audio output module, such as bone conduction headphones, is used to provide private audio prompts to the interviewer. The prompt generation and output module further includes: converting the prompts into a speech signal and outputting it to the interviewer through the audio output module.
[0060] In one embodiment, the real-time analysis engine further includes a question generation module; the question generation module is used to generate follow-up questions based on the keywords and prompts, and present them in an AR overlay manner on the AR display module of the smart glasses terminal.
[0061] In one embodiment, the multimodal fusion analysis module is specifically used to include:
[0062] The real-time received video and audio information is timestamped and aligned.
[0063] Micro-expression recognition and body language analysis are performed on the real-time received video information to extract video emotional features;
[0064] The real-time received audio information is converted from speech to text, and keywords and audio emotional features are extracted.
[0065] The semantic consistency analysis results are obtained by performing semantic consistency analysis on the keywords and the historical static data of the corresponding current candidates. The credibility analysis is then performed by combining the semantic consistency analysis results of the keywords with the video sentiment features and audio sentiment features of the corresponding timestamps. If the credibility is less than a certain value, a prompt message is generated based on the keywords.
[0066] In one embodiment, the real-time analysis engine further includes: a closed-loop feedback module: the closed-loop feedback module is used for:
[0067] After the prompt information is presented in an AR overlay manner on the AR display module of the smart glasses terminal, the system receives real-time video and audio information during the interview process.
[0068] Identify predefined interactive behaviors of the interviewer from the video and audio information; the interactive behaviors include: predefined voice command information and predefined gesture information;
[0069] Based on predefined voice command information and predefined gesture information, the interviewer determines the prompt information for the keyword and labels it. The labeling includes at least: confirmation, ignoring, and marking as important.
[0070] Based on the statistics of interactive behaviors marked by the tags, the output frequency and content of subsequent prompts are adjusted to optimize the prompt strategy.
[0071] In this embodiment, the interviewer can interact with AR prompts (such as confirming, ignoring, or marking important information) through predefined whispered voice commands or subtle gestures. The system records these interactions to optimize subsequent prompting strategies, forming a closed-loop learning process.
[0072] The synergistic effects and key advantages of this invention include:
[0073] (1) It realizes the paradigm shift from “humans adapting to machines” to “machines enhancing humans”: Through seamless triggering and intelligent information flow, the analyzed decision support information is pushed to the interviewer’s perception range at the best time and in the most natural way, greatly expanding their cognitive ability.
[0074] (2) It created a new paradigm of “enhanced interview”: interviewers can process multiple information streams at the same time, including auditory, visual and system intelligent prompts, and make more accurate and in-depth interview judgments that humans could not do alone before.
[0075] (3) Provides a seamless cross-scene experience: Based on the unified interactive terminal of smart glasses, this invention is perfectly applicable to offline face-to-face interviews and remote video interviews, providing a highly consistent and immersive interview assistance experience.
[0076] (4) Improved the robustness and practicality of the system: Through the multimodal identity recognition strategy, it ensures reliable startup under various practical conditions, thereby enhancing the commercial value of the invention.
[0077] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0078] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A collaborative interview assistance system based on smart glasses, characterized in that, The system includes: smart glasses terminal and interview management server; The smart glasses terminal includes: an image acquisition module, an audio acquisition module, an AR display module, a first communication module, and a local processing module; the smart glasses terminal is for interviewers to wear on their heads. The interview management server includes: a candidate information database, a real-time analysis engine, and a second communication module; the candidate information database stores historical static data of candidates; the historical static data includes: resume data and assessment test paper data; the interview management server is connected to the first communication module of the smart glasses terminal through the second communication module. The real-time analysis engine is configured to include: an identity recognition module, a multimodal fusion analysis module, and a prompt information generation and output module; The identity recognition module is used to receive video and audio information collected by the image acquisition module and audio acquisition module of the smart glasses terminal during the initial stage of the interview; extract features that can characterize the identity based on the video and audio information of the initial stage of the interview; and obtain historical static data of the corresponding current candidate from the candidate information database based on the features that can characterize the identity. The multimodal analysis module is used to continuously receive video and audio information of the current candidate's interview process from the image acquisition module and audio acquisition module of the smart glasses terminal in real time during the interview process; and to extract keywords and sentiment features based on the real-time received video and audio information, and to perform multimodal fusion analysis based on the keywords, sentiment features and the current candidate's historical static data to generate a credibility assessment result. The prompt information generation and output module is used to: generate contextualized prompt information related to the current dialogue context based on the credibility assessment results, and present the prompt information in the AR display module of the smart glasses terminal in an AR overlay manner.
2. The interview assistance and collaboration system based on smart glasses according to claim 1, characterized in that, The identity recognition module is specifically used for: During the initial stage of the interview, the system receives video and audio information from the image acquisition module and audio acquisition module of the smart glasses terminal. The facial image of the current candidate is collected from the video information at the initial stage of the interview. If the facial image is compared with the facial image in the current candidate information database, and the comparison is successful, the identity of the current candidate and the corresponding historical static data are determined. If the comparison fails or there is no facial image in the current candidate information database, then name keywords are extracted from the audio information of the initial stage of the interview, and the name keywords are matched with the name field in the candidate information database to determine the current candidate's identity and corresponding historical static data.
3. The interview assistance and collaboration system based on smart glasses according to claim 1, characterized in that, The smart glasses terminal includes an audio output module; the audio output module is used to provide audio prompts to the interviewer.
4. The interview assistance and collaboration system based on smart glasses according to claim 3, characterized in that, The audio output module is a bone conduction headphone.
5. The interview assistance and collaboration system based on smart glasses according to claim 3, characterized in that, The prompt information generation and output module further includes: The prompt information is converted into a voice signal and output to the interviewer through the audio output module.
6. The interview assistance and collaboration system based on smart glasses according to claim 1, characterized in that, The notification information includes: conflict warning notification and emotional state reminder notification.
7. The interview assistance and collaboration system based on smart glasses according to claim 1, characterized in that, The real-time analysis engine also includes a question generation module; the question generation module is used to generate follow-up questions based on the keywords and prompts, and present them in an AR overlay manner on the AR display module of the smart glasses terminal.
8. The interview assistance and collaboration system based on smart glasses according to claim 1, characterized in that, The multimodal fusion analysis module is specifically used for: The real-time received video and audio information is timestamped and aligned. Micro-expression recognition and body language analysis are performed on the real-time received video information to extract video emotional features; The real-time received audio information is converted from speech to text, and keywords and audio emotional features are extracted. The semantic consistency analysis results are obtained by performing semantic consistency analysis between the keywords and the historical static data of the corresponding current candidates. The credibility analysis is performed by combining the semantic consistency analysis results of the keywords with the video sentiment features and audio sentiment features of the corresponding timestamps. If the credibility is less than a certain value, a prompt message is generated based on the keywords.
9. The interview assistance and collaboration system based on smart glasses according to claim 1, characterized in that, The real-time analysis engine further includes: a closed-loop feedback module; the closed-loop feedback module is used for: After the prompt information is presented in an AR overlay manner on the AR display module of the smart glasses terminal, the system receives real-time video and audio information during the interview process. Identify predefined interactive behaviors of the interviewer from the video and audio information; the interactive behaviors include: predefined voice command information and predefined gesture information; Based on predefined voice command information and predefined gesture information, the interviewer determines the prompt information for the keyword and labels it. The labeling includes at least: confirmation, ignoring, and marking as important. Based on the statistics of interactive behaviors marked by the tags, the output frequency and content of subsequent prompts are adjusted to optimize the prompt strategy.