Intelligent interview method and device based on large language model, equipment and storage medium

By using an intelligent interviewing method based on a large language model, dynamic job profiles and interview cognitive maps are generated. By combining multimodal information analysis and three major language models, the problems of low efficiency and strong subjectivity in traditional interviews are solved, achieving more precise and scientific interviews and generating comprehensive and reliable evaluation reports.

CN122114877APending Publication Date: 2026-05-29FOSHAN HUAYUE COMPUTER CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FOSHAN HUAYUE COMPUTER CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Traditional interviews are inefficient, subjective in assessment, and lack standardization, making it impossible to create accurate profiles and conduct in-depth analysis. This leads to the omission of excellent candidates or the misselection of unqualified candidates, and the interview process lacks scientific rigor and comprehensiveness.

Method used

The intelligent interview method based on large language models generates dynamic job profiles and interview cognitive maps. It combines multimodal information analysis and recursive questioning mechanisms, and leverages the synergistic effect of three major language models to accurately determine in-depth exploration strategies and evaluation goals, generating a comprehensive evaluation report.

Benefits of technology

It improves the relevance and scientific rigor of interviews, reduces subjective bias, deeply uncovers candidates' true abilities, generates comprehensive and reliable evaluation results, reduces the workload of interviewers, and shortens the interview cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114877A_ABST
    Figure CN122114877A_ABST
Patent Text Reader

Abstract

The application relates to the field of intelligent interviews, and discloses an intelligent interview method and device based on a large language model, equipment and a storage medium, which are used for fusing a large language model to realize intelligent and accurate interviews. The method comprises the following steps: acquiring post descriptions and candidate resume information, generating a dynamic post portrait and an interview cognitive graph; relying on a pre-trained first large language model, combining a current state of the cognitive graph to determine a deep exploration strategy and a current deep evaluation target; generating interactive content through a pre-trained second large language model, receiving candidate text, acoustic and visual multi-modal response information; analyzing the multi-modal response to obtain corresponding features, and updating the interview cognitive graph; judging whether a cognitive convergence threshold is reached, recursively asking questions if the threshold is not reached, and promoting the interview or ending the interview if the threshold is reached; after the interview is ended, generating a main evaluation conclusion based on a final cognitive graph, questioning and debating through a pre-trained third large language model, and generating a comprehensive evaluation report containing core arguments of the pro and con sides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent interviewing, and more particularly to an intelligent interviewing method, apparatus, device, and storage medium based on a large language model. Background Technology

[0002] In digital recruitment scenarios, interviews are a core part of talent screening, and traditional methods suffer from significant pain points such as inefficiency, subjective evaluation, and lack of standardization. Traditional interviews rely on the interviewer's personal experience, which is easily influenced by subjective preferences and fatigue, leading to inconsistent evaluation standards for candidates for the same position. This results in a higher probability of overlooking excellent candidates or mistakenly selecting unqualified candidates, increasing recruitment costs for companies.

[0003] As companies expand their recruitment scale and the number of candidates surges, interviewers need to invest a significant amount of time in screening resumes, conducting interviews, and writing evaluation reports. This results in low interview efficiency per person, making it difficult to meet the demands of large-scale, rapid recruitment. Furthermore, traditional interviews often employ a fixed questioning process, failing to dynamically adjust follow-up questions based on the candidate's resume characteristics and job requirements, thus hindering a deep understanding of the candidate's true abilities and job suitability.

[0004] Existing intelligent interview technologies are mostly limited to fixed question banks and simple speech recognition. They lack the ability to build accurate profiles of positions and candidates, cannot achieve dynamic optimization and in-depth exploration of the interview process, and the evaluation conclusions are mostly based on single-dimensional data, lacking comprehensiveness and objectivity, making it difficult to support scientific decision-making by enterprises.

[0005] Therefore, existing technologies still need improvement and development. Summary of the Invention

[0006] This invention provides an intelligent interviewing method, apparatus, device, and storage medium based on a large language model, which is used to integrate a large language model to achieve intelligent and accurate interviewing.

[0007] The first aspect of this invention provides an intelligent interview method based on a large language model. The method includes: acquiring job description information and candidate resume information; generating a dynamic job profile and an interview cognitive graph based on the job description information and the candidate resume information; determining a deep exploration strategy and a current deep evaluation target for the current interview stage based on the current state of the interview cognitive graph using a pre-trained first large language model; generating interactive content for the candidate based on the deep exploration strategy using a pre-trained second large language model, and receiving multimodal response information from the candidate based on the interactive content; parsing the candidate's multimodal response information to obtain text features, acoustic features, and visual features. The interview cognitive graph is updated based on the text features, acoustic features, and visual features to obtain the updated cognitive graph. It is then determined whether the cognitive convergence threshold for the current depth assessment goal has been reached. If not, the updated cognitive graph replaces the original interview cognitive graph, and the process returns to executing the first pre-trained language model based on the interview cognitive graph to determine the current deep exploration strategy and current depth assessment goal for the current interview stage. When the cognitive convergence threshold for the current depth assessment goal is reached, the final cognitive graph is obtained, and a main assessment conclusion is generated based on the final cognitive graph. The main assessment conclusion is then challenged and debated using a pre-trained third language model to generate a comprehensive assessment report, which includes core arguments from both sides.

[0008] Preferably, the step of acquiring job description information and candidate resume information, and generating a dynamic job profile and interview cognitive graph based on the job description information and candidate resume information, includes: acquiring and preprocessing job description information and candidate resume information; performing deep semantic analysis on the preprocessed job description information and candidate resume information to extract job requirement elements and candidate ability elements; generating a dynamic job profile based on the extracted job requirement elements, the dynamic job profile including hard qualifications, responsibilities, abilities, organizational environment, and evaluation preference dimensions; performing preliminary mapping and gap analysis on the extracted candidate ability elements and dynamic job profile, and constructing an interview cognitive graph based on the results of the preliminary mapping and gap analysis, the interview cognitive graph including entity nodes, evidence nodes, logical nodes, and relational edges representing the membership, support, contradiction, triggering, or association relationships between representational nodes.

[0009] Preferably, the step of determining the in-depth exploration strategy and current in-depth assessment target for the current interview stage based on the current state of the interview cognitive graph and using a pre-trained first large language model includes: performing a situational diagnosis on the current state of the interview cognitive graph using the pre-trained first large language model to obtain a situational diagnosis result, which includes the sufficiency, consistency, and existence of unresolved contradictions of existing evidence under each ability dimension; combining the situational diagnosis result and the preset weights of each ability dimension in the dynamic job profile, calculating the current exploration urgency score of each ability dimension in the dynamic job profile using a multi-factor weighted algorithm; selecting the ability dimension with the highest current exploration urgency score as the current in-depth assessment target, and matching and generating an in-depth exploration strategy from a preset strategy library based on the attributes of the current in-depth assessment target and the type of information gap exposed in the interview cognitive graph, which includes in-depth stress testing, behavioral event tracing, or value probes.

[0010] Preferably, the step of generating interactive content for candidates based on the deep exploration strategy using a pre-trained second language model, and receiving candidate multimodal response information based on the interactive content, includes: extracting relevant context and current dialogue history from the interview cognitive graph; inputting the deep exploration strategy, the extracted relevant context, and the current dialogue history into the pre-trained second language model, so that the pre-trained second language model generates interactive content for candidates under policy constraints, the interactive content including specific situational details, natural language questions, or simulated scene descriptions; presenting the interactive content to the candidates, and receiving candidate multimodal response information based on the interactive content, the multimodal response information including textual information, acoustic information, and visual information.

[0011] Preferably, the step of parsing the candidate's multimodal response information to obtain text features, acoustic features, and visual features, and updating the interview cognitive graph based on the text features, acoustic features, and visual features, includes: for the candidate's multimodal response information, obtaining text transcription as text features through speech recognition, extracting intonation, speech rate, and pause features as acoustic features through acoustic analysis, and extracting facial action units and posture features as visual features through visual analysis; aligning the text features, acoustic feature sequences, and visual features on the time axis to obtain aligned multimodal features; performing cross-modal consistency reasoning based on the aligned multimodal feature sequences to determine the consistency relationship between the text features, acoustic feature sequences, and visual features, and obtaining a consistency reasoning result; based on the consistency reasoning result, creating evidence nodes in the interview cognitive graph and establishing support, weakening, or contradictory logical relationships between the created evidence nodes and related ability dimension nodes, or updating evidence nodes and adjusting the support, weakening, or contradictory logical relationships between the updated evidence nodes and related ability dimension nodes, to obtain an updated cognitive graph.

[0012] Preferably, the determination of whether the cognitive convergence threshold for the current depth assessment goal has been reached, and if not, replacing the interview cognitive graph with an updated cognitive graph, and returning to execute the determination of the current interview stage's depth exploration strategy and current depth assessment goal based on the interview cognitive graph using a pre-trained first language model, includes: determining whether the current depth assessment goal meets the requirements of sufficient evidence nodes, diverse evidence sources, and the absence of unresolved major logical contradictions; if yes, determining that the cognitive convergence threshold for the current depth assessment goal has been reached, determining the next depth assessment goal, or ending the interview; if not, based on the updated interview cognitive graph, locating the specific information gaps or logical contradictions that led to non-convergence; generating targeted follow-up questioning strategies based on the type of gaps or logical contradictions, including requesting specific details, requesting clarification of inconsistent statements, or introducing new stressful situations; and generating the next round of specific interaction content using a pre-trained second language model based on the follow-up questioning strategies, until the cognitive convergence threshold for the current depth assessment goal is reached.

[0013] Preferably, when the cognitive convergence threshold for the current depth assessment target is reached, a final cognitive map is obtained, and a main assessment conclusion is generated based on the final cognitive map. The main assessment conclusion is then challenged and debated using a pre-trained third language model to generate a comprehensive assessment report. The comprehensive assessment report includes core arguments for and against the assessment, including: when the cognitive convergence threshold for the current depth assessment target is reached, obtaining the final cognitive map, and calculating a quantitative score and generating a preliminary main assessment conclusion based on the evidence weights and consistency scores of each ability dimension in the final cognitive map; inputting the final cognitive map and the preliminary main assessment conclusion into the pre-trained third language model to find evidence supporting the opposite conclusion and generating opposing opinions, which include opposing arguments and corresponding evidence citations; inputting the opposing opinions into the pre-trained main assessment model to conduct multiple rounds of simulated debate between the pre-trained main assessment model and the pre-trained third language model, and recording the arguments and evidence during the debate to generate simulated debate records; and generating a comprehensive assessment report based on the simulated debate records, which includes candidate basic information, the main assessment conclusion, core arguments for and against the assessment, and recruitment recommendations.

[0014] A second aspect of this invention provides an intelligent interview device based on a large language model, comprising: a first generation module, configured to acquire job description information and candidate resume information, and generate a dynamic job profile and an interview cognitive graph based on the job description information and the candidate resume information; a determination module, configured to determine the current deep exploration strategy and the current deep evaluation target for the current interview stage based on the current state of the interview cognitive graph using a pre-trained first large language model; a second generation module, configured to generate interactive content for the candidate based on the deep exploration strategy using a pre-trained second large language model, and receive candidate multimodal response information based on the interactive content; and a parsing module, configured to parse the candidate multimodal response information to obtain text features, acoustic features, and visual features, and... The interview cognitive graph is updated based on the text features, acoustic features, and visual features to obtain an updated cognitive graph. A judgment module is used to determine whether the cognitive convergence threshold for the current depth assessment goal has been reached. If not, the updated cognitive graph replaces the interview cognitive graph, and the process returns to execute the first pre-trained language model based on the interview cognitive graph to determine the current interview stage's depth exploration strategy and current depth assessment goal. A third generation module is used to obtain the final cognitive graph when the cognitive convergence threshold for the current depth assessment goal is reached, and to generate a main assessment conclusion based on the final cognitive graph. The main assessment conclusion is then questioned and debated using a pre-trained third language model to generate a comprehensive assessment report, which includes core arguments from both sides.

[0015] A third aspect of the present invention provides an intelligent interview device based on a large language model, comprising: a memory and at least one processor, wherein the memory stores computer-readable instructions, and the memory and the at least one processor are interconnected via a circuit; the at least one processor invokes the computer-readable instructions in the memory to cause the intelligent interview device based on the large language model to perform the various steps of the intelligent interview method based on the large language model as described above.

[0016] A fourth aspect of the present invention provides a computer-readable storage medium storing computer-readable instructions that, when executed on a computer, cause the computer to perform the steps of the intelligent interview method based on a large language model as described above.

[0017] The technical solution provided by this invention generates a dynamic job profile and interview cognitive map by acquiring job descriptions and candidate resumes, achieving precise matching between jobs and candidates and breaking the limitations of traditional one-size-fits-all interviews. Furthermore, through the synergistic efforts of three pre-trained language models—the first accurately determines deep exploration strategies and evaluation goals, the second generates personalized interactive content, and the third verifies evaluation conclusions through questioning and debate—it significantly reduces human subjectivity bias and enhances the relevance and scientific rigor of interviews. Simultaneously, by integrating multimodal information parsing technology, it comprehensively captures candidates' textual, acoustic, and visual features, and combined with a recursive questioning mechanism, it deeply uncovers candidates' true abilities, avoiding superficial evaluations and ensuring comprehensive and reliable evaluation results. In addition, by judging cognitive convergence thresholds, it reasonably controls the interview pace, automating the entire interview process, significantly reducing the workload of interviewers and shortening the interview cycle. The final comprehensive evaluation report includes core arguments from both sides, clearly presenting the candidate's strengths and weaknesses, providing a comprehensive and reliable reference for recruitment decisions. It also possesses good job-company fit and can be widely applied to recruitment scenarios across various industries. Attached Figure Description

[0018] Figure 1 A flowchart illustrating the intelligent interview method based on a large language model provided in this embodiment of the invention; Figure 2 A schematic diagram of the structure of an intelligent interview device based on a large language model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an intelligent interview device based on a large language model, provided in an embodiment of the present invention. Detailed Implementation

[0019] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0020] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 An intelligent interview method based on a large language model, as described in this embodiment of the invention, includes: S101. Obtain job description information and candidate resume information, and generate dynamic job profile and interview cognitive map based on job description information and candidate resume information; S102. Based on the current state of the interview cognitive graph, determine the deep exploration strategy and current deep evaluation target for the current interview stage through the pre-trained first language model. S103. Based on the deep exploration strategy, the system generates interactive content for candidates through a pre-trained second language model and receives multimodal response information from candidates based on the interactive content. S104. Analyze the candidate's multimodal response information to obtain text features, acoustic features and visual features, and update the interview cognitive map based on the text features, acoustic features and visual features to obtain the updated cognitive map. S105. Determine whether the cognitive convergence threshold for the current depth assessment target has been reached. If not, replace the interview cognitive graph with an updated cognitive graph and return to execute the deep exploration strategy and current depth assessment target for the current interview stage determined by the first pre-trained language model based on the interview cognitive graph. S106. When the cognitive convergence threshold for the current depth evaluation target is reached, the final cognitive map is obtained, and the main evaluation conclusion is generated based on the final cognitive map. The main evaluation conclusion is questioned and debated through the pre-trained third language model, and a comprehensive evaluation report is generated, which includes the core arguments of both sides.

[0021] It is understood that the executing entity of this invention can be an intelligent interview device based on a large language model, or it can be a terminal or a server; no specific limitation is made here. This embodiment of the invention will be described using a server as an example.

[0022] This embodiment presents an intelligent interview method based on a large language model. By acquiring job descriptions and candidate resumes, it generates dynamic job profiles and interview cognitive maps, achieving precise matching between jobs and candidates, breaking the limitations of traditional one-size-fits-all interviews. Furthermore, through the synergistic efforts of three pre-trained language models—the first accurately determines deep exploration strategies and evaluation objectives, the second generates personalized interactive content, and the third verifies evaluation conclusions through questioning and debate—it significantly reduces human subjectivity bias and enhances the relevance and scientific rigor of the interview. Simultaneously, by integrating multimodal information parsing technology, it comprehensively captures the candidate's textual, acoustic, and visual features, combined with a recursive questioning mechanism, it deeply uncovers the candidate's true abilities, avoiding superficial evaluations and ensuring comprehensive and reliable assessment results. In addition, by using cognitive convergence threshold judgment, it reasonably controls the interview pace, automating the entire interview process, significantly reducing the interviewer's workload and shortening the interview cycle. The final comprehensive evaluation report includes core arguments from both sides, clearly presenting the candidate's strengths and weaknesses, providing a comprehensive and reliable reference for recruitment decisions. It also possesses good job-company fit and can be widely applied to recruitment scenarios across various industries.

[0023] In this embodiment, step S101 involves acquiring job description information and candidate resume information, and generating a dynamic job profile and interview cognitive graph based on the job description information and candidate resume information. This includes: acquiring and preprocessing job description information and candidate resume information; performing deep semantic analysis on the preprocessed job description information and candidate resume information to extract job requirement elements and candidate ability elements; generating a dynamic job profile based on the extracted job requirement elements, including dimensions of hard qualifications, responsibilities, abilities, organizational environment, and evaluation preferences; performing preliminary mapping and gap analysis on the extracted candidate ability elements and dynamic job profile, and constructing an interview cognitive graph based on the results of the preliminary mapping and gap analysis. The interview cognitive graph includes entity nodes, evidence nodes, logical nodes, and relational edges representing the membership, support, contradiction, triggering, or association relationships between representational nodes.

[0024] In this embodiment, the job description information includes the job title, responsibilities, job requirements (education, major, work experience, skills, abilities, etc.), and organizational environment (team size, work pace, corporate culture, etc.); the candidate resume information includes basic personal information, educational background, work experience, project experience, skill certificates, self-evaluation, etc.

[0025] The preprocessing stage mainly involves information cleaning and standardization, specifically including: removing invalid information (such as redundant advertisements in resumes and vague statements in job descriptions), standardizing information formats (such as classifying different expressions of project management capabilities into a unified project management dimension), correcting information errors (such as checking outliers in education and years of work experience), and extracting key information fragments (such as project achievements in resumes and core skill requirements in job descriptions) to ensure the accuracy of subsequent semantic parsing.

[0026] In this embodiment, the preprocessed job description information and candidate resume information are input into a semantic parsing model for deep semantic mining and element extraction. The semantic parsing model can be fine-tuned based on a pre-trained large language model. The extraction of job requirements elements mainly focuses on four core dimensions: hard qualifications (education, major, years of work experience, certificates, etc.), responsibilities and tasks (core work content, work objectives, etc.), abilities and qualities (communication skills, logical thinking, stress resistance, innovation ability, etc.), and organizational environment suitability (such as adaptability to a fast-paced work environment, whether it matches the corporate culture, etc.). The extraction of candidate ability elements corresponds to the job requirements elements, exploring the candidate's specific performance in each dimension (such as communication skills demonstrated in work experience, professional skills demonstrated in project experience, and professional qualities demonstrated in self-evaluation, etc.).

[0027] For example, regarding the job description's ability to independently manage the entire process of a project from initiation to completion, possessing strong problem-solving and cross-departmental communication skills, the responsibilities can be extracted as full-process project management, and the competencies as problem-solving and cross-departmental communication skills. Similarly, regarding the candidate's resume's achievements in leading the implementation of three provincial-level projects, coordinating five cross-departmental teams to solve core technical challenges, and achieving a 100% on-time delivery rate, the candidate's competencies can be extracted as project management skills (full-process), cross-departmental communication skills, and problem-solving skills, linked to specific evidence such as "leading three provincial-level projects, coordinating five cross-departmental teams, and achieving a 100% on-time delivery rate."

[0028] In this embodiment, a multi-dimensional dynamic job profile is constructed based on the extracted job requirement elements. This profile is not fixed and will be dynamically adjusted in subsequent interviews based on the candidate's actual performance and the refinement of job requirements, adjusting the weights and evaluation criteria of each dimension. The dynamic job profile specifically includes five dimensions, each described in detail below: Hard qualifications: Clearly define the minimum entry requirements for the position, such as a bachelor's degree or above, a computer-related major, more than 3 years of backend development experience, and holding a PMP certificate. This dimension is the core basis for veto or priority consideration. Job Responsibilities and Tasks: Detail the core work content and delivery goals of the position, such as being responsible for backend interface development and optimization, participating in project requirement review, resolving online system failures, and cooperating with the frontend team to complete integration testing, so as to provide a basis for subsequent evaluation of the candidate's job suitability; Competency and Qualification Dimension: Clearly define the core soft skills and hard skills required for the position, and distinguish between essential skills and bonus skills. For example, essential skills include: Java development skills, experience in using the Spring Boot framework, and logical thinking ability; bonus skills include: experience in microservice architecture and communication skills. Organizational environment dimension: Based on the actual situation of the enterprise, clarify the work environment and cultural adaptation requirements of the positions, such as Internet companies, fast-paced work mode, flat management, and emphasis on innovation and collaboration; Evaluation preference dimensions: Based on the company's recruitment needs, set the evaluation priority and weight for each dimension. For example, for technical positions, professional skills are evaluated first (weight 60%), followed by logical thinking ability (weight 20%), and finally communication ability (weight 20%).

[0029] In this embodiment, the interview cognitive graph includes a basic structure mapping from resume to job profile and hypotheses to be tested. The interview cognitive graph is a graph data structure, where node types include: Entity nodes: such as candidates, positions, and specific project experience.

[0030] Evidence points: such as the candidate's claim: I led the XX project; behavioral indicators: a firm tone of voice when making statements.

[0031] Logical nodes: such as assessment dimensions: leadership; capability requirements: cross-departmental communication.

[0032] Nodes are connected by directed edges, and the types of edge relationships include: membership (e.g., evidence belongs to a certain ability dimension), support / weakening (e.g., a behavioral indicator supports or weakens a claim), contradiction (e.g., two statements are inconsistent), triggering (e.g., a situation triggers a certain behavior), association (general association), etc.

[0033] In this embodiment, the job description and candidate resume are preprocessed to remove invalid information and standardize data format to ensure information quality. At the same time, deep semantic parsing is used to accurately extract job requirements and candidate ability elements to avoid information omissions or misinterpretations. Moreover, the generated dynamic job profile covers multiple dimensions such as hard qualifications and abilities, fully matching the actual needs of the job and possessing dynamic adaptability. In addition, the constructed interview cognitive graph clearly presents the gap between the candidate's abilities and job requirements, as well as the relationship between existing evidence, through nodes such as entities, evidence, and logic, and multiple relational edges, transforming abstract ability information into a visualized logical structure.

[0034] In this embodiment, step S102 involves determining the current in-depth exploration strategy and current in-depth assessment target for the current interview stage based on the current state of the interview cognitive graph using a pre-trained first large language model. This includes: performing a situational diagnosis on the current state of the interview cognitive graph using the pre-trained first large language model to obtain a situational diagnosis result, which includes the sufficiency, consistency, and existence of unresolved contradictions of existing evidence under each capability dimension; combining the situational diagnosis result and the preset weights of each capability dimension in the dynamic job profile, calculating the current exploration urgency score of each capability dimension in the dynamic job profile using a multi-factor weighted algorithm; selecting the capability dimension with the highest current exploration urgency score as the current in-depth assessment target, and matching and generating an in-depth exploration strategy from a preset strategy library based on the attributes of the current in-depth assessment target and the type of information gap exposed in the interview cognitive graph. The in-depth exploration strategy includes in-depth stress testing, behavioral event tracing, or value probes.

[0035] In this embodiment, the interview cognitive graph in the current state (initial state or updated state after the previous round of follow-up questions) is input into the pre-trained first language model. The pre-trained first language model has been fine-tuned with a large amount of interview scenario data and has situational diagnosis capabilities. It can comprehensively analyze each ability dimension in the cognitive graph and output situational diagnosis results.

[0036] In this embodiment, the first pre-trained large language model is a large language model fine-tuned with specific task instructions (such as fine-tuning based on Llama 3, GPT, or a domestic open-source platform). During training, a massive amount of interview cognitive graph situational data, job evaluation dimension data, and exploration urgency scoring cases are used as the training set. Through supervised fine-tuning (SFT), the model masters the core capabilities of situational diagnosis of the interview cognitive graph, exploration urgency calculation, and matching exploration strategies with evaluation goals. At the same time, a small amount of human feedback reinforcement learning (RLHF) is introduced to optimize the accuracy of strategy and goal determination.

[0037] In this embodiment, the situational diagnosis results include the sufficiency and consistency of existing evidence under each capability dimension, as well as whether there are any unresolved contradictions.

[0038] The sufficiency of existing evidence under each competency dimension refers to whether the existing evidence (from resume analysis and previous interview responses) under each competency dimension is sufficient to support the evaluation conclusion. For example, in the Java development dimension of professional skills, if there is only a statement of familiarity with Java in the resume without specific project cases, the evidence is insufficient.

[0039] Consistency of evidence refers to judging whether there are contradictions in the evidence across different competency dimensions. For example, if a candidate claims to have led a project development, but the resume does not mention the relevant project, or if the candidate stated in the previous interview that they participated in auxiliary work on the project, this constitutes inconsistency of evidence.

[0040] Unresolved contradictions refer to the points of conflict that exist in each ability dimension but have not yet been clarified. For example, in the work experience dimension, if the resume shows 2 years and the candidate states 3 years, without further clarification, it is considered an unresolved contradiction.

[0041] In this embodiment, the higher the current urgency score, the more priority is given to in-depth exploration of that dimension.

[0042] The core formula of the multi-factor weighted algorithm is: Investigation Urgency Score = (Insufficient Evidence Coefficient × Weight 1) + (Contradictory Evidence Coefficient × Weight 2) + (Unresolved Contradictory Evidence Coefficient × Weight 3) × Job Dimension Weight; where the Insufficient Evidence Coefficient, Contradictory Evidence Coefficient, and Unresolved Contradictory Evidence Coefficient are all values ​​between 0 and 1 (e.g., if the evidence is completely insufficient, the coefficient is 1; if the evidence is completely sufficient, the coefficient is 0). Weight 1, Weight 2, and Weight 3 are preset factor weights that can be adjusted according to the company's needs. For example, if priority is given to the points of contradiction, Weight 3 can be set to 0.5, Weight 1 to 0.3, and Weight 2 to 0.2. The job dimension weight is the preset weight of each ability dimension in the dynamic job profile, such as the professional skills dimension weight of a technical position being 0.6.

[0043] For example, if the weight of the professional skills dimension for a backend development position is 0.6, and the coefficients for insufficient evidence, contradiction, and unresolved contradiction are 0.8, 0.2, and 0.5 respectively, then the urgency score for this dimension is (0.8×0.3+0.2×0.2+0.3×0.5)×0.6=(0.24+0.04+0.15)×0.6=0.43×0.6=0.258. If the weight of the communication skills dimension is 0.2, and the coefficients for insufficient evidence, contradiction, and unresolved contradiction are 0, then its urgency score is (0.5×0.3+0×0.2+0×0.5)×0.2=0.15×0.2=0.03. In this case, the urgency of the professional skills dimension is higher, and it should be prioritized as the current in-depth assessment target.

[0044] In this embodiment, if multiple dimensions have the same score, the core requirement dimension of the job is selected as the current target based on the priority of the dynamic job profile.

[0045] Deep exploration strategies refer to the specific questioning methods and interaction logic adopted to uncover the current deep assessment target. These strategies are generated by a pre-trained primary language model based on the attributes of the current deep assessment target (such as hard skills and soft skills) and the type of information gaps revealed in the interview cognitive map (such as insufficient evidence or contradictions), matching from a pre-built strategy library. Deep exploration strategies include deep stress testing, behavioral event tracing, or value probes.

[0046] Deep stress testing is suitable for assessing candidates' resilience, problem-solving abilities, and professional skills proficiency. By posing challenging and complex questions, it observes candidates' response speed, logical clarity, and the rationality of their solutions. For example, regarding professional skills, a question might be, "If your Java program has a memory leak, how would you troubleshoot it? Please explain the steps in detail, including the tools used and the core logic." Regarding resilience, a question might be, "If the project deadline is 3 days early and two team members have to take temporary leave, how would you handle the situation?"

[0047] Behavioral event tracing is suitable for assessing a candidate's past work experience, practical skills, sense of responsibility, etc. Based on the STAR method (Context, Task, Action, Result), it requires candidates to describe specific work scenarios and behaviors to uncover their true abilities. For example, it asks, "Please describe a project you led that encountered significant difficulties. What was your task at the time? What actions did you take? What was the final result?" Through the candidate's description, relevant statements in the resume are verified, and supplementary evidence is added.

[0048] Values ​​probes are used to assess a candidate's values, professional qualities, and fit with corporate culture. By asking questions related to values ​​and career choices, they can observe the candidate's stance and attitude, such as "What do you think is most important at work? How would you handle a situation where you disagree with your leader?" and "What are your views on overtime? Under what circumstances would you choose to work overtime?"

[0049] In this embodiment, relevant context and dialogue history are extracted from the interview cognitive graph to ensure that the interactive content fits the interview process and the candidate's actual situation, avoiding disconnected questioning. Moreover, under the constraints of the deep exploration strategy, the second language model generates contextualized and personalized interactive content, including various forms such as contextual details and natural questioning, which can improve the naturalness of the interview and the candidate's participation. In addition, receiving candidates' textual, acoustic, and visual multimodal response information breaks through the limitations of traditional interviews that only focus on textual answers. It can comprehensively capture multi-dimensional information such as candidates' language expression, emotional state, and body language, making the interview interaction more human and targeted, and comprehensively collecting information related to candidates' abilities and qualities. This provides rich and comprehensive evidence support for subsequent accurate assessment and improves the accuracy of the assessment results.

[0050] In this embodiment, in step S103, based on the deep exploration strategy, interactive content for the candidate is generated through a pre-trained second language model, and multimodal response information of the candidate is received based on the interactive content. This includes: extracting relevant context and current dialogue history from the interview cognitive graph; inputting the deep exploration strategy, the extracted relevant context, and the current dialogue history into the pre-trained second language model, so that the pre-trained second language model generates interactive content for the candidate under policy constraints. The interactive content includes specific situational details, natural language questions, or simulated scene descriptions; presenting the interactive content to the candidate, and receiving multimodal response information of the candidate based on the interactive content. The multimodal response information includes text information, acoustic information, and visual information.

[0051] In this embodiment, the second pre-trained language model is a base model with strong natural language generation and scene adaptation capabilities (such as ERNIE3.5, LLaMA3). The training set includes interview dialogues for various positions, scenario-simulated questioning cases, and candidate response samples. Through a combination of supervised fine-tuning and prompt engineering, the model can generate interactive content that fits the interview scenario, is highly targeted, and is natural and fluent based on deep exploration strategies, context, and dialogue history. At the same time, the model's ability to avoid repeated questions and dynamically adjust the difficulty of questions is optimized.

[0052] In this embodiment, contextual information related to the current in-depth assessment objective (such as the job requirements for this dimension, the candidate's existing relevant evidence, and existing information gaps) and the current interview dialogue history (such as the previous round of questions and the candidate's answers) are extracted from the interview cognitive graph to ensure that the generated interactive content is coherent and targeted, and to avoid repeated questions or deviation from the assessment objective.

[0053] For example, if the current in-depth assessment objective is project management capability, and the in-depth exploration strategy is behavioral event tracing, then the extracted contextual information includes "Job Requirements: Ability to independently manage the entire project process," "Existing Evidence from the Candidate: Resume mentions leading one project," and "Information Gap: Specific project management processes, difficulties encountered, and solutions." The current dialogue history includes the previous question, "Do you have experience leading projects?" and the candidate's answer, "Yes, I led the development of a small project." Subsequent interaction content is generated based on this to ensure coherence.

[0054] In this embodiment, the interactive content takes various forms, mainly including questions about specific scenario details, natural language questions, and simulated scenario descriptions.

[0055] Specific scenario-based detailed questions, which address information gaps by asking specific questions that can guide candidates to fill in details, such as "What specific modules were involved in the small project you led? How many people were on the project team? What were your main responsibilities?"

[0056] Natural language questioning involves asking questions related to the assessment objectives in a natural and friendly tone, which is similar to the interactive experience of a human interview. For example, "During the project, did you encounter any situations where team members did not cooperate well? How did you solve them?"

[0057] Simulated scenario description involves constructing simulated scenarios related to the actual work of the position, requiring candidates to provide solutions on-site, and assessing their practical operational skills and emergency handling capabilities. For example, "Suppose you are currently in charge of a project, with 10 days left until delivery, and suddenly a major bug is discovered in the core module, making it impossible to complete the development on time. How would you handle this? Please explain your thought process and steps in detail."

[0058] It should be noted that when the second language model generates interactive content, it automatically avoids asking the same questions repeatedly (by combining the dialogue history) and adjusts the difficulty and tone of the questions according to the candidate's past response style. For example, if the candidate's answer is concise, the question will be more guiding; if the candidate's answer is detailed, the question will be more focused on the core gap.

[0059] In this embodiment, the generated interactive content is presented to the candidate through an interview terminal (such as a computer, mobile phone, or interview pod). The candidate can respond in various ways, including voice, video, and text. The system simultaneously receives the candidate's multimodal response information. The multimodal response information includes text information, acoustic information, and visual information.

[0060] Textual information includes the candidate's response content entered via text (e.g., typing answers to questions) and the text converted from speech to text (e.g., the text content of a speech response after speech recognition). Acoustic information includes the sound characteristics of the candidate's speech response, such as tone (high and low), speech rate (fast and slow), pauses (number and duration), volume (loudness), and tone (firm, hesitant, perfunctory). Visual information includes the image characteristics of the candidate's video response, such as facial expressions (smiling, frowning, hesitant, calm), facial movement units (eye contact, nodding, shaking head, gestures), and posture (sitting upright, bending over, bowing head).

[0061] In this embodiment, step S104 involves parsing the candidate's multimodal response information to obtain text features, acoustic features, and visual features. The interview cognitive graph is then updated based on these features to obtain an updated cognitive graph. This includes: for the candidate's multimodal response information, obtaining text transcription as text features through speech recognition; extracting intonation, speech rate, and pause features as acoustic features through acoustic analysis; and extracting facial action units and posture features as visual features through visual analysis. The text features, acoustic feature sequences, and visual features are aligned on the time axis to obtain aligned multimodal features. Cross-modal consistency reasoning is performed based on the aligned multimodal feature sequences to determine the consistency relationship between the text features, acoustic feature sequences, and visual features, resulting in a consistency reasoning result. Based on the consistency reasoning result, evidence nodes are created in the interview cognitive graph, and support, weakening, or contradictory logical relationships are established between the created evidence nodes and relevant ability dimension nodes. Alternatively, evidence nodes are updated, and the support, weakening, or contradictory logical relationships between the updated evidence nodes and relevant ability dimension nodes are adjusted, resulting in an updated cognitive graph.

[0062] In this embodiment, corresponding parsing methods are used for the received text information, acoustic information, and visual information to extract core features, thereby achieving the quantification and standardization of multimodal information.

[0063] Natural Language Processing (NLP) technology is used to parse the text information and extract core text features, including keywords (such as Java, project management, and cross-departmental collaboration), semantic similarity (the semantic match between the candidate's answer and the job requirements), logical coherence (the clarity of the answer), content completeness (whether the question was answered completely and whether key details were added), and accuracy (whether there are any errors or contradictions in the expression). For example, if a candidate answers, "I used Java to develop the backend interface and coordinated with the frontend and testing teams to complete the integration testing and ensure the stable operation of the interface," the extracted keywords include Java, backend interface, and cross-departmental coordination. The logical coherence score is 0.8 (out of 1), and the content completeness score is 0.9.

[0064] Speech signal processing technology is used to analyze acoustic information and extract core acoustic features, including speech rate (e.g., words per minute), intonation variance (reflecting the degree of intonation fluctuation; the larger the variance, the more fluctuating the intonation), pause features (e.g., the average number of pauses per sentence and the longest single pause duration), and tone features (judging tone as firm, hesitant, or perfunctory through parameters such as sound frequency and amplitude). For example, if a candidate answers questions with a fast speech rate (180 words per minute), few pauses (0.5 times per sentence on average), and a firm tone, it can be judged that they are familiar with the questions; if the speech rate is slow, there are many pauses (3 times per sentence on average), and the tone is hesitant, it may indicate that they are unfamiliar with the questions or are concealing information.

[0065] Computer vision technologies (such as facial recognition and posture recognition) are used to analyze visual information and extract core visual features, including facial expression features (such as recognizing smiles, frowns, and averted eye contact through facial motion units), posture features (such as upright posture, slouching, and frequent head-down posture), and eye contact features (such as the duration and frequency of eye contact with the camera). For example, when a candidate answers a question, frequent eye contact (more than 60% of the time spent in eye contact), upright posture, and a smile can reflect their confidence and composure; while averted eye contact, poor posture, and frowning may indicate nervousness, lack of confidence, or inaccurate answers.

[0066] In this embodiment, the interview cognitive graph is comprehensively updated based on the extracted text features, acoustic features, and visual features, resulting in an updated cognitive graph. The update mainly includes four aspects to ensure the real-time performance and completeness of the cognitive graph, specifically including: Supplementing Evidence Nodes: Keywords and core expressions from textual features, tone and speech rate from acoustic features, and facial expressions and postures from visual features are added as new evidence nodes to the cognitive graph and associated with corresponding ability dimension nodes. For example, mentioning a candidate's use of Java to develop backend interfaces is used as an evidence node, associated with the Java development dimension of professional skills, and marked as a supporting relationship; while avoiding eye contact and hesitant tone during the candidate's answers are used as evidence nodes, associated with the confidence dimension of communication skills, and marked as a weakening relationship.

[0067] Update logical relationships: Adjust the logical relationships between nodes based on the newly added evidence nodes. For example, if the candidate did not mention cross-departmental collaboration experience in their previous answer, but added that they coordinated the front-end and testing teams to complete joint debugging in this answer, then add a supporting relationship between the cross-departmental collaboration dimension of communication ability and the evidence node of coordinating cross-departmental teams; if the candidate's current answer contradicts their previous statement (e.g., they said they led a project before, but said they participated in a project this time), then add a contradictory relationship between the two evidence nodes and mark it as an unresolved contradiction.

[0068] Revise representation nodes: Based on acoustic and visual features, update the candidate's performance representation nodes, such as adding representation nodes like fluent and logical answers, and confidence and composure to the cognitive map, associating them with the corresponding ability dimensions, and providing a reference for subsequent evaluation.

[0069] Adjusting information gaps: Based on newly added evidence, determine whether previous information gaps have been filled. If filled, delete the corresponding information gap marker; if not fully filled, update the specific content of the information gap to provide a basis for the next round of follow-up questions. For example, if the previous information gap regarding Java development skills was the lack of specific project examples, and the candidate has now provided relevant examples, then that information gap is deleted; if the candidate only provided some details without specifying the frameworks and tools used, then the information gap is updated to "lack of experience using Java development frameworks and tools."

[0070] In this embodiment, the candidate's language logic, emotional state, and body language are synchronized through timeline alignment to achieve multimodal feature correlation. Then, cross-modal consistency reasoning is used to determine the consistency of information conveyed by each feature, avoiding evaluation bias caused by a single feature. Moreover, the interview cognitive graph is dynamically updated based on the reasoning results, creating or adjusting evidence nodes and their logical relationships with ability dimensions to achieve real-time accumulation and logical improvement of evidence. This allows the interview cognitive graph to dynamically iterate with the interview process, always maintaining an accurate mapping of the candidate's abilities. At the same time, through multimodal feature fusion and consistency verification, the reliability and comprehensiveness of the evidence are improved, laying a solid foundation for subsequent cognitive convergence judgment and evaluation conclusion generation.

[0071] In this embodiment, step S105 involves determining whether the cognitive convergence threshold for the current depth assessment goal has been reached. If not, the interview cognitive graph is replaced with an updated cognitive graph, and the process returns to executing the process of determining the current interview stage's depth exploration strategy and current depth assessment goal using a pre-trained first language model based on the interview cognitive graph. This includes: determining whether the current depth assessment goal meets the requirements of a sufficient number of evidence nodes, diverse evidence sources, and the absence of unresolved major logical contradictions; if not, the interview cognitive graph is replaced with an updated cognitive graph, and the process returns to executing the process of determining the current interview stage's depth exploration strategy and current depth assessment goal using a pre-trained first language model based on the interview cognitive graph; generating targeted follow-up questioning strategies based on the current interview stage's depth exploration strategy and current depth assessment goal, including requesting specific details, clarifying inconsistent statements, or introducing new pressure situations; and generating specific interaction content for the next round using a pre-trained second language model based on the follow-up questioning strategies, until the cognitive convergence threshold for the current depth assessment goal is reached.

[0072] In this embodiment, the cognitive convergence threshold is a preset standard for judging whether the current in-depth evaluation target has been sufficiently evaluated. It is set by the enterprise according to the recruitment needs, that is, judging whether the current in-depth evaluation target meets the requirements of the number of evidence nodes, the diversity of evidence sources, and the absence of unresolved major logical contradictions.

[0073] The number of evidence nodes meets the standard, meaning that the number of evidence nodes corresponding to the current in-depth assessment target meets the preset minimum requirement (e.g., at least 5 valid evidence nodes for the professional skills dimension and at least 3 valid evidence nodes for the soft skills dimension), and the evidence sources are diverse (e.g., there is evidence from the resume, as well as evidence from interview questions and answers, and multimodal feature evidence).

[0074] Diverse sources of evidence mean that the evidence points need to come from different channels to avoid the one-sidedness of a single source. For example, evidence in the dimension of professional skills should include skill descriptions in the resume, answers in the interview, the firmness of tone in acoustic features, and the confidence in visual features, etc., and should include at least two or more sources. The absence of unresolved major logical contradictions means that there are no unresolved major contradictions (such as contradictions in the description of core experience or core skills) between the evidence nodes corresponding to the current in-depth assessment target. If there are minor contradictions (such as inconsistencies in the details of the description) that do not affect the overall assessment conclusion, then the condition can be determined to be met.

[0075] By comparing the relevant information of the current in-depth assessment objective in the current interview cognitive map with the three conditions of the cognitive convergence threshold, it is determined whether the convergence requirements have been met. For example, if the current in-depth assessment objective is project management ability, and the preset minimum number of evidence nodes is 3, and there are currently 4 evidence nodes (1 from the resume, 2 from interview questions and answers, and 1 from visual features), with diverse evidence sources and no unresolved major contradictions, then the cognitive convergence threshold has been met. If there are only 2 evidence nodes, or if there are unresolved major contradictions between leading projects and participating projects, then the convergence threshold has not been met.

[0076] In this embodiment, if it is determined that the cognitive convergence threshold for the current depth assessment target has been reached, then one of the following two branches will be entered, which will be determined based on the overall progress of the interview.

[0077] Branch 1: If the current interview does not cover all core competency dimensions in the dynamic job profile, or if there are other assessment targets with higher urgency scores, then return to step S102, replace the interview cognitive map with an updated cognitive map, and return to execute the deep exploration strategy and current deep assessment target for the current interview stage determined by the first pre-trained language model based on the interview cognitive map, and start a new round of deep exploration and interaction.

[0078] Branch Two: If the current interview has covered all core competency dimensions in the dynamic job profile, and all core assessment objectives have reached the cognitive convergence threshold, with no unresolved major contradictions or information gaps, then the interview process will be terminated, and the process will proceed to the subsequent comprehensive assessment report generation stage.

[0079] In this embodiment, if it is determined that the cognitive convergence threshold for the current depth assessment target has not been reached, the specific reasons for the non-convergence are located based on the updated interview cognitive map (such as insufficient evidence, unresolved contradictions, or unfilled information gaps). Then, the process returns to step S102, redetermines the current depth assessment target and the depth exploration strategy for the current depth assessment target (adjusts the questioning direction and focuses on the reasons for non-convergence), and generates targeted interactive content through step S103 to trigger a new round of recursive follow-up questions. Step S105 is repeated until the cognitive convergence threshold is reached.

[0080] For example, if the current in-depth assessment target is project management ability, but it has not reached the convergence threshold because of an unresolved contradiction (the candidate states that they led a project, but this is not mentioned in their resume), then the in-depth exploration strategy is redefined as behavioral event tracing. This generates targeted interactive content, such as, "You just mentioned leading a project. What was the name of this project? What were the start and end dates? This information is not mentioned in your resume; could you please explain in detail?" After receiving the candidate's multimodal response information, the features are analyzed, the cognitive map is updated, and convergence is assessed again until the contradiction is resolved and the convergence threshold is reached.

[0081] It should be noted that the number of recursive follow-up questions can be preset to an upper limit (e.g., a maximum of 3 recursive follow-up questions for the same evaluation goal). If the upper limit is reached and convergence is still not achieved, a phased evaluation will be conducted based on the existing evidence, the reasons for non-convergence will be marked, and the process will move on to the next evaluation goal to avoid an excessively long interview process.

[0082] In this embodiment, the core conditions for cognitive convergence are clearly defined: sufficient evidence quantity, diverse sources, and no major logical contradictions, ensuring the rigor of the assessment. Simultaneously, if the convergence threshold is not reached, information gaps or logical contradictions are precisely located to avoid blindly pursuing further questions. Furthermore, targeted questioning strategies are generated based on the type of gap, and personalized interactive content is generated using the second language model to achieve targeted recursive questioning. In addition, multiple rounds of recursive questioning until the convergence threshold is reached ensure that each in-depth assessment target receives sufficient and reliable evidence, eliminating superficial questions encountered in interviews and ensuring that the core competency dimensions of the position are thoroughly explored. At the same time, the timely resolution of logical contradictions improves the accuracy and rigor of the interview cognitive map, providing a strong guarantee for the objectivity of the final assessment conclusion.

[0083] In this embodiment, in step S106, when the cognitive convergence threshold for the current depth assessment target is reached, the final cognitive map is obtained, and a main assessment conclusion is generated based on the final cognitive map. The main assessment conclusion is challenged and debated using a pre-trained third language model to generate a comprehensive assessment report. The comprehensive assessment report includes the core arguments of both sides, including: when the cognitive convergence threshold for the current depth assessment target is reached, the final cognitive map is obtained, and a quantitative score is calculated and a preliminary main assessment conclusion is generated based on the evidence weights and consistency scores of each ability dimension in the final cognitive map; the final cognitive map and the preliminary main assessment conclusion are input into the pre-trained third language model to find evidence supporting the opposite conclusion and generate opposing opinions, which include opposing arguments and corresponding evidence citations; the opposing opinions are input into the pre-trained main assessment model so that the pre-trained main assessment model and the pre-trained third language model conduct multiple rounds of simulated debate, and the arguments and evidence during the debate process are recorded to generate simulated debate records; based on the simulated debate records, a comprehensive assessment report is generated, which includes the core arguments of both sides.

[0084] In this embodiment, when the cognitive convergence threshold for the current depth assessment target is reached, all information in the final state interview cognitive graph is extracted, including evidence nodes, logical relationships, representation nodes, matching gaps, etc. of each ability dimension. Based on the preset assessment algorithm, the quantitative score of the candidate in each ability dimension (e.g., 80 points for professional skills and 75 points for communication skills) is calculated, and combined with the weight of the dynamic job profile, the comprehensive score is calculated to generate the preliminary conclusion of the main assessment.

[0085] In this embodiment, the core content of the preliminary conclusion of the main assessment includes the candidate's overall score, specific scores and brief descriptions of each ability dimension, the candidate's match with the position (e.g., 85% match, recommended to proceed to the second interview; 60% match, not recommended for hiring), core strengths (e.g., solid professional skills, rich project management experience), and core weaknesses (e.g., cross-departmental communication skills need improvement, lack of relevant case support). All conclusions must be linked to the corresponding evidence nodes in the cognitive map to ensure that they are based on evidence.

[0086] In this embodiment, the pre-trained third language model is a base model (such as ERNIE4.0 or Claude3) with strong logical debate, evidence mining, and refutation capabilities. The training set includes a large number of interview evaluation conclusions, debate cases for both sides, and correlation data between evidence and evaluation conclusions. Through a combination of adversarial training and supervised fine-tuning, the pre-trained third language model can accurately mine evidence in the interview cognitive graph that supports the opposite evaluation conclusions, generate reasonable opposing arguments, and have the ability to conduct multiple rounds of logical debate with the main evaluation model, ensuring the dialectical and objective nature of the evaluation conclusions.

[0087] In this embodiment, the third language model acts as the opposing side. Based on the initial conclusion of the main assessment, it mines evidence from the interview cognitive map to support the opposite conclusion, generating objections. These objections include clear counterarguments and corresponding evidence citations, avoiding unfounded rebuttals. For example, if the main assessment concludes that the candidate has excellent project management skills, the third language model might find evidence that the candidate has only led one small project, lacks experience in large-scale project management, and that the project delivery cycle exceeded expectations. This generates an objection arguing that the candidate's project management skills do not reach an excellent level, and cites corresponding evidence nodes.

[0088] In this embodiment, the objections generated by the third language model are input into the main evaluation model (the model responsible for generating the main evaluation conclusion). The main evaluation model, acting as the affirmative side, refutes the objections based on evidence from the interview cognitive map and proposes supplementary arguments to support the main evaluation conclusion. The third language model (the negative side) then further explores evidence in response to the affirmative side's rebuttal and proposes new objections. This process is repeated for multiple rounds of simulated debate (e.g., 2-3 rounds) until neither side can propose any new arguments.

[0089] The main assessment model is also based on general large language models such as ERNIE, and has been specially fine-tuned for interview assessment scenarios. The training set includes a massive amount of interview cognitive graphs, quantitative scoring cases of ability and assessment conclusion samples. It has the core ability to calculate quantitative scores based on graph evidence, generate preliminary assessment conclusions and refute opposing arguments.

[0090] In this embodiment, the entire process of multiple rounds of simulated debate is recorded in real time, including each argument of the affirmative and negative sides, the corresponding evidence cited (evidence nodes in the associated cognitive map), and the logical relationship between the arguments, forming a complete simulated debate record.

[0091] In this embodiment, based on the preliminary conclusions of the main assessment and the simulated debate record, the core arguments of both sides are integrated to generate a comprehensive assessment report, the core content of which includes the following five parts: Candidate Basic Information: Briefly introduce the candidate's name, education, work experience, applied position, and other core information; The main assessment conclusion should clearly define the candidate's overall score, job suitability, core strengths and weaknesses, and adopt the core content of the preliminary conclusion of the main assessment, making appropriate revisions based on the debate results. The affirmative side's core arguments: In the integrated mock debate, the affirmative side (main evaluation model) supported the main evaluation conclusion with the core arguments and corresponding evidence, such as the candidate's 100% on-time delivery rate of the project and the coordination of 5 cross-departmental teams to solve core problems, proving that the candidate has excellent project management capabilities. The opposing side's core arguments: Integrating the core arguments and corresponding evidence presented by the opposing side (the third language model) against the main evaluation conclusion in the simulated debate, such as the candidate's relatively small-scale projects, lack of experience in managing large projects, and unclear logic in his answers when facing stress tests, proving that his project management ability still has room for improvement. Recruitment Recommendations: Based on the comprehensive evaluation results, we provide specific recruitment recommendations, such as recommending that candidates proceed to the second interview to further assess their large-scale project management capabilities; if the match is high, we recommend hiring them; if their core competencies do not meet the job requirements, we do not recommend hiring them.

[0092] In this embodiment, a quantitative main assessment conclusion is generated based on the evidence weight and consistency score of the final interview cognitive map, ensuring the objectivity and quantifiability of the assessment. Simultaneously, the use of a third language model to challenge the main assessment conclusion, uncovering opposing evidence and generating dissenting opinions, breaks the limitations of a single assessment perspective. Furthermore, through multiple rounds of simulated debate between the main assessment model and the third language model, the arguments of both sides are comprehensively reviewed, ensuring the comprehensiveness of the assessment conclusion. In addition, the final comprehensive report includes core arguments from both sides, recruitment suggestions, etc., clearly presenting the assessment results and providing complete supporting evidence and decision-making references. By using dialectical thinking, it breaks the subjectivity and one-sidedness of traditional assessments, making the assessment conclusion more objective, comprehensive, and persuasive, providing a scientific and reliable basis for recruitment decisions.

[0093] The above describes the intelligent interview method based on a large language model in the embodiments of the present invention. The following describes the apparatus in the embodiments of the present invention. Please refer to [link / reference]. Figure 2 The implementation methods of the intelligent interview device based on a large language model in this invention include: The first generation module 201 is used to obtain job description information and candidate resume information, and generate a dynamic job profile and interview cognition map based on the job description information and candidate resume information. The determination module 202 is used to determine the deep exploration strategy and the current deep evaluation target for the current interview stage based on the current state of the interview cognitive graph and through the pre-trained first large language model. The second generation module 203 is used to generate interactive content for candidates based on the deep exploration strategy through a pre-trained second language model, and to receive candidate multimodal response information based on the interactive content. The parsing module 204 is used to parse the candidate's multimodal response information to obtain text features, acoustic features and visual features, and update the interview cognitive map based on the text features, the acoustic features and the visual features to obtain the updated cognitive map; The judgment module 205 is used to determine whether the cognitive convergence threshold for the current depth assessment target has been reached. If not, the updated cognitive graph is used to replace the interview cognitive graph, and the execution of the first pre-trained language model based on the interview cognitive graph to determine the current interview stage's depth exploration strategy and current depth assessment target is returned. The third generation module 206 is used to obtain the final cognitive map when the cognitive convergence threshold for the current depth evaluation target is reached, and generate the main evaluation conclusion based on the final cognitive map. The main evaluation conclusion is questioned and debated by the pre-trained third language model to generate a comprehensive evaluation report, which includes the core arguments of both sides.

[0094] In this embodiment, by acquiring job descriptions and candidate resumes, a dynamic job profile and interview cognitive map are generated, achieving precise matching between jobs and candidates, breaking the limitations of the traditional one-size-fits-all approach to interviews. Furthermore, through the collaborative efforts of three pre-trained language models—the first accurately determines deep exploration strategies and evaluation goals, the second generates personalized interactive content, and the third verifies evaluation conclusions through questioning and debate—human subjective bias is significantly reduced, enhancing the relevance and scientific rigor of the interview. Simultaneously, the integration of multimodal information parsing technology comprehensively captures the candidate's textual, acoustic, and visual features, combined with a recursive questioning mechanism, enabling in-depth exploration of the candidate's true abilities, avoiding superficial evaluations, and ensuring comprehensive and reliable evaluation results. Moreover, by using cognitive convergence threshold judgment, the interview pace is reasonably controlled, achieving full automation of the interview process, significantly reducing the interviewer's workload and shortening the interview cycle. The final comprehensive evaluation report includes core arguments from both sides, clearly presenting the candidate's strengths and weaknesses, providing a comprehensive and reliable reference for recruitment decisions. It also possesses good job-company fit and can be widely applied to recruitment scenarios across various industries.

[0095] Figure 2 The structure of the intelligent interview device based on a large language model shown does not constitute a limitation on the intelligent interview device based on a large language model, and can implement the steps of the intelligent interview method based on a large language model provided in the above-described method embodiments.

[0096] above Figure 2 The intelligent interview device based on a large language model in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The intelligent interview device based on a large language model in this embodiment of the invention will be described in detail from the perspective of hardware processing.

[0097] Figure 3This is a schematic diagram of the structure of an intelligent interview device based on a large language model according to an embodiment of the present invention. The device 300 can vary significantly due to different configurations or performance, and may include one or more central processing units (CPUs) 310 (e.g., one or more processors) and a memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 333 or data 332. The memory 320 and storage media 330 can be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown), each module including a series of instruction operations on the device 300. Furthermore, the processor 310 may be configured to communicate with the storage media 330 and execute the series of instruction operations in the storage media on the device 300.

[0098] Device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.

[0099] This invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the steps of an intelligent interview method based on a large language model.

[0100] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0101] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0102] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent interview method based on a large language model, characterized in that, The intelligent interview method based on a large language model includes: Obtain job description information and candidate resume information, and generate a dynamic job profile and interview cognitive map based on the job description information and candidate resume information; Based on the current state of the interview cognitive graph, the deep exploration strategy and the current deep evaluation target for the current interview stage are determined through the pre-trained first language model. Based on the aforementioned deep exploration strategy, interactive content for candidates is generated through a pre-trained second language model, and multimodal response information of candidates is received based on the interactive content. The candidate's multimodal response information is analyzed to obtain text features, acoustic features, and visual features. The interview cognitive map is then updated based on the text features, acoustic features, and visual features to obtain the updated cognitive map. Determine whether the cognitive convergence threshold for the current depth assessment target has been reached. If not, update the cognitive graph to replace the interview cognitive graph, and return to execute the first pre-trained language model based on the interview cognitive graph to determine the depth exploration strategy and the current depth assessment target for the current interview stage. When the cognitive convergence threshold for the current depth assessment target is reached, the final cognitive map is obtained, and the main assessment conclusion is generated based on the final cognitive map. The main assessment conclusion is then questioned and debated through a pre-trained third language model to generate a comprehensive assessment report, which includes the core arguments of both sides.

2. The intelligent interview method based on a large language model according to claim 1, characterized in that, The process of obtaining job description information and candidate resume information, and generating a dynamic job profile and interview cognitive map based on the job description information and candidate resume information, includes: Acquire and preprocess job description information and candidate resume information; Deep semantic analysis is performed on the preprocessed job description information and candidate resume information to extract job requirement elements and candidate ability elements. Based on the extracted job requirement elements, a dynamic job profile is generated, which includes hard qualifications, responsibilities, abilities and qualities, organizational environment and evaluation preferences. The extracted candidate competency elements and dynamic job profiles are preliminarily mapped and gap analyzed. Based on the results of the preliminary mapping and gap analysis, an interview cognitive graph is constructed. The interview cognitive graph includes entity nodes, evidence nodes, logical nodes, and relational edges representing the membership, support, contradiction, triggering, or association relationships between the nodes.

3. The intelligent interview method based on a large language model according to claim 1, characterized in that, Based on the current state of the interview cognitive graph, the deep exploration strategy and current deep evaluation target for the current interview stage are determined through a pre-trained first large language model, including: The current state of the interview cognitive graph is diagnosed by using a pre-trained first language model to obtain the situation diagnosis results, which include the sufficiency and consistency of existing evidence under each ability dimension and whether there are any unresolved contradictions. Combining the situation diagnosis results and the preset weights of each capability dimension in the dynamic job profile, the current urgency score of each capability dimension in the dynamic job profile is calculated using a multi-factor weighted algorithm. The current ability dimension with the highest urgency score is selected as the current in-depth assessment target. Based on the attributes of the current in-depth assessment target and the type of information gap exposed in the interview cognitive map, an in-depth exploration strategy is matched and generated from a pre-set strategy library. The in-depth exploration strategy includes in-depth stress testing, behavioral event tracing, or value probes.

4. The intelligent interview method based on a large language model according to claim 1, characterized in that, The process of generating interactive content for candidates based on the deep exploration strategy using a pre-trained second language model, and receiving candidate multimodal response information based on the interactive content, includes: Extract relevant context and current dialogue history from the interview cognitive graph; The deep exploration strategy, the extracted relevant context, and the current dialogue history are input into the pre-trained second language model so that the pre-trained second language model generates interactive content for the candidate under policy constraints. The interactive content includes specific situational details, natural language questions, or simulated scene descriptions. The interactive content is presented to the candidate, and multimodal response information of the candidate is received based on the interactive content. The multimodal response information includes text information, acoustic information, and visual information.

5. The intelligent interview method based on a large language model according to claim 1, characterized in that, The candidate's multimodal response information is parsed to obtain textual features, acoustic features, and visual features. Based on these features, the interview cognitive map is updated to obtain the updated cognitive map. For candidate multimodal response information, text transcription obtained through speech recognition is used as text features. Acoustic features are extracted from intonation, speech rate and pauses through acoustic analysis, and facial movement units and posture features are extracted from visual analysis. Align the text features, acoustic feature sequences, and visual features on the time axis to obtain aligned multimodal features; Cross-modal consistency inference is performed based on aligned multimodal feature sequences to determine the consistency relationship between text features, acoustic feature sequences and visual features, and to obtain consistency inference results; Based on the results of consistent reasoning, evidence nodes are created in the interview cognitive graph, and the supporting, weakening, or contradictory logical relationships between the created evidence nodes and the relevant ability dimension nodes are established, or the evidence nodes are updated and the supporting, weakening, or contradictory logical relationships between the updated evidence nodes and the relevant ability dimension nodes are adjusted to obtain an updated cognitive graph.

6. The intelligent interview method based on a large language model according to claim 1, characterized in that, The determination of whether the cognitive convergence threshold for the current depth assessment target has been reached is made. If not, the interview cognitive graph is updated and replaced, and the process of determining the current deep exploration strategy and current depth assessment target for the current interview stage using the pre-trained first language model based on the interview cognitive graph is returned. This includes: Determine whether the current in-depth assessment objectives meet the following criteria: sufficient number of evidence nodes, diverse sources of evidence, and absence of unresolved major logical contradictions. If not, the updated cognitive graph is used to replace the interview cognitive graph, and the process is returned to execute the first pre-trained language model based on the interview cognitive graph to determine the current interview stage's deep exploration strategy and current deep evaluation target. Based on the current in-depth exploration strategy and current in-depth assessment goals of the interview stage, generate targeted follow-up questioning strategies, including requesting specific details, requesting clarification of inconsistent statements, or introducing new stressful situations. Based on the aforementioned questioning strategy, the next round of specific interactive content is generated through a pre-trained second language model until the cognitive convergence threshold for the current depth evaluation target is reached.

7. The intelligent interview method based on a large language model according to claim 1, characterized in that, When the cognitive convergence threshold for the current depth assessment target is reached, the final cognitive map is obtained, and a main assessment conclusion is generated based on the final cognitive map. The main assessment conclusion is then challenged and debated using a pre-trained third language model, generating a comprehensive assessment report. This comprehensive assessment report includes core arguments from both sides, including: When the cognitive convergence threshold for the current depth assessment target is reached, the final cognitive map is obtained, and based on the evidence weights and consistency scores of each ability dimension in the final cognitive map, a quantitative score is calculated and a preliminary conclusion of the main assessment is generated. The final cognitive map and the preliminary conclusions of the main evaluation are input into the pre-trained third language model to find evidence supporting the opposite conclusions and generate objections, which include opposing arguments and corresponding evidence citations. The opposing opinions are input into the pre-trained master evaluation model, so that the pre-trained master evaluation model and the pre-trained third language model can conduct multiple rounds of simulated debate, and the arguments and evidence in the debate process are recorded to generate simulated debate records; Based on the simulated debate record, a comprehensive evaluation report is generated, which includes basic information about the candidate, the main evaluation conclusion, the core arguments of both sides, and recruitment recommendations.

8. An intelligent interview device based on a large language model, characterized in that, include: The first generation module is used to obtain job description information and candidate resume information, and generate a dynamic job profile and interview cognitive map based on the job description information and candidate resume information; The determination module is used to determine the deep exploration strategy and the current deep evaluation target for the current interview stage based on the current state of the interview cognitive graph and through the pre-trained first language model. The second generation module is used to generate interactive content for candidates based on the deep exploration strategy using a pre-trained second language model, and to receive candidate multimodal response information based on the interactive content. The parsing module is used to parse the candidate's multimodal response information to obtain text features, acoustic features, and visual features, and update the interview cognitive map based on the text features, acoustic features, and visual features to obtain the updated cognitive map; The judgment module is used to determine whether the cognitive convergence threshold for the current depth assessment target has been reached. If not, the updated cognitive graph is used to replace the interview cognitive graph, and the execution of the first pre-trained language model based on the interview cognitive graph to determine the current interview stage's depth exploration strategy and current depth assessment target is returned. The third generation module is used to obtain the final cognitive map when the cognitive convergence threshold for the current depth evaluation target is reached, and to generate the main evaluation conclusion based on the final cognitive map. The main evaluation conclusion is then questioned and debated by the pre-trained third language model to generate a comprehensive evaluation report, which includes the core arguments of both sides.

9. An intelligent interview device based on a large language model, characterized in that, It includes a memory and at least one processor, wherein the memory stores computer-readable instructions; The at least one processor invokes the computer-readable instructions in the memory to perform the steps of the intelligent interview method based on a large language model as described in any one of claims 1-7.

10. A computer-readable storage medium storing computer-readable instructions thereon, characterized in that, When the computer-readable instructions are executed by the processor, they implement the steps of the intelligent interview method based on a large language model as described in any one of claims 1-7.