Digital human adaptive interactive interview feedback method and system

CN122472726BActive Publication Date: 2026-09-29UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610971543.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-01
Publication Date
2026-09-29
Estimated Expiration
2046-07-01

AI Technical Summary

Technical Problem

1)大模型遗忘导致记忆断裂与逻辑不连贯

Benefits of technology

[0009]与现有技术相比,本发明所提供的方法通过构建全局动态共享记忆池与多角色智能体协同框架,实现了跨阶段的上下文状态继承与逻辑一致性,大幅提升了数字人面试的科学性、针对性与反馈的实际指导价值,对人工智能在复杂交互场景与高阶人才选拔领域的落地具有重要的实际应用意义,有助于辅助面试效果,提升企业招聘与高校招生筛选人才的能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122472726B_ABST
    Figure CN122472726B_ABST
Patent Text Reader

Abstract

The application discloses a digital human adaptive interactive interview feedback method and system, in the pre-interview and initial interaction stage, candidate labels are collected through multi-source interaction, the initial image of the candidate is constructed, and a global dynamic shared memory pool is initialized; based on the initial image, a multi-agent collaborative architecture is called to perform cross-stage interview interaction, logical coherence is maintained through horizontal breadth and vertical depth, and the global dynamic shared memory pool is dynamically summarized and updated; in the interview interaction deep digging stage, the RAG technology generated by combining retrieval enhancement and the historical records in the global dynamic shared memory pool is used to trigger adaptive dynamic follow-up questions and open high-order seminars; through fusion of dynamic scores in each stage of the interview, a forced structured high-dimensional evaluation feedback report is generated. The method realizes cross-stage context state inheritance and logical consistency, greatly improves the scientificity, pertinence and actual guiding value of the feedback of the digital human interview.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a digital human adaptive interactive interview feedback method and system. Background Technology

[0002] Digital human interview systems, as efficient recruitment and assessment tools, have been widely used in corporate screening and university mock interviews. They aim to assist human interviewers in automating the evaluation of candidates' comprehensive qualities and professional skills through multi-round interactive dialogues. Early interview systems were mostly based on fixed templates or question banks, resulting in a rigid question-and-answer process that struggled to handle real high-level interview scenarios (such as recruitment for senior technical positions and doctoral selection). Real high-level interviews are not single-point question-and-answer sessions, but rather dynamic, context-dependent, long-sequence interactive processes. The system needs to accurately remember the candidate's historical answer logic, knowledge gaps, and on-the-spot performance across multiple assessment stages, including resume feature mining and in-depth professional knowledge exploration, and generate targeted, in-depth follow-up questions accordingly to realistically probe the candidate's cognitive boundaries and thought coherence.

[0003] However, real-world interviews involve complex interactions across multiple stages. Most existing methods rely solely on a single large language model or a pre-defined rule engine to drive the interview process. While these methods achieve some degree of automated question answering and basic follow-up questions, they still suffer from the following significant shortcomings when dealing with lengthy, in-depth, and multimodal complex interview scenarios: 1) Large-scale model forgetting leads to memory fragmentation and logical incoherence. Existing technologies typically rely on a single large language model to generate sequential dialogues. In long, multi-round interviews, the continuously increasing dialogue history can easily exceed the model's context window limitations. This underlying architectural flaw can cause the model to suffer from "catastrophic forgetting," resulting in a disconnect between the preceding and following parts of the interview, and a lack of information linkage and logical coherence between different assessment stages. 2) The follow-up questioning mechanism lacks specificity and in-depth cognitive probing capabilities. Existing systems largely rely on static keyword matching or unconstrained free-flowing large models for question generation, lacking real-time quantitative assessment of the candidate's current answer in terms of "content coverage" and "logical chain completeness." The system cannot accurately pinpoint logical gaps or cognitive deficiencies in the candidate's answer, making it difficult to adaptively generate compelling and targeted in-depth follow-up questions, and consequently, to effectively measure the candidate's true cognitive boundaries.

[0004] 3) Lack of open-ended discussion capabilities in high-level professional scenarios. The generation logic of existing systems is mostly based on closed-domain question banks or the model's own general prior knowledge. When facing highly specialized scenarios such as scientific research and postgraduate entrance examinations, it can only remain at a superficial question-and-answer level, focusing on the correctness of knowledge points. The system lacks a mechanism to introduce external high-level long texts (such as cutting-edge academic papers) and extract key entities, making it unable to construct non-conclusion-oriented open-ended discussions. This makes it difficult to quantitatively assess candidates' higher-level potential, such as problem awareness and methodological transfer. 4) Interview feedback evaluation dimensions are singular and static. Most existing systems provide a one-time static summary of the dialogue text or a single-dimensional accuracy score after the interview, ignoring the time-series characteristics of the interview process. They fail to incorporate dynamic performance characteristics such as "performance fluctuation variance between stages" and "error correction ability under pressure and follow-up questioning" into the evaluation system, resulting in output feedback that lacks multi-dimensional interpretability and scientific guidance. 5) Lack of Personalized Customization and Precise Path Planning in the Pre-Interview "Cold Start" Phase: Existing systems typically use standardized opening lines or fixed initial question banks in the initial interaction phase, failing to combine the candidate's proactive intentions (such as target position) with passive resume information (such as structured resume features). This indiscriminate, uniform questioning approach prevents the system from accurately identifying the candidate's true background and makes it difficult to dynamically generate a personalized initial interview path, severely impacting the targeting of subsequent interactions.

[0005] In view of this, the present invention is hereby proposed. Summary of the Invention

[0006] The purpose of this invention is to provide a digital human adaptive interactive interview feedback method and system to solve the aforementioned technical problems in the prior art. The method described in this invention achieves cross-stage context state inheritance and logical consistency by constructing a global dynamic shared memory pool and a multi-role intelligent agent collaborative framework. This significantly improves the scientific rigor, relevance, and practical guidance value of digital human interviews, and has important practical application significance for the implementation of artificial intelligence in complex interactive scenarios and high-level talent selection.

[0007] The objective of this invention is achieved through the following technical solution: A digital human adaptive interactive interview feedback method, the method comprising: Step 1: In the pre-interview and initial interaction stage, collect candidate tags through multi-source interaction, build the initial profile of the candidate, and initialize the global dynamic shared memory pool. Step 2: Based on the initial profile, invoke the multi-agent collaborative architecture to execute cross-stage interview interactions. Maintain logical coherence through horizontal breadth and vertical depth mining, and dynamically summarize and update the global dynamic shared memory pool to ensure the context of different stages of the interview and prevent forgetting. Step 3: In the in-depth interview interaction stage, combine the retrieval enhancement generation RAG technology with the historical records in the global dynamic shared memory pool to trigger adaptive dynamic follow-up questions and open-ended high-level discussions. Step 4: Generate a mandatory structured, high-dimensional evaluation feedback report by integrating the dynamic scores from each stage of the interview.

[0008] A digital human adaptive interactive interview feedback system, the system comprising: The pre-interview path planning module is used to collect candidate tags through multi-source interaction during the pre-interview and initial interaction stages, build an initial candidate profile, and initialize a global dynamic shared memory pool. The cross-stage multi-agent control module is used to call the multi-agent collaborative architecture to execute cross-stage interview interactions based on the initial profile. It maintains logical coherence through horizontal breadth and vertical depth mining, and dynamically summarizes and updates the global dynamic shared memory pool to ensure the context of different stages of the interview and prevent forgetting. The Real-Time Dynamic Follow-Up Questions and Advanced Discussion Module is used to trigger adaptive dynamic follow-up questions and open-ended advanced discussions during the in-depth interview interaction phase by combining retrieval-enhanced RAG technology with historical records in a globally dynamic shared memory pool. The structured assessment report generation module is used to generate a mandatory structured, high-dimensional assessment feedback report by integrating dynamic scores from each stage of the interview.

[0009] Compared with existing technologies, the method provided by this invention achieves cross-stage context state inheritance and logical consistency by constructing a global dynamic shared memory pool and a multi-role intelligent agent collaborative framework. This significantly improves the scientific nature, relevance, and practical guidance value of digital human interviews. It has important practical application significance for the application of artificial intelligence in complex interaction scenarios and high-level talent selection, and helps to improve interview effectiveness and enhance the ability of enterprises to recruit and universities to screen talents. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of the digital human adaptive interactive interview feedback method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of the digital human adaptive interactive interview feedback system provided in an embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them, and do not constitute a limitation on the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the protection scope of the present invention.

[0013] First, the following explanations are provided for the terms that may be used in this article: The term "and / or" means that either or both can be achieved simultaneously. For example, X and / or Y means that it includes both "X" or "Y" as well as the three cases of "X and Y".

[0014] The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0015] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0016] The technical solution provided by this invention will be described in detail below. Contents not described in detail in the embodiments of this invention belong to prior art known to those skilled in the art. For example... Figure 1 The diagram shown is a flowchart of a digital human adaptive interactive interview feedback method provided in an embodiment of the present invention. The method includes: Step 1: In the pre-interview and initial interaction stage, collect candidate tags through multi-source interaction, build the initial profile of the candidate, and initialize the global dynamic shared memory pool. In this step, the data structure and concurrent processing mechanism of the global dynamic shared memory pool are specifically as follows: Firstly, for general multi-turn question-answering scenarios, a persistent memory unit based on a relational database is constructed. A connection to the relational database is established through a database connection component, and each turn of dialogue is stored as a structured persistent record. Each record includes a session identifier (session_id), a turn identifier (turn_id), a user identifier (user_id), a message role, and message content. The database layer automatically maintains the creation time (created_at) and update time (updated_at). Among them, the role is used to distinguish between user-side messages and agent-side messages; the turn_id is used to uniquely identify a single turn of dialogue; and the session_id is used to aggregate multiple turns of messages into the same session context. Based on the above structure, the system can restore the complete historical message sequence in chronological order, thereby realizing cross-round context splicing and continuous semantic reasoning; Then, when facing complex task scenarios such as mock interviews and paper discussions, a state snapshot memory unit based on JSON structure is constructed. Specifically, the resume is extracted using the Optical Character Recognition (OCR) method, the result is stored in the ocr_result field, and the multi-agent dialogue state snapshot is stored in the dialogue_record field. The dialogue_record field adopts a JSON array structure, and each element in the array represents a dialogue event or a state advancement event. The multi-agent dialogue state snapshot includes control fields: interview stage identifier (interviewer_state), background type (background), lab type (lab), major category (major), job category (job), paper discussion sub-state (paper_state), mentor identifier (teacher), research interest direction (interest), selected paper object (paper), candidate question pool (question_pool), current question index (current_question_index), whether follow-up questions have been triggered (followup_triggered), follow-up question index (followup_question_index), direction change count (direction_change_count), whether it is a follow-up question round (ex_query), and interview mode (interview_mode). Among them, the selected paper object is used to store paper metadata; the candidate question pool is used to store the candidate question pool automatically generated for paper discussion, including the question design intent. Through the above structure, the system not only remembers the historical question and answer text, but also the task flow state, question selection state, and follow-up question control state, forming a global dynamic shared memory pool that can be shared and accessed by multiple intelligent agent modules. For the concurrent processing mechanism of the global dynamic shared memory pool, a process consistency strategy based on task state machine and a two-layer truncation mechanism combining the business layer and template layer are adopted for ultra-long context control, wherein: The process consistency strategy based on task state machines is as follows: 1) For general multi-turn question-and-answer links, the atomicity of single write operations in relational databases is used to ensure basic consistency. Single-turn messages are used as the basic write unit, and incremental persistence of dialogue is achieved through session identifiers and turn identifiers. When the same turn identifier is written repeatedly, atomic updates are completed by using database unique key constraints in conjunction with insert conflict update semantics (e.g., ONDUPLICATE KEY UPDATE). 2) For complex task chains, a synchronous strategy of "reading the latest state snapshot -> appending new dialogues or state nodes to memory -> writing the updated complete JSON snapshot back to the database" is adopted, and each task advancement is an update of a complete state snapshot; In practice, when different stages of the intelligent agent module (such as the professional interview agent and the research experience agent) are woken up, they can obtain a unified understanding of the task stage by reading state fields such as interviewer_state and paper_state, which avoids state read and write conflicts and facilitates subsequent process backtracking and review.

[0017] The two-layer truncation mechanism combining the business layer and the template layer is as follows: 1) Business layer length control: Set an upper limit threshold for the length of the current user input (e.g., a maximum of 32,768 characters) to prevent abnormally long inputs from directly entering the model inference chain at the source; 2) Template layer window truncation: During the message assembly stage, the system prompts, historical conversation sequences, and current user input are concatenated in sequence. To adapt to different large models, a strategy combining sliding windows and left-side truncation is adopted.

[0018] For example, for Qwen-type templates, a reverse sliding window assembly algorithm is used to traverse historical question-answer pairs in reverse chronological order, and backfill the context starting from the most recent round. When the cumulative token length reaches the window limit max_window_size, the forward expansion stops. For Baichuan-type templates or general base classes, after retaining the token sequence formed by the most recent several rounds of historical messages, left truncation or tail pruning is performed to ensure that only the context sequence that is closest to the current time and has the most reference value for the current interaction is retained.

[0019] By using the data structure and concurrent processing mechanism of the global dynamic shared memory pool, we can save the general dialogue history in a structured way, express the stage state of collaborative tasks in the form of state snapshots, and achieve effective adaptation of large models and ultra-long windows without significantly increasing the complexity of reasoning.

[0020] In the specific implementation, during the pre-interview and initial interaction stage, candidates interact with the agent through daily dialogue boxes. The agent stores the conversation records in the auxiliary database and extracts core personality tags through the underlying algorithm and writes them into the global dynamic shared memory pool. At the same time, the agent combines the candidate's selected active tags (such as desired position and skill preferences) and the passive tags parsed from the resume to construct the candidate's initial profile. To construct a comprehensive, three-dimensional, and highly discriminative initial profile of candidates, and to avoid the biased information resulting from a single data source, the intelligent agent needs to extract features from three dimensions: subjective intentions, objective resumes, and implicit personality traits. The sources of candidate label features include: Actively select tags The candidate's preferred job position, preferred mentor, and core skills preferences, as voluntarily filled out; Passive parsing tags The intelligent agent extracts structured entity information related to educational background, professional skills, and project experience from candidates' resumes. Specifically, it uses OCR and natural language processing technologies to convert resumes into plain text and utilizes a pre-trained Named Entity Recognition (NER) model to extract entity information through sequence labeling. In the implementation, a default value fallback mechanism is set up to address information gaps during the recognition process: missing key fields are uniformly filled with a "to be detected" label, which is written as a state variable into the pre-constraint conditions so that targeted information completion questions can be automatically triggered during the cold start phase of the subsequent formal interview. Pre-interactive tags The long text records of conversations between candidates and agents are stored in the underlying auxiliary log database. The large model is called to perform semantic and sentiment analysis on the long text records of conversations. The large model outputs corresponding implicit labels, including labels such as strong willingness to communicate, extroverted personality, and high stress resistance, which are written into the core main database as pre-interaction labels. Finally, the three labels are vectorized, concatenated, and aggregated. The specific process is as follows: First, the discrete label text is converted into a unified structured descriptive text, for example, into a structured descriptive text of "Job Title: Algorithm Engineer; Skills: Python; Implicit Trait: Positive Communication Skills"; then, the structured descriptive text is input into a pre-trained text representation model for feature encoding, generating three sets of feature vectors; finally, the three sets of encoded feature vectors are concatenated and dimensionality-reduced according to their dimensions to form the candidate's global initial prior label set. , is represented as: ; Here, Concat() represents the vector concatenation operation, used to concatenate the actively selected label vectors. Passive parsing of label vectors and pre-interactive label vectors By performing head-to-tail concatenation along the feature dimension, a unified global initial prior label set is formed. ; The prior label set It is pre-written into the global dynamic shared memory pool and serves as a prerequisite constraint for the first stage of generating the underlying large model, dynamically generating exclusive initial entry points for different candidates.

[0021] Step 2: Based on the initial profile, invoke the multi-agent collaborative architecture to execute cross-stage interview interactions. Maintain logical coherence through horizontal breadth and vertical depth mining, and dynamically summarize and update the global dynamic shared memory pool to ensure the context of different stages of the interview and prevent forgetting. In this step, based on the initial profile, the resume interviewing agent lays out various basic questions horizontally according to the candidate's tags and job requirements to explore the breadth of the candidate's knowledge; and the professional interviewing agent conducts in-depth vertical analysis on specific answers to explore the depth of the candidate's knowledge. To prevent long-term question-and-answer sessions from causing "catastrophic forgetting" in large models, each agent at each stage extracts the personalized summary of the current stage, the deep logic of the candidates, and cognitive blind spots during handover, and writes them in a structured manner into the global dynamic shared memory pool. For the k-th stage, the agent extracts three core features based on the long text of the interaction at the current stage, including: deep logic. This refers to the thought process a candidate uses when analyzing complex problems and building solutions; and the knowledge points they have correctly grasped. This refers to the professional concepts and technical details that the candidate accurately states in their answers; cognitive blind spots. This refers to the candidate's weak areas of knowledge, such as answering incorrectly, deliberately avoiding questions, or expressing them vaguely. In terms of the specific extraction method, the large model is called and a preset instruction prompt is injected. The prompt first defines the large model as an interview review expert, then the complete dialogue record of the current stage is input into the large model, and finally the large model is forced to output three types of core features in JSON format.

[0022] At the end of the k-th stage, the agent uses an attention mechanism to reduce the dimensionality of the three core features, generating a personalized summary vector for the k-th stage. And update it to the global dynamic shared memory pool in real time. This process is represented as: ; in, This indicates the state of the global dynamic shared memory pool at the end of the previous interaction phase (or the previous round); This represents the latest global dynamic shared memory pool state obtained after incorporating the core features of the current stage; Update() represents the data update operation; When the agent in the (k+1)th stage initiates a deep question, it first reads the candidate's cognitive blind spots from the global dynamic shared memory pool. and deep logic .

[0023] The system adds dynamic performance tags based on the candidate's answers, including tags for clear logic or emotional tension when faced with follow-up questions. These dynamic performance tags are stored in a global dynamic shared memory pool to adjust the pressure level of subsequent questions.

[0024] The above mechanism ensures that even in the later stages of the interview, the agent can still remember the weaknesses shown by the candidate in the early stages and conduct coherent cross-dimensional discussions.

[0025] Step 3: In the in-depth interview interaction stage, combine the retrieval enhancement generation RAG technology with the historical records in the global dynamic shared memory pool to trigger adaptive dynamic follow-up questions and open-ended high-level discussions. In this step, during the in-depth interview interaction phase, when it is necessary to examine the candidate's true cognitive boundaries, the agent retrieves records of the candidate's incorrect answers from the global dynamic shared memory pool. This, combined with an external professional domain-specific retrieval-augmented generation (RAG) knowledge base, generates in-depth follow-up questions. The RAG is a real-time inference engine and retrieval-augmented generation technology. Specifically: Regarding the candidate's current answer The intelligent agent, based on the answer Two scores were obtained: Logical Integrity Score Content coverage score ; Among them, logical integrity score The acquisition method is as follows: The underlying reasoning engine performs dependency parsing and comparison between the answer text and the set standard problem-solving step tree. Through the pre-set thinking chain scoring Prompt of the large model, it performs normalized scoring from 0 to 1 on three dimensions: causal relationship, step coherence, and argument self-consistency. The weighted average of the three is taken to obtain the logical completeness score. The process is represented as: ; Where C, L, and S represent the causal relationship score, the step coherence score, and the argument consistency score after dependency parsing, respectively. These are preset weighting coefficients; Content coverage score The method of obtaining the information is as follows: The answer is extracted using Named Entity Recognition (NER) technology. The core entity word set in It also reads the set of knowledge points for the corresponding questions from the RAG knowledge base in the external professional field. The RAG is a real-time inference engine and retrieval enhancement generation technology; the content coverage score is obtained by calculating the semantic overlap and entity recall rate between the two. ,Right now: ; In obtaining the logical integrity score Content coverage score The answer is derived by combining the two indicators. Quality score , is represented as: ; in These are learnable or preset weighting coefficients; When quality score When the value falls below a set threshold, the innovative probing engine will be triggered; wherein, the threshold is preset based on the distribution of historical interview samples, and the preferred value range is [0.6, 0.8]. This innovative inquiry engine first retrieves cognitive blind spots recorded in the global dynamic shared memory pool. To address cognitive blind spots Input into the RAG knowledge base to retrieve relevant higher-order knowledge slices. ; Slice the higher-level knowledge retrieved Cognitive blind spots Current session context The inputs are combined and fed into a larger model as prompts, driving the model to generate a series of follow-up questions targeting the candidate's cognitive blind spots. The splicing formula is expressed as: ; in, The `SystemPrompt` function is used to define the role attributes, tone, and follow-up questioning logic constraints of the large model at the underlying level, acting as a "senior technical interviewer." `Concat()` represents a vector concatenation operation.

[0026] In the specific implementation, the structured Prompt template passed to the large model is: (as a senior technical interviewer) The candidate's answers in the previous conversation were superficial, and the system recorded that the candidate failed to grasp the following knowledge points in the previous stage: [ Please refer to the following authoritative knowledge base snippets: This generates a highly challenging series of follow-up questions, requiring cross-dimensional questioning of historical weaknesses to test the candidate's true technical depth. By feeding this structured Prompt into a large model, it can drive the model to generate targeted, innovative follow-up prompts.

[0027] For example, suppose a candidate answered a basic question incorrectly ("cache breakdown"). If the candidate's subsequent answer about "high-concurrency architecture" is superficial, the system will combine information from the RAG knowledge base regarding "advanced protection algorithms against cache avalanche" to directly raise a series of in-depth probing questions that incorporate past errors. This strategy generates multi-dimensional cross-validation questions by associating the candidate's historical forgetting points with the flaws in their current answers, effectively improving the accuracy of probing the candidate's true abilities. Traditional follow-up questions are often based on rigid matching of preset logic trees. This embodiment of the invention utilizes a real-time inference engine and the retrieval enhancement generation technology RAG to achieve an innovative and highly in-depth probing mechanism.

[0028] Furthermore, in high-level scenarios such as postgraduate admission and doctoral selection, standard Q&A alone cannot effectively assess a candidate's research potential. Therefore, this invention designs an open discussion module, and the specific process of triggering an open high-level discussion is as follows: First, based on the actively selected tags constructed in the initial stage, precise matching is performed in the RAG knowledge base to extract cutting-edge academic papers. To avoid candidates answering mechanically based on pre-set conclusions and using a masking mechanism to block direct conclusions, only the research background, core academic questions and research motivations are extracted and retained as discussion materials. The prompts ask the large model to introduce the research background to the candidate and guide the candidate to explore solutions; For example, "What challenges would you face if you applied this approach to one of your past projects?" In practice, the content template for this Prompt is defined as: "You are an academic interviewer. Please introduce the following research background to the candidate and guide them to explore possible solutions in an open-ended manner. Do not directly give the actual conclusions in the text."

[0029] The intelligent agent initiates questions using open-ended materials as a starting point, and, in conjunction with information from the global dynamic shared memory pool, guides candidates to conduct open-ended methodological discussions based on their research background. It quantitatively assesses the candidates' problem awareness and ability to transfer methods. Horizontal information refers to the breadth of the candidates' cross-domain knowledge, while vertical information refers to the depth of the logical chain and the depth of knowledge mastery during the in-depth exploration of specific technical points.

[0030] Step 4: Generate a mandatory structured, high-dimensional evaluation feedback report by integrating the dynamic scores from each stage of the interview.

[0031] In this step, at each stage of the interview, each agent independently scores the candidate in the background based on the candidate's real-time responses, forming a stage-by-stage score sequence. After the interview, the central control agent comprehensively calculates the scores at each stage, the variance of emotional expression fluctuations, and the ability to correct errors when questioned about them. It then uses a forced template to output information including the candidate's core strengths and weaknesses, key areas for improvement (combined with cognitive blind spots). A structured diagnostic assessment report that includes the cognitive boundaries revealed during follow-up questioning and the overall hiring preference ("strong recommendation", "conditional recommendation", etc.); among which, the error correction ability is calculated using a formula: ; in This represents a candidate's overall ability to correct errors when questioned about them, i.e., the dynamic error correction score; M represents the number of rounds of questioning. and , respectively, represent the logical accuracy scores before and after the error correction guidance; the subscript j indicates the j-th round of follow-up questioning.

[0032] The above mechanism effectively reduces invalid and redundant information, and enhances the relevance and guidance value of the assessment report.

[0033] In its specific implementation, the method further includes: For multi-agent cooperative scheduling and state machine transitions, a method combining prefix-tuning and reinforcement learning from human feedback (RLHF) is adopted to achieve lightweight multi-role adaptation and stable optimization of generation strategies. Specifically: In the prefix fine-tuning stage, for different stages in the multi-agent collaborative scenario, an independent learnable role prefix parameter matrix (i.e., a set of continuous, optimizable dense vectors used to guide the large model to present specific role prior knowledge) is set for each agent in each stage. The resume-based AI agent is responsible for verifying basic information and asking background questions based on the candidate's resume. The corresponding prefix parameter matrix is ​​as follows: The professional interview agent is responsible for assessing candidates' professional skills, project experience, and technical depth in a specific field. The corresponding prefix parameter matrix is ​​as follows: The higher-order academic agent is responsible for evaluating the candidate's research capabilities, academic vision, and ability to analyze complex theories; the corresponding prefix parameter matrix is ​​as follows. ; The corresponding prefix parameter matrix is ​​concatenated to the front of each multi-head attention layer in the Transformer model, and the corresponding attention calculation formula is updated as follows: ; in This represents the output (or feature representation) of the self-attention mechanism after incorporating the role-specific prefix parameters of each agent. and is the prefix parameter matrix for independent training of the agent at each stage; Q is the query vector, K is the key vector, V is the value vector; d is the vector dimension; is the scaling factor to prevent gradient vanishing; softmax is the normalization exponential function; the superscript T indicates the transpose sign; This method, by freezing the parameters of the main model and concatenating a small number of learnable vectors, can drive a large model to present corresponding role behaviors and questioning styles at different interview stages, effectively reducing the training overhead of multi-role adaptation. During the RLHF fine-tuning phase, the Proximal Policy Optimization (PPO) algorithm is used to design the reward function and optimize the dialogue generation strategy of the large model. The core loss function is... for: ; in The ratio of the probabilities of the new and old strategies is given. This indicates the dialogue generation strategy network parameters (or model weights) that need to be optimized during the RLHF fine-tuning phase of the large model. This is the estimated value of the dominance function; To truncate hyperparameters; Expressing expected value based on experience; This is a truncation function used to constrain the policy update magnitude within a certain range. Within the range; min() means taking the minimum value; The loss function pass The function constrains the magnitude of each policy update, aiming to maximize the expected policy advantage while maintaining the stability of the large model during multi-round dialogue generation training. To enable the system to accurately extract stage summaries and write them into a global shared memory pool, a supervised training dataset is constructed based on manually labeled "ideal memory summaries." The cross-entropy loss function guides the weight updates of the multilayer perceptron (MLP) after the feature extraction layer. for: ; in Preset baseline reference label; This represents the predicted probability distribution output by the Multilayer Perceptron (MLP); the subscript i indicates the index of the category; log() is the logarithmic function. By minimizing this loss function Large models can extract the core performance features of candidates from dialogues and write them into a structured global dynamic shared memory pool. To address the retrieval capabilities for dynamic follow-up questions and advanced discussions, a sparse-dense dual-path recall fusion mechanism is employed. This combines sparse retrieval based on the BM25 (Best Matching 25) algorithm with semantic retrieval based on high-dimensional dense vectors. The BM25 algorithm is a ranking function used in information retrieval to calculate the relevance between query terms and documents. The pre-trained parameters of the main model are kept frozen, and only the generation part of the query vector Query is fine-tuned to ensure that the generated retrieval query can both connect to the preceding memory context and accurately hit the professional content in the knowledge base. For a given query vector q and RAG knowledge base document d, the comprehensive retrieval score is calculated. Defined as: ; in The text literal matching score is based on Term Frequency-Inverse Document Frequency (TF-IDF) to ensure that technical terms are not missed due to semantic ambiguity. To characterize the high-dimensional dense vectors output by large models; Indicates cosine similarity; The fusion weights are dynamically adjustable.

[0034] Specifically, the high-dimensional dense vector It is a fixed-length numerical vector generated by extracting features from a text string using a pre-trained deep neural network model. This vector maps the semantic information of the text to a high-dimensional continuous space, making the cosine similarity of semantically similar texts in this space higher. This enables the accurate capture of synonyms and homonyms of technical terms, making up for the shortcomings of literal matching (TF-IDF) in semantic mining.

[0035] The aforementioned sparse-dense dual-path recall fusion mechanism takes into account the complementarity of lexical-level precise matching and semantic-level similarity retrieval, ensuring both accuracy and relevance of the follow-up content.

[0036] Based on the above method, embodiments of the present invention also provide a digital human adaptive interactive interview feedback system, such as... Figure 2 The diagram shown is a structural schematic of a digital human adaptive interactive interview feedback system provided in an embodiment of the present invention. The system includes: The pre-interview path planning module is used to collect candidate tags through multi-source interaction during the pre-interview and initial interaction stages, build an initial candidate profile, and initialize a global dynamic shared memory pool. The cross-stage multi-agent control module is used to call the multi-agent collaborative architecture to execute cross-stage interview interactions based on the initial profile. It maintains logical coherence through horizontal breadth and vertical depth mining, and dynamically summarizes and updates the global dynamic shared memory pool to ensure the context of different stages of the interview and prevent forgetting. The Real-Time Dynamic Follow-Up Questions and Advanced Discussion Module is used to trigger adaptive dynamic follow-up questions and open-ended advanced discussions during the in-depth interview interaction phase by combining retrieval-enhanced RAG technology with historical records in a globally dynamic shared memory pool. The structured assessment report generation module is used to generate a mandatory structured, high-dimensional assessment feedback report by integrating dynamic scores from each stage of the interview.

[0037] The specific implementation process of each module in the above system is described in the method implementation example.

[0038] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method.

[0039] It is worth noting that the contents not described in detail in the embodiments of the present invention belong to the prior art known to those skilled in the art.

[0040] In summary, the method described in this embodiment of the invention, by constructing a global dynamic shared memory pool and a multi-role intelligent agent collaborative framework, achieves cross-stage context state inheritance and logical consistency, significantly improving the scientific rigor, relevance, and practical guidance value of digital human interviews. In the field of artificial intelligence and virtual digital human interaction, it helps machines accurately detect the true cognitive boundaries of candidates, providing strong support for the application of AI interviews in high-level professional scenarios and the expansion of evaluation dimensions, and is of great significance to the development of AI interviews.

[0041] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.

Claims

1. A digital human adaptive interactive interview feedback method, characterized in that, The method includes: Step 1: In the pre-interview and initial interaction stage, collect candidate tags through multi-source interaction, build the initial profile of the candidate, and initialize the global dynamic shared memory pool. Specifically, the data structure and concurrent processing mechanism of the global dynamic shared memory pool are as follows: Firstly, for general multi-turn question-answering scenarios, a persistent memory unit based on a relational database is constructed. A connection to the relational database is established through a database connection component, and each turn of dialogue is stored as a structured persistent record. Each record includes a session identifier (session_id), a turn identifier (turn_id), a user identifier (user_id), a message role, and message content. The database layer automatically maintains the creation time (created_at) and update time (updated_at). Among them, the role is used to distinguish between user-side messages and agent-side messages; the turn_id is used to uniquely identify a single turn of dialogue; and the session_id is used to aggregate multiple turns of messages into the same session context. Then, when facing complex task scenarios such as mock interviews and paper discussions, a state snapshot memory unit based on JSON structure is constructed. Specifically, the resume is extracted using the optical character recognition (OCR) method, the result is stored in the ocr_result field, and the multi-agent dialogue state snapshot is stored in the dialogue_record field. The dialogue_record field uses a JSON array structure, and each element of the array represents a dialogue event or a state advancement event. The multi-agent dialogue state snapshot includes control fields: interview stage identifier (interviewer_state), background type (background), lab type (lab), major category (major), job category (job), paper discussion sub-state (paper_state), mentor identifier (teacher), research interest direction (interest), selected paper object (paper), candidate question pool (question_pool), current question index (current_question_index), whether follow-up questions have been triggered (followup_triggered), follow-up question index (followup_question_index), direction change count (direction_change_count), whether it is a follow-up question round (ex_query), and interview mode (interview_mode). Among them, the selected paper object is used to store paper metadata; the candidate question pool is used to store the candidate question pool automatically generated for paper discussion, including the question design intent. For the concurrent processing mechanism of the global dynamic shared memory pool, a process consistency strategy based on task state machine and a two-layer truncation mechanism combining the business layer and template layer are adopted for ultra-long context control, wherein: The process consistency strategy based on task state machines is as follows: 1) For general multi-turn question-and-answer links, the atomicity of single write operations in relational databases is used to ensure basic consistency. Single-turn messages are used as the basic write unit, and incremental persistence of dialogue is achieved through session identifiers and turn identifiers. When the same turn identifier is written repeatedly, atomic updates are completed by using database unique key constraints in conjunction with insert conflict update semantics. 2) For complex task chains, a synchronous strategy of "reading the latest state snapshot -> appending new dialogues or state nodes to memory -> writing the updated complete JSON snapshot back to the database" is adopted, and each task advancement is an update of a complete state snapshot; The two-layer truncation mechanism combining the business layer and the template layer is as follows: 1) Business layer length control: Set an upper limit threshold for the current user input length to prevent abnormally long inputs from directly entering the model inference chain at the source; 2) Template layer window truncation: During the message assembly stage, the system prompts, historical conversation sequences, and current user input are concatenated in sequence, and a strategy combining sliding window and left-side truncation is adopted; Step 2: Based on the initial profile, invoke the multi-agent collaborative architecture to execute cross-stage interview interactions. Maintain logical coherence through horizontal breadth and vertical depth mining, and dynamically summarize and update the global dynamic shared memory pool to ensure the context of different stages of the interview and prevent forgetting. Based on the initial profile, the resume interviewing agent will horizontally lay out various basic questions according to the candidate's tags and job requirements to explore the breadth of the candidate's knowledge; and the professional interviewing agent will vertically delve into specific answers to explore the depth of the candidate's knowledge. During the handover process, each agent extracts the personalized summary of the current stage, the candidate's deep logic, and cognitive blind spots, and writes them in a structured manner into a global dynamic shared memory pool. For the k-th stage, the agent extracts three core features based on the long text of the interaction at the current stage, including: deep logic. This refers to the thought process a candidate uses when analyzing complex problems and building solutions; and the knowledge points they have correctly grasped. This refers to the professional concepts and technical details that the candidate accurately states in their answers; cognitive blind spots. This refers to the candidate's weak areas of knowledge, such as answering incorrectly, deliberately avoiding questions, or expressing themselves vaguely. At the end of the k-th stage, the agent uses an attention mechanism to reduce the dimensionality of the three core features, generating a personalized summary vector for the k-th stage. And update it to the global dynamic shared memory pool in real time. This process is represented as: ; in, This indicates the state of the global dynamic shared memory pool at the end of the previous interaction phase; This represents the latest global dynamic shared memory pool state obtained after incorporating the core features of the current stage; Update() represents the data update operation; When the agent in the (k+1)th stage initiates a deep question, it first reads the candidate's cognitive blind spots from the global dynamic shared memory pool. and deep logic ; Step 3: In the in-depth interview interaction stage, combine the retrieval enhancement generation RAG technology with the historical records in the global dynamic shared memory pool to trigger adaptive dynamic follow-up questions and open-ended high-level discussions. The specific process for triggering open-ended high-level discussions is as follows: First, based on the actively selected tags constructed in the initial stage, precise matching is performed in the RAG knowledge base to extract cutting-edge academic papers. A masking mechanism is used to block direct conclusions, and only the research background, core academic questions and research motivations are extracted and retained as discussion materials. The prompts ask the large model to introduce the research background to the candidate and guide the candidate to explore solutions; The intelligent agent initiates questions using open-ended materials as a starting point, and, in conjunction with information from the global dynamic shared memory pool, guides candidates to conduct open-ended methodological discussions based on their research background, and quantitatively assesses the candidates' problem awareness and ability to transfer methods. Step 4: Generate a mandatory structured, high-dimensional evaluation feedback report by integrating the dynamic scores from each stage of the interview.

2. The digital human adaptive interactive interview feedback method according to claim 1, characterized in that, In step 1, during the pre-interview and initial interaction phase, the candidate interacts with the agent through a daily dialog box. The agent stores the conversation record in the auxiliary database and extracts core personality tags through the underlying algorithm and writes them into the global dynamic shared memory pool. At the same time, by combining the active tags selected by the candidate and the passive tags obtained from the resume analysis, an initial profile of the candidate is constructed. The sources of candidate label features include: Actively select tags The candidate's preferred job position, preferred mentor, and core skills preferences, as voluntarily filled out; Passive parsing tags The intelligent agent extracts structured entity information related to educational background, professional skills, and project experience from candidates' resumes. Specifically, it uses OCR and natural language processing technologies to convert the resume into plain text and utilizes a pre-trained named entity recognition model (NER) to extract entity information through sequence labeling. Pre-interactive tags The long text records of conversations between candidates and agents are stored in the underlying auxiliary log database. The large model is called to perform semantic and sentiment analysis on the long text records of conversations. The large model outputs corresponding implicit labels, including labels such as strong willingness to communicate, extroverted personality, and high stress resistance, which are written into the core main database as pre-interaction labels. Finally, the three types of labels are vectorized, concatenated, and aggregated. The specific process is as follows: First, the discrete label text is converted into a unified structured description text; then, the structured description text is input into a pre-trained text representation model for feature encoding, generating three sets of feature vectors; finally, the three sets of encoded feature vectors are concatenated and dimensionality-reduced according to their dimensions to form the candidate's global initial prior label set. , is represented as: ; Here, Concat() represents the vector concatenation operation, used to concatenate the actively selected label vectors. Passive parsing of label vectors and pre-interactive label vectors By performing head-to-tail concatenation along the feature dimension, a unified global initial prior label set is formed. ; The prior label set It is pre-written into the global dynamic shared memory pool and serves as a prerequisite constraint for the first stage of generating the underlying large model, dynamically generating exclusive initial entry points for different candidates.

3. The digital human adaptive interactive interview feedback method according to claim 1, characterized in that, In step 3, during the in-depth interview phase, when it is necessary to examine the candidate's true cognitive boundaries, the agent retrieves records of the candidate's incorrect answers from the global dynamic shared memory pool. This, combined with external professional domain retrieval enhancement generation (RAG) knowledge base, generates in-depth follow-up questions. Specifically, the RAG is a real-time reasoning engine and retrieval enhancement generation technology: Regarding the candidate's current answer The intelligent agent, based on the answer Two scores were obtained: Logical Integrity Score Content coverage score ; Among them, logical integrity score The acquisition method is as follows: The underlying reasoning engine performs dependency parsing and comparison between the answer text and the set standard problem-solving step tree. Through the pre-set thinking chain scoring Prompt of the large model, it performs normalized scoring from 0 to 1 on three dimensions: causal relationship, step coherence, and argument self-consistency. The weighted average of the three is taken to obtain the logical completeness score. The process is represented as: ; Where C, L, and S represent the causal relationship score, the step coherence score, and the argument consistency score after dependency parsing, respectively. These are preset weighting coefficients; Content coverage score The method of obtaining the information is as follows: The answer is extracted using Named Entity Recognition (NER) technology. The core entity word set in It also reads the set of knowledge points for the corresponding questions from the RAG knowledge base in the external professional field. The RAG is a real-time inference engine and retrieval enhancement generation technology; the content coverage score is obtained by calculating the semantic overlap and entity recall rate between the two. ,Right now: ; In obtaining the logical integrity score Content coverage score The answer is derived by combining the two indicators. Quality score , is represented as: ; in These are learnable or preset weighting coefficients; When quality score When the threshold is lower than a set threshold, the innovative probing engine will be triggered; wherein, the threshold is preset based on the distribution of historical interview samples; This innovative inquiry engine first retrieves cognitive blind spots recorded in the global dynamic shared memory pool. To address cognitive blind spots Input into the RAG knowledge base to retrieve relevant higher-order knowledge slices. ; Slice the higher-order knowledge retrieved Cognitive blind spots Current session context The inputs are combined and fed into a larger model as prompts, driving the model to generate a series of follow-up questions targeting the candidate's cognitive blind spots. The splicing formula is expressed as: ; in, The `--system-command-prompt` keyword is used to define the role attributes, tone, and follow-up questioning logic constraints of the large model at the underlying level, acting as a "senior technical interviewer." `Concat()` represents a vector concatenation operation.

4. The digital human adaptive interactive interview feedback method according to claim 1, characterized in that, In step 4, at each stage of the interview, each agent independently scores the candidate in the background based on the candidate's real-time responses, forming a stage-based score sequence. After the interview, the central control agent comprehensively calculates the scores at each stage, the variance of emotional expression fluctuations, and the ability to correct errors when questioned about them. It then outputs a structured diagnostic assessment report using a forced template, including the candidate's core strengths and weaknesses, key areas for improvement, and overall hiring preference. The error correction ability is calculated using a formula: ; in This represents a candidate's overall ability to correct errors when questioned about them, i.e., the dynamic error correction score; M represents the number of rounds of questioning. and , respectively, represent the logical accuracy scores before and after the error correction guidance; the subscript j indicates the j-th round of follow-up questioning.

5. The digital human adaptive interactive interview feedback method according to claim 1, characterized in that, The method further includes: for multi-agent cooperative scheduling and state machine transitions, a combination of prefix fine-tuning and reinforcement learning (RLHF) based on human feedback is used to achieve lightweight multi-role adaptation and stable optimization of generation strategies. Specifically: During the prefix fine-tuning phase, for different stages in the multi-agent collaborative scenario, an independent learnable role prefix parameter matrix is ​​set for each agent at each stage, where: The resume-based AI agent is responsible for verifying basic information and asking background questions based on the candidate's resume. The corresponding prefix parameter matrix is ​​as follows: The professional interview agent is responsible for assessing candidates' professional skills, project experience, and technical depth in a specific field. The corresponding prefix parameter matrix is ​​as follows: The higher-order academic agent is responsible for evaluating the candidate's research capabilities, academic vision, and ability to analyze complex theories; the corresponding prefix parameter matrix is ​​as follows. ; The corresponding prefix parameter matrix is ​​concatenated to the front of each multi-head attention layer in the Transformer model, and the corresponding attention calculation formula is updated as follows: ; in This represents the output of the self-attention mechanism after incorporating the role-specific prefix parameters of each agent. and is the prefix parameter matrix for independent training of the agent at each stage; Q is the query vector, K is the key vector, V is the value vector; d is the vector dimension; is the scaling factor to prevent gradient vanishing; softmax is the normalization exponential function; the superscript T indicates the transpose sign; During the RLHF fine-tuning phase, a proximal strategy is used to optimize the PPO algorithm's reward function, and the dialogue generation strategy of the large model is optimized, with the core loss function... for: ; in The ratio of the probabilities of the new and old strategies is given. This represents the dialogue generation strategy network parameters that need to be optimized during the RLHF fine-tuning phase of the large model. This is the estimated value of the dominance function; To truncate hyperparameters; Expressing expected value based on experience; This is a truncation function used to constrain the policy update magnitude within a certain range. Within the range; min() means taking the minimum value; A supervised training dataset was constructed based on manually annotated "ideal memory summaries." The weights of a multilayer perceptron (MLP) after the feature extraction layer were updated using a cross-entropy loss function. for: ; in Preset baseline reference label; This represents the predicted probability distribution output by the Multilayer Perceptron (MLP); the subscript i indicates the index of the category; log() is the logarithmic function. By minimizing this loss function Large models can extract the core performance features of candidates from dialogues and write them into a structured global dynamic shared memory pool. To enhance retrieval capabilities for dynamic follow-up questions and advanced discussions, a sparse-dense dual-path recall fusion mechanism is employed. This mechanism combines sparse retrieval based on the BM25 algorithm with semantic retrieval based on high-dimensional dense vectors. The BM25 algorithm is a ranking function used in information retrieval to calculate the relevance between query terms and documents. For a given query vector q and RAG knowledge base document d, the comprehensive retrieval score is used to determine the relevance. Defined as: ; in The text literal matching score is based on term frequency-inverse document frequency (TF-IDF) to ensure that technical terms are not missed due to semantic ambiguity. To characterize the high-dimensional dense vectors output by large models; Indicates cosine similarity; The fusion weights are dynamically adjustable.

6. A digital human adaptive interactive interview feedback system, characterized in that, The system includes: The pre-interview path planning module is used to collect candidate tags through multi-source interaction during the pre-interview and initial interaction stages, build an initial candidate profile, and initialize a global dynamic shared memory pool. Specifically, the data structure and concurrent processing mechanism of the global dynamic shared memory pool are as follows: Firstly, for general multi-turn question-answering scenarios, a persistent memory unit based on a relational database is constructed. A connection to the relational database is established through a database connection component, and each turn of dialogue is stored as a structured persistent record. Each record includes a session identifier (session_id), a turn identifier (turn_id), a user identifier (user_id), a message role, and message content. The database layer automatically maintains the creation time (created_at) and update time (updated_at). Among them, the role is used to distinguish between user-side messages and agent-side messages; the turn_id is used to uniquely identify a single turn of dialogue; and the session_id is used to aggregate multiple turns of messages into the same session context. Then, when facing complex task scenarios such as mock interviews and paper discussions, a state snapshot memory unit based on JSON structure is constructed. Specifically, the resume is extracted using the optical character recognition (OCR) method, the result is stored in the ocr_result field, and the multi-agent dialogue state snapshot is stored in the dialogue_record field. The dialogue_record field uses a JSON array structure, and each element of the array represents a dialogue event or a state advancement event. The multi-agent dialogue state snapshot includes control fields: interview stage identifier (interviewer_state), background type (background), lab type (lab), major category (major), job category (job), paper discussion sub-state (paper_state), mentor identifier (teacher), research interest direction (interest), selected paper object (paper), candidate question pool (question_pool), current question index (current_question_index), whether follow-up questions have been triggered (followup_triggered), follow-up question index (followup_question_index), direction change count (direction_change_count), whether it is a follow-up question round (ex_query), and interview mode (interview_mode). Among them, the selected paper object is used to store paper metadata; the candidate question pool is used to store the candidate question pool automatically generated for paper discussion, including the question design intent. For the concurrent processing mechanism of the global dynamic shared memory pool, a process consistency strategy based on task state machine and a two-layer truncation mechanism combining the business layer and template layer are adopted for ultra-long context control, wherein: The process consistency strategy based on task state machines is as follows: 1) For general multi-turn question-and-answer links, the atomicity of single write operations in relational databases is used to ensure basic consistency. Single-turn messages are used as the basic write unit, and incremental persistence of dialogue is achieved through session identifiers and turn identifiers. When the same turn identifier is written repeatedly, atomic updates are completed by using database unique key constraints in conjunction with insert conflict update semantics. 2) For complex task chains, a synchronous strategy of "reading the latest state snapshot -> appending new dialogues or state nodes to memory -> writing the updated complete JSON snapshot back to the database" is adopted, and each task advancement is an update of a complete state snapshot; The two-layer truncation mechanism combining the business layer and the template layer is as follows: 1) Business layer length control: Set an upper limit threshold for the current user input length to prevent abnormally long inputs from directly entering the model inference chain at the source; 2) Template layer window truncation: During the message assembly stage, the system prompts, historical conversation sequences, and current user input are concatenated in sequence, and a strategy combining sliding window and left-side truncation is adopted; The cross-stage multi-agent control module is used to call the multi-agent collaborative architecture to execute cross-stage interview interactions based on the initial profile. It maintains logical coherence through horizontal breadth and vertical depth mining, and dynamically summarizes and updates the global dynamic shared memory pool to ensure the context of different stages of the interview and prevent forgetting. Based on the initial profile, the resume interviewing agent will horizontally lay out various basic questions according to the candidate's tags and job requirements to explore the breadth of the candidate's knowledge; and the professional interviewing agent will vertically delve into specific answers to explore the depth of the candidate's knowledge. During the handover process, each agent extracts the personalized summary of the current stage, the candidate's deep logic, and cognitive blind spots, and writes them in a structured manner into a global dynamic shared memory pool. For the k-th stage, the agent extracts three core features based on the long text of the interaction at the current stage, including: deep logic. This refers to the thought process a candidate uses when analyzing complex problems and building solutions; and the knowledge points they have correctly grasped. This refers to the professional concepts and technical details that the candidate accurately states in their answers; cognitive blind spots. This refers to the candidate's weak areas of knowledge, such as answering incorrectly, deliberately avoiding questions, or expressing themselves vaguely. At the end of the k-th stage, the agent uses an attention mechanism to reduce the dimensionality of the three core features, generating a personalized summary vector for the k-th stage. And update it to the global dynamic shared memory pool in real time. This process is represented as: ; in, This indicates the state of the global dynamic shared memory pool at the end of the previous interaction phase; This represents the latest global dynamic shared memory pool state obtained after incorporating the core features of the current stage; Update() represents the data update operation; When the agent in the (k+1)th stage initiates a deep question, it first reads the candidate's cognitive blind spots from the global dynamic shared memory pool. and deep logic ; The Real-Time Dynamic Follow-Up Questions and Advanced Discussion Module is used to trigger adaptive dynamic follow-up questions and open-ended advanced discussions during the in-depth interview interaction phase by combining retrieval-enhanced RAG technology with historical records in a globally dynamic shared memory pool. The specific process for triggering open-ended high-level discussions is as follows: First, based on the actively selected tags constructed in the initial stage, precise matching is performed in the RAG knowledge base to extract cutting-edge academic papers. A masking mechanism is used to block direct conclusions, and only the research background, core academic questions and research motivations are extracted and retained as discussion materials. The prompts ask the large model to introduce the research background to the candidate and guide the candidate to explore solutions; The intelligent agent initiates questions using open-ended materials as a starting point, and, in conjunction with information from the global dynamic shared memory pool, guides candidates to conduct open-ended methodological discussions based on their research background, and quantitatively assesses the candidates' problem awareness and ability to transfer methods. The structured assessment report generation module is used to generate a mandatory structured, high-dimensional assessment feedback report by integrating dynamic scores from each stage of the interview.

7. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • LLM-based adaptive interview simulation system and device

    CN120744047A

  • Multi-stage LLM with unlimited context

    US12387050B1