Interview assessment methods, devices, computer equipment, storage media, and program products based on large language models.

By employing a multi-dimensional evaluation process based on a large language model, the problems of subjectivity and inefficiency in traditional interview-based assessments are solved, achieving intelligent, standardized, and highly efficient interview assessments, applicable to fields such as talent selection, psychological evaluation, and educational assessment.

CN122367290APending Publication Date: 2026-07-10XUANXING INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Traditional human interview-based assessment methods are characterized by strong subjectivity, low efficiency, low standardization, high cost, difficulty in achieving high efficiency and consistency, and lack of in-depth questioning capabilities, making it difficult to meet the needs of large-scale assessments.

Method used

We employ an interview assessment method based on a large language model. Through a multi-dimensional and multi-level assessment process, including assessments of completeness, credibility, and internalization of abilities, we combine a pre-set question bank with intelligent follow-up questions to achieve standardized and adaptive assessment.

Benefits of technology

It improves the accuracy and objectivity of interview assessments, reduces labor costs, and achieves intelligent and efficient assessments, making it suitable for fields such as talent selection, psychological assessment, and educational evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122367290A_ABST
    Figure CN122367290A_ABST
Patent Text Reader

Abstract

This application relates to an interview assessment method, apparatus, computer device, storage medium, and program product based on a large language model. It obtains initial interview content for a target subject, then uses a large language model to perform completeness assessment, credibility assessment, ability internalization assessment, and preliminary ability assessment. Based on the preliminary ability assessment results, it conducts personalized follow-up questions to obtain the target interview content. Finally, it uses the large language model to perform a comprehensive ability assessment to determine the overall ability assessment result for the target subject. Through multi-dimensional and multi-level analysis, it improves the accuracy and objectivity of interview assessment; by utilizing the semantic understanding capabilities of the large language model, it achieves intelligent follow-up questioning and adaptive assessment, improving the intelligence level of interview assessment. This not only significantly reduces manual costs but also improves the efficiency of interview assessment. It achieves intelligent, standardized, and highly efficient interview assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an interview assessment method, apparatus, computer device, computer-readable storage medium, and computer program product based on a large language model. Background Technology

[0002] In various fields such as talent selection, psychological assessment, educational evaluation, and clinical diagnosis, interview-based assessment has long been an important evaluation method, relying primarily on human interviewers to ask questions, observe, judge, and comprehensively score. However, traditional human interview-based assessment methods have many technical shortcomings, severely restricting their promotion and application in large-scale, high-efficiency, and highly consistent application scenarios.

[0003] In recent years, some automated assessment methods based on information technology have emerged, such as using a pre-set question bank to ask questions in a fixed order and combining keyword matching or simple rule engines to conduct preliminary analysis of the answers. However, these methods generally lack semantic understanding, contextual reasoning, and dynamic interaction capabilities. They cannot adjust questioning strategies based on the respondent's real-time responses and are essentially still at the "electronic questionnaire" stage, making it difficult to meet the needs of deep cognition and adaptive interview assessment. Summary of the Invention

[0004] Therefore, it is necessary to provide an interview assessment method, apparatus, computer equipment, computer-readable storage medium, and computer program product based on a large language model to address the aforementioned technical problems.

[0005] Firstly, this application provides an interview evaluation method based on a large language model, the method comprising: Obtain the initial interview content for the target audience, and use a large language model to evaluate the completeness of the initial interview content based on a pre-configured completeness evaluation standard to determine the completeness evaluation result. Based on the completeness assessment results, complete interview content that meets the completeness target is determined, and the credibility assessment of the complete interview content is performed by calling the large language model to determine the credibility assessment results. Based on the credibility assessment results, credible interview content that meets the credibility objectives is determined. The large language model is then invoked to perform a capability internalization assessment on the credible interview content, and the capability internalization assessment results are determined. Based on the internalization assessment results, determine the interview content to be assessed that meets the internalization goals, call the large language model to conduct a preliminary ability assessment on the interview content to be assessed, and determine the preliminary ability assessment results. A question template matching the preliminary ability assessment results is determined from a preset question bank, and target interview content is determined based on the interview content to be assessed and the question template. Based on the initial interview content, the improved interview content, the credible interview content, the interview content to be evaluated, and the target interview content, the large language model is invoked to perform a comprehensive ability assessment, and the comprehensive ability assessment result for the target object is determined.

[0006] In one embodiment, the completeness assessment criteria include a multi-dimensional assessment criteria constructed based on the STAR principle. Each dimension includes multiple completeness features, and each completeness feature has a corresponding first weight. The step of calling a large language model to assess the completeness of the initial interview content based on the pre-configured completeness assessment criteria and determining the completeness assessment result includes: for each dimension's multiple completeness features, calling the large language model to obtain the first score corresponding to each completeness feature of the initial interview content; and weighting the first score and first weight corresponding to each of the multiple completeness features in each dimension to obtain the completeness assessment result of the corresponding dimension.

[0007] In one embodiment, determining the completeness interview content that meets the completeness target based on the completeness assessment results includes: determining the initial interview content as completeness interview content that meets the completeness target when the completeness assessment results of each dimension of the initial interview content all reach the completeness threshold of the corresponding dimension; determining that there are completeness assessment results in the initial interview content that do not reach the completeness threshold of the corresponding dimension, using a large language model based on multiple completeness features of the dimension to obtain a first supplementary interview content for the target object; determining a first merged content of the initial interview content and the first supplementary interview content, using a large language model based on a pre-configured completeness assessment standard to perform a completeness assessment on the first merged content, and determining the completeness assessment results of each dimension; determining that the first merged content is completeness interview content that meets the completeness target when the completeness assessment results of each dimension of the first merged content all reach the completeness threshold of the corresponding dimension.

[0008] In one embodiment, the step of calling the large language model to evaluate the credibility of the improved interview content and determining the credibility evaluation result includes: based on multiple preset credibility features, calling the large language model to obtain the second score corresponding to each credibility feature of the improved interview content, and each credibility feature also has a corresponding second weight; and performing weighted processing on the second scores and second weights corresponding to the multiple credibility features to obtain the credibility evaluation result of the improved interview content.

[0009] In one embodiment, determining credible interview content that meets the credibility target based on the credibility assessment result includes: if the credibility assessment result of the improved interview content reaches a credibility threshold, determining the improved interview content as credible interview content that meets the credibility target; if the credibility assessment result of the improved interview content does not reach the credibility threshold, invoking a large language model based on the multiple credibility features to obtain second supplementary interview content for the target object; determining a second merged content of the improved interview content and the second supplementary interview content, invoking the large language model to perform a credibility assessment on the second merged content, and determining the corresponding credibility assessment result; if the credibility assessment result of the second merged content reaches a credibility threshold, determining the second merged content as credible interview content that meets the credibility target.

[0010] In one embodiment, the step of calling the large language model to perform a capability internalization assessment on the credible interview content and determining the capability internalization assessment result includes: based on multiple preset capability internalization features, calling the large language model to obtain the third score corresponding to each capability internalization feature of the credible interview content, and each capability internalization feature also has a corresponding third weight; and performing weighted processing on the third scores and third weights corresponding to the multiple capability internalization features to obtain the capability internalization assessment result of the credible interview content.

[0011] In one embodiment, determining the interview content to be evaluated that meets the capability internalization target based on the capability internalization assessment result includes: if the capability internalization assessment result of the credible interview content reaches the capability internalization threshold, determining the credible interview content as the interview content to be evaluated that meets the capability internalization target; if the capability internalization assessment result of the credible interview content does not reach the capability internalization threshold, invoking a large language model based on the multiple capability internalization features to obtain a third additional interview content for the target object; determining a third combined content of the credible interview content and the third additional interview content, invoking the large language model to perform capability internalization assessment on the third combined content, and determining the corresponding capability internalization assessment result; if the capability internalization assessment result of the third combined content reaches the capability internalization threshold, determining the third combined content as the interview content to be evaluated that meets the capability internalization target.

[0012] In one embodiment, determining a follow-up question template matching the preliminary ability assessment result from a preset question bank, and determining target interview content based on the interview content to be assessed and the follow-up question template, includes: determining the information content difference between the preliminary ability assessment result and the adjacent higher level; determining a matching follow-up question template from the preset question bank based on the information content difference and the preliminary ability assessment result; rewriting the follow-up question template by calling the large language model according to the interview content to be assessed to generate target follow-up questions; and obtaining the target interview content for the target object regarding the target follow-up questions.

[0013] In one embodiment, the step of invoking the large language model to perform a comprehensive ability assessment based on the initial interview content, the improved interview content, the credible interview content, the interview content to be evaluated, and the target interview content, and determining the comprehensive ability assessment result for the target object, includes: integrating the initial interview content, the improved interview content, the credible interview content, the interview content to be evaluated, and the target interview content to construct an interview record dataset; invoking the large language model to perform behavior recognition on the interview record dataset to obtain multiple identified behavioral events; and determining the comprehensive ability assessment result for the target object based on the matching degree between each behavioral event and a preset ability standard.

[0014] Secondly, this application also provides an interview assessment device based on a large language model, the device comprising: The completeness assessment module is used to obtain the initial interview content for the target object, call the large language model to assess the completeness of the initial interview content based on the pre-configured completeness assessment criteria, and determine the completeness assessment result. The credibility assessment module is used to determine the completeness interview content that meets the completeness target based on the completeness assessment result, call the large language model to perform credibility assessment on the completeness interview content, and determine the credibility assessment result. The capability internalization assessment module is used to determine credible interview content that meets the credibility target based on the credibility assessment results, call the large language model to perform capability internalization assessment on the credible interview content, and determine the capability internalization assessment results. The preliminary competency assessment module is used to determine the interview content to be assessed that meets the competency internalization goal based on the competency internalization assessment results, call the large language model to perform a preliminary competency assessment on the interview content to be assessed, and determine the preliminary competency assessment results. The target content determination module is used to determine the question template that matches the preliminary ability assessment results from a preset question bank, and to determine the target interview content based on the interview content to be assessed and the question template. The comprehensive ability assessment module is used to perform a comprehensive ability assessment by calling the large language model based on the initial interview content, the improved interview content, the credible interview content, the interview content to be assessed, and the target interview content, and to determine the comprehensive ability assessment result for the target object.

[0015] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in the first aspect.

[0016] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0017] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect above.

[0018] The aforementioned large language model-based interview assessment method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire initial interview content for the target subject; evaluate the completeness of the initial interview content based on pre-configured completeness assessment criteria using a large language model; determine the completeness assessment result; determine complete interview content that meets the completeness objective based on the completeness assessment result; evaluate the credibility of the complete interview content using a large language model; determine the credibility assessment result; determine credible interview content that meets the credibility objective based on the credibility assessment result; evaluate the internalization of the credible interview content using a large language model; determine the internalization of ... Through multi-dimensional and multi-level analysis, it improves the accuracy and objectivity of interview assessments; by utilizing the semantic understanding capabilities of large language models, it achieves intelligent follow-up questioning and adaptive assessment, enhancing the intelligence level of interview assessments. This not only significantly reduces manual costs but also improves the efficiency of interview assessments. It achieves intelligent, standardized, and highly efficient interview assessments. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a diagram illustrating the application environment of an interview assessment method based on a large language model in one embodiment. Figure 2 This is a flowchart illustrating an interview assessment method based on a large language model in one embodiment; Figure 3 This is a flowchart illustrating the steps for determining and refining interview content in one embodiment. Figure 4 This is a flowchart illustrating the steps for determining credible interview content in one embodiment; Figure 5 This is a flowchart illustrating the steps for determining the interview content to be evaluated in one embodiment; Figure 6 This is a structural block diagram of an interview assessment device based on a large language model in one embodiment; Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0022] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0023] One of the core problems with traditional interview-based assessments is their strong subjectivity. Significant differences in professional background, experience level, cognitive preferences, and even emotional state among interviewers lead to varying interpretations and evaluation standards for the same respondent's answers, resulting in a lack of objectivity and comparability in the assessment results, making it difficult to ensure consistency across interviewers and time periods. Secondly, inefficiency limits the large-scale application of traditional methods. A complete human interview typically takes tens of minutes or even longer and requires professionally qualified personnel, resulting in high labor costs and failing to meet the high-throughput processing needs of scenarios such as corporate recruitment, large-scale psychological screening, or educational assessment. Furthermore, low standardization further weakens the scientific rigor and fairness of the assessment. The lack of a unified questioning logic, scoring dimensions, and feedback mechanism makes the interview process susceptible to temporary interference, leading to arbitrary procedures and making it difficult to ensure that all respondents are assessed under the same conditions, thus affecting the credibility and validity of the assessment results. In addition, insufficient depth of follow-up questioning is another bottleneck that traditional interviews struggle to overcome. Even experienced interviewers may find it difficult to fully analyze the completeness, logic, and authenticity of a candidate's answers in a short period of time, which may lead to insufficient information gathering and thus affect the accuracy and depth of the assessment.

[0024] Based on this, this application provides an interview evaluation method based on a large language model, which achieves intelligent, standardized and efficient interview evaluation by combining a large language model with the interview evaluation.

[0025] The interview assessment method based on a large language model provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Specifically, terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0026] In one exemplary embodiment, such as Figure 2 As shown, an interview assessment method based on a large language model is presented, which is then applied to... Figure 1 Taking the terminal in the example, the specific steps may include: Step 202: Obtain the initial interview content for the target audience, and use the large language model to evaluate the completeness of the initial interview content based on the pre-configured completeness evaluation criteria to determine the completeness evaluation result.

[0027] The target audience can be the users currently undergoing the interview and evaluation. For example, the target audience could be candidates for talent selection, individuals to be evaluated in psychological assessments, or subjects being evaluated in educational assessments.

[0028] The initial interview content can be the target subject's initial response text obtained based on standard interview questions corresponding to the interview task. Interview tasks include, but are not limited to, talent selection, psychological assessment, educational evaluation, and clinical diagnosis. Understandably, different standard interview questions are pre-configured for different interview tasks.

[0029] The perfection assessment criteria are pre-configured standards used for perfection assessment. In this embodiment, the perfection assessment criteria may include multi-dimensional assessment criteria built based on the STAR principle. For example, the STAR principle may include assessment criteria in the dimensions of S (Situation), T (Task), A (Action), and R (Result). S, based on situational cognition theory, assesses the impact mechanism of the situation on behavior from aspects such as background complexity, environmental constraints, and time pressure. T, based on goal-setting theory, assesses the impact of task characteristics on performance from aspects such as goal clarity, task difficulty, and scope of responsibility. A, based on behaviorist theory and cognitive-behavioral theory, assesses the correspondence between behavior and ability from aspects such as strategy selection, execution process, and degree of innovation. R, based on results-oriented theory, assesses the value of result feedback for ability assessment from aspects such as result quantification, scope of impact, and sustainability.

[0030] Each dimension includes multiple completeness features, each with a corresponding primary weight. For example, the S dimension may include completeness features such as time characteristics, location characteristics, background complexity, environmental constraints, and the number of personnel involved. The T dimension may include completeness features such as goal clarity, task complexity, responsibility boundaries, success criteria, and time requirements. The A dimension may include completeness features such as the completeness of the action sequence, methodological innovation, resource utilization, collaboration level, and decision quality. The R dimension may include completeness features such as quantifiable results, scope of impact, achievement rate, side effect control, and verifiability.

[0031] In this embodiment, after obtaining the initial interview content for the target object, a large language model can be invoked based on pre-configured completeness assessment criteria to evaluate the completeness of the initial interview content and determine the completeness assessment result. Specifically, for multiple completeness features of each dimension, the large language model can be invoked to obtain the first score corresponding to each completeness feature of the initial interview content, and the first scores and first weights corresponding to multiple completeness features in each dimension are weighted to obtain the completeness assessment result for the corresponding dimension. The large language model can be implemented using Deepseek, Qwen, etc., and this embodiment does not limit it to this specific implementation.

[0032] For example, consider an assessment scenario of "evaluating a project manager's project management skills." If the standard interview question is "Please describe an experience where you successfully managed a complex project," and the initial interview content of the target audience is, "I previously led a large project with a team of over 20 people. The project was complex, involving collaboration across multiple departments. I organized many meetings, coordinated resources from all parties, and ultimately the project was completed on time, with the client very satisfied," then a large language model can be used to perform semantic analysis on the initial interview content and calculate the completeness assessment results for each dimension. For instance, using prompts, the large language model can be asked to analyze the completeness of features such as time characteristics, location characteristics, background complexity, environmental constraints, and the number of people involved along the S dimension, and return the completeness score (i.e., the first score) for each feature along with the reasoning. For example, the response text might mention the number of people involved, briefly mention time and background complexity, but completely omit location or environmental characteristics. If the scores for each feature in dimension S returned after analysis by the large language model are (60, 0, 60, 80), and the first weights corresponding to each feature are (0.25, 0.25, 0.25, 0.25), then the completeness assessment result for dimension S can be obtained as 0.25*60 + 0.25*0 + 0.25*60 + 0.25*80 = 50 points. Based on this, the completeness assessment results of the initial interview content in each dimension can be obtained. It is understood that the first weights of each feature can be flexibly configured based on the specific assessment scenario, and this embodiment does not limit this.

[0033] This embodiment assesses completeness based on the STAR principle and preset standards, thereby achieving standardization and normalization of interview evaluation.

[0034] Step 204: Based on the perfection assessment results, determine the perfect interview content that meets the perfection goals, call the large language model to conduct a credibility assessment on the perfect interview content, and determine the credibility assessment results.

[0035] The completeness target can be a pre-configured completeness condition, such as a completeness threshold corresponding to each dimension. Completed interview content refers to interview content that meets the completeness target. Credibility assessment can be the process of evaluating the credibility of the completed interview content based on multiple preset credibility features.

[0036] In this embodiment, the perfected interview content that meets the perfectedness target is determined based on the perfectedness assessment results. Then, the large language model is called to conduct a credibility assessment on the perfected interview content to obtain the credibility assessment results of the perfected interview content.

[0037] Specifically, based on multiple preset credibility features, a large language model can be invoked to obtain the second score corresponding to each credibility feature of the complete interview content. Each credibility feature also has a corresponding second weight. Then, the second scores and second weights corresponding to multiple credibility features are weighted to obtain the credibility assessment result of the complete interview content.

[0038] For example, consider an assessment scenario: "Evaluating a project manager's data analysis capabilities." If the standard interview question is, "Please describe an experience where you made a significant product decision through data analysis," then the refined interview content, such as, "Last year, the product I was responsible for encountered user churn. We immediately organized the team for in-depth data analysis and discovered that user retention had dropped significantly. We analyzed various data metrics and ultimately found that the new version's features were too complex. Therefore, we decisively decided to simplify the product functionality and redesign the user interface. This decision was extremely correct; after the product went live, user retention increased dramatically, growing by more than 50%. Company leaders were very satisfied with this decision and specifically praised our team. This experience made me deeply aware of the importance of data analysis," can then utilize a large language model to assess the credibility of the refined interview content and determine the credibility assessment result.

[0039] For example, if the preset credibility features include language consistency, detail density, and cognitive load, then prompts can be used to ask the large language model to return a credibility score (i.e., a second score) for each feature in terms of language consistency, detail density, and cognitive load, along with the reasoning behind the score. For instance, in analyzing the language consistency of a refined interview, the prompts would specify the need for analysis from the perspectives of time, person, and logic. Since the refined interview content contains instances of the interchangeable use of "I" and "we," if the large language model judges based on the reasoning "the boundary between individual and team contributions is unclear, potentially exaggerating individual roles," the final score for language consistency would be 50. For credibility analysis of the detail density dimension, the large language model is required to score based on the number of data indicators, the clarity of the analysis process, and the clarity of the decision-making process, and provide the reasoning behind the score. For credibility analysis of the cognitive load dimension, the large language model is required to judge the credibility of the response based on sentence length, pause words, and the number of corrections, identifying whether the narrative is simplified or overly fluent, and providing a score and the reasoning behind the score. This yields the second score corresponding to each credibility feature. Then, by weighting the second score using the second weight corresponding to each credibility feature, a more comprehensive credibility assessment result can be obtained. For example, if the weights for language consistency, detail density, and cognitive load are (0.4, 0.4, 0.2) and the second scores are (50, 60, 40), the weighted credibility assessment result is 0.4*50 + 0.4*60 + 0.2*40 = 52 points. By introducing credibility assessment, it is possible to effectively identify and process inaccurate information, thereby improving the credibility of the interview assessment.

[0040] Step 206: Based on the credibility assessment results, determine the credible interview content that meets the credibility objectives, call the large language model to conduct a capability internalization assessment on the credible interview content, and determine the capability internalization assessment results.

[0041] The credibility target can be a pre-configured credibility condition, such as a corresponding credibility threshold. Credible interview content refers to interview content that meets the credibility target. Capability internalization assessment can be a process of evaluating the degree of internalization of credible interview content based on multiple pre-defined capability internalization features.

[0042] In this embodiment, after determining the credible interview content that meets the credibility target based on the credibility assessment results, the large language model can then be invoked to perform a capability internalization assessment on the credible interview content in order to determine the capability internalization assessment results.

[0043] Specifically, based on multiple preset capability internalization features, a large language model is invoked to obtain the third score corresponding to each capability internalization feature of the credible interview content. Each capability internalization feature also has a corresponding third weight. Then, the third scores and third weights corresponding to the multiple capability internalization features are weighted to obtain the capability internalization evaluation result of the credible interview content.

[0044] For example, if the preset internalization features of multiple capabilities include behavioral universality, capability transferability, depth of reflection, and continuous improvement, analytical rules for behavioral universality analysis, capability transferability analysis, depth of reflection analysis, and evidence analysis for continuous improvement can be provided through prompts. For instance, the rules for depth of reflection analysis include the degree of description of the content of summarization and reflection. Then, the large language model can be asked to provide scores (i.e., third scores) and reasons for these scores in these four aspects. For example, if the credible interview content only mentions "a deep understanding of the importance of data analysis" but does not mention specific understanding or reflection content, the large model judges the depth of reflection to score 30 points (i.e., the third score). Based on this, the third score corresponding to each capability internalization feature can be obtained. Then, by weighting the third weight and third score corresponding to each capability internalization feature, the capability internalization assessment result of the credible interview content can be obtained. For example, if the weights for behavior universality, ability transferability, reflection depth, and continuous improvement are (0.25, 0.25, 0.25, 0.25) and the third score is (30, 50, 60, 40) respectively, then the weighted internalization assessment result is 45 points.

[0045] Step 208: Based on the results of the competency internalization assessment, determine the interview content to be assessed that meets the competency internalization objectives, call the large language model to conduct a preliminary competency assessment on the interview content to be assessed, and determine the preliminary competency assessment results.

[0046] The competency internalization goal can be a pre-configured competency internalization condition, such as a corresponding competency internalization threshold. The interview content to be evaluated is the interview content that meets the competency internalization goal. The preliminary competency assessment can be a process of preliminary competency grading based on preset competency performance standards.

[0047] In this embodiment, after determining the interview content to be evaluated that meets the ability internalization goals based on the ability internalization assessment results, the large language model can be invoked to perform a preliminary ability assessment on the interview content to be evaluated, so as to determine the preliminary ability assessment results. For example, based on the interview content to be evaluated that meets the ability internalization goals, the ability assessment engine of the large language model can be invoked, and a preliminary ability level can be determined according to the preset ability performance standards to obtain the preliminary ability level, i.e., the preliminary ability assessment result.

[0048] Step 210: Determine the question templates that match the preliminary ability assessment results from the preset question bank, and determine the target interview content based on the interview content to be assessed and the question templates.

[0049] The preset question bank can be a pre-configured question bank used for interview evaluation, i.e., a question bank used to ask questions during the interview evaluation process. The follow-up question template refers to a question template generated based on the preset question bank for asking additional questions. The target interview content is the final interview content determined based on the interview content to be evaluated and the follow-up question template.

[0050] In this embodiment, by determining a question template that matches the preliminary ability assessment results from a preset question bank, the target interview content can be determined based on the interview content to be assessed and the question template.

[0051] Step 212: Based on the interview content, use a large language model to conduct a comprehensive ability assessment and determine the comprehensive ability assessment results for the target subject.

[0052] The interview content can include initial interview content, improved interview content, credible interview content, interview content to be evaluated, and target interview content, that is, all interview content involved in the above-mentioned interview process.

[0053] Specifically, based on the initial interview content, the improved interview content, the credible interview content, the interview content to be evaluated, and the target interview content, a large language model can be invoked to conduct a comprehensive ability assessment, thereby determining the comprehensive ability assessment results for the target object.

[0054] The aforementioned interview assessment method based on a large language model involves several steps. First, initial interview content is obtained for the target audience. Then, based on pre-configured completeness assessment criteria, a large language model is invoked to assess the completeness of the initial interview content, determining the completeness assessment result. Based on the completeness assessment result, complete interview content that meets the completeness objective is determined. Next, the large language model is invoked to assess the credibility of the complete interview content, determining the credibility assessment result. Based on the credibility assessment result, credible interview content that meets the credibility objective is determined. Finally, the large language model is invoked to assess the internalization of the credible interview content, determining the internalization assessment result. Based on the internalization assessment result, interview content that meets the internalization objective is determined. Then, the large language model is invoked to conduct a preliminary ability assessment of the interview content, determining the preliminary ability assessment result. A follow-up question template matching the preliminary ability assessment result is then determined from a pre-set question bank. Based on the interview content to be assessed and the follow-up question template, target interview content is determined. Finally, based on the interview content, the large language model is invoked to conduct a comprehensive ability assessment to determine the comprehensive ability assessment result for the target audience. Through multi-dimensional and multi-level analysis, it improves the accuracy and objectivity of interview assessments; by utilizing the semantic understanding capabilities of large language models, it achieves intelligent follow-up questioning and adaptive assessment, enhancing the intelligence level of interview assessments. This not only significantly reduces manual costs but also improves the efficiency of interview assessments. It achieves intelligent, standardized, and highly efficient interview assessments.

[0055] In one exemplary embodiment, such as Figure 3 As shown, in step 204, the completeness interview content that meets the completeness target is determined based on the completeness assessment results, which may specifically include: Step 302: If the completeness assessment results of each dimension of the initial interview content have reached the completeness threshold of the corresponding dimension, the initial interview content is determined to be complete interview content that meets the completeness target.

[0056] Since perfecting interview content means ensuring that the completeness assessment results meet the completeness objectives, and the completeness assessment results include completeness scores across multiple dimensions, when the completeness objective is a completeness threshold, it includes the completeness thresholds corresponding to each dimension.

[0057] In this embodiment, the completeness assessment results of each dimension of the initial interview content can be compared with the completeness threshold of the corresponding dimension. If the completeness assessment results of each dimension reach the completeness threshold of the corresponding dimension, the initial interview content can be determined to be complete interview content that meets the completeness target.

[0058] Step 304: If it is determined that the completeness assessment results in the initial interview content do not reach the completeness threshold of the corresponding dimension, determine the first supplementary interview content.

[0059] Specifically, if it is determined that the completeness assessment result in the initial interview content does not reach the completeness threshold of the corresponding dimension, then the large language model is invoked based on multiple completeness features of that dimension to obtain the first supplementary interview content for the target object.

[0060] For example, after determining the completeness assessment results of the four dimensions of the STAR principle based on the initial interview content, a judgment can be made based on the completeness threshold of each dimension. If the completeness score of one or more dimensions is lower than the corresponding completeness threshold, a dimension is randomly selected, and the large language model is invoked to guide the follow-up questioning module for further questioning. For example, if the completeness threshold for dimension S is 65 points, and the current completeness score for that dimension is 50 points, then the score is lower than the threshold, and therefore follow-up questioning is required. The specific dimension that is not complete, the completeness score result, the specific reasons for the completeness judgment, and the follow-up questioning rules are input into the large language model through prompt words, and the large language model is asked to generate corresponding follow-up questions. Based on the follow-up questions, the supplementary response content from the target subject, i.e., the first supplementary interview content, is obtained.

[0061] Step 306: Determine the first combined content of the initial interview content and the first supplementary interview content, and determine the completeness assessment results of each dimension of the first combined content.

[0062] The first merged content is the interview content obtained by merging the initial interview content and the first supplementary interview content. Specifically, by determining the first merged content of the initial interview content and the first supplementary interview content, and using a large language model based on pre-configured completeness assessment criteria, the completeness assessment results of each dimension are determined.

[0063] Step 308: If the completeness assessment results of each dimension of the first merged content have reached the completeness threshold of the corresponding dimension, the first merged content is determined to be complete interview content that meets the completeness target.

[0064] Specifically, if the completeness assessment results of each dimension of the first merged content reach the corresponding completeness threshold, then the first merged content can be determined as complete interview content that meets the completeness target. If there are cases where the completeness assessment results do not reach the corresponding completeness threshold, then follow-up questions are added based on step 304 above, and the completeness of the four dimensions of the STAR principle is recalculated. Then, it is re-evaluated whether the completeness target is met, until the completeness target is met, and the corresponding complete interview content is obtained.

[0065] In this embodiment, the completeness of the four dimensions of the STAR principle can be recalculated for each round of responses, and then a reassessment can be made to determine whether the completeness target is met. Since the person being evaluated (the target audience) may supplement their responses with information not only for the follow-up question dimensions but also for other dimensions, recalculation and assessment can reduce the number of follow-up questioning rounds, thus avoiding redundant questioning and improving evaluation efficiency.

[0066] In one exemplary embodiment, such as Figure 4 As shown, in step 206, credible interview content that meets the credibility objective is determined based on the credibility assessment results, which may specifically include: Step 402: If the credibility assessment result of the improved interview content reaches the credibility threshold, the improved interview content is determined to be credible interview content that meets the credibility target.

[0067] Since credible interview content is interview content whose credibility assessment results meet the credibility objective, and the credibility objective can be a credibility threshold, after determining the credibility assessment results of the improved interview content, it can be compared with the credibility threshold. If the credibility assessment results of the improved interview content reach the credibility threshold, then the improved interview content can be determined as credible interview content that meets the credibility objective.

[0068] Step 404: If the credibility assessment result of the improved interview content does not reach the credibility threshold, determine the second supplementary interview content.

[0069] Specifically, if the credibility assessment result of the improved interview content does not reach the credibility threshold, a second supplementary interview content for the target object is obtained by calling a large language model based on multiple credibility features.

[0070] For example, after determining the corresponding credibility assessment result based on the improved interview content, if the corresponding credibility score is lower than the credibility threshold, the large language model is invoked based on multiple preset credibility features to guide the follow-up questioning module. For instance, if the credibility threshold is 60 points, and the current improved interview content has a credibility score of 50 points, then the score is lower than the threshold, and follow-up questioning is deemed necessary. The specific credibility score, the specific reasons for the credibility judgment, and the follow-up questioning rules are then input into the large language model via prompt words, requesting the large language model to generate corresponding follow-up questions. Based on the follow-up questions, the supplementary response content from the target audience, i.e., the second supplementary interview content, is obtained.

[0071] Step 406: Determine the second combined content of the improved interview content and the second supplementary interview content, and determine the credibility assessment result corresponding to the second combined content.

[0072] The second merged content is the interview content obtained by merging the improved interview content and the second supplementary interview content. Specifically, by determining the second merged content of the improved interview content and the second supplementary interview content, and by calling a large language model to evaluate the credibility of the second merged content, the corresponding credibility evaluation result is determined.

[0073] Step 408: If the credibility assessment result of the second merged content reaches the credibility threshold, the second merged content is determined to be credible interview content that meets the credibility target.

[0074] Specifically, if the credibility assessment result of the second merged content reaches the credibility threshold, then the second merged content can be determined as credible interview content that meets the credibility target. If its credibility assessment result does not reach the credibility threshold, follow-up questions are added based on step 404 above, and the corresponding credibility is recalculated. Then, the credibility target is reassessed until it is met, at which point the corresponding credible interview content is obtained. By introducing credibility analysis, it can effectively identify and process false information, thereby improving the credibility of the interview content.

[0075] In one exemplary embodiment, such as Figure 5 As shown, in step 208, the interview content to be assessed that meets the competency internalization objectives is determined based on the competency internalization assessment results. Specifically, this may include: Step 502: If the assessment result of the ability internalization of the credible interview content reaches the ability internalization threshold, the credible interview content is determined as the interview content to be assessed that meets the ability internalization goal.

[0076] Since the interview content to be evaluated is the interview content whose competency internalization assessment results meet the competency internalization goals, and the competency internalization goals can be competency internalization thresholds, after determining the competency internalization assessment results of credible interview content, it can be compared with the competency internalization threshold. If the competency internalization assessment results of credible interview content reach the competency internalization threshold, then the credible interview content can be determined as the interview content to be evaluated that meets the competency internalization goals.

[0077] Step 504: If the assessment results of the ability internalization of the credible interview content do not reach the ability internalization threshold, determine the third supplementary interview content.

[0078] Specifically, if the assessment result of the ability internalization of the credible interview content does not reach the ability internalization threshold, a third supplementary interview content for the target object is obtained by calling a large language model based on multiple ability internalization features.

[0079] For example, after determining the corresponding competency internalization assessment result based on credible interview content, if the corresponding competency internalization score is lower than the competency internalization threshold, the large language model is invoked based on multiple preset competency internalization features to guide the follow-up questioning module for further questioning. For instance, if the competency internalization threshold is 75 points, and the current credible interview content has a competency internalization score of 70 points, then the score is lower than the threshold, and therefore follow-up questioning is deemed necessary. The competency internalization score, the reasons for the score, and the follow-up questioning rules are then input into the large language model via prompt words, requiring the large language model to generate corresponding follow-up questions. Based on these follow-up questions, supplementary responses from the target audience are obtained, i.e., the third supplementary interview content.

[0080] Step 506: Determine the third combined content of the credible interview content and the third supplementary interview content, and determine the capability internalization assessment results corresponding to the third combined content.

[0081] The third merged content is the interview content obtained by merging the credible interview content and the third supplementary interview content. Specifically, by determining the third merged content of the credible interview content and the third supplementary interview content, and by calling the large language model to perform a capability internalization assessment on the third merged content, the corresponding capability internalization assessment result is determined.

[0082] Step 508: If the assessment result of the third merged content reaches the capability internalization threshold, the third merged content is determined as the interview content to be assessed to meet the capability internalization objective.

[0083] Specifically, if the assessment result of the third merged content's ability internalization reaches the ability internalization threshold, then the third merged content can be identified as the interview content to be assessed that meets the ability internalization objective.

[0084] If the internalization assessment result does not reach the internalization threshold, further follow-up questions are asked based on step 504 above, and the corresponding internalization score is recalculated. Then, it is reassessed whether the internalization goal is met, until the internalization goal is met, at which point the corresponding interview content to be evaluated is obtained. By assessing internalized capabilities and dynamically adjusting the questioning strategy based on real-time assessment results, the efficiency of information acquisition is improved.

[0085] In an exemplary embodiment, in step 210, a question template matching the preliminary ability assessment result is determined from a preset question bank, and target interview content is determined based on the interview content to be assessed and the question template. Specifically, this may further include: determining the information content difference between the preliminary ability assessment result and the adjacent higher level; determining a matching question template from the preset question bank based on the information content difference and the preliminary ability assessment result; rewriting the question template using a large language model according to the interview content to be assessed to generate target follow-up questions; and obtaining the target interview content for the target object regarding the target follow-up questions.

[0086] Specifically, firstly, by constructing a feature vector space for capability rating, establishing an evidence likelihood function, and calculating the posterior probability distribution in real time, the information difference ΔI(Li,Li+1)=|log2(P(Li|Et))-log2(P(Li+1|Et))| between the current capability rating Li (i.e., the preliminary capability assessment result) and the adjacent capability rating Li+1 (i.e., the adjacent next higher level) is calculated based on the principle of information theory, thereby generating the optimal probing strategy that maximizes the expected information gain.

[0087] For example, if the ability rating space is Ω={L1,L2,...,Ln}, where Li represents the i-th ability level, then each ability level Li corresponds to a feature vector: Li=[competency_vector,behavioral_indicators,evidence_requirements] Where: competency_vector represents the competency feature vector (with dimension k), behavioral_indicators represents the set of behavioral indicators, and evidence_requirements represents the evidence requirement weights.

[0088] For example, L2 = [(team collaboration = 0.8, goal setting = 0.7, resource allocation = 0.5, ...), the set of behavioral indicators = {taking initiative to assume responsibility, organizing team meetings, assigning clear tasks ...}, the weight of evidence = {key behavior × 0.6, result data × 0.4}], and all levels constitute the feature vector space Ω = {L1, L2, ..., Ln}.

[0089] The evidence likelihood function P(E|Li) measures the probability of observing current behavioral evidence E under the rank Li hypothesis. Specifically, it can be implemented by calculating the semantic matching degree between the behavioral event and the feature vectors of each rank using a large language model. For example, if an interviewee describes "I was responsible for assigning team tasks and completed the project on time," the matching degree with the L2 feature vector is 0.75, i.e., P(E|L2) = 0.75; and the matching degree with L3 is 0.30, i.e., P(E|L3) = 0.30.

[0090] The posterior probability distribution P(Li|E) is determined based on Bayes' theorem, such as P(Li|E) = P(E|Li) × P(Li) / P(E). Combined with the prior probabilities of each level (initially set to a uniform distribution or determined by preliminary rating results), the probabilities of each level are updated in real time to obtain the posterior distribution {P(L1|E), P(L2|E), P(L3|E)}. After each new answer is obtained, the posterior distribution can be updated to the prior for the next round, achieving dynamic Bayesian updating. The information content of the ability rating based on information theory is calculated as follows: I(Li)=-log2(P(Li|Evidence_current)) Where P(Li|Evidence_current) represents the posterior probability that the interviewee belongs to ability level Li under the current evidence conditions.

[0091] The calculation of the information content difference between adjacent levels is as follows: ΔI(Li,Li+1)=|I(Li)-I(Li+1)|=|log2(P(Li+1|E))-log2(P(Li|E))| The above can be simplified to: ΔI(Li,Li+1)=log2(P(Li|E) / P(Li+1|E)) Then, based on the principle of maximizing information content, and according to the information content calculation results, the ability level with the greatest difference is selected, and question templates that match the current ability assessment (i.e., the preliminary ability assessment results) are intelligently filtered from the preset question bank. A large language model is then invoked to personalize the filtered question templates according to the context of the interview content to be assessed, generating targeted follow-up questions. Finally, additional questions are asked to the target subject based on the target follow-up questions to obtain the target interview content regarding the target follow-up questions.

[0092] The ability level with the greatest difference can be determined in the following way: For example, the information content of each level Li under the current evidence Et can be calculated as I(Li) = -log2(P(Li|Et)), and the information content difference between adjacent levels (i.e., the current level Li and the adjacent previous level (Li+1)) is ΔI(Li,Li+1) = |log2(P(Li|Et) / P(Li+1|Et))|. The smaller the difference, the closer the probabilities between the two levels are, the more difficult it is to distinguish them, and the more necessary it is to ask follow-up questions. For example, if P(L2|Et) = 0.45 and P(L3|Et) = 0.40, then ΔI(L2,L3) = |log2(0.45 / 0.40)| ≈ 0.17 bits, indicating low discrimination and requiring follow-up questions on behavioral evidence at the L2 / L3 boundary. If P(L1|Et) = 0.10, then ΔI(L1,L2) = |log2(0.10 / 0.45)| ≈ 2.17 bits, indicating high discrimination and no need to follow up on the L1-L2 boundary. Therefore, by selecting the level pair with the smallest ΔI (i.e., the adjacent level with the closest posterior probability) among all adjacent level pairs (i.e., level Li and its adjacent previous level (Li+1)), the corresponding follow-up questions will bring the greatest information gain. Therefore, questions targeting this level boundary can be selected from the question bank first.

[0093] In one scenario, a large language model can be invoked to semantically match the obtained response text with preset ability standards. Based on the degree of matching between the behavioral performance identified in the response text and the behavioral indicators of each level, the most likely ability level range of the interviewee can be determined, serving as the initial grading basis for subsequent adaptive questioning. For example, taking "leadership" ability as an example, if the interviewee's answer reflects goal setting and team division of labor (entry-proficiency behavioral indicators), but does not reflect cross-departmental influence or strategic resource integration (proficiency-expert indicators), then the initial grading is set to the proficiency range, and subsequent follow-up questions will focus on behavioral evidence related to the proficiency level.

[0094] In an exemplary embodiment, in step 212, based on the interview content, a large language model is invoked to perform a comprehensive ability assessment to determine the comprehensive ability assessment result for the target object. Specifically, this may include: integrating the initial interview content, improved interview content, credible interview content, interview content to be assessed, and target interview content to construct an interview record dataset; invoking the large language model to perform behavior recognition on the interview record dataset to obtain multiple identified behavioral events; and determining the comprehensive ability assessment result for the target object based on the matching degree between each behavioral event and the preset ability standard.

[0095] The interview content includes initial interview content, refined interview content, credible interview content, interview content to be evaluated, and target interview content. Behavioral events can be events that include "verb + object + result".

[0096] Specifically, a complete interview record dataset is constructed by integrating the initial interview content, improved interview content, credible interview content, interview content to be evaluated, and target interview content obtained above.

[0097] Then, the large language model is invoked, and its behavioral event encoding engine automatically identifies and extracts all behavioral events from the interview content based on the encoding rule of "verb + object + result". The identification of behavioral events can be achieved using sentence boundary recognition and segmentation techniques. For example, semantic analysis can be performed based on the large language model to identify multiple behavioral events in the interview record dataset and output an independent set of behavioral events {S1, S2, ..., Sn}.

[0098] Then, through the ability assessment engine of the large language model, the matching degree between each behavioral event and the preset ability standards is calculated. These preset ability standards can be implemented through standardized modeling, and can include features across multiple dimensions. Taking "leadership ability" as an example: Competency Standards = { Skill Name: "Leadership" Ability Levels: [Beginner (1-2), Proficient (3-4), Expert (5-6), Specialist (7-8), Master (9-10)] Behavioral performance: [A specific set of behavioral indicators, including behavioral descriptions and weights] Contrarian Indicator: [Set of Negative Behavioral Manifestations] } By invoking the capability assessment engine of a large language model, the output set of behavioral events is matched with preset capability standards based on semantic similarity to obtain the matching degree of each behavioral event. For example, each behavioral event in the set can be scored against the preset capability standards to obtain the score, i.e., the matching degree, of each behavioral event.

[0099] Finally, based on the matching degree of each behavioral event, the target object's various ability indicators are quantitatively scored, thereby generating a comprehensive ability assessment result. The quantitative scoring logic for each ability is: Matching_Score = Σ(w i ×b i ), where w i b is the weighting coefficient for behavioral performance (including negative behavior). iEach behavioral performance (including negative behaviors) is assigned a quantitative score. The overall ability assessment result can be a weighted sum of the quantitative scores corresponding to each ability indicator. For example, the overall ability assessment result Total_Score = Σ(Wj × Matching_Score_j), where Wj is the preset weight coefficient of the j-th ability in the current ability standard (the sum of the weights of each ability is 1), and Matching_Score_j represents the quantitative score corresponding to the j-th ability indicator. The overall ability assessment result can also include score details and level conclusions for each ability.

[0100] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0101] Based on the same inventive concept, this application also provides a large language model-based interview assessment device for implementing the aforementioned large language model-based interview assessment method. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the large language model-based interview assessment device provided below can be found in the limitations of the large language model-based interview assessment method described above, and will not be repeated here.

[0102] In one exemplary embodiment, such as Figure 6 As shown, an interview assessment device based on a large language model is provided, including: a completeness assessment module 602, a credibility assessment module 604, a capability internalization assessment module 606, a preliminary capability assessment module 608, a target content determination module 610, and a comprehensive capability assessment module 612, wherein: The completeness assessment module 602 is used to obtain the initial interview content for the target object, call the large language model to perform a completeness assessment on the initial interview content based on the pre-configured completeness assessment criteria, and determine the completeness assessment result. The credibility assessment module 604 is used to determine the complete interview content that meets the completeness target based on the completeness assessment result, call the large language model to perform credibility assessment on the complete interview content, and determine the credibility assessment result. The capability internalization assessment module 606 is used to determine credible interview content that meets the credibility target based on the credibility assessment result, call the large language model to perform capability internalization assessment on the credible interview content, and determine the capability internalization assessment result. The preliminary ability assessment module 608 is used to determine the interview content to be assessed that meets the ability internalization goal based on the ability internalization assessment result, call the large language model to perform a preliminary ability assessment on the interview content to be assessed, and determine the preliminary ability assessment result. The target content determination module 610 is used to determine the question template that matches the preliminary ability assessment result from the preset question bank, and to determine the target interview content based on the interview content to be assessed and the question template. The comprehensive ability assessment module 612 is used to perform a comprehensive ability assessment by calling the large language model based on the initial interview content, the improved interview content, the credible interview content, the interview content to be assessed, and the target interview content, and to determine the comprehensive ability assessment result for the target object.

[0103] In an exemplary embodiment, the completeness assessment criteria include a multi-dimensional assessment criteria constructed based on the STAR principle. Each dimension includes multiple completeness features, and each completeness feature has a corresponding first weight. The completeness assessment module is specifically used to: for each dimension's multiple completeness features, call a large language model to obtain the first score corresponding to each completeness feature of the initial interview content; and perform weighted processing on the first score and first weight corresponding to each of the multiple completeness features in each dimension to obtain the completeness assessment result of the corresponding dimension.

[0104] In an exemplary embodiment, the credibility assessment module is further configured to: determine that the initial interview content is complete interview content that meets the completeness target if the completeness assessment results of each dimension of the initial interview content all reach the completeness threshold of the corresponding dimension; if the initial interview content contains completeness assessment results that do not reach the completeness threshold of the corresponding dimension, call a large language model based on multiple completeness features of the dimension to obtain the first supplementary interview content for the target object; determine the first merged content of the initial interview content and the first supplementary interview content, call a large language model based on a pre-configured completeness assessment standard to perform a completeness assessment on the first merged content, and determine the completeness assessment results of each dimension; if the first merged content contains completeness assessment results of each dimension that reach the completeness threshold of the corresponding dimension, determine that the first merged content is complete interview content that meets the completeness target.

[0105] In an exemplary embodiment, the credibility assessment module is further configured to: based on multiple preset credibility features, call a large language model to obtain the second score corresponding to each credibility feature of the improved interview content, and each credibility feature also has a corresponding second weight; perform weighted processing on the second scores and second weights corresponding to the multiple credibility features respectively to obtain the credibility assessment result of the improved interview content.

[0106] In an exemplary embodiment, the capability internalization assessment module is further configured to: determine that the improved interview content is credible interview content that meets the credibility target if the credibility assessment result of the improved interview content reaches the credibility threshold; if the credibility assessment result of the improved interview content does not reach the credibility threshold, call a large language model based on the multiple credibility features to obtain a second supplementary interview content for the target object; determine a second merged content of the improved interview content and the second supplementary interview content, call the large language model to perform a credibility assessment on the second merged content, and determine the corresponding credibility assessment result; if the credibility assessment result of the second merged content reaches the credibility threshold, determine that the second merged content is credible interview content that meets the credibility target.

[0107] In an exemplary embodiment, the capability internalization assessment module is further configured to: based on multiple preset capability internalization features, call a large language model to obtain the third score corresponding to each capability internalization feature of the credible interview content, and each capability internalization feature also has a corresponding third weight; perform weighted processing on the third scores and third weights corresponding to the multiple capability internalization features to obtain the capability internalization assessment result of the credible interview content.

[0108] In an exemplary embodiment, the preliminary capability assessment module is further configured to: determine, if the capability internalization assessment result of the credible interview content reaches the capability internalization threshold, determine the credible interview content as interview content to be assessed that meets the capability internalization target; if the capability internalization assessment result of the credible interview content does not reach the capability internalization threshold, invoke a large language model based on the multiple capability internalization features to obtain a third supplementary interview content for the target object; determine a third combined content of the credible interview content and the third supplementary interview content, invoke the large language model to perform capability internalization assessment on the third combined content, and determine the corresponding capability internalization assessment result; if the capability internalization assessment result of the third combined content reaches the capability internalization threshold, determine the third combined content as interview content to be assessed that meets the capability internalization target.

[0109] In an exemplary embodiment, the target content determination module is further configured to: determine the information content difference between the preliminary ability assessment result and the adjacent previous level; determine a matching question template from a preset question bank based on the information content difference and the preliminary ability assessment result; rewrite the question template by calling the large language model according to the interview content to be assessed to generate target follow-up questions; and obtain the target interview content of the target object for the target follow-up questions.

[0110] In an exemplary embodiment, the comprehensive ability assessment module is further configured to: integrate the initial interview content, the improved interview content, the credible interview content, the interview content to be assessed, and the target interview content to construct an interview record dataset; call the large language model to perform behavior recognition on the interview record dataset to obtain multiple identified behavioral events; and determine the comprehensive ability assessment result for the target object based on the matching degree between each behavioral event and the preset ability standard.

[0111] The modules in the aforementioned large language model-based interview assessment device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0112] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 7As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When executed by the processor, the computer program implements an interview assessment method based on a large language model. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0113] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0114] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0115] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0116] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0117] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0118] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0119] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0120] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An interview assessment method based on a large language model, characterized in that, The method includes: Obtain the initial interview content for the target audience, and use a large language model to evaluate the completeness of the initial interview content based on a pre-configured completeness evaluation standard to determine the completeness evaluation result. Based on the completeness assessment results, complete interview content that meets the completeness target is determined, and the credibility assessment of the complete interview content is performed by calling the large language model to determine the credibility assessment results. Based on the credibility assessment results, credible interview content that meets the credibility objectives is determined. The large language model is then invoked to perform a capability internalization assessment on the credible interview content, and the capability internalization assessment results are determined. Based on the internalization assessment results, determine the interview content to be assessed that meets the internalization goals, call the large language model to conduct a preliminary ability assessment on the interview content to be assessed, and determine the preliminary ability assessment results. A question template matching the preliminary ability assessment results is determined from a preset question bank, and target interview content is determined based on the interview content to be assessed and the question template. Based on the initial interview content, the improved interview content, the credible interview content, the interview content to be evaluated, and the target interview content, the large language model is invoked to perform a comprehensive ability assessment, and the comprehensive ability assessment result for the target object is determined.

2. The method according to claim 1, characterized in that, The perfection assessment criteria include a multi-dimensional assessment criteria based on the STAR principle. Each dimension includes multiple perfection features, and each perfection feature has a corresponding first weight. The completeness assessment of the initial interview content is performed using a large language model based on pre-configured completeness assessment criteria, and the completeness assessment result is determined, including: For each dimension with multiple completeness features, a large language model is invoked to obtain the first score corresponding to each completeness feature of the initial interview content; The first score and first weight corresponding to multiple perfection features in each dimension are weighted to obtain the perfection evaluation result of the corresponding dimension.

3. The method according to claim 2, characterized in that, The step of determining the completeness interview content that meets the completeness objectives based on the completeness assessment results includes: If the completeness assessment results of each dimension of the initial interview content reach the corresponding completeness threshold, the initial interview content is determined to be complete interview content that meets the completeness target. If it is determined that the completeness assessment result in the initial interview content does not reach the completeness threshold of the corresponding dimension, the large language model is invoked based on multiple completeness features of the dimension to obtain the first supplementary interview content for the target object. The first merged content of the initial interview content and the first supplementary interview content is determined. Based on the pre-configured completeness assessment criteria, the large language model is called to assess the completeness of the first merged content and determine the completeness assessment results of each dimension. If the completeness assessment results of each dimension of the first merged content are determined to meet the completeness threshold of the corresponding dimension, the first merged content is determined to be complete interview content that meets the completeness target.

4. The method according to claim 1, characterized in that, The process of calling the large language model to perform a credibility assessment on the refined interview content and determining the credibility assessment result includes: Based on multiple preset credibility features, a large language model is invoked to obtain the second score corresponding to each credibility feature of the complete interview content, and each credibility feature also has a corresponding second weight. The credibility assessment results of the improved interview content are obtained by weighting the second score and second weight corresponding to the multiple credibility features.

5. The method according to claim 4, characterized in that, The process of determining credible interview content that meets the credibility objective based on the credibility assessment results includes: If the credibility assessment result of the improved interview content reaches the credibility threshold, the improved interview content is determined to be credible interview content that meets the credibility target. If the credibility assessment result of the improved interview content does not reach the credibility threshold, a large language model is invoked based on the multiple credibility features to obtain a second supplementary interview content for the target object. The second merged content of the improved interview content and the second supplementary interview content is determined, and the credibility of the second merged content is evaluated by calling the large language model to determine the corresponding credibility evaluation result. If the credibility assessment result of the second merged content reaches the credibility threshold, the second merged content is determined to be credible interview content that meets the credibility target.

6. The method according to claim 1, characterized in that, The process of calling the large language model to perform a capability internalization assessment on the credible interview content and determining the capability internalization assessment results includes: Based on multiple preset capability internalization features, a large language model is invoked to obtain the third score corresponding to each capability internalization feature of the credible interview content. Each capability internalization feature also has a corresponding third weight. The third score and third weight corresponding to multiple internalization features of the ability are weighted to obtain the internalization assessment result of the credible interview content.

7. The method according to claim 6, characterized in that, The step of determining the interview content to be assessed based on the capability internalization assessment results, which meets the capability internalization objectives, includes: If the assessment result of the ability internalization of the credible interview content reaches the ability internalization threshold, the credible interview content is determined to be the interview content to be assessed that meets the ability internalization goal. If the assessment result of the ability internalization of the credible interview content does not reach the ability internalization threshold, a third supplementary interview content for the target object is obtained by calling a large language model based on the multiple ability internalization features. The third merged content of the credible interview content and the third additional interview content is determined, and the large language model is invoked to perform a capability internalization assessment on the third merged content to determine the corresponding capability internalization assessment result. If the assessment result of the third merged content reaches the capability internalization threshold, the third merged content is determined to be the interview content to be assessed that meets the capability internalization objective.

8. The method according to claim 1, characterized in that, The step of determining a follow-up question template from a pre-set question bank that matches the preliminary ability assessment results, and determining target interview content based on the interview content to be assessed and the follow-up question template, includes: Determine the information content difference between the preliminary capability assessment results and the adjacent higher level; Based on the differences in information content and the preliminary ability assessment results, a matching question template is determined from the preset question bank; Based on the interview content to be evaluated, the large language model is invoked to rewrite the question follow-up template and generate target follow-up questions; Obtain the content of the target interview regarding the follow-up questions asked by the target.

9. The method according to claim 1, characterized in that, Based on the initial interview content, the refined interview content, the credible interview content, the interview content to be evaluated, and the target interview content, the large language model is invoked to perform a comprehensive ability assessment, determining the comprehensive ability assessment result for the target object, including: The initial interview content, the improved interview content, the credible interview content, the interview content to be evaluated, and the target interview content are integrated to construct an interview record dataset; The large language model is invoked to perform behavior recognition on the interview transcript dataset, resulting in multiple identified behavioral events; Based on the degree of matching between each behavioral event and the preset capability standard, a comprehensive capability assessment result for the target object is determined.

10. An interview assessment device based on a large language model, characterized in that, The device includes: The completeness assessment module is used to obtain the initial interview content for the target object, call the large language model to assess the completeness of the initial interview content based on the pre-configured completeness assessment criteria, and determine the completeness assessment result. The credibility assessment module is used to determine the completeness interview content that meets the completeness target based on the completeness assessment result, call the large language model to perform credibility assessment on the completeness interview content, and determine the credibility assessment result. The capability internalization assessment module is used to determine credible interview content that meets the credibility target based on the credibility assessment results, call the large language model to perform capability internalization assessment on the credible interview content, and determine the capability internalization assessment results. The preliminary competency assessment module is used to determine the interview content to be assessed that meets the competency internalization goal based on the competency internalization assessment results, call the large language model to perform a preliminary competency assessment on the interview content to be assessed, and determine the preliminary competency assessment results. The target content determination module is used to determine the question template that matches the preliminary ability assessment results from a preset question bank, and to determine the target interview content based on the interview content to be assessed and the question template. The comprehensive ability assessment module is used to perform a comprehensive ability assessment by calling the large language model based on the initial interview content, the improved interview content, the credible interview content, the interview content to be assessed, and the target interview content, and to determine the comprehensive ability assessment result for the target object.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.