Resume processing method based on large model and electronic device

By using a large model to identify, correct, and complete resume information, a resume knowledge graph is generated and evaluated from multiple dimensions. This solves the problems of low efficiency in manual screening and difficulty in iterating rule-based systems, thereby improving the accuracy and automation of resume screening.

CN121481492BActive Publication Date: 2026-06-12CHIA TAI TIANQING PHARMA GRP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHIA TAI TIANQING PHARMA GRP CO LTD
Filing Date
2026-01-07
Publication Date
2026-06-12

Smart Images

  • Figure CN121481492B_ABST
    Figure CN121481492B_ABST
Patent Text Reader

Abstract

The present disclosure provides a large model-based resume processing method and an electronic device, and relates to the technical field of large models and agents. The method comprises: identifying structured first resume information from a to-be-processed resume file; inputting the first resume information and an error correction prompt word as prompt information into a large model for error correction processing, and generating a resume knowledge graph from output structured second resume information, the error correction processing comprising a conflict information arbitration determination operation; inputting the resume knowledge graph, the first resume information and an information completion prompt word as prompt information into the large model for information completion processing to obtain structured third resume information, the information completion processing comprising a quantitative conversion operation on ambiguous expressions; and based on the third resume information, evaluating the to-be-processed resume file in multiple preset evaluation dimensions to obtain a resume evaluation result. The method can automatically extract accurate resume information and perform multi-dimensional analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, specifically to the field of large model and intelligent agent technology, and particularly to a resume processing method and electronic device based on a large model. Background Technology

[0002] In the field of human resources recruitment, faced with a massive number of resumes in various formats (such as PDFs, scans, and images), traditional screening methods mainly rely on manual reading or keyword matching systems based on fixed rules. However, manual screening is inefficient, subjective, and prone to errors due to fatigue. On the other hand, rule-based automated systems make it difficult to achieve both efficiency and quality in the screening process due to the difficulty in iterating and updating the rules in a timely manner, thus failing to meet the needs of modern efficient and precise recruitment. Summary of the Invention

[0003] This disclosure presents a resume processing method, apparatus, electronic device, computer-readable storage medium, and computer program product based on a large model.

[0004] In a first aspect, embodiments of this disclosure propose a resume processing method based on a large model, comprising: identifying structured first resume information from a resume file to be processed; inputting the first resume information and preset error correction prompts as prompts into a preset large model for error correction processing, and generating a resume knowledge graph from the output structured second resume information in the form of a knowledge graph; wherein, the error correction processing includes arbitration and judgment operations for conflicting information; inputting the resume knowledge graph, the first resume information, and preset information completion prompts as prompts into the large model for information completion processing to obtain structured third resume information; wherein, the information completion processing includes quantification and conversion operations for fuzzy expressions; and evaluating the resume file to be processed on multiple preset evaluation dimensions based on the third resume information to obtain a resume evaluation result.

[0005] Secondly, embodiments of this disclosure propose a resume processing apparatus based on a large model, comprising: a first resume information recognition unit configured to identify structured first resume information from a resume file to be processed; an error correction processing and knowledge graph generation unit configured to input the first resume information and preset error correction prompts as prompts into a preset large model for error correction processing, and generate a resume knowledge graph from the output structured second resume information in the form of a knowledge graph; wherein the error correction processing includes an arbitration judgment operation for conflicting information; an information completion unit configured to input the resume knowledge graph, the first resume information, and preset information completion prompts as prompts into the large model for information completion processing to obtain structured third resume information; wherein the information completion processing includes a quantification conversion operation for fuzzy expressions; and a multi-dimensional evaluation unit configured to evaluate the resume file to be processed on multiple preset evaluation dimensions based on the third resume information to obtain a resume evaluation result.

[0006] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the large-model-based resume processing method as described in the first aspect.

[0007] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer to implement the large-model-based resume processing method as described in the first aspect.

[0008] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the steps of the resume processing method based on a large model as described in the first aspect.

[0009] The resume processing solution based on a large model disclosed herein firstly identifies structured first resume information directly from the resume files to be processed, providing a preliminary standardized data foundation for subsequent in-depth processing and improving the initial processing capability for multi-format source files. Secondly, by combining the first resume information with error correction prompts and inputting it into the large model for error correction processing, and specifically including arbitration and judgment operations for conflicting information, it can effectively identify and resolve potential contradictions, ambiguities, or redundant information in the resume, thereby outputting more accurate and consistent second resume information. This second resume information is then used to generate a resume knowledge graph in the form of a knowledge graph, enabling the internal connections of information to be structured. Standardized representation lays a solid foundation for subsequent deep semantic processing. Next, by inputting the resume knowledge graph, initial resume information, and information completion prompts into a large model for information completion processing, including a specific quantification conversion operation for ambiguous expressions, the model can intelligently complete, refine, and standardize simplified, colloquial, and other vague descriptions in the resume, resulting in more complete and semantically clearer third-party resume information. This significantly improves the usability and comparability of the resume content. Finally, based on this high-quality, structured third-party resume information, evaluation is conducted across multiple preset evaluation dimensions, generating more objective, comprehensive, and interpretable resume evaluation results. This solution, through two targeted applications of the large model at key stages (error correction and completion), effectively overcomes the two major bottlenecks of inaccurate original information and semantic ambiguity. Ultimately, it achieves a simultaneous improvement in the accuracy, automation, and evaluation depth of resume screening, and also realizes automated and intelligent processing from the original resume file to the final evaluation result.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0012] Figure 1 This is an exemplary system architecture to which this disclosure can be applied;

[0013] Figure 2 A flowchart illustrating a resume processing method based on a large model, provided for embodiments of this disclosure;

[0014] Figure 3 A flowchart illustrating a method for arbitrating conflict information, provided as an embodiment of this disclosure;

[0015] Figure 4A flowchart illustrating a method for compensating cross-regional information and aligning different representations of the same entity, provided for embodiments of this disclosure;

[0016] Figure 5 A flowchart illustrating a method for quantization conversion of fuzzy representations provided in this embodiment of the disclosure;

[0017] Figure 6 A flowchart illustrating a scoring method for different preset evaluation dimensions, provided for embodiments of this disclosure;

[0018] Figure 7 A structural block diagram of a resume processing device based on a large model provided in this disclosure embodiment;

[0019] Figure 8 This is a schematic diagram of the structure of an electronic device suitable for performing a resume processing method based on a large model, provided as an embodiment of the present disclosure. Detailed Implementation

[0020] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0021] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0022] Figure 1 An exemplary system architecture 100 is shown, to which embodiments of the large-model-based resume processing method, apparatus, electronic device, and computer-readable storage medium of this disclosure can be applied.

[0023] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0024] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include resume screening applications, image conversion applications, and instant messaging applications.

[0025] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.

[0026] Server 105 can provide various services through its built-in applications. Taking a resume screening application that provides resume processing services for massive amounts of resumes to enterprise managers as an example, server 105 can achieve the following effects when running this resume screening application: First, it receives resume files to be processed from terminal devices 101, 102, and 103 via network 104, and identifies structured first resume information from the resume files to be processed; then, it inputs the first resume information and preset error correction prompts as prompts into a preset large model for error correction processing, and generates a resume knowledge graph in the form of a knowledge graph based on the output structured second resume information. This error correction processing includes arbitration and judgment operations for conflicting information; next, it inputs the resume knowledge graph, the first resume information, and preset information completion prompts as prompts into the large model for information completion processing, obtaining structured third resume information. This information completion processing includes quantitative conversion operations for fuzzy expressions; finally, based on the third resume information, it evaluates the resume file to be processed on multiple preset evaluation dimensions to obtain the resume evaluation result.

[0027] It should be noted that, in addition to being temporarily obtained from terminal devices 101, 102, and 103 via network 104, the resume files to be processed can also be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (for example, when starting to process previously stored resume processing tasks), it can choose to directly obtain this data from locally. In this case, the exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.

[0028] Because complex analysis of resume files requires significant computing resources and power, the resume processing methods based on large models provided in the subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the resume processing device based on large models is also generally located within the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also perform the aforementioned calculations performed by the server 105 through their installed resume screening applications, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the resume screening application determines that its terminal device has strong computing power and abundant remaining computing resources, the terminal device can perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the resume processing device based on large models can also be located within terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.

[0029] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0030] Please refer to Figure 2 , Figure 2 A flowchart of a resume processing method based on a large model provided for embodiments of this disclosure, wherein process 200 includes the following steps:

[0031] Step 201: Identify the structured first resume information from the resume file to be processed;

[0032] This step is intended for the implementation of a large-model-based resume processing method (e.g., Figure 1The server 105 shown, or a resume screening application installed on server 105, identifies structured first resume information from the resume files to be processed. The aim is to initially transform the raw, varied resume files into a structured data format that a machine can initially understand and process. The "resume file to be processed" is a broader concept, encompassing various electronic document formats or audio / video file formats actually received by users in recruitment scenarios, such as text-selectable PDF documents, image files in pure image formats (e.g., JPG, PNG), images or PDF files generated from scanned paper documents, and audio / video files. After receiving these files, the server, for document-formatted files, can first identify these visual information such as layout and font styles, as these are often crucial for understanding the structure of the resume content, and then perform text information recognition to obtain structured first resume information; for audio / video files, it can first perform speech-to-text processing or take screenshots and then extract text from the images.

[0033] Specifically, the execution entity does not perform simple Optical Character Recognition (OCR) when executing this step. Traditional OCR outputs are often continuous, undifferentiated text streams, losing the original text's position, grouping, and visual hierarchy on the page. This step requires producing "structured" information, so the execution entity must also perform "layout awareness." In practice, the execution entity can call or have a built-in document image analysis and understanding module. This module first parses the file to be processed; if it is a PDF, it extracts its page images and possible text layer information; if it is an image, it processes the image directly. Subsequently, this module can perform layout analysis on the page image... Computer vision technology is used to identify different logical regions. These logical regions are logical blocks determined by a combination of visual features (such as font size, bolding, alignment, and spacing) and text content semantics (such as the presence of title-like words like "work experience" and "project experience"). Examples include "personal information title area," "educational background list area," and "detailed work experience area." Simultaneously or after identifying these regions, the aforementioned execution entity can perform text detection and recognition within each region to ensure that the text in the image is accurately converted into digital text. Crucially, the binding relationship between each text fragment and its corresponding logical region, as well as the relative position or order of the text within the region, are preserved.

[0034] Through the above operations, the "first resume information" generated by the executing entity is a preliminary form of structured data. This structured data does not refer to a database record with precisely filled standard fields such as "name," "company," and "position," but rather an intermediate representation. It typically manifests as a hierarchical, labeled data structure, such as a list or tree structure, where each node represents an identified section area (which may have a section type label, such as "section_heading" or "item_list"), and the nodes are associated with the ordered text content identified within that area. For example, the executing entity might output a structure containing a block labeled "work_experience," under which several "job_entry" sub-blocks are arranged chronologically, each containing identified lines of raw text. This structure preserves the visual organization logic of the resume document, providing crucial information for subsequent steps to distinguish which text belongs to the same experience, which are titles, and which are detailed descriptions. In short, "First Resume Information" has transformed from unstructured document images into a set of programmable text that retains the original document organization logic, completing the first key data morphological change in the information intelligent processing flow.

[0035] One implementation method, including but not limited to, is as follows: First, a visual Transformer model is used to identify different page regions in the resume file to be processed; then, a differentiable binarization network is used to determine the corresponding text within each page region; finally, structured first resume information is generated based on the membership relationship between the text within each page region and its respective page region.

[0036] Specifically, the aforementioned execution entity first invokes the Vision Transformer (VIT) model to perform a global analysis of the resume document image. VIT is a deep learning model that applies the successful Transformer architecture from natural language processing to computer vision tasks. Its technical principle lies in segmenting the input image into a series of fixed-size image patches and linearly embedding each patch into a vector sequence. Through a multi-layered self-attention mechanism within the model, these vectors dynamically capture the long-range dependencies between all patches in the image, thereby achieving a global understanding of the image content. In the specific implementation of this step, the execution entity preprocesses the resume document to be processed into a uniform-sized image and inputs it into the pre-trained VIT model. This model, through its learned prior knowledge of document structure, can output the recognition results of different semantic regions in the image. For example, it accurately distinguishes the "Personal Basic Information Title Area," "Education Background List Area," "Work Experience Details Area," "Skills Block," and "Project Experience Description Area," etc. These regions are given by the model in the form of bounding box coordinates and region type labels, constituting a machine interpretation of the resume's logical structure.

[0037] Next, based on the obtained page layout region segmentation, a Differentiable Binarization Network (DBNet) is further used to perform refined text detection and recognition on each identified page layout region. DBNet is a model specifically designed for text detection in complex scenes. Its core technology lies in introducing a differentiable binarization operation, enabling the model to learn end-to-end how to accurately predict whether each pixel belongs to a text region, and it has excellent adaptability to text lines of arbitrary shapes (such as curved or tilted). In practice, the aforementioned execution entity inputs image slices of each page layout region output by the VIT model into DBNet. DBNet first outputs a probability map representing the probability that each pixel belongs to text, and then generates a binary mask of the text region through its differentiable binarization module, finally outputting a fine-grained bounding box at the level of each text line or word. Subsequently, the aforementioned execution entity can call or integrate an optical character recognition engine to perform character recognition on the image content within these detected text boxes, converting the image into a string of text. At this point, each page area is associated with a set of text content it contains and its relative positions within the area.

[0038] Step 202: Input the first resume information and the preset error correction prompts together as prompt information into the preset large model for error correction processing, and generate a resume knowledge graph in the form of a knowledge graph from the output structured second resume information;

[0039] Building upon step 201, this step aims to have the aforementioned executing entity input the first resume information and a pre-designed error-correction prompt as prompt information into a pre-designed large model for error correction processing, outputting structured second resume information. This second resume information is then used to generate a resume knowledge graph. The purpose is to deeply verify, correct, and integrate the previously identified "structured first resume information" to generate a more accurate, consistent, and logically coherent "second resume information," ultimately organizing it into a standardized "resume knowledge graph." To achieve this, the aforementioned executing entity combines the "first resume information" with a pre-designed "error-correction prompt" as prompt information, inputting it into a pre-designed large model (also known as a large language model or multimodal large model). This "error-correction prompt" is actually a complex text template containing specific task descriptions, role settings, and domain knowledge constraints. Its function is to precisely guide the large model to understand and execute a series of pre-designed error-correction sub-tasks. Furthermore, the powerful contextual understanding and reasoning capabilities of the large model automatically process various complex problems in the original extracted information, thereby significantly improving the data quality foundation for subsequent processes.

[0040] Specifically, the "error correction process" can cover multiple levels of error correction operations, such as firstly including arbitration and determination of conflicting information. This is to address the fact that the same entity (such as the employer company) may appear in different locations in a resume under different names (e.g., "XX Group" and "XX Treasure"), leading to factual conflicts or ambiguities. Therefore, when the aforementioned executing entity performs arbitration judgment through the large model, it can instruct the large model to simulate a "reviewer or arbitrator" with domain knowledge. This first identifies these semantically similar candidate entities and, combined with the layout area labels carried in the "first resume information" (e.g., whether the information comes from the "Work Experience" title area or the "Project Description" body area), assigns initial confidence weights to candidate entities from different sources. Furthermore, the aforementioned executing entity can also guide the large model to establish temporal logic constraints, that is, align each candidate entity with the time information identified in the resume (e.g., the start and end years of employment). If a certain entity's statement cannot find a reasonable position on the timeline, its credibility will be significantly reduced. Finally, the large model integrates spatial weights and temporal logic, and calls upon its inherent common sense knowledge (e.g., knowing that "XX Treasure" is a business under "XX Group"), thereby normalizing the judgment of conflicting entities and outputting the most credible and standard entity statement as the arbitration result.

[0041] Of course, this error correction process can also include compensation operations for cross-regional information and alignment operations for different representations of the same entity, in order to further eliminate problems in the first resume information. Specific explanations will be provided in detail in subsequent embodiments. After completing the above error correction process, the executing entity can generate a resume knowledge graph from the corrected and unified "structured second resume information" output by the large model, in the form of a "knowledge graph." This means organizing it according to a "entity-relationship-entity" triple structure (for example, constructing a triple from the elements "Zhang San," "employed at," and "XX Group"; this transformation process can be guided by a predefined knowledge graph schema, which specifies common entity and relation types in the resume domain). This transforms linear, paragraph-style text information into a structured network data model with interconnected nodes and edges. This transformation not only allows the inherent relationships of information to be explicitly represented but also lays a crucial data foundation for subsequent semantic retrieval, relational reasoning, and deep analysis based on graph structures.

[0042] To drive the large model to perform a series of predetermined and complex text analysis and correction tasks, the "correction prompt word" is not a simple instruction, but a structured text template that explicitly explains the "correction operation description information" to be performed. For example, it instructs the model to "identify and resolve contradictions and inconsistencies in the resume regarding company name, job timeline, or skill descriptions." When the aforementioned entity sends this set of prompt information to the large model, the model's primary task is to parse the prompt, extracting and understanding the specific "target correction operation to be performed." One implementation method, including but not limited to: using the large model to determine the target correction operation based on the correction operation description information contained in the prompt word; and using the large model to perform the target correction operation on the first resume information. That is, the prompt word should at least contain correction operation description information so that the aforementioned entity can use the large model to determine the target correction operation to be performed, such as a set of sub-tasks like "conflict arbitration," "information compensation," and "terminology alignment."

[0043] To ensure that the large model can perform error correction in a manner consistent with professional domain knowledge, the error correction prompt can also include first role setting information and first domain knowledge constraint information. This further controls the large model to assume the first role corresponding to the first role setting information, and controls the large model to determine the error correction constraint information to be followed when performing the target error correction operation based on the first domain knowledge constraint information. Thus, during the execution phase, the large model can be controlled to assume the first role and perform the target error correction operation on the first resume information under the constraints of the error correction constraint information. This implementation further refines the design of "error correction prompts" by including "first role setting information" and "first domain knowledge constraint information." When these elements are included in the provided error correction prompts, the guidance and control over the large model becomes more precise and powerful. "First role setting information" instructions, such as "Please play the role of a senior human resources data governance expert," essentially inject a virtual "personality" or perspective into the large model. Technically, this provides a strong contextual prior, making the model more inclined to mimic the role's professional knowledge background, thinking patterns, and expression habits when generating responses, thereby enhancing the professionalism and rationality of error correction decisions. "First domain knowledge constraint information," on the other hand, provides more specific operational boundaries and rules. For example, "When resolving conflicts, please prioritize information located under the 'Work Experience' heading area and refer to the common parent-subsidiary relationship diagrams of Chinese internet companies." This is equivalent to adding filters and signposts to the large model's reasoning process, ensuring that its output does not deviate from business common sense and predetermined data governance standards.

[0044] In practical terms, when constructing a complete prompt message, the aforementioned executor can smoothly assemble "first resume information" (usually embedded in structured text formats such as JSON), role setting instructions, domain knowledge constraint clauses, and specific error correction operation descriptions into a coherent natural language instruction according to a predetermined template format. The executor then sends this instruction by calling the large model service through an application programming interface (API). Upon receiving the instruction, the large model undergoes an internal instruction following and role-playing process. It first "places" itself into the designated "first role," then interprets the "first domain knowledge constraint information" based on that role's cognitive framework, deriving the "error correction constraint information" that must be followed when performing the error correction operation. Ultimately, the text response produced by the large model strictly adheres to the role identity, and its conflict resolution solutions, information compensation suggestions, or terminology alignment results clearly reflect the guidance of domain constraints. For example, when a resume contains both "XX Group" and "XX Bao," the large-scale model, positioned as a "human resources data governance expert" and constrained by relevant corporate relationships, is more likely to output a standardized judgment such as "Based on the organizational relationship, 'XX Bao' should be uniformly merged under its parent company entity 'XX Group'," rather than a vague or incorrect suggestion. In this way, the aforementioned implementing entity effectively disciplines and guides the large-scale model's powerful general-purpose language capabilities onto the specific and professional task of resume information correction, achieving accuracy and reliability in automated processing.

[0045] Step 203: Input the resume knowledge graph, the first resume information, and the preset information completion prompts into the large model for information completion processing to obtain the structured third resume information;

[0046] Building upon step 202, this step involves the aforementioned executing entity inputting the resume knowledge graph, the first resume information, and preset information completion prompts into the large model for information completion processing, resulting in structured third resume information. The aim is to deeply enhance the information that has already undergone preliminary error correction and structured organization, thereby addressing common issues in resumes such as colloquialisms, ambiguity, and missing key details. Specifically, through the semantic understanding and reasoning capabilities of the large model, incomplete descriptions are transformed into standardized, quantifiable, and evaluable expressions.

[0047] Among them, the structured knowledge graph provides the large model with a clear, machine-readable context about the candidate's skills, experience and the relationships between them; the first resume information retains the text content identified from the original document and organized by page area, which will serve as supplementary original corpus to ensure that the original expressions that may not be fully structured but contain details, directly extracted by character recognition technology, are not lost during the completion process; the preset information completion prompts are the key control instructions that guide the entire completion process. Their content will not only indicate the task of "performing information completion", but can also define the specific rules, output format and professional context to be followed for the completion operation.

[0048] One of the core processing operations in this information completion process is the "quantification transformation of vague expressions." This refers to converting descriptions in resumes that lack specific technical details, action boundaries, and measurable results, such as "responsible for the backend system," "improved performance," and "participated in big data projects," into standardized sentences rich in clear technology stacks, specific action verbs, business scenarios, and quantifiable indicators. At the technical level, this relies on the powerful contextual reasoning, common-sense knowledge fusion, and instruction-following capabilities of the large model. The aforementioned execution entity can combine the resume knowledge graph (providing structured relational context), the initial resume information (providing raw text details), and completion prompts (providing transformation rules and format constraints) to form a composite prompt information input model. When interpreting this prompt, the large model simultaneously considers structured relational constraints (e.g., knowing from the knowledge graph that this person possesses "Spark" skills), the original vague text (e.g., "worked on recommendation systems"), and the hard requirements in the prompts (e.g., "must use the STAR method to describe and supplement the quantifiable results"), thereby performing comprehensive reasoning.

[0049] At the practical level, when implementing this step, the aforementioned executing entity can pre-store or dynamically generate "information completion prompts" optimized for different job types in the system. For example, when processing a specific resume, the aforementioned executing entity first retrieves the corresponding knowledge graph data and the first resume information from storage, and serializes them in a predetermined format (e.g., describing the knowledge graph in JSON-LD format and listing the first resume information in the form of text paragraphs). Then, these serialized data are combined with the completion prompt template to form a complete and coherent natural language instruction, which is sent to the large model service through the API. After receiving this composite instruction, the large model will execute its internal information completion logic. For example, when encountering the vague statement "I've developed a recommendation system," the large model retrieves relevant skill entities from the context (e.g., it learns from the knowledge graph that the candidate entity is associated with the skills "collaborative filtering" and "Spark"). Based on the constraints of "using strong verbs" and "supplementing quantitative indicators" in the prompts, it performs logical chain reasoning: the subject is the candidate, the object is the recommendation system, the available technologies are collaborative filtering and Spark, the reasonable action is "development" or "implementation," and the possible business goal is "personalized recommendation." This leads to the inference of a reasonable quantitative improvement direction, such as "click-through rate." Ultimately, the model generates "structured third-party resume information," such as "I developed a recommendation system using the Spark framework and collaborative filtering algorithm, increasing user content click-through rate by 15%." This output is not only professionally worded and clearly structured but also contains quantitative dimensions that can be directly used by subsequent evaluation modules, achieving a fundamental improvement from vague descriptions to calculable and comparable structured information. Through this step, the aforementioned executing entity significantly enhances the semantic density and evaluation usability of the resume information, laying a solid data foundation for subsequent objective and accurate intelligent scoring.

[0050] To drive the large model to perform semantic-level deep enhancement and reconstruction of the information that has been initially corrected and structured, the "information completion prompt" is not a simple instruction, but a structured text template that clearly explains the "information completion operation description information" to be completed to the large model. When the aforementioned execution entity sends this combined prompt information to the large model, the primary task of the large model is to parse this prompt, extract and understand what the "target information operation to be performed" is. One implementation method, including but not limited to, is as follows: using the large model to determine the target information completion operation to be performed based on the information completion operation description information contained in the information completion prompt; and using the large model to perform the target information completion operation on the resume knowledge graph and the first resume information. In other words, the aforementioned execution entity first relies on the "information completion operation description information" contained in the "information completion prompt" to guide the large model in determining the "target information completion operation that needs to be performed." This description information is a clear task definition, such as "Please identify and complete the missing technical details, quantitative achievements, and business background in all work experiences and project descriptions in the resume, and transform colloquial expressions into professional statements." When the large model receives this compound prompt, it internally parses the instructions and extracts one or more specific, executable sets of operation objectives from the natural language description, such as "supplementing the core technology stack for each project experience," "replacing vague verbs with strong verbs that conform to the STAR method," and "inferring a reasonable quantitative range for achievements lacking numbers," etc. This process constrains and guides the large model's open text generation capabilities toward solving the specific problem of incomplete and unprofessional resume information.

[0051] To make the completion operation more accurate and in line with professional domain expectations, the information completion prompts can also include second role setting information and second domain knowledge constraint information. In this case, the large model station can be additionally controlled to occupy the second role corresponding to the second role setting information, and the information completion constraint information that the large model should follow when performing the target information completion operation based on the second domain knowledge constraint information can be controlled. This information supplementary constraint information includes: the output format template of the completed information. Then the execution phase is adaptively changed to: controlling the large model station to occupy the second role, and performing the target information completion operation on the resume knowledge graph and the first resume information under the constraints of the information completion constraint information.

[0052] This implementation further allows embedding "second role setting information" and "second domain knowledge constraint information" within the "information completion prompts." When the prompts provided by the aforementioned executing entity contain these elements, its control over the behavior of the large model reaches a more granular level. The second role setting information and second domain knowledge constraint information are similar to the first role setting information and first domain knowledge constraint information mentioned above, differing only in the target audience. Furthermore, the "second domain knowledge constraint information" provides specific operational guidelines and boundaries, such as "during completion, the backend project must explicitly mention the framework and database used, the frontend project must explain the technology stack and compatibility requirements, and all quantitative indicators should be reasonable and conform to industry common sense." A key constraint is the "output format template for the completed information," for example, requiring the model to strictly adhere to the sentence structure "[action verb] based on [technology stack] [project type], achieving [quantitative indicator] [business value]." This ensures that different resumes, after completion, have structural consistency and comparability in their descriptions, thus facilitating subsequent automated evaluation.

[0053] In practical terms, when executing this step, the aforementioned entity can assemble all the elements into a complete prompt according to a predetermined logic. For example, it can serialize the "resume knowledge graph" into a descriptive text (such as listing the main triples), and then organize the "first resume information" into paragraph-based plain text. Next, it smoothly combines this data with phrases containing role settings, domain constraints, and operational descriptions into a coherent natural language prompt, seemingly directed at an expert. This prompt is then sent to the large model service via its integrated API. When interpreting this complex prompt, the large model first "positions itself," simulating the thought process of the designated expert during its internal reasoning. Then, it parses and internalizes the domain constraints and output templates, treating them as rules that must be followed when generating text. Ultimately, the text output by the large model possesses both professional depth and strictly adheres to the preset format and content requirements. For example, when the model encounters a knowledge graph entry that links "Python" and "data analysis" skills but only states "some data analysis experience" in the initial resume, the model, positioned as a "technical architect" and constrained by templates, will generate a standardized and concise statement such as, "Led a sales data trend analysis project using Python and Pandas libraries, built a predictive model, and improved monthly report generation efficiency by 40%." Through this series of designs, the general capabilities of the large model are transformed into a stable, reliable, and professional automated resume enhancement tool.

[0054] Step 204: Based on the third resume information, evaluate the resume file to be processed on multiple preset evaluation dimensions to obtain the resume evaluation results.

[0055] Building upon step 203, this step aims to have the aforementioned implementing entity conduct a final quantitative and comprehensive evaluation of the "structured third-party resume information" obtained after the in-depth processing and enhancement of the preceding steps, in order to obtain the resume evaluation results. In other words, it aims to transform the candidate's abilities and potential contained in a resume into a set of objective and comparable quantitative indicators, thereby providing a direct and reliable basis for recruitment decisions.

[0056] In executing this step, the "third-party resume information" relied upon by the aforementioned executing entity is a high-quality structured data that has undergone error correction, completion, and standardization. It contains clear, complete, and semantically unambiguous descriptions of the candidate's experience, skills, and achievements. The executing entity combines this information with a pre-designed "scoring prompt," forming an instruction input to the large model. This "scoring prompt" acts as the examiner's "scoring criteria" and "scoring rules," clearly defining the "multiple preset evaluation dimensions" that need to be assessed. Specifically, in some other implementations of this embodiment, these dimensions can be concretized as "professional fit," "project complexity," "clarity of expression," and "development potential," as well as other similar evaluation angles. Each dimension corresponds to an examination of different aspects of the candidate: professional fit focuses on the degree of alignment between skills and job requirements; project complexity measures the technical depth and scale of past experience; clarity of expression assesses the logical coherence and information presentation quality of the resume itself; and development potential infers future possibilities based on growth trajectory and background. The server clearly conveys the definitions and judgment criteria of these evaluation dimensions to the large model through the prompt.

[0057] At the technical level, the aforementioned implementing entity can input "third-party resume information" along with "scoring prompts" containing multi-dimensional scoring rules into a large model. Essentially, this constrains and guides the powerful semantic understanding and reasoning capabilities of the large model into a structured evaluation task. Upon receiving this composite instruction, the large model will follow the guidance of the prompts, sequentially conducting in-depth analysis of the "third-party resume information" for each preset evaluation dimension. For example, when evaluating "professional matching," the model will extract the skill list from the resume and perform semantic comparison and overlap calculations with the target job responsibilities that may be implicitly or explicitly given in the prompts. When evaluating "project complexity," the model will analyze elements such as the technology stack in the project description, the scale of the business problem solved, and the applicant's role within it. For each dimension, the large model will output a "raw score" based on its analysis. This score is typically a value within a predetermined range (e.g., 0-10 points), reflecting the model's quantitative judgment of the level of that dimension.

[0058] In practical terms, the aforementioned implementing entity can pre-configure "scoring prompt word" templates optimized for different job types within the system. When a resume needs to be evaluated, the server first retrieves the corresponding "third-party resume information," preparing it in structured text format (such as JSON or a specific formatted natural language summary). This information is then combined with the scoring prompt word templates to form a complete, task-specific natural language prompt, which is sent to the large model service via the application programming interface (API). After processing, the large model returns a text response containing the original scores for each dimension. The implementing entity parses this response to extract the numerical values ​​corresponding to each dimension. Further weighted calculations can then be performed. For example, the weight coefficients for each evaluation dimension (such as professional matching degree, project complexity, etc.) can be pre-stored or dynamically configured. The original scores for each dimension are then multiplied by their corresponding weight coefficients, and finally, all these weighted scores are summed to obtain a comprehensive "resume scoring result." This mathematical calculation process integrates the discrete judgments of the large model on each dimension into a unified, sortable total score. At the same time, it reflects the relative importance of different dimensions in the final decision through weight allocation, thus completing the final transformation from semantic information to actionable decision data.

[0059] One specific implementation method is as follows: inputting the third-party resume information and preset scoring prompts into the large model as prompt information, and obtaining the original scores output for each preset evaluation dimension; wherein, the scoring prompts include multiple preset evaluation dimensions, including: professional matching degree, project complexity, expression clarity and development potential; using the large model, the original scores of each preset evaluation dimension are weighted according to the corresponding preset weights, and the resume score result is obtained by the sum of the weighted scores.

[0060] The resume processing method based on a large model provided in this disclosure effectively overcomes the two major bottlenecks of inaccurate original information and semantic ambiguity by applying the large model twice in key stages (error correction and completion). Ultimately, it achieves simultaneous improvement in the accuracy, automation, and evaluation depth of resume screening, and also realizes automated and intelligent processing from the original resume file to the final evaluation result.

[0061] To deepen your understanding of how to use large models for error correction, please refer to [link / reference]. Figure 3 , Figure 3 A flowchart of a method for arbitrating conflict information provided in this disclosure embodiment aims to provide a specific implementation of the arbitration and determination operation for conflict information included in the target error correction operation, wherein process 300 includes the following steps:

[0062] Step 301: Use the arbitration decision module in the large model to identify multiple semantically similar candidate entity information from different page areas in the first resume information;

[0063] Initiating this operation first relies on the generation of the initial "first resume information," which includes not only text content but also page area tags for each text segment. Based on this, this step involves the aforementioned executing entity calling the "arbitration and judgment module" within its integrated, task-optimized large language model. This module scans and clusters semantically similar entity representations across the entire text. Specifically, the executing entity inputs all the "first resume information" into this module, which utilizes its built-in semantic understanding capabilities to calculate the semantic similarity between different entity names, identifying different textual representations referring to the same or highly related things as "multiple candidate entity information." For example, it identifies "XX Group," "XX Treasure," and "XXX GROUP" as a group of candidates. Technically, this relies on the common-sense knowledge about vocabulary, aliases, and company organizational structures acquired by the large model during training.

[0064] Step 302: Use the arbitration decision module to determine the page area weight based on the confidence level of the page area to which each candidate entity information belongs;

[0065] This step aims to guide the arbitration module into the weighting stage, whereby the aforementioned implementing entity assigns an initial "confidence level" to different page areas according to preset rules. In practice, this can involve storing a mapping table internally. For example, areas under the main headings "Work Experience" or "Education Background" might be assigned high confidence, while areas within the main text of "Self-Evaluation" or "Project Details" might be assigned relatively lower confidence. The arbitration module then assigns a corresponding "page area weight" based on which area each candidate entity's information was extracted from. This step reflects a thorough consideration of the information source's authority; information in the formal resume section is generally more reliable than supplementary descriptions.

[0066] Step 303: Use the arbitration decision module to determine the weight adjustment strategy based on whether the information of each candidate entity is associated with time information;

[0067] This step aims to further introduce temporal logic as a key corrective factor by the aforementioned implementing entity. Specifically, the arbitration module checks whether each candidate entity's information is associated with explicitly identified time information in the resume. For example, in the text "At XX Group (2019-2022)," the entity "XX Group" is explicitly associated with the time period "2019-2022." The implementing entity will then execute a "weight adjustment strategy": for candidate entities with such associated time information, their weight is increased, as this enhances their empirical validity as a specific piece of resume experience; conversely, for entity statements without explicit time information, their weight is decreased, as they may only be background mentions or general statements. This strategy is based on the fundamental logic that valid work or educational experience in a resume must occur within a specific time period, effectively distinguishing between "factual statements" and "background information."

[0068] Step 304: Using the arbitration judgment module, determine the comprehensive weight of each candidate entity information based on the page area weight and weight adjustment strategy to which each candidate entity information belongs;

[0069] After taking into account the weights of different page areas and adjusting for time relevance, the arbitration module will calculate a "comprehensive weight" for each candidate entity. This process can be a simple weighted sum or a more complex function calculation.

[0070] Step 305: Using the arbitration decision module, determine the target candidate entity information as the arbitration decision result of the conflict information based on the comprehensive weight of each candidate entity information.

[0071] Typically, the candidate entity with the highest weight will be selected as the arbitration result, which is the final standardized representation of the conflict information.

[0072] This embodiment provides a specific implementation method for "arbitration and judgment operation of conflicting information" through steps 301-305, aiming to solve the problem of textual differences or conflicts in different areas of a resume that point to the same real entity. For example, a resume may state "XX Group" in the "Work Experience" section and mention "XX Treasure" in the "Project Description" section. Although the two are different in text, they are highly related in semantics. The execution entity can automatically resolve such conflicts by executing a series of logical steps provided in this embodiment, simulating the reasoning process of a professional reviewer, and outputting the most credible and standard entity description, thereby ensuring the consistency and accuracy of subsequent information processing.

[0073] To further understand the other possible error correction operations, in addition to the above embodiments, please refer to [link to relevant documentation]. Figure 4 , Figure 4A flowchart of a method for compensating cross-regional information and aligning different representations of the same entity, provided for embodiments of this disclosure, is provided. The flowchart 400 includes the following steps:

[0074] Step 401: For the compensation operation of cross-regional information included in the target error correction operation, use the cross-regional compensation module in the large model to identify each entity in the first resume information;

[0075] Step 402: Use the cross-regional compensation module to determine the correlation between each identified entity and different page areas, and supplement the original first resume information of the target entity that should belong to the target page area with entity information based on the correlation.

[0076] In steps 401-402, the compensation operation for cross-regional information performed by the aforementioned executing entity is triggered based on a practical problem: key entity information in the resume (such as employer company name, university graduated from) may not appear completely in its most expected logical area (such as under the headings "Work Experience" or "Education Background"), but is scattered in other areas (such as "Self-Evaluation" or "Project Details"). To address this problem, the aforementioned executing entity invokes a specific functional unit in the large model, namely the "Cross-Regional Compensation Module." This module, based on the inherent powerful named entity recognition capability of the large model, performs a global scan of the entire "first resume information" (including all page areas and their text), identifying all entities such as names of people, organizations, locations, and skills. Identification is not the end point; the key is to determine the "relationship between each entity and different page areas."

[0077] At the technical level, the aforementioned execution entity can guide the large model to perform semantic and logical reasoning through designed prompts: analyzing whether the context (e.g., "XXX Bus") of a certain entity appearing in the "Self-Evaluation" section explicitly points to a specific work experience when it appears in the text. If such a strong association exists, even if the company name field in the target "Work Experience" area is empty, it will be determined that the entity "should belong" to the target area, and "entity information supplementation" will be performed, that is, the target entity "XXX Bus" will be filled back into the employer field of the corresponding work experience. This process achieves automatic aggregation and completion of information, significantly improving the completeness and accuracy of the data.

[0078] Step 403: For the alignment operation of different representations of the same entity contained in the target error correction operation, use the representation alignment module in the large model to perform hierarchical parsing of each sentence in the first resume information, and determine the different representations of the same entity at the same level based on the obtained hierarchical parsing results.

[0079] Step 404: Use the representation alignment module to unify different representations of the same entity into the target representation.

[0080] In steps 403-404, the "alignment operation for different representations of the same entity" performed by the aforementioned execution entity is used to solve another common problem: the same objective thing may be referred to in different text forms in different positions of the resume (such as the technology stack "K8S" and "Kubernetes", or the company name "XX Group" and "XX Treasure"). To eliminate this ambiguity, the aforementioned execution entity can call the "representation alignment module" in the large model. This module first performs hierarchical parsing on each sentence in the "first resume information".

[0081] At the technical level, this is typically achieved by guiding a large model to perform deep semantic understanding, similar to dependency parsing, to identify the grammatical modification relationships and subordinate structures between words in a sentence. For example, when parsing "proficient in K8S, familiar with the operation and maintenance of the Kubernetes containerization platform," the module can understand that "K8S" and "Kubernetes containerization platform" are in parallel object positions in the sentence, both serving as objects of "proficiency" and "familiarity," thus determining that they are "at the same level" and refer to the same core technology in the current context. Based on this hierarchical parsing result, the module can accurately "identify different expressions of the same entity at the same level." Finally, the aforementioned execution entity can guide the module to perform standardized mapping, that is, "unify different expressions of the same entity into the target expression." In practice, the aforementioned execution entity can pre-define or introduce an industry-standard terminology mapping rule through prompt words, such as specifying that "K8S" and "Kubernetes" are unified as "Kubernetes." Based on this rule, the module replaces all identified synonyms with unique standardized terms, thereby ensuring the consistency of terminology throughout the document and removing vocabulary-level obstacles for subsequent content-based accurate matching and evaluation.

[0082] The solution provided in this embodiment systematically repairs the dispersion and ambiguity defects in the original resume information caused by layout and expression habits by sequentially performing the above-mentioned cross-regional compensation and expression alignment operations. This not only enhances the information consistency and logical self-consistency within a single resume, but also makes the information between different resumes comparable, substantially improving the data quality that serves as the cornerstone of subsequent intelligent analysis and evaluation.

[0083] To deepen your understanding of how to use large models for information completion, please refer to [link to relevant documentation]. Figure 5 , Figure 5The flowchart of a method for quantizing and converting fuzzy representations provided in this embodiment of the present disclosure aims to provide a specific implementation of the quantization and conversion operation of fuzzy representations included in the target information completion operation. The process 500 includes the following steps:

[0084] Step 501: Use the quantization transformation module in the large model to identify unstructured fuzzy description information in the first resume information that does not appear in the resume knowledge graph;

[0085] The inputs for initiating this step include two key data sources: one is the already constructed "resume knowledge graph," which is a structured network organized in the form of "entity-relationship-entity" triples; the other is the "first resume information" as the original corpus, which is a collection of texts divided by page area.

[0086] Building upon this, the aforementioned executing entity invokes the "quantization transformation module" within the larger model. The first step of this module is to perform difference comparison and fuzzy identification, specifically, "identifying unstructured fuzzy descriptive information in the first resume information not appearing in the resume knowledge graph." Technically, the executing entity guides this module to compare the encoded structured entities and relationships in the knowledge graph with the original text segments in the "first resume information." Text fragments that cannot find corresponding nodes or edges in the knowledge graph and whose expression lacks specific verbs, technical terms, or quantification results are judged by the module as "fuzzy descriptive information." For example, the knowledge graph may already contain a record (candidate, proficient in, Python), but a project description might contain the sentence "did some data processing with Python." The phrase "did some data processing" lacks specific actions, tools, and results details, thus being identified as a target fuzzy description.

[0087] Step 502: Use the quantization conversion module to convert the fuzzy description information into a standardized description using strong verbs based on the STAR rule;

[0088] After identifying the target text, the aforementioned execution entity guides the quantitative conversion module to perform the core operation of "converting ambiguous descriptive information into standardized descriptions using strong verbs based on the STAR method." The STAR (Situation, Task, Action, and Result) method is a framework for structuring descriptions of behaviors and outcomes, with its core being the clarification of the situation, task, action, and result. In practice, the aforementioned execution entity can command the module to perform the following operations on the ambiguous description using the "action verb guidance pool" and "dimensional constraint operator" built into the prompts: First, infer the core action to be strengthened based on the context, replacing "done" with specific verbs such as "designed," "analyzed," or "built"; second, complete the specific technology stack or tools involved in the action, which may be obtained from other parts of the knowledge graph (such as skill nodes) or from other sentences in the same description; finally, infer or attach a reasonable and measurable business output or result for this action. For example, regarding the statement "I've done some data processing with Python," the module might combine the "Pandas" library and "sales reports" background mentioned elsewhere in the resume to transform it into "I used Python and the Pandas library to clean, summarize, and analyze trends in historical sales data, generating visualized weekly reports to support departmental decision-making." Essentially, this process guides the large model through a "semantic instantiation" reasoning process, moving from a generalized description to concrete examples.

[0089] Step 503: Use the quantification conversion module to supplement the resume knowledge graph with standardized descriptions.

[0090] After the transformation is complete, the aforementioned executing entity will enter the information fusion stage, namely, "supplementing the resume knowledge graph with standardized descriptions." At this point, the newly generated standardized descriptions are no longer vague text blocks, but clear semantic units containing explicit subjects, actions, objects, and outputs. The executing entity will instruct the quantitative transformation module or a dedicated graph updater to deconstruct and map this new unit onto the knowledge graph schema. For example, from the description "Used Python and Pandas libraries to clean, summarize, and analyze trends in historical sales data...", "Python," "Pandas," "data cleaning," and "trend analysis" can be extracted as skill or activity entities, establishing a "mastered" or "performed" relationship with the candidate entity. Simultaneously, "supported departmental decision-making" can be considered an attribute or output result of this relationship. In this way, vague text with low value density, originally outside the structured graph, is transformed into a series of new, high-quality knowledge nodes and relationship edges, organically integrated into the ever-expanding resume knowledge graph. This greatly enriches the semantic content and evaluation dimensions of the graph, providing a complete and high-quality data foundation for ultimately achieving accurate, multi-dimensional intelligent scoring.

[0091] This embodiment performs deep semantic transformation and integration on unstructured text existing in the "first resume information" that is not yet covered in the "resume knowledge graph" generated after the previous error correction process, through steps 501-503. Specifically, it targets the colloquial, vague, and lacking-specific-detail descriptions commonly found in resumes, transforming them into standardized, quantifiable, and evaluable expressions through the specialized functions of a large language model, and finally integrating them into a unified knowledge graph structure, thereby completing the final sublimation of candidate information from "raw mentions" to "structured knowledge".

[0092] For step 502, a specific implementation method can be as follows: use the quantization conversion module to determine the contextual supplementary information related to the fuzzy description information in the first resume information; use the quantization conversion module to complete the logical reasoning chain of the fuzzy description information according to the contextual supplementary information to obtain the supplemented fuzzy description information; use the quantization conversion module to convert the supplemented fuzzy description information into a standardized description using strong verbs according to the STAR rule.

[0093] In the initial state of this embodiment, the target "fuzzy description information" to be converted is clearly defined. The quantification conversion module first identifies contextual supplementary information related to the fuzzy description information in the first resume information. This first resume information, as a rich original corpus, contains all text identified from various sections of the resume. The quantification conversion module analyzes the semantics of the sentence or paragraph containing the fuzzy description information, and then searches for other text fragments related to it in terms of topic, entity, or technical field throughout the entire "first resume information." For example, for the fuzzy description "responsible for backend development," the module might retrieve "Java, Spring Cloud, MySQL" from the resume's skills list area and keywords such as "high concurrency" and "payment system" from other project descriptions. This retrieved information collectively constitutes valuable "contextual supplementary information," providing factual basis for subsequent reasoning.

[0094] After obtaining contextual clues, the aforementioned execution entity guides the module into the core reasoning stage. This involves completing the logical reasoning chain based on supplementary contextual information to obtain the supplemented fuzzy description. This operation aims to simulate the thought process of human experts, connecting fragmented clues into a logically complete narrative. In practice, the quantification module executes a series of internal reasoning steps based on the inherent causal and common-sense reasoning capabilities of the large model. For example, it associates "responsible for backend development" (fuzzy action) with "Java, Spring Cloud" (contextual technology) and "payment system" (contextual business), inferring that the candidate's likely specific responsibility is "developing the payment system backend using Java and Spring Cloud frameworks." Essentially, this process constructs a logical reasoning chain from "general role" to "specific technical action" and "clear business object," filling in the key missing links in the original fuzzy description and forming a more semantically rich and definite "supplemented fuzzy description."

[0095] Finally, a standardized output transformation is performed, which involves "converting the supplemented fuzzy description into a standardized description using strong verbs according to the STAR method." At this point, the "supplemented fuzzy description" already contains the necessary elements, but may still exist in declarative language. The STAR method provides a structured expression framework that requires the description to include context, task, action, and result. It can force modules to use a predefined "strong verb guide pool" (such as replacing "do" with "design," "refactor," or "improve") through prompts, and organize the language according to the template of "action + technology / tool ​​+ business goal + quantified result." For example, the aforementioned inference "developing the payment system backend using Java and Spring Cloud framework," combined with the business value that may be inferred from the context (such as "supporting peak transaction volumes"), can be transformed into: "Designed and refactored a payment backend system based on Java and Spring Cloud microservice architecture, successfully supporting a peak transaction volume of 100,000 transactions per second."

[0096] This embodiment focuses on solving a core and delicate operation in "information completion processing": how to systematically transform a highly generalized, vague descriptive statement (e.g., "system performance has been optimized") into a specific, professional, and evaluable standardized description. Through its integrated quantification conversion module (a functional unit within the large language model specifically designed for this task), a three-stage deep semantic reasoning and reconstruction process is executed to ensure the accuracy and practicality of the completed information. The final conversion outputs a standardized description that conforms to professional expression standards and contains quantifiable evaluation dimensions, thus achieving a qualitative leap from vague colloquial language to precise structured knowledge, significantly improving the feasibility and accuracy of subsequent intelligent evaluation.

[0097] Based on any of the above embodiments, to further understand how to achieve a comprehensive evaluation across multiple dimensions based on third-party resume information, please refer to [link to relevant documentation]. Figure 6 , Figure 6 A flowchart of a scoring method for different preset evaluation dimensions provided in this disclosure embodiment is included in process 600, comprising the following steps:

[0098] Step 601: For professional matching, use a large model to extract skill information from the third resume information, calculate the degree of overlap between the skill information and the job description information of the target position, and determine the original score of professional matching based on the degree of overlap.

[0099] Regarding "professional matching," the aforementioned implementing entity will guide the large model to perform an analysis focused on skill alignment. This entity will provide the large model with "third-party resume information" (which includes completed and standardized skill descriptions) along with "job description information for the target position." Technically, the large model acts as a semantic matching engine. It first systematically "extracts skill information" from the "third-party resume information," typically a structured skill list or technical keywords scattered throughout the job description. Then, the large model semantically "calculates the degree of overlap," meaning it understands the core competencies required for each responsibility in the job description and compares their semantic similarity with the resume skills. It considers the breadth, depth, and relevance of skills to the core job requirements. For example, for a "backend development engineer" position, the overlap weight of mastering "Java" and "microservice architecture" is far higher than mastering only "Python scripting." Finally, based on the large model's overall assessment of the matching depth, it maps to a predefined scoring range (e.g., 0-10 points), thus "determining the original score for professional matching."

[0100] Step 602: For project complexity, use a large model to extract project experience information from the third resume information, including project scale description, technical depth description and historical role description, and determine the original score of project complexity based on the project experience information.

[0101] Regarding "project complexity," the aforementioned implementing entities shifted their evaluation focus to the technical content and scale of the candidate's past experience. They instructed the large-scale model to specifically "extract project experience information" from the "third-party resume information" and guided the model to examine it from three sub-dimensions: "project scale description" (e.g., number of users involved, data volume, team size), "technical depth description" (e.g., the sophistication of the technologies used, the technical difficulties solved, and the complexity of the architectural design), and "historical role description" (e.g., whether the candidate was a project leader, core developer, or participant). In practice, the large-scale model analyzes each project description, identifies quantitative indicators and technical terms, and infers the project's complexity level based on built-in common sense or domain knowledge. For example, a project that "led the design of a distributed system supporting billions of users" would receive high ratings in all three sub-dimensions. The server-wide integrated model then transforms the analysis results of these sub-dimensions into a raw score representing the overall "project complexity."

[0102] Step 603: Regarding clarity of expression, use a large model to determine the degree of difference between the third resume information and the first resume information, and determine the structural rationality of the resume content based on the structural information of the first resume information, and determine the original score of clarity of expression based on the degree of difference and structural rationality.

[0103] Regarding "clarity of expression," the evaluation criteria of the aforementioned implementing entities uniquely include the writing and organization quality of the resume document itself. This dimension is achieved through two calculable indicators: "the degree of difference between the third resume information and the first resume information" and "the structural rationality of the first resume information." The former assesses the extent of information completion and transformation. The principle is that a resume with clear original expression requires less completion, resulting in a low degree of difference between the two; conversely, if the original resume (first resume information) is vague and disorganized, the "third resume information" formed after system transformation will differ significantly from it. The server calculates this semantic or structural difference using a large model. The latter directly evaluates the logical structure of the "first resume information," such as whether the chapters are complete, the timeline is clear, and the descriptions are orderly. The server combines these two indicators, and the large model judges the resume candidate's logical thinking and communication skills, outputting a "raw score for clarity of expression."

[0104] Step 604: For development potential, use a large model to extract growth information from the third resume information, including educational background, historical job change trajectory and cross-domain experience, and determine the original score of development potential based on the growth information.

[0105] Regarding "development potential," the aforementioned implementing entities conduct predictive analysis based on history and trends. The server guides the large model to extract key elements constituting "growth information" from "third-party resume information," primarily including: "educational background" (e.g., school, major), "historical job change trajectory" (e.g., promotion speed, company platform changes), and "cross-domain experience" (e.g., whether the candidate possesses multi-skilled abilities or cross-disciplinary project experience). By analyzing these elements, the large model constructs a comprehensive profile of the candidate's growth potential, adaptability, and learning ability. For example, a rapid promotion trajectory from engineer to technical lead, or a background possessing both technical and product experience, typically indicates high potential. Based on the large model's assessment of this profile, the server ultimately "determines the original score for development potential."

[0106] Through the independent and in-depth scoring process across the four dimensions mentioned above, the implementing entities generated multi-faceted and interpretable raw evaluation data for each resume, providing a solid basis for subsequent weighted summarization and final decision-making.

[0107] This embodiment performs detailed quantitative analysis on the processed "third-party resume information" across multiple preset evaluation dimensions through steps 601-604, thereby generating an original score for each dimension. This process is a key step in intelligent evaluation. The aforementioned execution entity does not make a holistic, fuzzy judgment on the resume, but rather uses the deep semantic analysis capabilities of a large language model to decompose the comprehensive evaluation into four independently calculable and logically clear sub-tasks: "professional matching degree," "project complexity," "clarity of expression," and "development potential," and performs standardized scoring operations on each sub-task.

[0108] Based on the previous embodiment, when it is found that the original score of any preset evaluation dimension is lower than the score lower limit preset for the corresponding preset evaluation dimension, the resume score result of the corresponding resume file to be processed is attached with a target identifier, and then the resume file to be processed with the target identifier attached is sent to the target storage location according to the preset sending path.

[0109] In this embodiment, the initial input consists of the pre-calculated raw scores for each dimension, and a pre-configured "score lower limit" for each dimension. This "score lower limit" is a configurable threshold parameter representing the minimum acceptable level of ability for that dimension. For example, for a technical R&D position, the score lower limit for "professional matching" might be set to 7 points, while the lower limit for "clarity of expression" might be set to 5 points. These lower limits are typically pre-defined by domain experts based on the core requirements of the position and stored in the server's configuration system. The core operation performed by the aforementioned execution entity is a logical judgment: sequentially traversing each preset evaluation dimension and checking whether its corresponding raw score is strictly "lower" than the score lower limit set for that dimension. This judgment process is deterministic; once any dimension is found to be unsatisfactory, the subsequent processing flow is triggered. When the triggering condition is met—that is, when any dimension's score is lower than its lower limit—the aforementioned execution entity will perform a marking operation on the "resume scoring result" of this "resume file to be processed." "Adding a target identifier" is the specific manifestation of this operation. In terms of technical implementation, this identifier can be a status label written to a specific field in the database (such as status = 'rejected_by_threshold'), or it can be a specific attribute marker attached to the scoring result data structure (such as flag: {critical_fail: true, failed_dimension: 'professional matching degree'}). The core function of this identifier is that in any subsequent processing or display stage, the system or other interfaces can quickly identify that this resume was rejected because it violated a hard rule, thus eliminating the need to consider its total score or other dimensions of performance.

[0110] After the identification is attached, the aforementioned execution entity then performs automated sorting. The "preset sending path" and "target storage location" are predefined processing rules by the system. In practice, this might mean that the server moves or copies all relevant data of the resume (original file, parsed information, scoring results, and rejection identifier) ​​from a database partition or message queue designated for "pending processing" or "pending review" to a dedicated storage area designated for "rejected" or "talent pool reserves." Another implementation involves calling an internal or external message notification interface to generate a log or notification containing the resume ID and rejection reason, and sending it to a designated monitoring system or administrator interface.

[0111] The automated "judgment-marking-routing" process provided in this embodiment ensures that unqualified resumes are removed from the main processing pipeline in a timely and accurate manner, greatly improving the efficiency of batch processing and the consistency of decision-making, while retaining complete audit trail records. In other words, the core logic of the solution provided in this embodiment lies in implementing a "one-vote veto" mechanism based on hard thresholds. When a candidate has a significant weakness in a core competency dimension, regardless of their performance in other aspects, they are usually directly excluded from further consideration.

[0112] Building upon any of the above embodiments, to address the rigidity of fixed, absolute scoring thresholds when faced with natural fluctuations in resume quality across different recruitment batches, a dynamic watermark based on the relative performance of resumes within the same batch can be introduced. This allows the evaluation system to intelligently adapt to the actual competitive landscape of each batch, thereby making more reasonable and fairer screening decisions. One specific implementation could be: setting the upper limit of the scores for multiple resume files in the same processing batch under the same preset evaluation dimension as the maximum value of the original scores and the lower limit as the minimum value of the original scores.

[0113] The premise for the aforementioned executing entity to perform this operation is that it has already completed the preliminary evaluation of all resumes within the same processing batch. A "processing batch" refers to a collection of resumes received and processed by the executing entity within a specific time period (such as a recruitment cycle or a working day) or for a specific position. The executing entity has already calculated the original score for each resume in this batch on the same preset evaluation dimensions, such as "professional matching" and "project complexity." At this point, the executing entity does not immediately compare these original scores with a preset absolute score (such as 7.5 points). Instead, it first performs statistical analysis within the batch: for each preset evaluation dimension, the server iterates through all resumes in the batch to find the maximum and minimum scores for that dimension. According to the rules of this step, this maximum value is defined as the upper limit of the score for that dimension within the batch, and the minimum value is defined as the lower limit. Technically, this is essentially performing extreme value statistics before data normalization or standardization. Its purpose is to capture the actual distribution range of the candidates' performance in that dimension in the current batch, establishing a benchmark for subsequent relative evaluation.

[0114] In practice, the aforementioned execution entity maintains a set of resume ratings for the current batch in memory or cache. When a batch of resumes is processed, a post-processing thread is started to perform a fast aggregation query for each evaluation dimension. For example, for the "project complexity" dimension, it might find that the original ratings of the 100 resumes in this batch are distributed between 4.5 and 9.2. The execution entity would then record 9.2 as the upper limit for the "project complexity" rating of this batch and 4.5 as the lower limit. This dynamically determined upper and lower limit replaces the fixed threshold and becomes the new basis for subsequent "veto" or "selection based on merit" logic. For example, the system can be configured to "eliminate resumes with a certain dimension rating in the lowest 10% of the batch" or "select resumes with a certain dimension rating in the highest 20% of the batch". In this way, the evaluation criteria can "rise with the tide" or "adapt to market conditions" as the overall level of the batch increases. For example, in a group of resumes filled with elites, an 8 might only be considered average (close to the lower limit), while in a group of resumes with relatively little experience, a 7 might be considered top-tier (close to the upper limit).

[0115] This mechanism greatly enhances the system's adaptability and fairness in dealing with markets of varying talent density, ensuring that the screening results are always based on the relative competitiveness of the current candidate pool, rather than absolute scores detached from context. This simulates how experienced recruiters, after reviewing a large number of resumes, naturally form intuitive judgments about the "average level" and "excellent threshold" of the candidates, and then transform these intuitions into automated, quantifiable decision-making rules.

[0116] Based on any of the above embodiments, corresponding recommended interview questions can also be generated based on the resume evaluation results. That is, when the above-mentioned executing entity attempts to perform this operation, the input is the already calculated comprehensive score and detailed analysis data of each dimension. The executing entity uses these structured evaluation conclusions to drive its integrated large language model to automatically generate a series of targeted interview questions that match the candidate's specific profile. This extends the output of the screening system from "judgmental scores" to "operational tools," assisting interviewers in conducting more efficient and in-depth follow-up assessments.

[0117] From a technical perspective, this step relies on the powerful conditional text generation and logical reasoning capabilities of the large language model. The aforementioned execution entity uses the "resume evaluation results" (including scores for each dimension and a summary of possible key evidence) as the core context, combined with a specific question generation prompt, and inputs them into the large model. This prompt instructs the large model to act as a "senior interviewer" and generate questions based on the following logic: First, it delves deeper into strengths; for example, if the "project complexity" score is high, it instructs the model to generate questions that explore the project's technical details, architectural decisions, and challenge-overcoming processes. Second, it verifies and explores weaknesses; for example, if the "clarity of expression" score is low, it instructs the model to generate questions that require the candidate to clearly retell or summarize a particular experience to assess their on-site communication skills. Finally, it designs forward-looking questions based on potential; for example, for a high score in "development potential," it can generate questions about career planning and the willingness to learn new fields. In this way, the generated interview questions and resume evaluation conclusions form a tight logical loop.

[0118] In practical terms, after completing the evaluation, the aforementioned execution entity can initiate a post-processing thread. This thread serializes the evaluation results into a structured natural language description, such as: "Candidate Zhang San, overall score 85. Strengths: Professional matching (9 / 10), skill stack highly matches the target position; Project complexity (8 / 10), experience leading a medium-sized distributed system. Dimensions to be examined: Clarity of expression (6 / 10), the original resume description is relatively brief." The execution entity concatenates this description with a preset "interview question generation prompt" template to form a complete instruction and sends it to the large model. Upon receiving the instruction, the large model, based on its built-in interview knowledge and understanding of the context, outputs a set of specific questions, such as: "1. Please elaborate on the biggest technical challenge you encountered in the XX distributed system and the solution you led. 2. Your resume mentions proficiency in Spring Cloud. Please compare the similarities and differences between Spring Cloud and Dubbo in microservice governance, and explain the reasons for your choice based on your projects. 3. Please give me a clear introduction to the project for which you made the greatest contribution in three minutes."

[0119] Furthermore, the aforementioned implementing entity can also attempt to enhance the explanatory power of the evaluation results, guiding the large model to provide "metacognitive" output regarding the evaluation process itself. This means not only providing scores but also generating evidence chains and explanatory text supporting those scores. In practical implementation, the implementing entity can add additional instructions to the scoring prompts, requiring the large model to include a brief "scoring basis" alongside the raw score for each dimension. This basis must directly cite or summarize specific content from the "third-party resume information." For example, a score of 9 for "professional matching" might be explained as: "The candidate's resume clearly lists all core technology stacks required for the position, including Java, Kafka, and Redis, and provides application descriptions in two major projects." The implementing entity can then package this explanatory text along with the scoring results and recommended interview questions into the final evaluation report. This enhancement significantly improves the transparency, credibility, and acceptability of the scoring results, enabling recruiters to clearly understand the system's scoring logic rather than facing an untraceable "black box" score, thereby enhancing the practical value and user trust of the entire intelligent screening system.

[0120] To enhance understanding, this disclosure also provides a specific implementation scheme in light of a particular deficiency in existing technology:

[0121] Suppose that Company A has designed an intelligent resume screening system for its publicly advertised "Senior Backend Development Engineer" position. This system is deployed on the company's internal server, and its core workflow consists of the following three stages:

[0122] The first stage is multimodal resume parsing. The system receives raw resume files in various formats and uses a layout-aware OCR engine that integrates Visual Transformer (ViT) and Differentiable Binarization Network (DBNet) to not only recognize the text but also understand its layout structure and hierarchical relationships, outputting preliminary structured text information with region labels.

[0123] The second stage is structured information fusion and enhancement. The system performs in-depth error correction, completion and standardization on the initially extracted information, and constructs a unified resume knowledge graph.

[0124] The third stage is large-scale intelligent scoring and decision-making. Based on the completed high-quality structured information, the system performs multi-dimensional quantitative scoring under the drive of the large model, and automatically generates screening conclusions and interview suggestions.

[0125] The following is a detailed explanation of the system's workflow through a specific resume processing case.

[0126] Phase 1: Multimodal Resume Analysis and Preliminary Structuring

[0127] The system received a PDF resume submitted by applicant "Zhang San". The resume had a rather complex layout, with the "Work Experience" and "Project Experience" sections interleaved.

[0128] Layout-aware OCR extraction: The server first converts the PDF into a high-resolution image, then calls the ViT model for global layout analysis. The ViT model divides the image into multiple small blocks and accurately identifies several semantic regions and their boundaries, such as "personal information," "educational background," "work experience," "project experience," and "skills," through a self-attention mechanism. Next, for each identified region, the server calls the DBNet model for refined text detection and recognition. DBNet, through its differentiable binarization operation, accurately segments the text lines within the region, correctly handling even slight tilts or curves in the text. Finally, the system outputs a set of texts with region affiliation labels, namely "first resume information." For example, it clearly knows that the text "2019.07 - 2022.08" and "responsible for payment system backend development" belong to the same "work experience" entry region.

[0129] Post-processing and conflict arbitration: The system detected conflicting employer information in the resume, with one entry mentioning "XX Group" and another in the project description stating "responsible for… at XXbao…". The system then initiated a weighted voting-based conflict resolution logic: First, it assigned credibility weights to different areas (e.g., the "Work Experience" area had a higher weight than the "Project Description" area); second, it checked whether the entities were associated with credible timestamps ("XX Group" was associated with a clear employment period of 2019-2022); finally, it invoked a large language model as the arbitrator, combining common business knowledge ("XXbao" is a business under "XX Group") to normalize the entity, determining the standardized entity to be "XX Group". Simultaneously, the system performed a cross-regional scan using named entity recognition, discovering that the "Self-Evaluation" section mentioned "accumulated experience in large-scale system architecture at XXX Group," while a certain company information was missing from the "Work Experience" section. The system automatically filled this information back into the corresponding experience field through semantic association. After this step, the system obtained a "second resume" that corrected the contradictions and filled in the missing information.

[0130] Phase Two: Structured Information Fusion and Semantic Enhancement

[0131] The system converts the "second resume information" into a unified "resume knowledge graph" schema. For example, it generates triples: (Zhang San, employed at, XX Group), (Zhang San, proficient in, Java), (XX Group, time period, 2019-2022). However, the graph can only cover structured, explicit information and cannot effectively represent vague descriptions in the original resume, such as "handled a recommendation system".

[0132] Information completion guided by the system's prompts: The system initiates the information completion process. It inputs the constructed "resume knowledge graph" (as the first input), the original "first resume information" text (as the second input, providing contextual details), and a carefully designed "completion prompt" into the large language model. The prompt includes: (a) role setting: "You are a senior technical recruitment expert"; (b) action verb guidance: requiring the use of strong verbs such as "design," "development," and "optimization"; (c) output template constraints: must include "[action verb] based on [technology stack] [project], realizing [quantitative indicators] [business value]"; (d) domain knowledge: associating the skills "Spark" and "collaborative filtering" already identified in the resume.

[0133] Semantic Transformation and Graph Update: Guided by prompts, the large model performs logical reasoning chain completion on the fuzzy description "I've got a recommendation system working": identifying the subject (Zhang San), the object (recommendation system), the associated technology (Spark, collaborative filtering), inferring reasonable actions (R&D), and speculating on business value (personalized distribution). The final output is a standardized description: "Based on the Spark platform, I developed a collaborative filtering recommendation system, achieving accurate personalized content distribution and a 15% increase in click-through rate." The system then deconstructs this new description, treating "Spark," "collaborative filtering," and "recommendation system" as new entities or attributes, updating and enriching the original "resume knowledge graph," forming a complete and standardized "third resume information."

[0134] Phase 3: Large-scale intelligent scoring and automated decision-making

[0135] The system is now evaluating candidate Zhang San based on high-quality "third-party resume information". The target position is "Senior Backend Development Engineer", and the job description (JD) emphasizes experience in high-concurrency and distributed systems.

[0136] Multi-dimensional quantitative scoring: The system inputs "third-party resume information" and a set of "scoring prompts" into the large model. The scoring prompts define four core dimensions and their weights (e.g., professional matching degree 0.4, project complexity 0.3, clarity of expression 0.2, development potential 0.1), and provide detailed scoring criteria for each dimension (e.g., a 0-10 scale and corresponding descriptions). The large model acts as the evaluator, analyzing each dimension: for "professional matching degree," it compares Zhang San's skill stack (Java, SpringCloud, Redis, etc.) with the job requirements and gives a high score; for "project complexity," it analyzes the scale and technical depth of his "payment system" project and gives a relatively high score; for "clarity of expression," it compares the information differences before and after completion to evaluate the structural logic of the original resume; for "development potential," it analyzes his education and career advancement trajectory. Assume the original output score is [9, 8, 7, 8].

[0137] Comprehensive Calculation and Dynamic Threshold Determination: The server performs a weighted calculation based on preset weights to obtain a comprehensive score. Simultaneously, the system employs a "dynamic watermark" strategy: Having processed 100 resumes for the same position this week, the system automatically calculates the distribution of scores across all dimensions for this batch. Assuming the lowest possible score (lower limit) for "professional matching" is 6 points and the highest possible score (upper limit) is 9.5 points, Zhang San scores 9 points in this dimension, placing him among the top performers in the batch. Based on the urgency of the recruitment (e.g., "urgent"), the system sets the acceptance threshold at a "professional matching" score greater than the 85th percentile of the batch (assumed to be 8.5 points). Zhang San's score of 9 points exceeds this threshold.

[0138] Decision Generation and Derivative Output: Since Zhang San's scores across all dimensions meet the requirements and his overall score is high, the system marks him as "Recommended for Interview." Furthermore, the system automatically generates targeted recommended interview questions, such as: "Please elaborate on your design scheme for handling high-concurrency transactions in the XX Group payment system and the challenges you encountered." This question directly stems from the specific content of his high-scoring dimension, "Project Complexity." Simultaneously, the system can attach explanatory notes to the scoring results, such as "Professional Matching 9 points: The candidate's core skills in Java, Spring Cloud, Redis, etc., perfectly match the JD requirements and have been applied in relevant large-scale projects."

[0139] As demonstrated by the above embodiments, this system effectively overcomes information extraction errors from complex PDFs / scanned documents through ViT+DBNet combined layout-aware OCR; it transforms ambiguous expressions into standardized semantics that can be quantified and evaluated by integrating knowledge graphs and prompting engineering into a large-scale model; and it achieves objective, interpretable, and batch-adaptable intelligent evaluation through large-scale model scoring based on multi-dimensional prompt word design and dynamic thresholds. Ultimately, the system can automatically output a clearly structured and accurate competency assessment report and interview action guide from a poorly formatted and ambiguous original resume, significantly improving the efficiency, accuracy, and fairness of the recruitment process.

[0140] Building upon any of the above embodiments, in an embodiment designed to optimize system performance, the executing entity not only performs single inference iterations but also deploys a multi-task joint training framework for continuous learning. Specifically, the executing entity prepares a batch of training data during initialization or periodic updates. Each data point includes: the original resume file, manually verified and corrected entity information (as annotations for the information extraction task), and comprehensive scores and interview decisions from senior HR personnel (as annotations for the scoring task). The server employs a shared underlying representation neural network architecture, where the OCR post-processing module (responsible for entity recognition and normalization) and the large model scoring module (responsible for dimensional scoring) share intermediate layer encodings of resume semantics. During training, the executing entity calculates a composite loss function: this function simultaneously includes the cross-entropy loss for information extraction (measuring the accuracy of entity recognition) and the mean squared error loss for scoring prediction (measuring how close the score is to the manually labeled data). By simultaneously optimizing these two tasks through backpropagation, the shared encoding layer learns feature representations that are valuable for resume screening, thereby significantly improving the accuracy and generalization ability of both modules and achieving a synergistic optimization effect where more accurate error correction leads to more precise scoring.

[0141] In another embodiment aimed at improving the model's discriminative power, the aforementioned execution entity can also utilize a contrastive learning strategy for self-supervised or weakly supervised training. The execution entity can first construct specific sample pairs: two resumes of the same candidate, one before and one after polishing for the same position, as positive sample pairs; and a resume submitted for "Java Backend Engineer" and another for "Marketing Specialist" as negative sample pairs. During model training, the execution entity inputs the "third resume information" from the resumes into an encoder, converting it into a feature vector. The design goal of the contrastive learning loss function is to reduce the distance between the feature vectors of positive sample pairs while increasing the distance between the feature vectors of negative sample pairs. Specifically, this loss function (such as the InfoNCE loss) is integrated into the aforementioned composite loss function as an important component. Furthermore, the composite loss function explicitly includes: a scoring regression loss (calculating the mean squared error between the model's predicted score and the human score) and a classification prediction loss (calculating the cross-entropy between the model's "suggested interview" probability and the actual interview result). The execution entity can combine these three loss terms through a weighted summation and use this total loss to guide the update of the model parameters. This mechanism enables the system to learn from a large amount of resume data that does not explicitly label specific scores but whose job affiliation and relative quality are known, and to continuously evolve its judgment criteria.

[0142] Further reference Figure 7 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a resume processing device based on a large model, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0143] like Figure 7As shown, the resume processing device 700 based on a large model in this embodiment may include: a first resume information recognition unit 701, an error correction processing and knowledge graph generation unit 702, an information completion unit 703, and a multi-dimensional evaluation unit 704. The system includes the following components: a first resume information recognition unit 701, configured to identify structured first resume information from the resume file to be processed; an error correction and knowledge graph generation unit 702, configured to input the first resume information and preset error correction prompts as prompts into a preset large model for error correction processing, and generate a resume knowledge graph from the output structured second resume information in the form of a knowledge graph; wherein the error correction processing includes arbitration and judgment operations for conflicting information; an information completion unit 703, configured to input the resume knowledge graph, the first resume information, and preset information completion prompts as prompts into a large model for information completion processing, obtaining structured third resume information; wherein the information completion processing includes quantification and conversion operations for fuzzy expressions; and a multi-dimensional evaluation unit 704, configured to evaluate the resume file to be processed on multiple preset evaluation dimensions based on the third resume information, and obtain a resume evaluation result.

[0144] In this embodiment, the specific processing and technical effects of the first resume information recognition unit 701, the error correction and knowledge graph generation unit 702, the information completion unit 703, and the multi-dimensional evaluation unit 704 in the large model-based resume processing device 700 can be referred to respectively. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.

[0145] In some optional implementations of this embodiment, the first resume information recognition unit 701 is further configured to:

[0146] Use a visual Transformer model to identify different layout areas in the resume file to be processed;

[0147] Differentiable binarization networks are used to determine the text within different page layout regions;

[0148] Based on the hierarchical relationship between the text in each region and its corresponding page area, generate structured first resume information.

[0149] In some other optional implementations of this embodiment, the error correction processing and knowledge graph generation unit 702 includes an error correction processing subunit configured to input the first resume information and preset error correction prompts as prompt information into a preset large model for error correction processing. The error correction processing subunit includes:

[0150] The target error correction operation determination module is configured to use a large model to determine the target error correction operation to be performed based on the error correction operation description information contained in the error correction prompt words.

[0151] The error correction execution module is configured to perform target error correction operations on the first resume information using a large model.

[0152] In some other optional implementations of this embodiment, the error correction processing subunit may further include:

[0153] The first role positioning control module is configured to respond to the error correction prompt, which also includes first role setting information and first domain knowledge constraint information, and control the large model to stand in the first role corresponding to the first role setting information.

[0154] The error correction constraint information determination module is configured to control the large model to determine the error correction constraint information that should be followed when performing the target error correction operation based on the first domain knowledge constraint information;

[0155] Correspondingly, the error correction execution module is further configured as follows:

[0156] Control the large model to take the first role, and perform the target error correction operation on the first resume information under the constraint of error correction constraint information.

[0157] In some other optional implementations of this embodiment, the error correction operation execution module includes:

[0158] The arbitration decision execution submodule is configured to handle the arbitration decision of conflicting information included in the target error correction operation. It utilizes the arbitration decision module in the large model to identify multiple semantically similar candidate entity information identified from different page regions in the first resume information. The arbitration decision module then determines the page region weight based on the confidence level of each candidate entity information's respective page region. It further determines a weight adjustment strategy based on whether each candidate entity information is associated with time information; this strategy includes increasing the weight of candidate entity information associated with time information and decreasing the weight of candidate entity information not associated with time information. Finally, the arbitration decision module determines the comprehensive weight of each candidate entity information based on its page region weight and the weight adjustment strategy. Finally, based on the comprehensive weight of each candidate entity information, the arbitration decision module determines the target candidate entity information as the arbitration result for conflicting information.

[0159] In some other optional implementations of this embodiment, the error correction process further includes a compensation operation for cross-regional information, and the error correction operation execution module includes:

[0160] The compensation operation execution submodule is configured to perform compensation operations on cross-regional information included in the target error correction operation. It uses the cross-regional compensation module in the large model to identify each entity in the first resume information; it uses the cross-regional compensation module to determine the correlation between each identified entity and different page areas, and supplements the original part of the first resume information in the target page area with entity information of the target entities that should belong to the target page area.

[0161] In some other optional implementations of this embodiment, the error correction process further includes an alignment operation for different representations of the same entity, and the error correction operation execution module includes:

[0162] The alignment execution submodule is configured to perform alignment operations on different representations of the same entity contained in the target error correction operation. It uses the representation alignment module in the large model to perform hierarchical parsing on each sentence in the first resume information and determines different representations of the same entity at the same level based on the obtained hierarchical parsing results. The representation alignment module then unifies the different representations of the same entity into the target representation.

[0163] In some other optional implementations of this embodiment, the information completion unit 703 includes:

[0164] The target information completion operation determination subunit is configured to use the large model to determine the target information completion operation to be performed based on the information completion operation description information contained in the information completion prompt words.

[0165] The target information completion operation execution subunit is configured to perform target information completion operations on the resume knowledge graph and the first resume information using the large model.

[0166] In some other optional implementations of this embodiment, the information completion unit 703 may further include:

[0167] The second role positioning control subunit is configured to respond to the information completion prompts, which also include second role setting information and second domain knowledge constraint information, and control the large model to position itself in the second role corresponding to the second role setting information.

[0168] The information completion constraint information determination subunit is configured to control the large model to determine the information completion constraint information to be followed when performing the target information completion operation based on the second domain knowledge constraint information; wherein, the information completion constraint information includes: the output format template of the completed information;

[0169] Correspondingly, the target information completion operation execution subunit is further configured as follows:

[0170] The control large model station plays a second role and performs target information completion operations on the resume knowledge graph and the first resume information under the constraints of information completion constraints.

[0171] In some other optional implementations of this embodiment, the target information completion operation execution subunit includes:

[0172] The fuzzy description information determination module is configured to perform a quantization transformation operation on fuzzy expressions included in the target information completion operation. It uses the quantization transformation module in the large model to determine unstructured fuzzy description information in the first resume information that does not appear in the resume knowledge graph.

[0173] The standardized description conversion module is configured to use the quantization conversion module to convert fuzzy description information into standardized descriptions using strong verbs based on the STAR rule;

[0174] The supplementary module is configured to use the quantization transformation module to supplement the resume knowledge graph with standardized descriptions.

[0175] In some other optional implementations of this embodiment, the standardized description conversion module is further configured to:

[0176] The quantification conversion module is used to determine contextual supplementary information related to the fuzzy description information in the first resume information;

[0177] The quantization conversion module is used to complete the logical reasoning chain of the fuzzy description information based on the contextual supplementary information, so as to obtain the supplemented fuzzy description information.

[0178] The quantization conversion module uses the STAR method to convert the supplemented fuzzy description information into a standardized description using strong verbs.

[0179] In some other optional implementations of this embodiment, the multi-dimensional evaluation unit 704 includes:

[0180] The original score output sub-units for each evaluation dimension are configured to input the third resume information and preset score prompts as prompts into the large model to obtain the original scores output for each preset evaluation dimension. The score prompts include multiple preset evaluation dimensions, including: professional matching degree, project complexity, clarity of expression and development potential.

[0181] The weighted processing unit is configured to use a large model to weight the original scores of each preset evaluation dimension according to the corresponding preset weights, and obtain the resume score result as the sum of the weighted scores.

[0182] In some other optional implementations of this embodiment, the original score output subunits of each evaluation dimension are further configured as follows:

[0183] Regarding professional matching, a large model is used to extract skill information from the third resume information and calculate the degree of overlap with the job description information of the target position, and the original score of professional matching is determined based on the degree of overlap.

[0184] To assess project complexity, a large model is used to extract project experience information from third-party resumes, including descriptions of project scale, technical depth, and historical responsibilities and roles. Based on this project experience information, an initial score for project complexity is determined.

[0185] Regarding clarity of expression, a large model is used to determine the degree of difference between the third resume information and the first resume information, and the structural rationality of the resume content is determined based on the structural information of the first resume information. The original score of clarity of expression is determined based on the degree of difference and structural rationality.

[0186] To assess development potential, a large model is used to extract growth information from third-party resumes, including educational background, historical job change trajectory, and cross-disciplinary experience, and an original score for development potential is determined based on this growth information.

[0187] In some other optional implementations of this embodiment, the resume processing device 700 based on a large model may further include:

[0188] The identifier appending unit is configured to append a target identifier to the resume scoring result of the corresponding resume file in response to the existence of an original score for any preset evaluation dimension that is lower than the score lower limit preset for the corresponding preset evaluation dimension.

[0189] The filtering and sending unit is configured to send the resume file to be processed, which is attached with the target identifier, to the target storage location according to the preset sending path.

[0190] In some other optional implementations of this embodiment, the upper limit of the scores of multiple resume files to be processed in the same batch under the same preset evaluation dimension is the maximum value of the original scores, and the lower limit of the scores is the minimum value of the original scores.

[0191] In some other optional implementations of this embodiment, the resume processing device 700 based on a large model may further include:

[0192] The recommended interview question generation unit is configured to generate corresponding recommended interview questions based on the resume evaluation results.

[0193] This embodiment exists as a device embodiment corresponding to the above method embodiment. The resume processing device based on the large model provided in this embodiment effectively overcomes the two major bottlenecks of inaccurate original information and semantic ambiguity by applying the large model twice in key stages (error correction and completion). Ultimately, it achieves simultaneous improvement in the accuracy, automation level and evaluation depth of resume screening, and also realizes automated and intelligent processing from the original resume file to the final evaluation result.

[0194] According to embodiments of this disclosure, this disclosure also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the resume processing method based on a large model as described in any of the above embodiments.

[0195] According to embodiments of this disclosure, this disclosure also provides a readable storage medium storing computer instructions that enable a computer to implement the large-model-based resume processing method described in any of the above embodiments when executed.

[0196] According to embodiments of this disclosure, this disclosure also provides a computer program product that, when executed by a processor, can implement the resume processing method based on a large model as described in any of the above embodiments.

[0197] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0198] like Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0199] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0200] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a large-model-based resume processing method. For example, in some embodiments, the large-model-based resume processing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the large-model-based resume processing method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform a resume processing method based on a large model by any other suitable means (e.g., by means of firmware).

[0201] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0202] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0203] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0204] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0205] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0206] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0207] According to the technical solution of this disclosure, by applying the large model twice in key stages (error correction and completion), the two major bottlenecks of inaccurate original information and semantic ambiguity are effectively overcome. Ultimately, the accuracy, automation and evaluation depth of resume screening are improved simultaneously, and the automated and intelligent processing from the original resume file to the final evaluation result is realized.

[0208] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0209] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A resume processing method based on a large model, characterized in that, include: Identify the structured first resume information from the resume files to be processed; The first resume information and preset error correction prompts are input together as prompts into a preset large model for error correction processing, including: using the large model to determine the target error correction operation to be performed based on the error correction operation description information contained in the error correction prompts; for the arbitration judgment operation on conflicting information included in the target error correction operation, using the arbitration judgment module in the large model to determine multiple semantically similar candidate entity information identified from different page areas in the first resume information; using the arbitration judgment module to determine the page area weight based on the confidence level of the page area to which each candidate entity information belongs; using the arbitration judgment module to determine the weight adjustment strategy based on whether each candidate entity information is associated with time information; wherein, the weight adjustment strategy includes: The weights of candidate entity information associated with time information are increased, and the weights of candidate entity information not associated with time information are decreased. The arbitration decision module determines the comprehensive weight of each candidate entity information based on the page area weight and weight adjustment strategy to which each candidate entity information belongs. The arbitration decision module then determines the target candidate entity information as the arbitration result for conflict information based on the comprehensive weight of each candidate entity information. The error correction prompt includes first role setting information and first domain knowledge constraint information, controlling the large model to position itself according to the first role setting information and to perform the target error correction operation on the first resume information under the constraints of the error correction constraint information to be followed when performing the target error correction operation, as determined by the first domain knowledge constraint information. The structured second resume information output is used to generate a resume knowledge graph in the form of a knowledge graph. The error correction process includes an arbitration judgment operation for conflicting information. The large model is used to determine the required information completion operation based on the information completion operation description information contained in the preset information completion prompts. For the quantization conversion operation of fuzzy expressions included in the information completion operation, the quantization conversion module in the large model is used to identify unstructured fuzzy description information in the first resume information not appearing in the resume knowledge graph. The quantization conversion module is used to determine the semantics of the sentence or paragraph containing the fuzzy description information, and to identify contextual supplementary information in the first resume information that is related to the semantics in terms of topic, entity, or technical field. The quantization conversion module is then used to complete the logical reasoning chain of the fuzzy description information based on the contextual supplementary information to obtain the supplemented fuzzy description information. Finally, the quantization conversion module is used to process the supplemented fuzzy description information according to the STAR method. The information is converted into a standardized description using strong verbs. This standardized description is generated according to a preset template using a predefined strong verb guidance pool through a prompt word enforcement module. The preset template includes actions, technologies / tools, business goals, and quantified results. The standardized description is then used by the quantification conversion module to supplement the resume knowledge graph, resulting in structured third resume information. The information completion prompts include second role setting information and second domain knowledge constraint information. This controls the large model to be positioned in the second role corresponding to the second role setting information. Under the constraints of the information completion constraint information that should be followed when performing the target information completion operation based on the second domain knowledge constraint information, the target information completion operation is performed on the resume knowledge graph and the first resume information. The information completion constraint information includes the output format template of the completed information. Based on the third resume information, the resume file to be processed is evaluated on multiple preset evaluation dimensions to obtain the resume evaluation result; The step of inputting the first resume information and preset error correction prompts as prompts into a preset large model for error correction processing further includes: for the compensation operation of cross-region information included in the target error correction operation, using the cross-region compensation module in the large model to identify each entity in the first resume information; using the cross-region compensation module to determine the correlation between each identified entity and different page areas, and supplementing the original part of the first resume information in the target page area with entity information of the target entities that should belong to the target page area according to the correlation.

2. The method according to claim 1, characterized in that, The process of identifying structured first resume information from the resume file to be processed includes: The visual Transformer model is used to identify different layout areas in the resume file to be processed. Differentiable binarization networks are used to determine the text within different page layout regions; Based on the hierarchical relationship between the text in each region and its corresponding page area, structured first resume information is generated.

3. The method according to claim 1, characterized in that, The step of inputting the first resume information and preset error correction prompts together as prompt information into a preset large model for error correction processing also includes: For the alignment operation of different representations of the same entity included in the target error correction operation, the representation alignment module in the large model is used to perform hierarchical parsing on each sentence in the first resume information, and different representations of the same entity at the same level are determined based on the obtained hierarchical parsing results; the representation alignment module is used to unify the different representations of the same entity into the target representation.

4. The method according to any one of claims 1-3, characterized in that, The process of evaluating the resume file to be processed based on the third resume information across multiple preset evaluation dimensions to obtain resume evaluation results includes: The third resume information and the preset scoring prompts are input into the large model as prompt information to obtain the original scores output for each preset evaluation dimension; wherein, the scoring prompts include multiple preset evaluation dimensions, including: professional matching degree, project complexity, expression clarity and development potential. The original scores of each preset evaluation dimension are weighted according to the corresponding preset weights using the large model, and the resume score result is obtained by summing the weighted scores.

5. The method according to claim 4, characterized in that, The third resume information and preset rating prompts are input into the large model as prompt information to obtain the original scores output for each preset evaluation dimension, including: Regarding the professional matching degree, the large model is used to extract skill information from the third resume information, calculate the degree of overlap between the skill information and the job description information of the target position, and determine the original score of professional matching degree based on the degree of overlap. Regarding the project complexity, the large model is used to extract project experience information, including project scale description, technical depth description, and historical role description, from the third resume information, and the original score of project complexity is determined based on the project experience information. Regarding the clarity of expression, the large model is used to determine the degree of difference between the third resume information and the first resume information, and the structural rationality of the resume content is determined based on the structural information of the first resume information. The original score of the clarity of expression is determined based on the degree of difference and the structural rationality. Based on the development potential, the large model is used to extract growth information, including educational background, historical job change trajectory and cross-domain experience, from the third resume information, and an original score for development potential is determined based on the growth information.

6. The method according to claim 5, characterized in that, Also includes: In response to the existence of an original score for any of the preset evaluation dimensions being lower than the lower limit of the score set for the corresponding preset evaluation dimension, a target identifier is appended to the resume scoring result of the corresponding resume file to be processed. The resume file to be processed, which includes the target identifier, is sent to the target storage location via a preset sending path.

7. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the resume processing method based on a large model as described in any one of claims 1-6.