Resume processing method based on large model and electronic equipment
Patent Information
- Application Number
- HK42026126826
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-01-06
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Application Publication Number (43) Application Publication Date (21) Application Number 202610018060.9 (22) Application Date 2026.01.07 (71) Applicant: Chia Tai Tianqing Pharmaceutical Group Co., Ltd. Address: No. 369, Yuzhou South Road, Lianyungang City, Jiangsu Province, 222062 (72) Inventors: Xie Chengrun, Cao Fenze, Wu Zhaoxuan, Zhang Xiaojun (51) Int.Cl. G06Q 10 / 1053 (2023.01) G06N 5 / 04 (2023.01) G06N 5 / 045 (2023.01) G06N 5 / 025 (2023.01) G06F 40 / 16 (2020.01) G06F 40 / 232 (2020.01) G06F 40 / 30 (2020.01) (54) Invention Title: Resume Processing Method and Electronic Device Based on Large Model (57) Abstract: This invention discloses a resume processing method and electronic device based on a large model, relating to the fields of large model and intelligent agent technology. The method includes: identifying structured first resume information from a resume file to be processed; inputting the first resume information and error correction prompts as prompts into a large model for error correction processing, and generating a resume knowledge graph from the output structured second resume information, wherein the error correction processing includes arbitration judgment operation for conflicting information; inputting the resume knowledge graph, the first resume information, and information completion prompts as prompts into a large model for information completion processing to obtain structured third resume information, wherein the information completion processing includes quantitative conversion operation for fuzzy expressions; and evaluating the resume file to be processed on multiple preset evaluation dimensions based on the third resume information to obtain resume evaluation results. This method can automatically extract accurate resume information and perform multi-dimensional analysis. Claims 3 pages, Description 26 pages, Drawings 5 pages, CN 121481492 A 2026.02.06 CN 1 21 48 14 92 A 1. A resume processing method based on a large model, characterized in that it includes: identifying structured first resume information from a resume file to be processed; inputting the first resume information and preset error correction prompts as prompts into a preset large model for error correction processing, and generating a resume knowledge graph from the output structured second resume information in the form of a knowledge graph, wherein the error correction processing includes arbitration judgment operation for conflicting information; using the large model to determine the information completion operation to be performed based on the information completion operation description information contained in the preset information completion prompts; for the quantization conversion operation of fuzzy expressions included in the information completion operation, using the quantization conversion module in the large model to convert the first resume information that does not appear in the resume knowledge graph.The method involves: identifying unstructured, fuzzy descriptive information; using the quantization conversion module based on the STAR rule to convert the fuzzy descriptive information into standardized descriptions using strong verbs; using the quantization conversion module to supplement the resume knowledge graph with the standardized descriptions to obtain structured third resume information; and evaluating the resume file to be processed on multiple preset evaluation dimensions based on the third resume information to obtain a resume evaluation result. 2. The method according to claim 1, wherein identifying structured first resume information from the resume file to be processed includes: using a visual Transformer model to identify different page regions in the resume file to be processed; using a differentiable binarization network to determine the corresponding text within each page region; and generating structured first resume information based on the membership relationship between the text within each page region and its respective page region. 3. The method according to claim 1, characterized in that, the step of inputting the first resume information and preset error correction prompts together as prompt information into a preset large model for error correction processing includes: using the large model to determine the target error correction operation to be performed based on the error correction operation description information contained in the error correction prompts; for the arbitration judgment operation on conflicting information included in the target error correction operation, using the arbitration judgment module in the large model to determine multiple candidate entity information that are semantically similar and identified from different page areas in the first resume information; using the arbitration judgment module to determine the page area weight based on the confidence level of the page area to which each candidate entity information belongs; using the arbitration judgment module to determine the weight adjustment strategy based on whether each candidate entity information is associated with time information; wherein, the weight adjustment strategy includes: increasing the weight of candidate entity information associated with time information and decreasing the weight of candidate entity information not associated with time information; using the arbitration judgment module to determine the comprehensive weight of each candidate entity information based on the page area weight to which each candidate entity information belongs and the weight adjustment strategy; The arbitration determination module determines the target candidate entity information as the arbitration determination result of conflict information based on the comprehensive weight of each candidate entity information. 4. The method according to claim 3, characterized in that, the step of inputting the first resume information and the preset error correction prompt words as prompt information into the preset large model for error correction processing, further includes: for the compensation operation of cross-region information included in the target error correction operation, using the cross-region compensation module in the large model to identify each entity in the first resume information; using the cross-region compensation module to determine the correlation between each identified entity and different page areas, and according to the correlation, the target entity that should belong to the target page area is assigned to the target page area.The method supplements entity information in the original portion of the first resume information in the target layout area; or, for the alignment operation of different expressions of the same entity included in the target error correction operation, the expression alignment module in the large model performs hierarchical parsing on each sentence in the first resume information, and determines different expressions of the same entity at the same level based on the obtained hierarchical parsing results; and uses the expression alignment module to unify the different expressions of the same entity into the target expression. 5. The method according to claim 1, wherein the step of using the quantization conversion module to convert the fuzzy description information into a standardized description using strong verbs based on the STAR rule includes: using the quantization conversion module to determine contextual supplementary information related to the fuzzy description information in the first resume information; using the quantization conversion module to complete the logical reasoning chain of the fuzzy description information according to the contextual supplementary information to obtain the supplemented fuzzy description information; and using the quantization conversion module to convert the supplemented fuzzy description information into a standardized description using strong verbs according to the STAR rule. 6. The method according to any one of claims 1-5, characterized in that, the step of evaluating the resume file to be processed on multiple preset evaluation dimensions based on the third resume information to obtain a resume evaluation result includes: inputting the third resume information and preset scoring prompts as prompt information into the large model to obtain original scores output for each of the preset evaluation dimensions; wherein, the scoring prompts include multiple preset evaluation dimensions, and the multiple preset evaluation dimensions include: professional matching degree, project complexity, expression clarity and development potential; and using the large model to weight the original scores of each of the preset evaluation dimensions according to corresponding preset weights to obtain a resume score result consisting of the sum of the weighted scores. 7. The method according to claim 6, characterized in that, the third resume information and preset scoring prompts are input into the large model as prompt information to obtain original scores output for each preset evaluation dimension, including: For the professional matching degree, using the large model to extract skill information from the third resume information, calculating the degree of overlap with the job description information of the target position, and determining the original score of professional matching degree based on the degree of overlap; For the project complexity, using the large model to extract project experience information including project scale description, technical depth description, and historical job role description from the third resume information, and determining the original score of project complexity based on the project experience information; For the clarity of expression, using the large model to determine the degree of difference between the third resume information and the first resume information, and determining the structural rationality of the resume content based on the structural information of the first resume information, and determining the original score of the professional matching degree based on the degree of overlap;The method of claim 7 further comprises: in response to the existence of an original score for the clarity of expression based on the degree of difference and the rationality of the structure; and for the development potential, using the large model to extract growth information including educational background, historical job change trajectory and cross-domain experience from the third resume information, and determining the original score for the development potential based on the growth information. 8. The method of claim 7 further comprises: in response to the existence of an original score for any of the preset evaluation dimensions being lower than a pre-set score lower limit for the corresponding preset evaluation dimension, attaching a target identifier to the resume score result of the corresponding resume file to be processed; and sending the resume file to be processed with the attached target identifier to a target storage location according to a preset sending path. 9. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the resume processing method based on the large model according to any one of claims 1-8. Claims 3 / 3 Page 4 CN 121481492 A Resume Processing Method and Electronic Device Based on Large Model Technical Field
[0001] This disclosure relates to the field of data processing technology, specifically to the field of large model and intelligent agent technology, and particularly to a resume processing method and electronic device based on a large model. Background Art
[0002] In the field of human resource recruitment, facing a massive amount of resumes in various formats (such as PDFs, scanned copies, and images), traditional screening methods mainly rely on manual reading or keyword matching systems based on fixed rules. However, manual screening is inefficient, subjective, and prone to errors due to fatigue. On the other hand, rule-based automated systems are difficult to iterate and update in a timely manner, making it difficult to achieve both efficiency and quality in the screening process, and thus failing to meet the needs of modern efficient and accurate recruitment. Summary of the Invention
[0003] Embodiments of this disclosure propose a resume processing method, apparatus, electronic device, computer-readable storage medium, and computer program product based on a large model.
[0004] In a first aspect, embodiments of this disclosure propose a resume processing method based on a large model, comprising: identifying structured first resume information from a resume file to be processed; inputting the first resume information and preset error correction prompts as prompts into a preset large model for error correction processing, and generating a resume knowledge graph from the output structured second resume information in the form of a knowledge graph; wherein, the error correction processing includes an arbitration judgment operation for conflicting information; inputting the resume knowledge graph, the first resume information, and preset information completion prompts as prompts into the large model for information completion processing to obtain structured third resume information; wherein, the information completion processing includes a quantization conversion operation for fuzzy expressions;Regarding the third resume information, the resume file to be processed is evaluated on multiple preset evaluation dimensions to obtain a resume evaluation result.
[0005] In a second aspect, embodiments of this disclosure propose a resume processing device based on a large model, comprising: a first resume information recognition unit, configured to identify structured first resume information from the resume file to be processed; an error correction processing and knowledge graph generation unit, configured to input the first resume information and preset error correction prompts as prompts into a preset large model for error correction processing, and generate a resume knowledge graph in the form of a knowledge graph from the output structured second resume information; wherein, the error correction processing includes an arbitration judgment operation for conflicting information; an information completion unit, configured to input the resume knowledge graph, the first resume information, and preset information completion prompts as prompts into the large model for information completion processing to obtain structured third resume information; wherein, the information completion processing includes a quantification conversion operation for fuzzy expressions; and a multi-dimensional evaluation unit, configured to evaluate the resume file to be processed on multiple preset evaluation dimensions based on the third resume information to obtain a resume evaluation result.
[0006] In a third aspect, embodiments of the present disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the large-model-based resume processing method as described in the first aspect.
[0007] In a fourth aspect, embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing computer instructions for enabling a computer to implement the large-model-based resume processing method as described in the first aspect. Specification 1 / 26 pages 5 CN 121481492 A
[0008] In a fifth aspect, embodiments of the present disclosure provide a computer program product including a computer program, the computer program being executed by a processor to implement the steps of the large-model-based resume processing method as described in the first aspect.
[0009] The resume processing scheme based on a large model provided in this disclosure firstly identifies structured first resume information directly from the resume file to be processed, providing a preliminary standardized data foundation for subsequent in-depth processing and improving the initial processing capability for multi-format source files; secondly, by combining the first resume information with error correction prompts and inputting it into the large model for error correction processing, and specifically including arbitration judgment operations for conflicting information, it can effectively identify and resolve potential contradictions, ambiguities, or redundant information in the resume, thereby outputting more accurate and consistent second resume information, which is then used to generate a resume knowledge graph in the form of a knowledge graph, so that the internal relationship of information can be represented in a structured and standardized manner, providing a foundation for subsequent processing.Deep semantic processing laid a good foundation; then, by inputting the resume knowledge graph, first resume information and information completion prompts into the large model for information completion processing, and specifically including quantitative conversion operations for vague expressions, it can intelligently complete, refine and standardize the abbreviated, colloquial and other vague descriptions in the resume, and obtain third resume information with more complete information and clearer semantics, thereby significantly improving the usability and comparability of resume content; finally, based on this high-quality, structured third resume information, evaluation is carried out on multiple preset evaluation dimensions, which can generate more objective, comprehensive and interpretable resume evaluation results. This solution effectively overcomes the two major bottlenecks of inaccurate original information and semantic ambiguity through two targeted applications (error correction and completion) of the large model in key links, and finally achieves simultaneous improvement in the accuracy, automation and evaluation depth of resume screening, and also realizes automated and intelligent processing from the original resume file to the final evaluation result.
[0010] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Brief Description of the Drawings
[0011] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: FIG1 is an exemplary system architecture to which this disclosure can be applied; FIG2 is a flowchart of a large-model-based resume processing method provided by an embodiment of this disclosure; FIG3 is a flowchart of a method for arbitrating conflict information provided by an embodiment of this disclosure; FIG4 is a flowchart of a method for compensating cross-regional information and aligning different representations of the same entity provided by an embodiment of this disclosure; FIG5 is a flowchart of a method for quantifying and converting fuzzy representations provided by an embodiment of this disclosure; FIG6 is a flowchart of a scoring method for different preset evaluation dimensions provided by an embodiment of this disclosure; FIG7 is a structural block diagram of a large-model-based resume processing device provided by an embodiment of this disclosure; FIG8 is a structural schematic diagram of an electronic device suitable for performing a large-model-based resume processing method provided by an embodiment of this disclosure. Detailed Description of Embodiments
[0012] Exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0013] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solutions of this disclosure all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0014] FIG1 shows an exemplary system architecture 100 of an embodiment of the resume processing method, apparatus, electronic device, and computer-readable storage medium based on a large model that can be applied to this disclosure.
[0015] As shown in FIG1, the system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0016] Users can use the terminal devices 101, 102, and 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications for information communication between the terminal devices 101, 102, 103 and the server 105 can be installed, such as resume screening applications, image conversion applications, instant messaging applications, etc.
[0017] The terminal devices 101, 102, 103 and the server 105 can be hardware or software. When the terminal devices 101, 102, 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers, etc.; when the terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above, and can be implemented as multiple software or software modules, or as a single software or software module, without specific limitations. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; when the server is software, it can be implemented as multiple software or software modules, or as a single software or software module, without specific limitations.
[0018] Server 105 can provide various services through its built-in applications. Taking a resume screening application that provides resume processing services for massive amounts of resumes to enterprise managers as an example, server 105 can achieve the following effects when running this resume screening application: First, it receives resume files to be processed from terminal devices 101, 102, and 103 via network 104, and identifies structured first resume information from the resume files to be processed; then, it inputs the first resume information and preset error correction prompts as prompt information into a preset large model for error correction processing, and generates a resume knowledge graph in the form of a knowledge graph based on the output structured second resume information. This error correction processing includes arbitration and judgment operations for conflicting information; next, it uses the resume knowledge graph, the first resume information, and preset information completion prompts as prompts.Information is input into the large model for information completion processing to obtain structured third resume information. This information completion processing includes quantitative conversion operations for fuzzy expressions. Finally, based on the third resume information, the resume file to be processed is evaluated on multiple preset evaluation dimensions to obtain the resume evaluation result.
[0019] It should be noted that the resume file to be processed can be temporarily obtained from the terminal devices 101, 102, and 103 through the network 104, or it can be pre-stored locally on the server 105 in various ways. Therefore, when the server 105 detects that this data has been stored locally (for example, when starting to process the previously stored resume processing task), it can choose to directly obtain this data from the local storage. In this case, the exemplary system architecture 100 may not include the terminal devices 101, 102, and 103 and the network 104.
[0020] Since complex analysis of resume files requires a large amount of computing resources and strong computing power, the resume processing method based on a large model provided in the subsequent embodiments of this disclosure is generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the resume processing device based on a large model is also generally set in the server 105. However, it should also be noted that when the terminal devices 101, 102, and 103 also have sufficient computing power and computing resources, the terminal devices 101, 102, and 103 can also complete the various calculations performed by the server 105 through the resume screening applications installed on them, and thus output the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, but the resume screening application determines that the terminal device has strong computing power and a lot of remaining computing resources, the terminal device can be allowed to perform the above calculations, thereby appropriately reducing the computing pressure on server 105. Correspondingly, the resume processing device based on the large model can also be set in terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also not include server 105 and network 104.
[0021] It should be understood that the number of terminal devices, networks, and servers in FIG1 is merely illustrative. Depending on the implementation needs, there can be any number of terminal devices, networks, and servers.
[0022] Please refer to Figure 2, which is a flowchart of a resume processing method based on a large model provided in this embodiment of the present disclosure. The process 200 includes the following steps: Step 201: Identify structured first resume information from the resume file to be processed; This step aims to have the execution entity of the resume processing method based on a large model (e.g., server 105 shown in Figure 1, or a resume screening application installed on server 105) identify structured first resume information from the resume file to be processed.The goal is to initially transform raw resume files of varying formats into a structured data format that can be understood and processed by a machine. This "resume file to be processed" is a broader concept encompassing various electronic document formats and audio / video file formats actually received by users in recruitment scenarios. Examples include PDF documents with optional text, image files (such as JPG and PNG), images or PDF files scanned from paper documents, and audio / video files. Upon receiving these files, the server, for document-formatted files, which often contain complex layouts, font styles, and other visual information crucial for understanding the resume's structure, can first identify this visual information before recognizing the text information to obtain structured initial resume information. For audio / video files, speech-to-text conversion or screenshotting followed by text extraction from the image can be performed first.
[0023] Specifically, when the above-mentioned execution entity performs this step, it does not perform simple optical character recognition (OCR), because the output of traditional OCR is often a continuous and undifferentiated stream of text, which loses the position, grouping and visual hierarchy of the original text on the page. However, this step requires the production of "structured" information, so the above-mentioned execution entity must also perform "layout awareness". In practice, the aforementioned execution entity can call or integrate a document image analysis and understanding module. This module first parses the file to be processed. If it is a PDF, it extracts the page images and possible text layer information; if it is an image, it processes the image directly. Subsequently, the module can perform layout analysis on the page images, identifying different logical regions through computer vision technology. These logical regions are logical blocks determined based on visual features (such as font size, bolding, alignment, and spacing) and semantics of text content (such as the presence of heading terms like "work experience" or "project experience"). Examples include "personal information title area," "educational background list area," and "detailed work experience area." Simultaneously or after identifying the regions, the aforementioned execution entity can perform text detection and recognition within each region to ensure that the text in the image is accurately converted into digital text, and crucially, to preserve the binding relationship between each text fragment and its corresponding logical region, as well as the relative position or order of the text within the region.
[0024] Through the above operations, the "first resume information" generated by the executing entity is a preliminary structured data. This structured data does not refer to a database record that has been precisely filled with standard fields such as "name," "company," and "position," but rather an intermediate representation. It is usually presented as a hierarchical, labeled data structure, such as a list or tree structure, where each node represents an identified page area (which may be accompanied by an area type label, such as...).The nodes “section_heading” and “item_list” are associated with ordered text content identified within that area. For example, the execution entity might output a structure containing a section description block labeled “work_experience” (page 4 / 26, CN 121481492 A). This block contains several “job_entry” sub-blocks arranged chronologically, each containing identified lines of original text. This structure preserves the visual organization logic of the resume document, providing crucial information for subsequent steps to distinguish which text belongs to the same experience, which are titles, and which are detailed descriptions. In short, “first resume information” transforms an unstructured document image into a programmable collection of text that retains the original document organization logic, completing the first key data morphological change in the intelligent information processing flow.
[0025] One implementation method, including but not limited to, is as follows: First, the visual Transformer model is used to identify different page regions in the resume file to be processed; then, a differentiable binarization network is used to determine the corresponding text within each page region; finally, based on the membership relationship between the text within each page region and its respective page region, structured first resume information is generated.
[0026] Specifically, the above-mentioned execution entity first calls the Vision Transformer (VIT) model to perform global analysis on the image of the resume document. VIT is a deep learning model that applies the successful Transformer architecture in natural language processing to computer vision tasks. Its technical principle is that it divides the input image into a series of fixed-size image blocks and linearly embeds each image block into a vector sequence. Through the multi-layer self-attention mechanism inside the model, these vectors can dynamically capture the long-range dependencies between all blocks in the image, thereby achieving a global understanding of the image content. In the specific practice of this step, the above-mentioned execution entity preprocesses the resume file to be processed into an image of uniform size, and then inputs it into the pre-trained VIT model. The model, through its learned prior knowledge of document structure, can output the recognition results of different semantic regions in the image. For example, it can accurately distinguish the "personal basic information title area", "educational background list area", "work experience details area", "skills block" and "project experience description area", etc. These regions are given by the model in the form of bounding box coordinates and region type labels, which constitute the machine interpretation of the logical structure of the resume layout.
[0027] Then, based on the obtained layout region division, a differentiable binarization network is further used.The DBNet network performs refined text detection and recognition on each identified page region. DBNet is a model specifically designed for text detection in complex scenes. Its core technology lies in introducing a differentiable binarization operation, enabling the model to learn end-to-end how to accurately predict whether each pixel belongs to a text region, and exhibiting excellent adaptability to text lines of arbitrary shapes (such as curved or tilted). In practice, the execution entity inputs image slices of each page region output by the VIT model into DBNet. DBNet first outputs a probability map representing the probability that each pixel belongs to text, then generates a binary mask of the text region through its differentiable binarization module, finally outputting fine-grained bounding boxes at the level of each text line or word. Next, the execution entity can call or integrate an optical character recognition engine to perform character recognition on the image content within these detected text boxes, converting the image into a string of text. Thus, each page region is associated with a set of text content it contains and its relative positions within the region.
[0028] Step 202: Input the first resume information and the preset error correction prompts as prompts into the preset large model for error correction processing, and generate a resume knowledge graph from the output structured second resume information in the form of a knowledge graph; Based on step 201, this step aims to have the above-mentioned execution entity input the first resume information and the preset error correction prompts as prompts into the preset large model for error correction processing, output structured second resume information, and then generate a resume knowledge graph from the second resume information in the form of a knowledge graph. The purpose is to perform in-depth verification, correction and integration of the previously identified "structured first resume information" to generate a more accurate, consistent and logically clear "second resume information", and finally organize it into the standardized form of "resume knowledge graph". To achieve this goal, the aforementioned executing entity combines the "first resume information" with a pre-designed "error correction prompt" to form the prompt information, which is then input into a preset large model (also known as a large language model or a multimodal large model). The "error correction prompt" is actually a complex text template containing specific task descriptions, role settings, and domain knowledge constraints. Its function is to precisely guide the large model to understand and execute a series of preset error correction sub-tasks. This allows the large model to automatically process various complex problems in the original extracted information through its powerful contextual understanding and reasoning capabilities, thereby significantly improving the data quality foundation of subsequent processes.
[0029] Specifically, this "error correction processing" can cover multiple levels of error correction operations. For example, it can first include arbitration and judgment operations for conflicting information. This is to consider that the same entity (such as the employer company) may appear in different locations in the resume under different names.When conflicts arise (e.g., between "XX Group" and "XX Treasure"), leading to factual conflicts or ambiguities, the aforementioned executing entity can instruct the large model to simulate a "reviewer or arbitrator" with domain knowledge. This model first identifies semantically similar candidate entities and, combined with the layout area labels carried in the "first resume information" (e.g., whether the information comes from the "Work Experience" title area or the "Project Description" body area), assigns initial confidence weights to candidate entities from different sources. Furthermore, the executing entity can guide the large model to establish temporal logic constraints, aligning each candidate entity with the time information identified in the resume (e.g., start and end years of employment). If an entity's statement cannot find a reasonable position on the timeline, its credibility will be significantly reduced. Finally, the large model integrates spatial weights and temporal logic, and invokes its inherent common sense knowledge (e.g., knowing that "XX Treasure" is a business under "XX Group"), to normalize the conflicting entities and output the most credible and standard entity statement as the arbitration result.
[0030] Of course, the error correction process can also include compensation operations for cross-regional information and alignment operations for different representations of the same entity, in order to further eliminate problems in the first resume information. Specific explanations will be provided in detail in subsequent embodiments. After completing the above error correction process, the executing entity can generate a resume knowledge graph from the corrected and unified "structured second resume information" output by the large model, in the form of a "knowledge graph." That is, it organizes the information according to a "entity-relationship-entity" triple structure (for example, constructing the three elements "Zhang San," "employed at," and "XX Group" into a triple, which can be guided by a predefined knowledge graph schema, which specifies common entity and relationship types in the resume domain). This transforms linear, paragraph-style text information into a structured network data model with interconnected nodes and edges. This transformation not only allows the inherent relationships of information to be explicitly represented but also lays a crucial data foundation for subsequent semantic retrieval, relational reasoning, and deep analysis based on graph structures.
[0031] To drive the large model to perform a series of predetermined and complex text analysis and correction tasks, the "correction prompt word" is not a simple instruction, but a structured text template that clearly explains to the large model the "error correction operation description information" that needs to be completed. For example, it instructs the model to "identify and resolve contradictions and inconsistencies in the resume regarding company name, job timeline, or skill descriptions." When the aforementioned executing entity sends this combined prompt information to the large model, the primary task of the large model is to parse this prompt and extract and understand what the "target error correction operation" specifically is. One implementation method, including but not limited to, is: using the large model to determine the error correction operation description information contained in the error correction prompt word.The target error correction operation to be performed is determined; the target error correction operation is performed on the first resume information using the large model. That is, the error correction prompt should contain at least the error correction operation description information, so that the above-mentioned execution subject can determine the target error correction operation to be performed by the large model according to the error correction operation description information contained therein, such as the set of sub-tasks such as "conflict arbitration", "information compensation" and "term alignment".
[0032] In order to ensure that the large model can perform error correction in a way that conforms to professional domain knowledge, the error correction prompt can also include the first role setting information and the first domain knowledge constraint information, so as to further control the large model to stand in the first role corresponding to the first role setting information, and control the large model to determine the error correction constraint information to be followed when performing the target error correction operation according to the first domain knowledge constraint information, so that the large model can be controlled to stand in the first role and perform the target error correction operation on the first resume information under the constraint of the error correction constraint information during the execution stage. This implementation further refines the design of "error correction prompts" by including "first role setting information" and "first domain knowledge constraint information." When these elements are included in the provided error correction prompts, the guidance and control over the large model becomes more precise and powerful. "First role setting information" instructions, such as "Please play the role of a senior human resources data governance expert," essentially inject a virtual "personality" or perspective into the large model. Technically, this provides a strong contextual prior, making the model more inclined to mimic the role's professional knowledge background, thinking patterns, and expression habits when generating responses, thereby enhancing the professionalism and rationality of error correction decisions. "First domain knowledge constraint information," on the other hand, provides more specific operational boundaries and rules. For example, "When resolving conflicts, please prioritize information located under the 'Work Experience' heading area and refer to the common parent-subsidiary relationship diagrams of Chinese internet companies." This adds filters and signposts to the large model's reasoning process, ensuring that its output does not deviate from business common sense and predetermined data governance standards.
[0033] In practical terms, when constructing complete prompt information, the aforementioned executing entity can smoothly assemble the "first resume information" (usually embedded in structured text formats such as JSON), role setting instructions, domain knowledge constraint clauses, and specific error correction operation descriptions into a coherent natural language instruction according to a predetermined template format. Subsequently, the aforementioned executing entity calls the large model service through the application programming interface and sends this instruction. After receiving the instruction, the large model will undergo an internal instruction following and role-playing process. It will first "place" itself into the designated "first role," and then interpret the "first domain knowledge constraint information" based on the cognitive framework of that role, deriving from it the necessary steps to perform the error correction operation.The "following of error-correction constraints" ultimately manifests in the fact that the text responses produced by the large model strictly adhere to the role's identity, and that the conflict resolution, information compensation suggestions, or terminology alignment results clearly reflect the guidance of domain constraints. For example, when a resume contains both "XX Group" and "XX Bao," the large model, positioned as a "human resources data governance expert" and constrained by relevant corporate relationships, is more likely to output a standardized judgment such as "Based on corporate organizational relationships, 'XX Bao' should be uniformly merged under its parent company entity 'XX Group'," rather than a vague or erroneous processing suggestion. In this way, the aforementioned implementing entities effectively discipline and guide the large model's powerful general language capabilities onto the specific and professional task track of resume information error correction, achieving accuracy and reliability in automated processing.
[0034] Step 203: Input the resume knowledge graph, the first resume information, and the preset information completion prompts as prompts into the large model for information completion processing to obtain structured third resume information; Based on step 202, this step aims to have the above-mentioned execution entity input the resume knowledge graph, the first resume information, and the preset information completion prompts as prompts into the large model for information completion processing to obtain structured third resume information. The purpose is to deeply enhance the information that has already undergone preliminary error correction and structured organization, thereby solving the common problems in resumes such as colloquialism, ambiguity, and missing key details. That is, through the semantic understanding and reasoning ability of the large model, incomplete descriptions are transformed into standardized and quantifiable evaluation expressions.
[0035] Among them, the structured knowledge representation of knowledge graphs provides the large model with a clear, machine-readable context about the candidate's skills, experience and the relationships between them; the first resume information retains the text content identified from the original document and organized by page area, which will serve as supplementary original corpus to ensure that the original expressions that may not be fully structured but contain details, which are directly extracted by character recognition technology, will not be lost during the completion process; the preset information completion prompt words are the key control instructions that guide the entire completion process. Their content will not only indicate the task of "to perform information completion", but also define the specific rules, output format and professional context to be followed in the completion operation.
[0036] One of the core processing operations of this information completion process is the "quantification transformation of vague expressions." This refers to converting descriptions in resumes, such as "responsible for the backend system," "improved performance," and "participated in big data projects," which lack specific technical details, action boundaries, and measurable results, into standardized sentences rich in clear technology stacks, specific action verbs, business scenarios, and quantitative indicators. At the technical principle level, this will rely on the powerful contextual reasoning and common sense knowledge possessed by the large model.The ability to integrate and follow instructions. The aforementioned executor can combine the resume knowledge graph (providing structured relational context), the first resume information (providing raw text details), and completion prompts (providing transformation rules and format constraints) to form a composite prompt information input model. When interpreting this prompt, this model will simultaneously consider the structured relational constraints (e.g., knowing from the knowledge graph that this person possesses the "Spark" skill), the raw ambiguous text (e.g., "has worked on recommendation systems"), and the hard requirements in the prompts (e.g., "must use the STAR method to describe and supplement the quantitative results"), thereby performing comprehensive reasoning.
[0037] In practical terms, when the above-mentioned execution entity implements this step, it can store or dynamically generate "information completion prompts" optimized for different job types in the system in advance. For example, when processing a specific resume, the above-mentioned execution entity first calls the corresponding knowledge graph data and the first resume information from the storage, and serializes them in a predetermined format (for example, describing the knowledge graph in JSON-LD format and listing the first resume information in the form of text paragraphs). Then, it combines these serialized data with the completion prompt template to form a complete and coherent natural language instruction, which is sent to the large model service through the API. After receiving this composite instruction, the large model will execute the internal information completion logic. For example, when encountering the vague statement "I've developed a recommendation system," the large model retrieves relevant skill entities from the context (e.g., it learns from the knowledge graph that the candidate entity is associated with the skills "collaborative filtering" and "Spark"). Based on the constraints of "using strong verbs" and "supplementing quantitative indicators" in the prompts, it performs logical chain reasoning: the subject is the candidate, the object is the recommendation system, the available technologies are collaborative filtering and Spark, the reasonable action is "development" or "implementation," and the possible business goal is "personalized recommendation." This leads to the inference of a reasonable quantitative improvement direction, such as "click-through rate." Ultimately, the model generates "structured third-party resume information," such as "I developed a recommendation system using the Spark framework and collaborative filtering algorithm, increasing user content click-through rate by 15%." This output is not only professionally worded and clearly structured but also contains quantitative dimensions that can be directly used by subsequent evaluation modules, achieving a fundamental improvement from vague descriptions to calculable and comparable structured information. Through this step, the aforementioned executing entity significantly enhances the semantic density and evaluation usability of the resume information, laying a solid data foundation for subsequent objective and accurate intelligent scoring.
[0038] To drive the large model to perform semantic-level deep enhancement and reconstruction of the information that has been initially corrected and structured, the "information completion prompt" is not a simple instruction, but a structured text template that clearly explains to the large model the "information completion operation description information" that needs to be completed. When the aforementioned execution entity sends this combination of prompt information...After receiving the prompt, the primary task of the large model is to parse the prompt, extract and understand the specific "target information operation to be performed." One possible implementation method, including but not limited to: using the large model to determine the target information completion operation based on the information completion operation description information contained in the information completion prompt; and using the large model to perform the target information completion operation on the resume knowledge graph and the first resume information. That is, the aforementioned execution entity first relies on the "information completion operation description information" contained in the "information completion prompt" to guide the large model in determining the "target information completion operation to be performed." This description information is a clear task definition, such as "Please identify and complete the missing technical details, quantitative achievements, and business background in all work experiences and project descriptions in the resume, transforming colloquial expressions into professional statements." When the large model receives this compound prompt, it internally parses the instructions, extracting one or more specific, executable sets of operation targets from the natural language description, such as "supplementing the core technology stack for each project experience," "replacing vague verbs with strong verbs conforming to the STAR method," and "inferring a reasonable quantitative range for achievements lacking numbers," etc. This process constrains and directs the open text generation capabilities of the large model towards addressing the specific problem of incomplete and unprofessional resume information. (Instruction manual, page 8 / 26, CN 121481492 A
[0039] ) To make the completion operation more accurate and in line with professional domain expectations, the information completion prompt can also include second role setting information and second domain knowledge constraint information. In this case, the large model can be additionally controlled to be positioned in the second role corresponding to the second role setting information, and the large model can be controlled to determine the information completion constraint information to be followed when performing the target information completion operation based on the second domain knowledge constraint information. This information completion constraint information includes: the output format template of the completed information. Then, the execution phase is adaptively changed to: controlling the large model to be positioned in the second role, and performing the target information completion operation on the resume knowledge graph and the first resume information under the constraint of the information completion constraint information.
[0040] This implementation further allows embedding "second role setting information" and "second domain knowledge constraint information" in the "information completion prompts". When the prompts provided by the aforementioned executing entity contain these elements, its control over the behavior of the large model will reach a more refined level. The second role setting information and second domain knowledge constraint information are similar to the first role setting information and first domain knowledge constraint information mentioned above, except for some differences caused by different target objects. Moreover, the "second domain knowledge constraint information" provides specific operation guidelines and boundaries, such as "When completing, the backend project must clearly mention the framework and database used, and the frontend project must explain the technology stack and compatibility requirements, and all quantitative indicators...""It should be reasonable and in line with industry common sense," and a key constraint is the "output format template of the completed information," for example, requiring the model to strictly follow the sentence structure of "[action verb] based on [technology stack] [project type], realizing [quantitative indicators] [business value]." This ensures that the descriptions of different resumes are structurally consistent and comparable after completion, thus facilitating subsequent automated evaluation.
[0041] In practical terms, when executing this step, the aforementioned entity can assemble all the elements into a complete prompt according to a predetermined logic. For example, it can serialize the "resume knowledge graph" into a descriptive text (such as listing the main triples), and then organize the "first resume information" into paragraph-based plain text. Next, it smoothly combines this data with phrases containing role settings, domain constraints, and operational descriptions into a coherent natural language prompt, seemingly a task assignment to an expert. This prompt is then sent to the large model service via its integrated API. When interpreting this complex prompt, the large model first "positions itself," simulating the thought process of the designated expert during its internal reasoning. Then, it parses and internalizes the domain constraints and output templates, using them as rules to be followed when generating text. Ultimately, the text output by the large model possesses both professional depth and strictly adheres to the preset format and content requirements. For example, When the model encounters a description in the knowledge graph that associates "Python" and "data analysis" skills but only states "did some data analysis" in the first resume information, the model, positioned as a "technical architect" and constrained by the template, will generate a standardized and neat statement such as "led a sales data trend analysis project using Python and Pandas libraries, built a predictive model, and improved the efficiency of monthly report generation by 40%." Through this series of designs, the general capabilities of the large model are transformed into a stable, reliable, and professional automated resume information enhancement tool.
[0042] Step 204: Based on the third resume information, evaluate the resume file to be processed on multiple preset evaluation dimensions to obtain the resume evaluation result.
[0043] Based on step 203, this step aims to have the above-mentioned executing entity perform the final ability quantification and comprehensive evaluation of the "structured third resume information" obtained after the deep processing and enhancement of the aforementioned steps to obtain the resume evaluation result. That is, to realize the purpose of transforming the candidate's ability and potential contained in a resume into a set of objective and comparable quantitative indicators, thereby providing a direct and reliable basis for recruitment decisions.
[0044] In executing this step, the "third-party resume information" relied upon by the aforementioned executing entity is a high-quality structured data that has undergone error correction, completion, and standardization. It contains clear, complete, and semantically unambiguous descriptions of the candidate's experience, skills, and achievements. The executing entity combines this information with a pre-designed "scoring prompt" to form an instruction.The input is sent to the large model. The "scoring prompt" here acts as the examiner's "scoring criteria" and "scoring rules," and its content specification (page 9 / 26, CN 121481492 A) clearly defines the "multiple preset evaluation dimensions" that need to be assessed. Specifically, in some other implementations of this embodiment, these dimensions can be concretized as "professional fit," "project complexity," "clarity of expression," and "development potential," as well as other similar evaluation angles. Each dimension corresponds to an examination of different aspects of the candidate: professional fit focuses on the degree of alignment between skills and job requirements; project complexity measures the technical depth and scale of past experience; clarity of expression assesses the logic and information presentation quality of the resume itself; and development potential infers future possibilities based on growth trajectory and background. The server clearly conveys the definitions and judgment criteria of these evaluation dimensions to the large model through the prompt.
[0045] At the technical principle level, the aforementioned executing entity can input the "third resume information" along with "scoring prompts" containing multi-dimensional scoring rules into the large model. Essentially, this constrains and guides the powerful semantic understanding and reasoning capabilities of the large model into a structured evaluation task. Upon receiving this composite instruction, the large model will follow the guidance of the prompts and conduct in-depth analysis of the "third resume information" for each preset evaluation dimension. For example, when evaluating "professional matching," the model will extract the skill list from the resume and perform semantic comparison and overlap calculations with the target job responsibilities that may be implicitly or explicitly given in the prompts. When evaluating "project complexity," the model will analyze the technology stack in the project description, the scale of the business problems solved, and the individual's role within it. For each dimension, the large model will output a "raw score" based on its analysis. This score is usually a value within a predetermined range (e.g., 0-10 points), reflecting the model's quantitative judgment of the level of that dimension.
[0046] In practical terms, the aforementioned execution entity can pre-configure "scoring prompt word" templates optimized for different job types in the system. When a resume needs to be evaluated, the server first calls the corresponding "third-party resume information," prepares it in structured text form (such as JSON or a specific format of natural language summary), and then combines this information with the scoring prompt word template to form a complete, task-specific natural language prompt, which is sent to the large model service through the application programming interface. After processing, the large model returns a text response containing the original scores for each dimension. The aforementioned execution entity parses this response to extract the corresponding values for each dimension, and then can further enter the weighted calculation stage. For example, the weight system of each evaluation dimension (such as professional matching degree, project complexity, etc.) can be pre-stored internally or dynamically configured.The system calculates the original scores of each dimension obtained from the analysis, multiplies them by the corresponding weight coefficients, and finally sums all these weighted scores to obtain a comprehensive "resume scoring result". This mathematical calculation process integrates the discrete judgments of each dimension by the large model into a unified, sortable total score, and at the same time reflects the relative importance of different dimensions in the final decision through weight allocation, thus completing the final transformation from semantic information to actionable decision data.
[0047] A specific implementation method can be: inputting the third resume information and the preset scoring prompt words as prompt information into the large model to obtain the original scores output on each preset evaluation dimension; wherein, the scoring prompt words include multiple preset evaluation dimensions, including: professional matching degree, project complexity, expression clarity and development potential; using the large model to weight the original scores of each preset evaluation dimension according to the corresponding preset weights, the resume scoring result is obtained by the sum of the weighted scores.
[0048] The resume processing method based on a large model provided in this embodiment effectively overcomes the two major bottlenecks of inaccurate original information and semantic ambiguity by applying the large model twice in key stages (error correction and completion). Ultimately, it achieves simultaneous improvement in the accuracy, automation and evaluation depth of resume screening, and also realizes automated and intelligent processing from the original resume file to the final evaluation result.
[0049] To deepen the understanding of how to use a large model for error correction, please refer to Figure 3. Figure 3 is a flowchart of a method for arbitration and determination of conflict information provided by an embodiment of this disclosure. It aims to provide a specific implementation method for arbitration and determination of conflict information included in the target error correction operation. The process 300 includes the following steps: Specification 10 / 26 Page 14 CN 121481492 A Step 301: Use the arbitration and determination module in the large model to determine multiple candidate entity information that are semantically similar and identified from different page areas in the first resume information; Starting this operation first depends on the generation of the "first resume information" in the early stage. This information not only includes text content, but also binds the page area label to which each text fragment belongs. On this basis, in this step, the above-mentioned execution subject calls the "arbitration and determination module" in the large language model that it integrates and has been optimized for specific tasks. In this step, the module scans and clusters semantically similar entity expressions in the whole text. Specifically, the aforementioned execution entity inputs all the "first resume information" into this module. This module utilizes its built-in semantic understanding capabilities to calculate the semantic similarity between different entity names, identifying different textual expressions referring to the same or highly related things as "multiple candidate entity information." For example, it identifies "XX Group," "XX Treasure," and "XXX GROUP" as a group of candidates. At the technical level, this relies on a large model.The common knowledge about vocabulary, aliases, and company organizational structure acquired during training.
[0050] Step 302: Use the arbitration judgment module to determine the page area weight based on the confidence level of the page area to which each candidate entity information belongs; This step aims to guide the arbitration judgment module into the weight judgment stage by the above-mentioned execution entity, that is, to assign an initial "confidence level" to different page areas according to preset rules. In practice, it can store a mapping table internally. For example, it stipulates that the area under the main title of "Work Experience" or "Education Background" has a high confidence level, while the area in the main text of "Self-evaluation" or "Project Details" has a relatively low confidence level. The arbitration judgment module can then assign the corresponding "page area weight" to each candidate entity information based on which area it was extracted from. The operation of this step is used to reflect the full consideration of the authority of the information source, that is, the information in the formal resume is usually more reliable than the supplementary description.
[0051] Step 303: Utilize the arbitration decision module to determine the weight adjustment strategy based on whether each candidate entity's information is associated with time information. This step aims to further introduce time logic as a key corrective factor by the aforementioned executing entity. That is, the arbitration decision module will check whether each candidate entity's information is associated with the time information explicitly identified in the resume. For example, in the text "In XX Group (2019-2022)," the entity "XX Group" is explicitly associated with the time interval "2019-2022." The aforementioned executing entity will implement the "weight adjustment strategy": for such candidate entities with associated time information, their weight will be increased because this enhances their empirical validity as a specific resume; conversely, for entity statements without explicit time information, their weight will be decreased because they may only be background mentions or general statements. This strategy is based on the fundamental logic that valid work or educational experiences in a resume must occur within a specific time period, effectively distinguishing between "factual statements" and "background information."
[0052] Step 304: Using the arbitration decision module, determine the comprehensive weight of each candidate entity information based on the page area weight and weight adjustment strategy to which each candidate entity information belongs; After integrating the page area weights and adjusting for time correlation, the arbitration decision module will calculate a "comprehensive weight" for each candidate entity. This process can be a simple weighted summation or a more complex function calculation.
[0053] Step 305: Using the arbitration decision module, determine the target candidate entity information as the arbitration decision result for conflict information based on the comprehensive weight of each candidate entity information.
[0054] Typically, the candidate entity with the highest weight will be selected as the arbitration result, i.e., the final standardized representation of the conflict information.
[0055] This embodiment provides a specific "arbitration decision operation for conflict information" through steps 301-305.The implementation method aims to solve the problem of discrepancies or conflicts in textual descriptions of the same real-world entity identified from different sections of a resume. For example, a resume might state "XX Group" in the "Work Experience" section and mention "XX Treasure" in the "Project Description" section. Although the two are different in text, they are highly semantically related. The aforementioned execution entity can automatically resolve such conflicts by executing a series of logical steps provided in this embodiment, simulating the reasoning process of a professional reviewer, and outputting the most reliable and standard entity descriptions, thereby ensuring the consistency and accuracy of subsequent information processing.
[0056] Based on the above embodiments, to deepen the understanding of other possible error correction operations, please also refer to Figure 4. Figure 4 is a flowchart of a method for compensation operation of cross-regional information and alignment operation of different descriptions of the same entity provided by the present disclosure. The process 400 includes the following steps: Step 401: For the compensation operation of cross-regional information included in the target error correction operation, the cross-regional compensation module in the large model is used to identify each entity in the first resume information; Step 402: The cross-regional compensation module is used to determine the correlation between each identified entity and different page areas, and according to the correlation, the target entity that should belong to the target page area is used to supplement the original part of the first resume information in the target page area with entity information; In steps 401-402, the triggering of the compensation operation of cross-regional information performed by the above-mentioned executing entity is based on a practical problem: that is, the key entity information in the resume (such as the employer company name, the graduating school) may not appear completely in its most expected logical area (such as under the title of "Work Experience" or "Education Background"), but is scattered in other areas (such as "Self-evaluation" or "Project Details"). To address this issue, the aforementioned executing entity invokes a specific functional unit within the large model, namely the "cross-regional compensation module." This module, based on the large model's inherent powerful named entity recognition capability, performs a global scan of the entire "first resume information" (including all page areas and their text), identifying all entities such as names, organization names, locations, and skills. Identification is not the end goal; the key is to determine the "relationship between each entity and different page areas."
[0057] At the technical principle level, the aforementioned executing entity can guide the large model to perform semantic and logical reasoning by designing prompts: analyzing whether the context (e.g., "XXX Bus") of an entity appearing in the "self-evaluation" paragraph explicitly points to a specific work experience when it appears in the context of the text. If such a strong association exists, even if the company name field in the target "work experience" area is empty, it will be determined that the entity "should belong" to the target area, and "entity information supplementation" will be performed, that is, the target entity "XXX Bus" will be backfilled into the employer field of the corresponding work experience. This process realizes information...The automatic aggregation and completion significantly improve the integrity and accuracy of the data.
[0058] Step 403: For the alignment operation of different expressions of the same entity included in the target error correction operation, the expression alignment module in the large model is used to perform hierarchical parsing of each sentence in the first resume information, and the different expressions of the same entity at the same level are determined according to the obtained hierarchical parsing results; Step 404: The expression alignment module is used to unify the different expressions of the same entity into the target expression.
[0059] In steps 403-404, the "alignment operation of different expressions of the same entity" performed by the above-mentioned execution subject is used to solve another common problem: that is, the same objective thing may be referred to in different text forms in different positions of the resume (such as the technology stack "K8S" and "Kubernetes", or the company name "XX Group" and "XX Treasure"). In order to eliminate this ambiguity, the above-mentioned execution subject can call the "expression alignment module" in the large model. This module first performs hierarchical parsing of each sentence in the "first resume information".
[0060] At the technical principle level, this is usually achieved by guiding a large model to perform deep semantic understanding similar to dependency parsing, in order to identify the grammatical modification relationship and the main-subordinate structure between words in a sentence. For example, when parsing "proficient in K8S, familiar with the operation and maintenance of Kubernetes containerization platform", the module can understand that "K8S" and "Kubernetes containerization platform" are in parallel object positions in the sentence, and together serve as the objects of "proficient in" and "familiar with", thus determining that they are "at the same level" and point to the same core technology in the current context. Based on this hierarchical parsing result, the module can accurately "determine different expressions of the same entity at the same level". Finally, the above-mentioned execution entity can guide the module to perform standardized mapping, that is, "unify the different expressions of the same entity into the target expression". In specific practice, the above-mentioned execution entity can pre-set or introduce an industry standard terminology mapping rule through prompt words, such as specifying that "K8S" and "Kubernetes" are unified as "Kubernetes". Based on this rule, the module replaces all identified synonyms with unique standardized terms, thereby ensuring the consistency of terminology throughout the document and removing vocabulary-level obstacles for subsequent content-based accurate matching and evaluation.
[0061] The solution provided in this embodiment systematically repairs the dispersion and ambiguity defects in the original resume information caused by layout and expression habits by sequentially performing the above-mentioned cross-regional compensation and expression alignment operations. This not only enhances the information consistency and logical self-consistency within a single resume, but also makes the information between different resumes comparable, substantially improving the data quality that serves as the cornerstone of subsequent intelligent analysis and evaluation.
[0062] To deepen the understanding of how to use a large model for information completion processing, please also refer to Figure 5. Figure 5 is a flowchart of a method for quantification conversion of fuzzy expressions provided by an embodiment of this disclosure. It aims to provide a specific implementation method for quantification conversion of fuzzy expressions included in the target information completion operation. Its process 500 includes the following steps: Step 501: Use the quantification conversion module in the large model to determine the unstructured fuzzy description information in the part of the first resume information that does not appear in the resume knowledge graph; The input when starting this step includes two key data sources: one is the already constructed "resume knowledge graph", which is a structured network organized in the form of "entity-relationship-entity" triples; the other is the "first resume information" as the original corpus, that is, the text set divided by page area.
[0063] On this basis, the above-mentioned execution subject calls the "quantification conversion module" in the large model. The first step of this module is to perform difference comparison and fuzzy recognition, that is, "determine the unstructured fuzzy description information in the part of the first resume information that does not appear in the resume knowledge graph". In terms of technical principles, the aforementioned execution entity guides the module to compare the structured entities and relationships encoded in the knowledge graph with the original text paragraphs in the "first resume information". Those text fragments that cannot find corresponding nodes or edges in the knowledge graph and whose expression lacks specific verbs, technical terms, or quantitative results will be judged by the module as "fuzzy description information". For example, there may already be a record (candidate, mastery, Python) in the knowledge graph, but in a certain project description, there is a sentence like "I have done some data processing with Python". The phrase "I have done some data processing" is identified as a target fuzzy description because it lacks specific actions, tools, and results details.
[0064] Step 502: Use the quantization conversion module to convert the fuzzy description information into a standardized description using strong verbs based on the STAR rule; After determining the target text, the aforementioned execution entity guides the quantization conversion module to perform the core operation of "converting the fuzzy description information into a standardized description using strong verbs based on the STAR rule". The STAR (Situation, Task, Action, and Result) framework is a structured description of behaviors and outcomes, with its core being the clear definition of situation, task, action, and result. In practice, the executing entity can use the "action verb guidance pool" and "dimensional constraint operator" built into the prompts to command the module to perform the following operations on ambiguous descriptions: First, infer the core action to be reinforced based on the context, replacing "done" with specific verbs such as "designed," "analyzed," or "built"; second, complete the specific technology stack or tools involved in the action, which may be drawn from other parts of the knowledge graph (such as skill nodes) or the same segment.The description is obtained from other sentences; finally, a reasonable and measurable business output or result is inferred or attached to this action. For example, for "doing some data processing with Python", the module may combine the "Pandas" library and "sales report" background mentioned elsewhere in the resume to transform it into "using Python and Pandas library to clean, summarize and trend analyze historical sales data, generate visual weekly reports, and support departmental decision-making". This process is essentially guiding the large model to perform a "semantic instantiation" reasoning from generalized description to specific instance. Manual 13 / 26 pages 17 CN 121481492 A
[0065] Step 503: Use the quantization conversion module to supplement the resume knowledge graph with the standardized description.
[0066] After the conversion is completed, the above-mentioned execution subject will enter the information fusion stage, that is, "supplement the resume knowledge graph with the standardized description". At this time, the newly generated standardized description is no longer a vague text block, but a clear semantic unit containing a clear subject, action, object and output. The aforementioned execution entity instructs the quantification transformation module or a dedicated graph updater to deconstruct and map this new unit onto the knowledge graph schema. For example, from the description "Using Python and Pandas libraries to clean, summarize, and analyze trends in historical sales data...", "Python," "Pandas," "data cleaning," and "trend analysis" can be extracted as skill or activity entities, establishing a "mastered" or "performed" relationship with the candidate entity. Simultaneously, "supported departmental decision-making" can be considered an attribute or output of this relationship. In this way, previously vague text with low value density, existing outside the structured graph, is transformed into a series of new, high-quality knowledge nodes and relationship edges, organically integrated into the ever-expanding resume knowledge graph. This greatly enriches the graph's semantic content and evaluation dimensions, providing a complete and high-quality data foundation for ultimately achieving accurate, multi-dimensional intelligent scoring.
[0067] In this embodiment, steps 501-503 perform deep semantic transformation and integration on the unstructured text in the "first resume information" that is not yet covered in the "resume knowledge graph" generated after the previous error correction process. That is, specifically targeting the colloquial, vague and lacking specific details in resumes, the specialization function of the large language model is used to transform them into standardized and quantifiable expressions, and finally integrate them into a unified knowledge graph structure, thereby completing the final sublimation of candidate information from "raw mentions" to "structured knowledge".
[0068] For step 502, a specific implementation method can be: using the quantization transformation module to determine the contextual supplementary information related to the vague description information in the first resume information; using the quantization transformation module to transform the model according to the contextual supplementary information.The fuzzy description information is supplemented by logical reasoning to obtain the supplemented fuzzy description information; the quantization conversion module is used to convert the supplemented fuzzy description information into a standardized description using strong verbs according to the STAR rule.
[0069] In the initial state of the embodiment, the target "fuzzy description information" to be converted has been clearly defined. The quantization conversion module first determines the contextual supplementary information related to the fuzzy description information in the first resume information. The first resume information is a rich original corpus containing all text identified from various page areas of the resume. The quantization conversion module analyzes the semantics of the sentence or paragraph where the fuzzy description information is located, and then searches for other text fragments related to it in terms of topic, entity or technical field in the entire "first resume information". For example, for the fuzzy description "responsible for backend development", the module may retrieve "Java, Spring Cloud, MySQL" from the skills list area of the resume, and retrieve keywords such as "high concurrency" and "payment system" from other project descriptions. The retrieved information together constitutes valuable "contextual supplementary information", providing factual basis for subsequent reasoning.
[0070] After obtaining contextual clues, the aforementioned execution entity guides the module into the core reasoning stage, that is, to complete the logical reasoning chain of the vague description information based on the supplementary contextual information, and obtain the supplemented vague description information. This operation aims to simulate the thought process of human experts, connecting fragmented clues into a logically complete narrative. In practice, the quantification conversion module will execute a series of internal reasoning steps based on the causal and common-sense reasoning ability inherent in the large model. For example, it will associate "responsible for backend development" (vague action) with "Java, Spring Cloud" (contextual technology) and "payment system" (contextual business), and infer that the candidate's possible specific responsibility is "to develop the backend of the payment system using Java and Spring Cloud framework". This process is essentially to construct a logical reasoning chain from "general role" to "specific technical action" and "clear business object", so that the original vague description is filled with key missing links, forming a "supplemented vague description information" with richer semantics and more certain information.
[0071] Finally, a standardized output transformation is performed, that is, "converting the supplemented fuzzy description information into a standardized description using strong verbs according to the STAR method." At this point, the "supplemented fuzzy description information" already contains the necessary elements, but may still exist in descriptive language. The STAR method provides a structured expression framework that requires the description to include context, task, action, and result. It can force the module to use a predefined "strong verb guide pool" through prompt words (such as replacing "do" with "design")."Reconstruction" and "Improvement"), and organize the language according to the template of "Action + Technology / Tools + Business Goals + Quantifiable Results". For example, the aforementioned inference "using Java and Spring Cloud framework to develop the payment system backend", combined with the business value that may be inferred from the context (such as "supporting transaction peaks"), is transformed into: "Designed and reconstructed a payment backend system based on Java and Spring Cloud microservice architecture, successfully supporting a peak transaction processing of 100,000 transactions per second".
[0072] That is, this embodiment focuses on solving a core and delicate operation in "information completion processing", namely, how to systematically transform a highly generalized and lacking detailed "fuzzy description information" (such as "optimized system performance") into a specific, professional and assessable "standardized description". Through its integrated quantization conversion module (i.e., the functional unit in the large language model specifically designed to perform such tasks), a three-stage deep semantic reasoning and reconstruction process was executed to ensure the accuracy and practicality of the supplementary information. Through the final conversion, a "standardized description" that conforms to professional expression norms and contains quantifiable evaluation dimensions was output, thereby completing the qualitative change from vague colloquial language to precise structured knowledge, greatly improving the feasibility and accuracy of subsequent intelligent evaluation.
[0073] Based on any of the above embodiments, in order to deepen the understanding of how to achieve a comprehensive evaluation of multiple evaluation dimensions based on third resume information, please refer to Figure 6. Figure 6 is a flowchart of a scoring method for different preset evaluation dimensions provided by the embodiments of this disclosure. Its process 600 includes the following steps: Step 601: For professional matching degree, the large model is used to extract skill information from the third resume information and calculate the degree of overlap with the job description information of the target position, and determine the original score of professional matching degree based on the degree of overlap; For "professional matching degree", the above-mentioned execution entity will guide the large model to perform an analysis focused on skill alignment. The aforementioned implementing entity provides the "third-party resume information" (which includes completed and standardized skill descriptions) along with the "job description information for the target position" to the large model. Technically, the large model acts as a semantic matching engine. It first systematically "extracts skill information" from the "third-party resume information," which is typically a structured skill list or technical keywords scattered throughout the job description. Then, the large model semantically "calculates the degree of overlap," meaning it understands the core competencies required for each responsibility in the job description and compares their semantic similarity with the resume skills. It also considers the breadth, depth, and relevance of the skills to the core job requirements. For example, for a "backend development engineer" position, the overlap weight of mastering "Java" and "microservice architecture" is far higher than mastering only "Python scripting."Finally, based on the overall judgment of the matching depth by the large model, it is mapped to a predefined scoring range (such as 0-10 points), thereby "determining the original score of professional matching degree".
[0074] Step 602: Regarding project complexity, the large model is used to extract project experience information, including project scale description, technical depth description and historical role description, from the third resume information, and the original score of project complexity is determined based on the project experience information; Regarding "project complexity", the evaluation focus of the above-mentioned execution subject shifts to the technical content and scale of the candidate's past experience. The above-mentioned execution subject instructs the large model to specifically "extract project experience information" from the "third resume information" and guides the model to examine it from three sub-dimensions: "project scale description" (such as the number of users involved, the amount of data, the team size), "technical depth description" (such as the advancement of the technology used, the technical difficulties solved, the complexity of the architecture design), and "historical role description" (such as serving as the person in charge, core developer or participant). In specific practice, the large model will parse each project description, identify the quantitative indicators and technical terms in it, and infer the complexity level of the project based on the built-in common sense or domain knowledge. For example, a project that "led the design of a distributed system supporting hundreds of millions of traffic" will receive high ratings in all three sub-dimensions. The server comprehensive model analyzes these sub-dimensions and converts them into an original score representing the overall "project complexity".
[0075] Step 603: For clarity of expression, the large model is used to determine the degree of difference between the third resume information and the first resume information, and the structural rationality of the resume content is determined based on the structural information of the first resume information. The original score of clarity of expression is determined based on the degree of difference and structural rationality. For "clarity of expression", the evaluation object of the above-mentioned execution subject uniquely includes the writing and organization quality of the resume document itself. This dimension is achieved through two calculable indicators: "the degree of difference between the third resume information and the first resume information" and "the structural rationality of the first resume information". The former assesses the extent of information completion and transformation. The principle is that a resume with clear original expression requires less completion and the difference between the two is low; conversely, if the original resume (first resume information) is vague and messy, the "third resume information" formed after system transformation will be very different from it. The server calculates this semantic or structural difference through a large model. The latter directly assesses the logical structure of the "first resume information", such as whether the chapters are complete, whether the timeline is clear, and whether the description is orderly. The server combines these two indicators, and the large model judges the resume candidate's logical thinking and communication skills, and outputs the "original score of expression clarity".
[0076] Step 604: For development potential, the large model is used to extract information including educational background, history, etc. from the third resume information.Growth information, including job change trajectory and cross-domain experience, is used to determine the original score of development potential.
[0077] Regarding "development potential," the aforementioned implementing entity performs predictive analysis based on history and trends. The server guides the large model to extract key elements constituting "growth information" from "third resume information," mainly including: "educational background" (such as school and major), "historical job change trajectory" (such as promotion speed and company platform change trends), and "cross-domain experience" (such as whether they possess composite skills or cross-disciplinary project experience). By analyzing these elements, the large model constructs a comprehensive profile of the candidate's growth potential, adaptability, and learning ability. For example, a rapid promotion trajectory from engineer to technical leader, or a background with both technical and product experience, usually indicates high potential. Based on the large model's evaluation of this profile, the server ultimately "determines the original score of development potential."
[0078] Through the above four independent and in-depth scoring process, the aforementioned implementing entity generates multi-angle, interpretable original evaluation data for each resume, providing a solid basis for subsequent weighted synthesis and final decision-making.
[0079] In this embodiment, steps 601-604 perform detailed quantitative analysis on the processed "third resume information" across multiple preset evaluation dimensions, thereby generating an original score for each dimension. This process is a key step in intelligent evaluation. The aforementioned execution entity does not perform a holistic, fuzzy evaluation of the resume, but rather decomposes the comprehensive evaluation into four independently calculable and logically clear sub-tasks—"professional matching degree," "project complexity," "clarity of expression," and "development potential"—through the deep semantic analysis capabilities of a large language model, and performs standardized scoring operations on each sub-task.
[0080] Based on the previous embodiment, when it is found that the original score of any preset evaluation dimension is lower than the pre-set lower limit for the corresponding preset evaluation dimension, a target identifier is attached to the resume score result of the corresponding resume file to be processed, and then the resume file to be processed with the target identifier attached is sent to the target storage location according to a preset sending path.
[0081] The input in the initial state of this embodiment is the calculated original scores of each dimension and the pre-configured "score lower limit" for each dimension. This "lower limit" is a configurable threshold parameter representing the minimum acceptable level of ability in that dimension. For example, for a technical R&D position, the lower limit for "professional matching" might be set at 7 points, while the lower limit for "clarity of expression" might be set at 5 points. These lower limits are typically predefined by domain experts based on the core requirements of the position and stored in the server's configuration system. The core operation performed by the aforementioned execution entity is a logical judgment: sequentially traversing each preset evaluation dimension and checking whether its corresponding original score is strictly "lower" than the score set for that dimension (see page 20 of the instruction manual on page 16 / 26).The lower limit of the score for CN 121481492 A. This judgment process is deterministic; once any dimension fails to meet the requirements, the subsequent processing flow is triggered. When the triggering condition is met—that is, when any dimension's score is lower than its lower limit—the aforementioned executing entity will perform a marking operation on the "resume scoring result" of this "resume file to be processed." "Attaching a target identifier" is the specific manifestation of this operation. Technically, this identifier can be a status label written to a specific field in the database (e.g., status = 'rejected_by_threshold'), or it can be a specific attribute marker attached to the scoring result data structure (e.g., flag: {critical_fail: true, failed_dimension: 'professional matching degree'}). The core function of this identifier is that, in any subsequent processing or display stage, the system or other interfaces can quickly identify that this resume has been eliminated due to violating hard rules, thus eliminating the need to consider its total score or performance in other dimensions.
[0082] After the identification is attached, the aforementioned execution entity immediately performs automated sorting. The "preset sending path" and "target storage location" are predefined processing rules of the system. In practice, this may mean that the server moves or copies all relevant data of the resume (original file, parsed information, scoring results, and elimination identifier) from a database partition or message queue used for "pending processing" or "pending review" to a dedicated storage area used for "eliminated" or "talent pool reserve". Another implementation is to call an internal or external message notification interface to generate a log or notification containing the resume ID and elimination reason, and send it to the designated monitoring system or administrator interface.
[0083] The automated "judgment-marking-routing" process provided in this embodiment ensures that unqualified resumes can be removed from the main processing pipeline in a timely and accurate manner, greatly improving the efficiency of batch processing and the consistency of decision-making, while retaining complete audit trail records. The core logic of the solution provided in this embodiment is to implement a "one-vote veto" mechanism based on a hard threshold. When a candidate has a significant weakness in a certain core competency dimension, regardless of their performance in other aspects, they will usually be directly excluded from further consideration.
[0084] Based on any of the above embodiments, in order to solve the rigidity problem exposed by fixed and absolute scoring thresholds when facing the natural fluctuations in resume quality in different recruitment batches, a dynamic water level based on the relative performance of resumes within the same batch can be introduced. This allows the evaluation system to intelligently adapt to the actual competitive situation of each batch, thereby making a more reasonable and fairer screening decision. A specific implementation method can be: multiple resume files to be processed in the same batch are processed in the same...The upper limit of the score under the preset evaluation dimension is the maximum value of the original score, and the lower limit of the score is the minimum value of the original score.
[0085] That is, the premise for the above-mentioned execution entity to perform this operation is that it has completed the preliminary evaluation of all resumes in the same processing batch. "Processing batch" refers to a set of resumes that the above-mentioned execution entity receives and processes in a specific time period (such as a recruitment cycle or a working day) or for a specific position. The above-mentioned execution entity has calculated the original score of each resume in this batch on the same preset evaluation dimension such as "professional matching degree" and "project complexity". At this time, the above-mentioned execution entity does not immediately compare these original scores with a preset absolute score (such as 7.5 points), but first performs statistical analysis within the batch: for each preset evaluation dimension, the server will traverse the original scores of all resumes in this batch on that dimension and find the maximum and minimum values. According to the rules of this step, this maximum value is defined as the upper limit of the score of that dimension in this batch, and the minimum value is defined as the lower limit of the score. In terms of technical principles, this is essentially performing extreme value statistics before data normalization or standardization. Its purpose is to capture the actual distribution range of the current batch of candidates' performance in this dimension, and to establish a benchmark for subsequent relative evaluation.
[0086] In practical terms, the above-mentioned execution entity maintains a set of scoring data for the resumes in this batch in memory or cache. When a batch of resumes is processed, a post-processing thread will be started to perform a fast aggregation query for each evaluation dimension. For example, for the "project complexity" dimension, it may find that the original scores of the 100 resumes in this batch are distributed between 4.5 and 9.2. Then the above-mentioned execution entity will record 9.2 as the upper limit of the "project complexity" score for this batch and 4.5 as the lower limit of the score. This dynamically determined upper and lower limits replace the fixed threshold and become the new basis for the "one-vote veto" or "merit-based admission" logic in the following specification 17 / 26 pages 21 CN 121481492 A. For example, the system can be configured to "eliminate resumes with scores in the lowest 10% of a certain dimension within the batch," or "select resumes with scores in the highest 20% of a certain dimension within the batch." In this way, the evaluation criteria can "rise with the tide" or "follow the market" as the overall level of the batch increases. For example, in a batch of resumes with a high concentration of experienced candidates, an 8 might only be considered average (close to the lower limit), while in a batch of resumes with relatively little experience, a 7 might be considered top-tier (close to the upper limit).
[0087] This mechanism greatly enhances the system's adaptability and fairness in dealing with markets with varying talent densities, ensuring that the screening results are always based on the relative competitiveness of the current candidate group, rather than absolute scores detached from context. This simulates how a senior recruiter, after reviewing a large number of resumes, will naturally form an "average level" and "excellent threshold" for the candidates in that batch.The intuitive judgment is transformed into an automated, quantifiable decision rule.
[0088] Based on any of the above embodiments, corresponding recommended interview questions can also be generated based on the resume evaluation results. That is, when the above-mentioned execution subject attempts to perform this operation, the input is the already calculated comprehensive score and detailed analysis data of each dimension. The above-mentioned execution subject uses these structured evaluation conclusions to drive its integrated large language model to automatically generate a series of targeted interview questions that match the specific profile of the candidate, thereby extending the output of the screening system from "judgmental score" to "operational tool" to assist interviewers in conducting more efficient and in-depth follow-up examinations.
[0089] From a technical principle perspective, the implementation of this step relies on the powerful conditional text generation and logical reasoning capabilities of the large language model. The aforementioned execution entity uses the "resume evaluation results" (including scores for each dimension and a summary of possible key evidence) as the core context, combined with a dedicated question generation prompt, and inputs them into the large model. This prompt instructs the large model to act as a "senior interviewer" and generate questions based on the following logic: First, it delves deeper into the strengths dimension. For example, if the "project complexity" score is high, it instructs the model to generate questions aimed at exploring the project's technical details, architectural decisions, and challenge-overcoming processes. Second, it verifies and explores the weaknesses dimension. For example, if the "clarity of expression" score is low, it instructs the model to generate questions that require the candidate to clearly retell or summarize a certain experience to assess their on-site communication skills. Finally, it designs forward-looking questions based on the potential dimension. For example, for a high score in "development potential," questions about career planning and willingness to learn new fields can be generated. In this way, the generated interview questions and resume evaluation conclusions form a close logical loop.
[0090] In practical terms, after completing the evaluation, the aforementioned execution entity can start a post-processing thread. This thread serializes the evaluation results into a structured natural language description, such as: "Candidate Zhang San, overall score 85. Strengths: Professional matching (9 / 10), skill stack highly matches the target position; Project complexity (8 / 10), has experience leading medium-sized distributed systems. Dimensions to be examined: Clarity of expression (6 / 10), the original resume description is relatively brief." The aforementioned execution entity concatenates this description with a preset "interview question generation prompt" template to form a complete instruction and sends it to the large model. After receiving the instruction, the large model will output a set of specific questions based on its built-in interview knowledge and understanding of the context, such as: "1. Please elaborate on the biggest technical challenge you encountered in the XX distributed system and the solution you led. 2. Your resume mentions proficiency in Spring Cloud. Please compare Spring Cloud and Dubbo in microservice governance."Similarities and differences, and explain the reasons for your selection in conjunction with your project. 3. Please give me a clear introduction to the project in which you have made the greatest contribution in three minutes.
[0091] Furthermore, the above-mentioned executing entity can also try to enhance the explanatory power of why the evaluation results are given, that is, guide the large model to output "meta-cognition" of the evaluation process itself, that is, not only give a score, but also generate evidence chain and explanatory text to support the score. In specific implementation, the above-mentioned executing entity can add additional instructions to the scoring prompts, requiring the large model to attach a short "scoring basis" while outputting the original score of each dimension. The basis should be directly quoted or summarized from the specific content in the "third resume information". For example, if a score of 9 is given for "professional matching degree", the explanation may be: "The candidate's resume clearly lists all the core technology stacks such as Java, Kafka, and Redis required for the position, And there are application descriptions in both main projects. The above-mentioned implementing entities can package these explanatory texts together with the scoring results and recommended interview questions into a final evaluation report. This enhanced function greatly improves the transparency, credibility and acceptability of the scoring results, enabling recruiters to clearly understand the logic of the system's scoring, rather than facing an untraceable "black box" score, thereby enhancing the practical value and user trust of the entire intelligent screening system.
[0092] To deepen understanding, this disclosure also provides a specific implementation scheme in conjunction with a specific prior art defect: Assume that for the "Senior Backend Development Engineer" recruitment position released by Company A, Company A has designed the following intelligent resume screening system. The intelligent resume screening system is deployed on the company's internal server, and its core workflow is divided into the following three stages: The first stage is multimodal resume parsing. The system receives raw resume files in various formats and, through the integration of Visual Transformer (ViT) and Differentiable Binarization Network (DBNet) layout-aware OCR engine, not only recognizes the text, but also understands its layout structure and hierarchical relationship, and outputs preliminary structured text information with regional labels; The second stage is structured information fusion and enhancement. The system performs in-depth error correction, completion, and standardization on the initially extracted information, and constructs a unified resume knowledge graph. The third stage is large-scale intelligent scoring and decision-making. Based on the completed high-quality structured information, the system performs multi-dimensional quantitative scoring under the drive of the large model, and automatically generates screening conclusions and interview suggestions.
[0093] The following is a detailed description of the system's working process through a specific resume processing case.
[0094] First stage: Multimodal resume parsing and preliminary structuring The system receives a PDF resume submitted by applicant "Zhang San". The resume has a relatively complex layout, with the "Work Experience" and "Project Experience" sections interleaved.
[0095] Page Layout Aware OCR Extraction: The server first converts the PDF into a high-resolution image, and then calls the ViT model to perform global page layout analysis. The ViT model divides the image into multiple small blocks and accurately identifies multiple semantic regions and their boundaries, such as "personal information", "educational background", "work experience", "project experience", and "skills", through a self-attention mechanism. Then, for each identified region, the server calls the DBNet model to perform refined text detection and recognition. DBNet accurately segments the text lines within the region through its differentiable binarization operation, and can correctly process even texts with slight tilt or curvature. Finally, the system outputs a set of texts with region affiliation labels, namely "first resume information". For example, it clearly knows that the two texts "2019.07-2022.08" and "responsible for payment system backend development" belong to the same "work experience" entry region.
[0096] Post-processing and Conflict Arbitration: Since one part of the resume mentions "XX Group" and another part mentions "responsible for XX Bao..." in the project description, the system detects conflicting information about the employer. At this point, the system initiates a conflict resolution logic based on weighted voting: First, it assigns confidence weights to different regions (e.g., the "Work Experience" region has a higher weight than the "Project Description" region); second, it checks whether the entities are associated with credible timestamps ("XX Group" is associated with a clear period of employment from 2019 to 2022); finally, it invokes a large language model as an arbitrator, combining common business knowledge ("XX Treasure" is a business under "XX Group") to normalize the entities, determining that the standardized entity is "XX Group". Simultaneously, the system performs a cross-regional scan using named entity recognition, discovering that the "Self-Evaluation" section mentions "accumulated experience in large-scale system architecture at XXX Group," while a certain company information is missing from the "Work Experience" section. The system automatically fills this information back into the corresponding experience field through semantic association. After this step, the system obtains a "second resume" that corrects the contradiction and completes the missing information.
[0097] Second Stage: Structured Information Fusion and Semantic Enhancement Specification 19 / 26 pages 23 CN 121481492 A The system converts the "second resume information" into a unified "resume knowledge graph" schema. For example, it generates triples: (Zhang San, employed at, XX Group), (Zhang San, proficient in, Java), (XX Group, time period, 2019-2022). However, the graph can only cover structured explicit information and cannot effectively represent vague descriptions in the original resume such as "handled a recommendation system".
[0098] Information completion guided by the prompting project: The system starts the information completion process. It takes the constructed "resume knowledge graph" (as the first input), the original "first resume information" text (as the second input, providing contextual details) and aA carefully designed "completeness prompt" is input into the large language model. The prompt includes: (a) role setting: "You are a senior technical recruitment expert"; (b) action verb guidance: requiring the use of strong verbs such as "design", "development", and "optimization"; (c) output template constraints: must include "[action verb] based on [technology stack] [project], realizing [quantitative indicators] [business value]"; (d) domain knowledge: associated with the skills "Spark" and "collaborative filtering" already identified in the resume.
[0099] Semantic conversion and graph update: Under the guidance of the prompt, the large model performs logical reasoning chain completion on the fuzzy description "finished a recommendation system": identifying the subject (Zhang San), object (recommendation system), associated technology (Spark, collaborative filtering), inferring reasonable actions (development), and speculating business value (personalized distribution). Finally, the standardized description is output: "Based on the Spark platform, a collaborative filtering recommendation system was developed, realizing accurate distribution of personalized content, and increasing the click-through rate by 15%". The system then deconstructs this new description, using "Spark," "collaborative filtering," and "recommendation system" as new entities or attributes, updating and enriching the original "resume knowledge graph," forming a complete and standardized "third resume information."
[0100] Third stage: The large model intelligent scoring and automated decision-making system now evaluates candidate Zhang San based on the high-quality "third resume information." The target position is "senior backend development engineer," and its job description (JD) emphasizes experience in high-concurrency and distributed systems.
[0101] Multi-dimensional quantitative scoring: The system inputs the "third resume information" and a set of "scoring prompts" into the large model. The scoring prompts define four core dimensions and their weights (e.g., professional matching degree 0.4, project complexity 0.3, expression clarity 0.2, development potential 0.1), and provide detailed scoring criteria for each dimension (e.g., 0-10 scale and corresponding descriptions). The large model acts as the evaluator, analyzing each dimension: In terms of "professional matching degree", it compares Zhang San's skill stack (Java, Spring Cloud, Redis, etc.) with the job requirements and gives a high score; in terms of "project complexity", it analyzes the scale description and technical depth of his "payment system" project and gives a relatively high score; in terms of "clarity of expression", it compares the information difference before and after completion and evaluates the structural logic of the original resume; in terms of "development potential", it analyzes his education and career advancement trajectory. Assume the original output score is [9, 8, 7, 8].
[0102] Comprehensive calculation and dynamic threshold judgment: The server performs weighted calculation according to preset weights to obtain a comprehensive score. At the same time, the system enables the "dynamic water level" strategy: 100 resumes for the same position have been processed this week, and the system automatically calculates the distribution of scores for each dimension in this batch. Assume the lowest score (lower limit of score) for "professional matching degree" in the batch is 6 points, and the highest score (upper limit of score) is 6 points.9.5 points. Zhang San scored 9 points in this dimension, which is among the top in the batch. The system sets the admission line based on the urgency of the recruitment (e.g., "urgent") as "professional matching degree" being greater than the 85th percentile of the batch (assumed to be 8.5 points). Zhang San's 9 points exceed this line.
[0103] Decision generation and derived output: Since Zhang San's scores in all dimensions meet the standard and his comprehensive score is high, the system marks him as "recommended for interview". In addition, the system automatically generates targeted recommended interview questions, such as: "Please elaborate on your design scheme for handling high-concurrency transactions in the XX Group payment system and the challenges you encountered". This question is directly derived from the specific content of his high-scoring dimension "project complexity". Meanwhile, the system can add explanatory notes to the scoring results, such as "Professional matching score 9 points: The candidate's core skills in Java, Spring Cloud, Redis, etc., are a perfect match with the JD requirements and have been applied in relevant large-scale projects." As seen in the above examples, this system effectively overcomes information extraction errors in miscellaneous PDFs / scanned documents by using ViT+DBNet combined with layout-aware OCR; it transforms ambiguous expressions into standardized semantics that can be quantified and evaluated by integrating knowledge graphs and prompting engineering into a large model for completion; and it achieves objective, interpretable, and batch-adaptable intelligent evaluation through a large model scoring based on multi-dimensional prompt word design and dynamic thresholds. Ultimately, the system can automatically output a clearly structured and accurate competency assessment report and interview action guide from a poorly formatted and ambiguous original resume, significantly improving the efficiency, accuracy, and fairness of the recruitment process.
[0104] Based on any of the above embodiments, in an embodiment that optimizes system performance, the execution subject not only performs single inference but also deploys a multi-task joint training framework to achieve continuous learning. That is, the execution subject can prepare a batch of training data during initialization or periodic updates. Each piece of data includes: the original resume file, entity information that has been manually verified and corrected (as annotation for the information extraction task), and the comprehensive score and interview decision given by senior HR (as annotation for the scoring task). The server is designed with a neural network architecture that shares the underlying representation, in which the OCR post-processing module (responsible for entity recognition and normalization) and the large model scoring module (responsible for dimensional scoring) share the intermediate layer encoding of resume semantics. During training, the execution subject calculates a composite loss function: this function simultaneously includes the cross-entropy loss of information extraction (measuring the accuracy of entity recognition) and the mean squared error loss of scoring prediction (measuring the degree to which the score is close to the manual annotation). By optimizing both tasks simultaneously through backpropagation, the shared encoding layer can learn feature representations that are of common value for resume screening, thereby significantly improving the accuracy and generalization ability of both modules and achieving a collaborative optimization effect where the more accurate the error correction, the more precise the scoring.
[0105] In another embodiment aimed at improving the model's discriminative power, the aforementioned execution entity can also utilize a contrastive learning strategy for self-supervised or weakly supervised training. The execution entity can first construct specific sample pairs: two resumes of the same candidate before and after polishing for the same position are considered positive sample pairs; a resume submitted for "Java Backend Engineer" and a resume submitted for "Marketing Specialist" are considered negative sample pairs. During model training, the execution entity inputs the "third resume information" of the resumes into an encoder, converting it into a feature vector. The design goal of the contrastive learning loss function is to reduce the distance between the feature vectors of positive sample pairs while increasing the distance between the feature vectors of negative sample pairs. Specifically, this loss function (such as the InfoNCE loss) is integrated into the aforementioned composite loss function as an important component. Furthermore, the composite loss function explicitly includes: scoring regression loss (calculating the mean squared error between the model's predicted score and the human score) and classification prediction loss (calculating the cross-entropy between the model's "suggested interview" probability and the actual interview result). The execution entity can combine these three loss terms through weighted summation and use this total loss to guide the update of the model parameters. This mechanism enables the system to learn from a large amount of resume data that does not explicitly label specific scores but whose job affiliation and relative quality are known, and to continuously evolve its discrimination criteria.
[0106] Referring further to FIG7, as an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a resume processing device based on a large model. This device embodiment corresponds to the method embodiment shown in FIG2, and the device can be specifically applied to various electronic devices.
[0107] As shown in FIG7, the resume processing device 700 based on a large model in this embodiment may include: a first resume information recognition unit 701, an error correction processing and knowledge graph generation unit 702, an information completion unit 703, and a multi-dimensional evaluation unit 704. The system includes a first resume information recognition unit 701, configured to identify structured first resume information from the resume file to be processed; an error correction processing and knowledge graph generation unit 702, configured to input the first resume information and preset error correction prompts as prompts into a preset large model for error correction processing, and generate a resume knowledge graph from the output structured second resume information in the form of a knowledge graph; wherein the error correction processing includes arbitration judgment operations for conflicting information; an information completion unit 703, configured to input the resume knowledge graph, the first resume information, and preset information completion prompts as prompts into a large model for information completion processing, obtaining structured third resume information; wherein the information completion processing includes quantification conversion operations for fuzzy expressions; and a multi-dimensional evaluation unit 704, configured to evaluate the third resume information in multiple dimensions.The resume file to be processed is evaluated on a preset evaluation dimension to obtain the resume evaluation result.
[0108] In this embodiment, the specific processing of the first resume information recognition unit 701, the error correction and knowledge graph generation unit 702, the information completion unit 703, and the multi-dimensional evaluation unit 704 in the large model-based resume processing device 700 and the resulting technical effects can be referred to the relevant descriptions of steps 201-204 in the corresponding embodiment of Figure 2, which will not be repeated here.
[0109] In some optional implementations of this embodiment, the first resume information recognition unit 701 is further configured to: use a visual Transformer model to identify different page areas in the resume file to be processed; use a differentiable binarization network to determine the corresponding text in the different page areas; and generate structured first resume information according to the membership relationship between the text in each area and its respective page area.
[0110] In some other optional implementations of this embodiment, the error correction processing and knowledge graph generation unit 702 includes an error correction processing subunit configured to input the first resume information and preset error correction prompts as prompt information into a preset large model for error correction processing. The error correction processing subunit includes: a target error correction operation determination module, configured to use the large model to determine the target error correction operation to be performed based on the error correction operation description information contained in the error correction prompts; and an error correction operation execution module, configured to use the large model to perform the target error correction operation on the first resume information.
[0111] In some other optional implementations of this embodiment, the error correction processing subunit may further include: a first role positioning control module, configured to control the large model to position itself in the first role corresponding to the first role positioning information in response to the error correction prompt word further including first role setting information and first domain knowledge constraint information; an error correction constraint information determination module, configured to control the large model to determine the error correction constraint information to be followed when performing the target error correction operation based on the first domain knowledge constraint information; correspondingly, the error correction operation execution module is further configured to: control the large model to position itself in the first role and perform the target error correction operation on the first resume information under the constraint of the error correction constraint information.
[0112] In some other optional implementations of this embodiment, the error correction operation execution module includes: an arbitration judgment operation execution submodule, configured to, for the arbitration judgment operation on conflict information included in the target error correction operation, use the arbitration judgment module in the large model to determine multiple candidate entity information that are semantically similar and identified from different page areas in the first resume information; use the arbitration judgment module to determine the page area weight based on the confidence level of the page area to which each candidate entity information belongs; use the arbitration judgment module to determine the weight adjustment strategy based on whether each candidate entity information is associated with time information; wherein, the weight adjustment strategy includes: increasing the weight of the page area associated with time information.The weight of candidate entity information is adjusted, and the weight of candidate entity information that is not associated with time information is reduced. The arbitration judgment module determines the comprehensive weight of each candidate entity information according to the weight of the page area to which each candidate entity information belongs and the weight adjustment strategy. The arbitration judgment module determines the arbitration judgment result of the target candidate entity information as conflict information according to the comprehensive weight of each candidate entity information.
[0113] In some other optional implementations of this embodiment, the error correction process also includes a compensation operation for cross-region information. The error correction operation execution module includes: a compensation operation execution submodule, which is configured to identify each entity in the first resume information using the cross-region compensation module in the large model for the compensation operation for cross-region information included in the target error correction operation; determine the correlation between each identified entity and different page areas using the cross-region compensation specification 22 / 26 page 26 CN 121481492 A module, and supplement the original part of the first resume information of the target page area with entity information according to the correlation.
[0114] In some other optional implementations of this embodiment, the error correction process further includes alignment operations for different representations of the same entity. The error correction operation execution module includes: an alignment execution submodule, configured to perform hierarchical parsing of each sentence in the first resume information using the representation alignment module in the large model, and determine different representations of the same entity at the same level based on the obtained hierarchical parsing results; and use the representation alignment module to unify the different representations of the same entity into the target representation.
[0115] In some other optional implementations of this embodiment, the information completion unit 703 includes: a target information completion operation determination subunit, configured to determine the target information completion operation to be performed based on the information completion operation description information contained in the information completion prompt words using the large model; and a target information completion operation execution subunit, configured to perform the target information completion operation on the resume knowledge graph and the first resume information using the large model.
[0116] In some other optional implementations of this embodiment, the information completion unit 703 may further include: a second role positioning control subunit, configured to control the large model to be positioned in the second role corresponding to the second role setting information in response to the information completion prompt word also including second role setting information and second domain knowledge constraint information; an information completion constraint information determination subunit, configured to control the large model to determine the information completion constraint information to be followed when performing the target information completion operation based on the second domain knowledge constraint information; wherein, the information completion constraint information includes: the output format template of the completed information; correspondingly, the target information completion operation execution subunit is further configured to: control the large model to be positioned in the second role, and under the constraint of the information completion constraint information, perform the operation on the resume knowledge graph and the second role setting information;A target information completion operation is performed on the resume information.
[0117] In some other optional implementations of this embodiment, the target information completion operation execution subunit includes: a fuzzy description information determination module, configured to determine unstructured fuzzy description information in the first resume information that does not appear in the resume knowledge graph by using the quantization conversion module in the large model for the fuzzy description quantization conversion operation included in the target information completion operation; a standardized description conversion module, configured to convert the fuzzy description information into a standardized description using strong verbs by using the quantization conversion module based on the STAR rule; and a supplementation module, configured to supplement the resume knowledge graph with the standardized descriptions by using the quantization conversion module.
[0118] In some other optional implementations of this embodiment, the standardized description conversion module is further configured to: determine contextual supplementary information related to the fuzzy description information in the first resume information by using the quantization conversion module; complete the logical reasoning chain of the fuzzy description information by using the quantization conversion module based on the contextual supplementary information to obtain the supplemented fuzzy description information; and convert the supplemented fuzzy description information into a standardized description using strong verbs by using the quantization conversion module according to the STAR rule.
[0119] In some other optional implementations of this embodiment, the multi-dimensional evaluation unit 704 includes: a subunit for outputting original scores for each evaluation dimension, configured to input the third resume information and the preset scoring prompt words (page 23 / 26, CN 121481492 A) as prompt information into the large model to obtain the original scores output for each preset evaluation dimension; wherein, the scoring prompt words include multiple preset evaluation dimensions, including: professional matching degree, project complexity, expression clarity and development potential; and a weighted processing unit, configured to use the large model to weight the original scores of each preset evaluation dimension according to the corresponding preset weights to obtain the resume score result, which is the sum of the weighted scores.
[0120] In some other optional implementations of this embodiment, the original score output subunits of each evaluation dimension are further configured as follows: For professional matching degree, a large model is used to extract skill information from the third resume information, calculate the degree of overlap between the skill information and the job description information of the target position, and determine the original score of professional matching degree based on the degree of overlap; For project complexity, a large model is used to extract project experience information, including project scale description, technical depth description and historical job role description, from the third resume information, and determine the original score of project complexity based on the project experience information; For clarity of expression, a large model is used to determine the degree of difference between the third resume information and the first resume information, and the structural rationality of the resume content is determined based on the structural information of the first resume information, and the original score of clarity of expression is determined based on the degree of difference and structural rationality.Regarding development potential, a large model is used to extract growth information, including educational background, historical job change trajectory, and cross-domain experience, from the third resume information, and the original score of development potential is determined based on the growth information.
[0121] In some other optional implementations of this embodiment, the resume processing device 700 based on the large model may further include: an identifier attaching unit, configured to attach a target identifier to the resume scoring result of the corresponding resume file to be processed in response to the existence of an original score of any preset evaluation dimension being lower than a preset score lower limit for the corresponding preset evaluation dimension; and a screening and sending unit, configured to send the resume file to be processed with the target identifier attached to a preset sending path to a target storage location.
[0122] In some other optional implementations of this embodiment, the upper limit of the scores of multiple resume files to be processed in the same processing batch under the same preset evaluation dimension is the maximum value of the original score, and the lower limit of the scores is the minimum value of the original score.
[0123] In some other optional implementations of this embodiment, the resume processing device 700 based on the large model may further include: a recommended interview question generation unit, configured to generate corresponding recommended interview questions based on the resume evaluation results.
[0124] This embodiment exists as a device embodiment corresponding to the above method embodiment. The resume processing device based on the large model provided in this embodiment effectively overcomes the two major bottlenecks of inaccurate original information and semantic ambiguity through two targeted applications (error correction and completion) of the large model in key links. Finally, it realizes the simultaneous improvement of the accuracy, automation level and evaluation depth of resume screening, and also realizes the automated and intelligent processing from the original resume file to the final evaluation result.
[0125] According to the embodiments of this disclosure, this disclosure also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to at least one processor; wherein, the memory stores instructions that can be executed by at least one processor, which are executed by at least one processor to enable the resume processing method based on the large model described in any of the above embodiments when the at least one processor is executed.
[0126] According to the embodiments of this disclosure, this disclosure also provides a readable storage medium, which stores computer instructions, which are used to enable a computer to execute and implement the resume processing method based on the large model described in any of the above embodiments.
[0127] According to embodiments of the present disclosure, the present disclosure also provides a computer program product that, when executed by a processor, can implement the large-model-based resume processing method described in any of the above embodiments.
[0128] FIG8 shows a schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure.The device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0129] As shown in FIG8, the device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the device 800 may also be stored in the RAM 803. The computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0130] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0131] Computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 801 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 801 performs the various methods and processes described above, such as resume processing methods based on large models. For example, in some embodiments, the large-model-based resume processing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by computing unit 801, one or more steps of the large-model-based resume processing method described above can be performed. Alternatively, in other embodiments, computing unit 801 can...It can be configured to perform a large-model-based resume processing method by any other suitable means (e.g., by means of firmware).
[0132] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), payload programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transfer data and instructions to the storage system, the at least one input device, and the at least one output device.
[0133] Program code for implementing the methods of this disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a standalone software package, or entirely on a remote machine or server.
[0134] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0135] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) for displaying information to the user.The system and technology described herein can be implemented in a computing system that includes a background component (e.g., a data server), a middleware component (e.g., an application server), a front-end component (e.g., a user computer with a graphical user interface or web browser through which the user can interact with the implementation of the system and technology described herein), or any combination of such background, middleware, or front-end components. Components of the system can be interconnected by digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0137] The computer system may include clients and servers. Clients and servers are generally far apart and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of traditional physical hosts and virtual private servers (VPS, Virtual Private Server) services, such as high management difficulty and weak business scalability.
[0138] According to the technical solution of the embodiments of this disclosure, by applying the large model twice in key links (error correction and completion), the two major bottlenecks of inaccurate original information and semantic ambiguity are effectively overcome, and the accuracy, automation and evaluation depth of resume screening are improved simultaneously. It also realizes the automated and intelligent processing from the original resume file to the final evaluation result.
[0139] It should be understood that various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially or in different orders, as long as the expected results of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0140] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, combinations, sub-combinations, and substitutions made under this disclosure are not permitted.Modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall all be included within the scope of protection of this disclosure. Instruction Manual 26 / 26 Page 30 CN 121481492 A Figure 1 Figure 2 Instruction Manual Drawing 1 / 5 Page 31 CN 121481492 A Figure 3 Instruction Manual Drawing 2 / 5 Page 32 CN 121481492 A Figure 4 Figure 5 Instruction Manual Drawing 3 / 5 Page 33 CN 121481492 A Figure 6 Figure 7 Instruction Manual Drawing 4 / 5 Page 34 CN 121481492 A Figure 8 Instruction Manual Drawing 5 / 5 Page 35 CN 121481492 A Abstract This invention provides a resume processing method based on a large model and electronic equipment, and relates to the technical field of large models and agents. The method comprises the following steps: identifying structured first resume information from a resume file to be processed; inputting the first resume information and the error correction prompt words as prompt information into a large model for error correction processing, and generating a resume knowledge map from the output structured second resume information, wherein the error correction processing includes arbitration judgment operation on conflicting information; Inputting the resume knowledge map, the first resume information and the informationcompletion prompt words together as prompt information into a large model for information completion processing, and obtaining structured third resume information, wherein the information completion processing includes the quantitative conversion operation of fuzzy expressions; Based on the third resume information, the resume files to be processed are evaluated in a plurality of preset evaluation dimensions to obtain resume evaluation results. This method can automatically extract accurate resume information and conduct multi-dimensional analysis.
Claims
1. A resume processing method based on a large model, characterized in that, include: Identify the structured first resume information from the resume files to be processed; The first resume information and the preset error correction prompts are input into the preset large model as prompt information for error correction processing. The structured second resume information output is used to generate a resume knowledge graph in the form of a knowledge graph. The error correction processing includes arbitration and judgment operations for conflicting information. The large model is used to determine the information completion operation to be performed based on the information completion operation description information contained in the preset information completion prompt words; for the quantization conversion operation of fuzzy expression contained in the information completion operation, the quantization conversion module in the large model is used to determine the unstructured fuzzy description information in the part of the first resume information that does not appear in the resume knowledge graph; The quantization conversion module uses the STAR method to convert the fuzzy description information into a standardized description using strong verbs. The standardized description is supplemented into the resume knowledge graph using the quantization conversion module to obtain structured third resume information; Based on the third resume information, the resume file to be processed is evaluated on multiple preset evaluation dimensions to obtain the resume evaluation result.
2. The method according to claim 1, characterized in that, The process of identifying structured first resume information from the resume file to be processed includes: The visual Transformer model is used to identify different layout areas in the resume file to be processed. Differentiable binarization networks are used to determine the text within different page layout regions; Based on the hierarchical relationship between the text in each region and its corresponding page area, structured first resume information is generated.
3. The method according to claim 1, characterized in that, The step of inputting the first resume information and preset error correction prompts together as prompt information into a preset large model for error correction processing includes: The large model is used to determine the target error correction operation to be performed based on the error correction operation description information contained in the error correction prompt words; For the arbitration and determination operation on conflict information included in the target error correction operation, the arbitration and determination module in the large model is used to determine multiple candidate entity information with similar semantics that are identified from different page areas in the first resume information. The arbitration determination module uses the confidence level of the page area to which each candidate entity information belongs to determine the page area weight. The arbitration determination module uses the time information associated with each candidate entity to determine a weight adjustment strategy; wherein the weight adjustment strategy includes: increasing the weight of candidate entity information associated with time information and decreasing the weight of candidate entity information not associated with time information. The arbitration determination module uses the page area weight and weight adjustment strategy to determine the comprehensive weight of each candidate entity information. The arbitration determination module uses the comprehensive weight of each candidate entity information to determine the target candidate entity information as the arbitration determination result of the conflict information.
4. The method according to claim 3, characterized in that, The step of inputting the first resume information and preset error correction prompts together as prompt information into a preset large model for error correction processing also includes: For the compensation operation on cross-regional information included in the target error correction operation, the cross-regional compensation module in the large model is used to identify each entity in the first resume information; the cross-regional compensation module is used to determine the correlation between each identified entity and different page areas, and based on the correlation, the target entities that should belong to the target page area are used to supplement the original part of the first resume information in the target page area with entity information; or For the alignment operation of different representations of the same entity included in the target error correction operation, the representation alignment module in the large model is used to perform hierarchical parsing on each sentence in the first resume information, and different representations of the same entity at the same level are determined based on the obtained hierarchical parsing results; the representation alignment module is used to unify the different representations of the same entity into the target representation.
5. The method according to claim 1, characterized in that, The process of using the quantization conversion module to convert the fuzzy description information into a standardized description using strong verbs based on the STAR rule includes: The quantization conversion module is used to determine contextual supplementary information related to the fuzzy description information in the first resume information; The quantization conversion module completes the logical reasoning chain of the fuzzy description information based on the context supplementary information to obtain the supplemented fuzzy description information; The quantization conversion module is used to convert the supplemented fuzzy description information into a standardized description using strong verbs according to the STAR rule.
6. The method according to any one of claims 1-5, characterized in that, The process of evaluating the resume file to be processed based on the third resume information across multiple preset evaluation dimensions to obtain resume evaluation results includes: The third resume information and the preset scoring prompts are input into the large model as prompt information to obtain the original scores output for each preset evaluation dimension; wherein, the scoring prompts include multiple preset evaluation dimensions, including: professional matching degree, project complexity, expression clarity and development potential. The original scores of each preset evaluation dimension are weighted according to the corresponding preset weights using the large model, and the resume score result is obtained by summing the weighted scores.
7. The method according to claim 6, characterized in that, The third resume information and preset rating prompts are input into the large model as prompt information to obtain the original scores output for each preset evaluation dimension, including: Regarding the professional matching degree, the large model is used to extract skill information from the third resume information, calculate the degree of overlap between the skill information and the job description information of the target position, and determine the original score of professional matching degree based on the degree of overlap. Regarding the project complexity, the large model is used to extract project experience information, including project scale description, technical depth description, and historical role description, from the third resume information, and the original score of project complexity is determined based on the project experience information. Regarding the clarity of expression, the large model is used to determine the degree of difference between the third resume information and the first resume information, and the structural rationality of the resume content is determined based on the structural information of the first resume information. The original score of the clarity of expression is determined based on the degree of difference and the structural rationality. Based on the development potential, the large model is used to extract growth information, including educational background, historical job change trajectory and cross-domain experience, from the third resume information, and an original score for development potential is determined based on the growth information.
8. The method according to claim 7, characterized in that, Also includes: In response to the existence of an original score for any of the preset evaluation dimensions being lower than the lower limit of the score set for the corresponding preset evaluation dimension, a target identifier is appended to the resume scoring result of the corresponding resume file to be processed. The resume file to be processed, which includes the target identifier, is sent to the target storage location via a preset sending path.
9. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the resume processing method based on a large model as described in any one of claims 1-8.