Resume analysis method, resume analysis device and computer storage medium
By converting resume documents into image format and utilizing geometric relationship models and text classification models, the problem of disordered text order in resume documents was solved, achieving highly accurate structured information extraction.
Patent Information
- Application Number
- CN202410540974.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing resume parsing methods are prone to text order disorder when dealing with diverse and complex resume documents, leading to inaccurate extraction of structured fields.
By acquiring the resume document to be parsed and converting it into image format, a pre-trained resume document geometric relationship model is used to extract text block association groups, combine identical text blocks, obtain structural tags based on geometric layout relationships, and use a text classification model to extract entities and obtain parsed entities.
It effectively avoids the problem of text line disorder, improves the accuracy and consistency of extracting structured information from resume documents, and reduces the problems of insufficient generalization and low accuracy caused by keyword matching and regular expression rules.
Smart Images

Figure CN120874807A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of text processing technology, and in particular to a resume parsing method, a resume parsing device, and a computer storage medium. Background Technology
[0002] In human resource recruitment, resume documents in formats such as HTML, DOC, DOCX, PDF, and images have become the mainstream. Resume parsing refers to extracting structured field information such as basic information, work experience, project experience, and education from resume documents. Resume parsing is a crucial foundation for higher-level applications such as resume screening, searching, and job matching.
[0003] Text extracted from resume documents often exhibits an issue of disordered text order. Typical text extraction tools scan resume documents from top to bottom and left to right to extract text. However, resume documents come in various formats and have complex layouts, with many being two-column or even multi-column layouts. In such cases, disordered text order directly affects the extraction of structured fields. Summary of the Invention
[0004] To address the aforementioned technical problems, this application proposes a resume parsing method, a resume parsing device, and a computer storage medium.
[0005] To address the aforementioned technical problems, this application proposes a resume parsing method, which includes:
[0006] Retrieve the resume document to be parsed;
[0007] The resume document is input into a pre-trained resume document geometric relationship model to extract the text block association groups of the resume document;
[0008] By associating different text blocks that have the same text blocks, we can obtain text fragments;
[0009] The structural labels of each text block are obtained based on the geometric layout relationship of each text association group;
[0010] The text fragments are categorized according to the structural tags;
[0011] Entity extraction is performed on the text fragments according to each structural tag to obtain the parsed entities of the resume document.
[0012] The structural tags include titles, subheadings, body text, and / or others.
[0013] The step of extracting entities from the text fragments according to each structural tag to obtain the parsed entities of the resume document includes:
[0014] The text fragments are classified using a pre-trained text classification model to obtain semantic labels for each text fragment;
[0015] Entity extraction is performed from each structural tag of the text fragment according to the semantic tags to obtain the parsed entities of the resume document.
[0016] The semantic tags include basic information, work experience, education experience, project experience, personal evaluation and / or others.
[0017] The geometric layout relationship of the text block association group includes the relationship between several characters and / or the relationship between several text blocks.
[0018] The resume parsing method further includes:
[0019] Obtain the resume document to be trained;
[0020] Based on the resume document to be trained, input tags are obtained, wherein the input tags include text block association relationships, text block labels, and text block coordinates;
[0021] The input labels and the resume document to be trained are input into the resume document geometric relationship model for training.
[0022] The input label also includes each character in the text block and its coordinates.
[0023] The resume parsing method further includes, after obtaining the resume document to be parsed, the method further includes:
[0024] Convert resume documents in different formats into image-based resume documents.
[0025] To address the aforementioned technical problems, this application also proposes a resume parsing device, which includes a resume acquisition module, a geometric relationship module, a fragment combination module, and a resume parsing module; wherein,
[0026] The resume acquisition module is used to acquire the resume document to be parsed;
[0027] The geometric relationship module is used to input the resume document into a pre-trained resume document geometric relationship model and extract the text block association groups of the resume.
[0028] The fragment combination module is used to combine different text block association groups that have the same text block to obtain text fragments;
[0029] The resume parsing module is used to obtain the structural tags of each text block based on the geometric layout relationship of each text association group; classify each text fragment according to the structural tags; and extract entities from the text fragments according to each structural tag to obtain the parsed entities of the resume document.
[0030] To address the aforementioned technical problems, this application also proposes a resume parsing device, which includes a memory and a processor coupled to the memory; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the resume parsing method described above.
[0031] To address the aforementioned technical problems, this application also proposes a computer storage medium for storing program data, which, when executed by a computer, is used to implement the aforementioned resume parsing method.
[0032] Compared with existing technologies, the beneficial effects of this application are as follows: the resume parsing method obtains the resume document to be parsed; the resume document is input into a pre-trained resume document geometric relationship model to extract the text block association groups of the resume document; different text block association groups with the same text blocks are combined to obtain text fragments; the structural labels of each text block are obtained based on the geometric layout relationship of each text association group; each text fragment is classified according to the structural labels; entity extraction is performed on the text fragments according to each structural label to obtain the parsed entities of the resume document. Through the above resume parsing method, the geometric layout relationship of the resume document is established, and the relationship between each block in the resume document is associated, which can avoid the problem of text line disorder in text extraction in existing resume parsing schemes. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] in:
[0035] Figure 1 This is a flowchart illustrating an embodiment of the resume parsing method provided in this application;
[0036] Figure 2 This is a flowchart illustrating the overall process of the resume parsing method provided in this application;
[0037] Figure 3 This is a schematic diagram illustrating the process of establishing the layout relationships of the resume document provided in this application;
[0038] Figure 4 This is a schematic diagram of the input and output format of the geometric relationship model of the resume document provided in this application;
[0039] Figure 5 This is a schematic diagram of the structure of an embodiment of the resume parsing device provided in this application;
[0040] Figure 6 This is a schematic diagram of another embodiment of the resume parsing device provided in this application;
[0041] Figure 7 This is a schematic diagram of the structure of an embodiment of the computer storage medium provided in this application. Detailed Implementation
[0042] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0043] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0044] Existing resume parsing methods mainly include two steps:
[0045] 1. Text extraction.
[0046] Use different text extraction tools to extract text, as well as rich text information such as text size, position, and color, from resume documents in formats such as HTML, DOC, DOCX, PDF, and images. For example, use pdfminer to extract text from PDF files and use OCR tools to extract text from image files.
[0047] 2. Extraction of structured fields.
[0048] Using text analysis and mining methods, structured fields are extracted from the text obtained in step 1. This step can usually be further divided into two sub-steps: First, the text extracted in step 1 is divided into blocks using dictionaries, regular expression rules, or machine learning methods; then, target fields are extracted from different text blocks using tools such as rules, dictionaries, or named entity recognition models.
[0049] Existing resume parsing methods suffer from insufficient generalization in their resume block segmentation approaches and are prone to errors. Most existing methods rely on commonly used title dictionaries and regular expression rules for resume block segmentation. However, personalized formatting and diverse textual expressions are common in resume documents, rendering manually compiled and maintained dictionaries and rules ineffective.
[0050] Existing resume analysis methods often suffer from unclear target text boundaries when extracting structured fields, leading to the extraction of incorrect fields. The work experience section of a resume includes company names, employment periods, and job responsibilities. Job responsibilities may include company names, such as "worked in cross-border e-commerce, Amazon e-commerce operations." Using dictionaries or named entity recognition models, "Amazon" might be extracted as the company name.
[0051] Therefore, taking advantage of the distinct layout of resume documents, a resume parsing method combining geometric layout relationships is proposed.
[0052] Please refer to details. Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating an embodiment of the resume parsing method provided in this application. Figure 2 This is a schematic diagram of the overall process of the resume parsing method provided in this application.
[0053] The resume parsing method of this application is applied to a resume parsing device, which can be a server, a terminal device, or a system in which the server and the terminal device cooperate with each other. Accordingly, the various parts of the resume parsing device, such as each unit, subunit, module, and submodule, can all be set in the server, all in the terminal device, or separately in the server and the terminal device.
[0054] Furthermore, the aforementioned server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software modules, such as software or software modules used to provide distributed server functionality, or as a single software program or software module; no specific limitations are made here.
[0055] like Figure 1 As shown, the specific steps are as follows:
[0056] Step S11: Obtain the resume document to be parsed.
[0057] In this embodiment of the application, after the resume parsing device obtains the resume document to be parsed, it needs to standardize the resume document format.
[0058] Specifically, the resume parsing device of this application selects to convert resume documents of different formats into image formats. For example, for doc and docx formats, it converts them into pdf format using the Python toolkit win32com, and then converts the pdf file into image format using the Python toolkit (fitz, etc.); for pdf format, it directly uses the Python toolkit (fitz, etc.) to convert the pdf file into image format; for html format, it converts the html file into image format using the Python toolkit imgkit.
[0059] Furthermore, the resume document format conversion tools involved in this application are not limited to Fitz, IMGKit, Win32Com, etc.
[0060] Step S12: Input the resume document into the pre-trained resume document geometric relationship model and extract the text block association groups of the resume document.
[0061] In this embodiment, the resume parsing device inputs the resume document after unifying the resume document format into a pre-trained resume document geometric relationship model to extract the text block association groups in the resume document.
[0062] Specifically, the geometric relationship model for resume documents involved in this application can employ models such as GeolayoutLM. During training, the geometric relationship model learns the geometric relationships between resume text fragments, including but not limited to: relative positional relationships, directional relationships, and collinear relationships. By learning the geometric layout relationships of resume documents, the layout relationships of resume documents can be effectively managed.
[0063] Furthermore, the geometric relationship model for resume documents involved in this application is not limited to GeolayoutLM, but can be other multimodal layout models, such as LayoutLMv3, combined with geometric relationships. The predefined set of structural tags is not limited to {title, subtitle, body text, others}, and can be adjusted according to the specific resume layout.
[0064] Please refer to details. Figure 3 , Figure 3 This is a schematic diagram illustrating the process of establishing the layout relationships of the resume document provided in this application.
[0065] like Figure 3As shown, the category label set of the resume document geometric relationship model is a predefined label set that can cover different components in the resume document. For example, the predefined label set S = {title, subtitle, body text, others}.
[0066] The text block association group output in step S12 may contain multiple category labels. Specifically, the input and output of the resume document geometric relationship model are both TextBoxList = [TB1, TB2, ..., TBi, TBn]. For details on the format of TBi, please refer to [link to relevant documentation]. Figure 4 , Figure 4 This is a schematic diagram of the input and output format of the geometric relationship model of the resume document provided in this application.
[0067] like Figure 4 In the given example, the text block association group includes the text "Work Experience" and the text "xxx Company". In the input / output format of the text block association group, `from` and `linking` indicate a link from text block 0 to text block 1, `id` represents the text block number, `text` represents the text of the text block, `box` represents the coordinate information of the text block, `label` represents the category of the text block, and the `words` field represents the coordinate information of each character block and character within the text block. In this application, `label` ∈ S, and `label` is called the structural label of the text block.
[0068] The category label for the text "Work Experience" is "Title", and the category label for the text "xxx Company" is "Subtitle".
[0069] Step S13: Combine different text blocks with the same text block into a group to obtain a text fragment.
[0070] In the embodiments of this application, according to Figure 3 As shown in the resume document layout relationship establishment process, when dividing text block association groups, it is possible to divide text in the same area into different text block association groups, for example... Figure 3 In the work experience section, the text blocks consisting of titles and subheadings, as well as the text blocks consisting of subheadings and body text, can be considered the same text segment. Specifically, the text blocks consisting of titles and subheadings in the work experience section, and the text blocks consisting of subheadings and body text in the work experience section, both contain the same text blocks, namely the subheadings in the work experience section. Therefore, these two text block groups can be combined into a text segment.
[0071] Therefore, different text block association groups with the same text blocks refer to different text block association groups that include at least some of the same text content.
[0072] Specifically, the resume parsing device, based on the resume geometric layout relationship established in step S12, categorizes the resumes according to... Figure 3 The document structure tree shown divides text visual blocks, i.e., text fragments. This application can divide text visual blocks using heuristic rules, i.e., by setting rules through linking to connect text blocks. For example, if there are linking rules: [0,1] and [1,2], they can be combined into [0,1,2] through recursive intersection [1]. The text of related text blocks is then merged. This yields m text content fragments T1, T2, T3, ..., T m .
[0073] Step S14: Obtain the structural labels of each text block based on the geometric layout relationship of each text association group.
[0074] In this embodiment, the resume parsing device obtains the structural tags of each text block based on the geometric layout relationship of each text association group; the geometric layout relationship of the text block association group includes the relationship between several texts and / or the relationship between several text blocks; the text fragments are classified according to the structural tags; entity extraction is performed on the text fragments according to the classification of each structural tag to obtain the parsed entity of the resume. The structural tags involved in this application include, but are not limited to: title, subtitle, body text and / or others.
[0075] Step S15: Classify the text fragments according to their structural tags.
[0076] In this embodiment, the resume parsing device further utilizes a pre-trained text classification model to classify the various text fragments and obtain semantic tags for each text fragment; it then extracts entities from the various structural tags of the text fragments according to the semantic tags to obtain the parsed entities of the resume. The semantic tags involved in this application include, but are not limited to: basic information, work experience, education experience, project experience, personal evaluation, and / or others.
[0077] Specifically, the resume parsing device uses a text classification model (BERT+softmax) to classify each text segment in step S13, obtaining a semantic label (semantic_label) for each text segment. The text semantic label is a predefined set of labels. Based on the resume content, the text semantic label set can be defined as: semantic_label = {basic information, work experience, education experience, project experience, personal evaluation, other}.
[0078] Furthermore, the text classification model involved in this application can also use other text classification models, such as TextCNN.
[0079] Step S16: Extract entities from the text fragments according to each structural tag to obtain the parsed entities of the resume document.
[0080] In this embodiment, the resume parsing device uses the Named Entity Recognition (GlobalPointer) method to extract structured fields from different types of text segments. For example, it extracts the company name, employment period, etc., from the "subheading" block of the "Work Experience" heading block.
[0081] Furthermore, the named entity recognition method involved in this application can also use other named entity recognition methods, such as UIE, BERT+BiLSTM+CRF, etc.
[0082] In this embodiment, the resume parsing method obtains the resume document to be parsed; inputs the resume document into a pre-trained resume document geometric relationship model to extract text block association groups; combines different text block association groups with the same text blocks to obtain text fragments; obtains the structural labels of each text block based on the geometric layout relationship of each text association group; classifies each text fragment according to the structural labels; and extracts entities from the text fragments according to each structural label to obtain the parsed entities of the resume document. By establishing the geometric layout relationship of the resume document and associating the relationships between various blocks in the resume, the method avoids the text line disorder problem that occurs in existing resume parsing schemes. This application first divides the text information blocks of the resume using geometric layout relationships, and then uses a text classification algorithm to classify the text information blocks with high accuracy. Furthermore, because the information blocks have rich contextual information, only a BERT-based classification model is needed to accurately classify the information blocks, without requiring a large amount of manual features.
[0083] This application establishes the geometric layout relationship of the resume document and associates the relationships between the blocks in the resume, thereby avoiding the problem of text line disorder in the text extraction of existing resume parsing solutions.
[0084] This application utilizes the geometric layout of resume documents, merges text blocks using heuristic rules, and divides the text blocks using a text classification model to obtain semantic tags for the text blocks. This avoids the problems of insufficient generalization and low accuracy caused by keyword matching and regular expression rules.
[0085] This application utilizes geometric layout structure tags and text block semantic tags to define the recognition scope of the named entity recognition method. This allows for targeted extraction of fields within the target range, avoiding interference from irrelevant text.
[0086] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0087] To implement the above resume parsing method, this application also proposes a resume parsing device, which can be found in the following details. Figure 5 , Figure 5 This is a schematic diagram of an embodiment of the resume parsing device provided in this application.
[0088] The resume parsing device 300 in this embodiment includes a resume acquisition module 31, a geometric relationship module 32, a fragment combination module 33, and a resume parsing module 34.
[0089] The resume acquisition module 31 is used to acquire the resume document to be parsed.
[0090] The geometric relationship module 32 is used to input the resume document into a pre-trained resume document geometric relationship model and extract the text block association groups of the resume document.
[0091] The fragment combination module 33 is used to combine different text block association groups with the same text block to obtain text fragments.
[0092] The resume parsing module 34 is used to obtain the structural tags of each text block based on the geometric layout relationship of each text association group; classify each text fragment according to the structural tags; and extract entities from the text fragments according to each structural tag to obtain the parsed entities of the resume document.
[0093] To implement the above resume parsing method, this application also proposes another resume parsing device, please refer to [link / reference needed]. Figure 6 , Figure 6 This is a schematic diagram of another embodiment of the resume parsing device provided in this application.
[0094] The resume parsing device 400 in this embodiment includes a processor 41, a memory 42, an input / output device 43, and a bus 44.
[0095] The processor 41, memory 42, and input / output device 43 are respectively connected to the bus 44. The memory 42 stores program data, and the processor 41 is used to execute the program data to implement the resume parsing method described in the above embodiment.
[0096] In this embodiment, processor 41 can also be referred to as a CPU (Central Processing Unit). Processor 41 may be an integrated circuit chip with signal processing capabilities. Processor 41 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor, or processor 41 can be any conventional processor.
[0097] This application also provides a computer storage medium; please refer to the following: Figure 7 , Figure 7 This is a schematic diagram of a computer storage medium according to an embodiment of the present application. The computer storage medium 600 stores a computer program 61, which, when executed by a processor, is used to implement the resume parsing method of the above embodiment.
[0098] When the embodiments of this application are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0099] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A resume parsing method, characterized in that, The resume parsing method includes: Retrieve the resume document to be parsed; The resume document is input into a pre-trained resume document geometric relationship model to extract the text block association groups of the resume document; By associating different text blocks that have the same text blocks, we can obtain text fragments; The structural labels of each text block are obtained based on the geometric layout relationship of each text association group; The text fragments are categorized according to the structural tags; Entity extraction is performed on the text fragments according to each structural tag to obtain the parsed entities of the resume document.
2. The resume parsing method according to claim 1, characterized in that, The structure tags include headings, subheadings, body text, and / or others.
3. The resume parsing method according to claim 1 or 2, characterized in that, The step of extracting entities from the text fragments according to each structural tag to obtain the parsed entities of the resume document includes: The text fragments are classified using a pre-trained text classification model to obtain semantic labels for each text fragment; Entity extraction is performed from each structural tag of the text fragment according to the semantic tags to obtain the parsed entities of the resume document.
4. The resume parsing method according to claim 3, characterized in that, The semantic tags include basic information, work experience, education experience, project experience, personal evaluation and / or others.
5. The resume parsing method according to claim 1, characterized in that, The geometric layout relationship of the text block association group includes the relationship between several characters and / or the relationship between several text blocks.
6. The resume parsing method according to claim 1, characterized in that, The resume parsing method also includes: Obtain the resume document to be trained; Based on the resume document to be trained, input tags are obtained, wherein the input tags include text block association relationships, text block labels, and text block coordinates; The input labels and the resume document to be trained are input into the resume document geometric relationship model for training.
7. The resume parsing method according to claim 6, characterized in that, The input label also includes each character in the text block and its coordinates.
8. The resume parsing method according to claim 1, characterized in that, After obtaining the resume document to be parsed, the resume parsing method further includes: Convert resume documents in different formats into image-based resume documents.
9. A resume parsing device, characterized in that, The resume parsing device includes a resume acquisition module, a geometric relationship module, a fragment combination module, and a resume parsing module; wherein... The resume acquisition module is used to acquire the resume document to be parsed; The geometric relationship module is used to input the resume document into a pre-trained resume document geometric relationship model and extract the text block association groups of the resume document; The fragment combination module is used to combine different text block association groups that have the same text block to obtain text fragments; The resume parsing module is used to obtain the structural tags of each text block based on the geometric layout relationship of each text association group; classify each text fragment according to the structural tags; and extract entities from the text fragments according to each structural tag to obtain the parsed entities of the resume document.
10. A resume parsing device, characterized in that, The resume parsing device includes a memory and a processor coupled to the memory; The memory is used to store program data, and the processor is used to execute the program data to implement the resume parsing method as described in any one of claims 1 to 8.
11. A computer storage medium, characterized in that, The computer storage medium is used to store program data, which, when executed by the computer, is used to implement the resume parsing method as described in any one of claims 1 to 8.