Resume data analysis method and device, electronic equipment and storage medium

By identifying industry categories in resume files and using deep learning models to combine industry knowledge graphs for resume analysis, the shortcomings of traditional methods in identifying professional vocabulary in different industries and analyzing non-document format resumes are solved, and higher parsing accuracy and comprehensiveness are achieved.

CN120218048APending Publication Date: 2025-06-27QIAN JIN NETWORK INFORMATION TECH SHANGHAI LTD

Patent Information

Application Number
CN202510203708.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing resume analysis methods are difficult to accurately identify professional vocabulary in different industries, which has affected the accuracy of resume content extraction. Especially under the diverse methods of providing resumes, it is difficult for traditional methods to effectively analyze resumes in non-document formats.

Method used

By identifying the industry categories in the resume file, determining the corresponding multimodal resume data and the deep learning model obtained by training in the industry knowledge graph, analyzing the resume text and auxiliary feature data, and generating structured resume data.

Benefits of technology

It improves the accuracy and comprehensiveness of resume data analysis, can effectively extract information from multimodal data, enhances the recognition ability of professional vocabulary in different industries, and is suitable for diverse resume formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218048A_ABST
    Figure CN120218048A_ABST
Patent Text Reader

Abstract

The invention relates to a resume data analysis method and device, electronic equipment and a storage medium, and the method comprises the steps: collecting a resume file of a candidate, the resume file comprising a resume file in a document format and / or a resume file in a non-document format; extracting a resume text and auxiliary feature data from the resume file; according to the resume text, identifying an industry category involved by the content of the resume file; a corresponding resume analysis model is determined according to the industry category, and the resume analysis model is a deep learning model obtained through training based on multi-mode resume data of the industry category and an industry knowledge graph; and based on the resume text and the auxiliary feature data, utilizing the resume analysis model to obtain structured resume data, wherein the structured resume data comprises a plurality of resume fields and field contents. According to the method, the accuracy and comprehensiveness of resume analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a resume parsing method, apparatus, electronic device, and storage medium. Background Art

[0002] In the fields of talent recruitment and human resource management, the analysis and processing of resumes is one of the core tasks. Especially for current recruitment platforms that concentrate a large amount of job seeker information and recruiter information, resume parsing is the basis for services such as resume screening, job recommendation, and online interviews. Moreover, the accuracy and comprehensiveness of resume parsing directly affect the quality of these services.

[0003] Currently, the targets of resume parsing and processing are usually resumes in text format or image format. When the resume is in image format, the text content in the picture is extracted as editable text through technologies such as OCR, and then natural language processing technology is used to extract information from the resume text to form structured information including personal basic information, educational experience, work experience, skills, etc. The extraction of information mainly includes two schemes. One scheme is to use a large number of complex text parsing rules; the other scheme is to classify the resume text through deep learning algorithms. However, both of these schemes have their respective limitations, resulting in less than ideal resume parsing effects, such as a large amount of manual work, inaccurate key information extracted due to a large amount of information loss, such as out-of-order, duplicate, missing, etc.

[0004] Based on these problems, the Chinese patent with the publication number CN113743052B and the invention name "A Resume Layout Analysis Method and Apparatus Integrating Multi-Modalities" and the Chinese patent application with the publication number CN118314594A and the invention name "Resume Information Extraction Method, Apparatus, Device, and Storage Medium" provide a scheme for determining the category of corresponding text content by combining the natural language information of the text and the position information or layout information of the text. Since the classification process refers to the position information of the text, the accuracy of content classification is improved compared to the scheme that only processes text information.

[0005] However, with the changes in the forms of employment, job hunting, and recruitment, the ways of providing resumes have become diversified. The documents serving as resumes are no longer just documents in standard formats such as Word or PDF, but also other forms of files, such as multimedia format files like audio and video, and even other professional format files. These non-document format resumes can provide richer information beyond language, and these are the information that recruiters hope to know more. For example, for a language worker, such as a training teacher, although their text resume will provide relevant descriptions about their work ability, such as education background, work experience, achievements, etc., the tone and intonation in the audio features can reflect the affinity as a training teacher, and this trait is a very important evidence for their good work performance or ability, but usually such traits do not appear in the text resume.

[0006] In addition, resumes and resume contents in different industries usually vary greatly, and the currently commonly used resume information extraction methods or resume parsing methods are difficult to accurately identify professional vocabulary in different industries, thus affecting the accuracy of resume content extraction. Summary of the Invention

[0007] In view of the technical problems existing in the prior art, the present invention proposes a resume data parsing method, device, electronic device, and storage medium to improve the accuracy of resume data parsing.

[0008] To solve the above technical problems, according to one aspect of the present invention, the present invention provides a resume data parsing method, including the following steps:

[0009] Collect resume files of candidates, where the resume files include resume files in document format and / or non-document format;

[0010] Extract resume text and auxiliary feature data from the resume files;

[0011] Identify the industry category involved in the content of the resume file according to the resume text;

[0012] Determine the corresponding resume parsing model according to the industry category, where the resume parsing model is a deep learning model trained based on multi-modal resume data and industry knowledge graph of the industry category; and

[0013] Obtain structured resume data based on the resume text and auxiliary feature data by using the resume parsing model, where the structured resume data includes multiple resume fields and field contents.

[0014] Optionally, when the resume file is a resume file in document format, the auxiliary feature data extracted from the resume file includes one or more of text style, text position, table corresponding to the text, and image.

[0015] Optionally, when the resume file in document format includes an image, it further includes: extracting one or more of text information, image description in text form, and visual image feature data from the image.

[0016] Optionally, the non-document format resume file is one or more of picture format, audio format, and video format.

[0017] Optionally, when the resume file is in picture format, the resume text extracted from the resume file is the resume text obtained by recognizing the text image in the picture and the table description and / or image description in text form; the auxiliary feature data includes visual image feature data.

[0018] Optionally, when the resume file is in audio format, the resume text extracted from the resume file includes the resume text obtained by speech recognition; the auxiliary feature data includes one or more of voice, intonation, high-frequency keywords, stress, weak sound, specific events, and emotion types.

[0019] Optionally, when the resume file is in video format, the resume text extracted from the resume file includes the resume text obtained by speech recognition; the auxiliary feature data includes one or more of the voice, intonation, high-frequency keywords, stress, weak sound, specific events, expressions, body movements, and emotion types of the people in the video.

[0020] Optionally, after extracting the resume text and auxiliary feature data from the resume file, it further includes:

[0021] Converting the auxiliary feature data into tag information based on the auxiliary feature type; and

[0022] Using the tag information to mark the corresponding resume text block.

[0023] Optionally, the step of identifying the industry category involved in the content of the resume file according to the resume text includes:

[0024] Extracting industry keywords from the resume text; and

[0025] Querying the industry keyword library based on the industry keywords to determine the matching industry category;

[0026] wherein, the corresponding industry keywords are stored in the industry keyword library by industry category; or

[0027] Input the resume text into an industry classification model, where the industry classification model is a deep learning model trained with industry keywords as training data; and

[0028] Obtain the industry category involved in the content of the resume file through the output of the industry classification model.

[0029] Optionally, the steps of obtaining structured resume data using the resume parsing model based on the resume text and auxiliary feature data include:

[0030] Fuse the auxiliary feature data into the resume text to obtain a new resume text;

[0031] Input the new resume text into the resume parsing model; and

[0032] Output structured resume data after the processing of the resume parsing model.

[0033] Optionally, before fusing the auxiliary feature data into the resume text, it includes:

[0034] Query the text segment corresponding to the auxiliary feature data from the resume text; and

[0035] In response to querying the text segment corresponding to the auxiliary feature data from the resume text, verify the authenticity and / or consistency of the content of the text segment based on the auxiliary feature data.

[0036] Optionally, the steps of fusing the auxiliary feature data into the resume text include:

[0037] Convert the auxiliary feature data into a first description text based on the auxiliary feature type;

[0038] Query whether the corresponding second description text is included in the resume text;

[0039] In response to the corresponding second description text being included in the resume text, compare the semantic ranges of the first description text and the second description text;

[0040] In response to the semantic range of the second description text being greater than or equal to that of the first description text, retain the second description text; in response to the semantic ranges of the first description text and the second description text intersecting, merge the first description text and the second description text and remove the text with duplicate semantics; and

[0041] In response to the corresponding second description text not being included in the resume text, merge the first description text into the resume text.

[0042] Optionally, when multiple formats of resume files are included in the resume file, the resume text and auxiliary feature data obtained based on each resume file respectively are used to obtain the corresponding structured resume data by using the resume parsing model; and

[0043] The structured resume data obtained from each resume file is fused based on a fusion strategy to obtain the final structured resume data.

[0044] Optionally, after the resume text and auxiliary feature data are obtained for each resume file, cross-validation is performed on the resume texts obtained from different resume files, and the inconsistent contents are marked; correspondingly, in the obtained structured resume data, the corresponding resume fields are marked; when the structured resume data obtained from each resume file is fused, the content of the resume field with the mark is retained.

[0045] Optionally, before the resume text or the new resume text is input into the resume parsing model, sentence segmentation is performed through a sentence segmentation model to obtain segmented text;

[0046] The segmented text is input into the resume parsing model; and

[0047] After being processed by the resume parsing model, structured resume data is output.

[0048] Optionally, the resume data parsing method further includes:

[0049] Recruitment requirement information is obtained, and multiple recruitment feature data are extracted from the recruitment requirement information, where the recruitment feature data includes recruitment feature fields and field contents;

[0050] Correspondingly, based on the resume text, recruitment feature data, and auxiliary feature data, structured resume data is obtained by using a resume parsing model; where the resume parsing model is a deep learning model trained based on the multi-modal resume data of the industry category, recruitment feature data, and industry knowledge graph, and the deep learning model is used to extract fields and field contents that match the industry category and recruitment features from the multi-modal resume data.

[0051] According to another aspect of the present invention, the present invention also provides a resume data parsing device, including:

[0052] A file collection module configured to collect the resume files of candidates, where the resume files include resume files in document format and / or non-document format;

[0053] A data extraction module configured to extract resume text and auxiliary feature data from the resume files;

[0054] An industry category recognition module, configured to recognize the industry category involved in the resume file content according to the resume text; and

[0055] A parsing module, configured to determine a corresponding resume parsing model according to the industry category, and obtain structured resume data by using the resume parsing model based on the resume text and auxiliary feature data. The structured resume data includes multiple resume fields and field contents. Among them, the resume parsing model is a deep learning model trained based on multi-modal resume data of the industry category and an industry knowledge graph.

[0056] According to another aspect of the present invention, the present invention also provides an electronic device, including a processor and a memory. A computer program instruction set is stored on the memory. When the processor executes the computer program instruction set on the memory, the foregoing resume data parsing method is implemented.

[0057] According to another aspect of the present invention, the present invention also provides a computer-readable storage medium. A computer program instruction set is stored on the computer-readable storage medium. When the computer program instruction set is executed by a processor, the foregoing resume data parsing method is implemented.

[0058] The present invention can extract multi-modal data from a resume file. Through the processing of multi-modal data, not only can the parsing accuracy be improved, but also information other than text information can be extracted. In addition, through the application of the industry knowledge graph, the accuracy and comprehensiveness of resume parsing are further increased. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Next, the preferred embodiments of the present invention will be further described in detail with reference to the drawings, where:

[0060] Figure 1 is a flowchart of a resume data parsing method according to an embodiment of the present invention;

[0061] Figure 2 is a flowchart of a method for training a resume parsing model according to an embodiment of the present invention;

[0062] Figure 3 is a schematic diagram of the principle of a resume parsing model according to an embodiment of the present invention;

[0063] Figure 4 is a flowchart of a method for constructing an industry knowledge graph according to an embodiment of the present invention;

[0064] Figure 5 is a flowchart of a method for fusing auxiliary feature data into a resume text according to an embodiment of the present invention;

[0065] Figure 6Schematic diagram of the principle of a resume parsing model according to another embodiment of the present invention;

[0066] Figure 7 Block diagram of the principle of a resume data parsing device according to an embodiment of the present invention;

[0067] Figure 8 Block diagram of the principle of a first application system according to an application embodiment of the present invention;

[0068] Figure 9 Block diagram of the principle of a second application system according to an application embodiment of the present invention; and

[0069] Figure 10 Schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] In the following detailed description, reference may be made to the accompanying drawings that form a part hereof and that show, by way of illustration, specific embodiments in which the invention may be practiced. In the drawings, like reference numerals describe substantially similar components in different views. The specific embodiments of the present invention have been described in sufficient detail below to enable those of ordinary skill in the art with relevant knowledge and technology to implement the technical solutions of the present invention. It should be understood that other embodiments may be utilized or structural, logical, or electrical changes may be made to the embodiments of the present invention.

[0072] Figure 1 Flowchart of a resume data parsing method according to an embodiment of the present invention. The resume data parsing method in this embodiment includes the following steps:

[0073] Step S11, collect the resume files of candidates, where the resume files include resume files in document format and / or non-document format.

[0074] Step S12, extract resume text and auxiliary feature data from the resume files.

[0075] Step S13, identify the industry category involved in the content of the resume files according to the resume text.

[0076] Step S14: Determine the corresponding resume parsing model according to the industry category, where the resume parsing model is a deep learning model trained based on the multimodal resume data and industry knowledge graph of the industry category.

[0077] Step S15: Obtain structured resume data by using the resume parsing model based on the resume text and auxiliary feature data. The structured resume data includes multiple resume fields and field contents that match the industry category.

[0078] When parsing a resume, the present invention can not only parse traditional document-format resumes, but also parse non-document-format resumes, such as picture-format resumes, audio-format resumes, and video-format resumes, so as to cope with the increasingly diverse resume-providing methods currently. Among them, the document-format resume can be a document generated by word processing software such as word or WPS, or a PDF-format document. Usually, a resume document will adopt corresponding editing styles for different parts and may also insert tables, images, etc. For a picture-format resume, it is usually a picture taken of a document. An audio-format resume is, for example, a recording of a candidate, and the candidate orally states his / her basic information, educational experience, work situation, etc. in the recording. A video-format resume is similar to an audio-format resume and is a video recorded of a candidate, and the candidate orally states his / her basic information, educational experience, work situation, etc. in the video. Therefore, in step S11, when collecting the resume files of candidates, document-format resumes or picture-format resumes are collected. For example, when a candidate provides a resume by email or other means, the resume file is obtained from the email. Another example is to read the resume file from the database of the recruitment platform according to the candidate's ID. In addition, audio-format or video-format resumes can also be collected by other means. For example, when a recruiter conducts a telephone interview or a remote video interview with a candidate, the telephone recording and the screencast file of the remote video interview obtained in these scenarios can be used as supplementary files for the candidate's resume. Of course, a candidate can also record an audio-format or video-format resume by himself / herself and provide it to the recruitment platform or enterprise by email.

[0079] In step S12, the present invention extracts resume text and auxiliary feature data from resume files in various formats. For example, for resume files in document format, the text can be directly read from the document to obtain the corresponding text. For example, the text is read line by line to obtain the resume text. In addition, when reading the text in the document, auxiliary feature data such as text style, text position, tables and images corresponding to the text are also obtained simultaneously. For example, the entire document is regarded as an operable object model, and the document structure is divided into nodes represented by a tree structure. For example, the entire document is used as the root node, a section or paragraph in the document is used as the second-level node, and a fragment in the paragraph or section is used as the third-level node. Different nodes have corresponding attributes, and these attributes together constitute the document style. Therefore, the text, text style, position information, etc. can be extracted by traversing the document nodes level by level. When a table appears in the document, the relevant attributes of the table, such as the number of rows, number of columns, row height, column width, etc., can also be obtained. For another example, the extraction of the aforementioned text, style, position, etc. can also be completed using some existing models, such as the LayoutLM (Layout Language Model) model and its series of models or the PaddleOCR model.

[0080] When the document includes images, the OCR (Optical Character Recognition) technology can also be used to convert the text in the images into text to extract the text information, or generate a text description of the image in text form. For example, based on the degree certificate image therein, a text description such as "This is a degree certificate" is generated. Further, the text description and the text recognized from the image can be used as a complete text content. For example: "This is a degree certificate: Name: XXX, School: XXXXX...". etc. In a specific embodiment, it can be implemented through a trained neural network model. In addition, for the images in the document, in addition to extracting the text information, some visual image feature data are also generated, such as color features (such as color histograms, color sets, etc.), texture features (such as obtaining feature parameters such as the fineness and direction of the texture by using the texture feature analysis method of the gray-level co-occurrence matrix), shape features (such as contour features or region features), spatial relationship features (such as the position in the current document, the distance from other objects, etc.). The present invention adopts a rich text extraction method for resume files in document format, so other information besides the text can be extracted.

[0081] When the resume file is in picture format, similar to the processing method of the images in the aforementioned document, the resume text obtained by recognizing the text image in the picture using technologies such as OCR, if there are tables and images, text form table descriptions and / or image descriptions can also be obtained, and at the same time, the aforementioned visual image feature data are also generated.

[0082] When the resume file is in audio format, through speech technologies such as Automatic Speech Recognition (ASR), the speech introduction of the candidate in the audio can be converted into text, and at the same time, auxiliary feature data such as speech, intonation, speech rate, high-frequency keywords, stress, weak stress, specific events, and emotion types are extracted. Specifically, different auxiliary features can be obtained through existing models for various special tasks. For example, a dedicated intonation model can identify auxiliary features such as intonation, stress, or weak stress in this section of audio, and a model based on the emotion recognition task can obtain the emotion in this section of audio, such as happy, sad, angry, anxious, etc. Some events and specific information other than speech in this section of audio can also be obtained based on a specific event parsing model, such as the duration, frequency, and other related feature information of laughter and the event of sneezing or coughing. The speech rate, high-frequency words, etc. of the candidate can be obtained through statistical means. The models for the above-mentioned various tasks can be determined according to actual needs and the application habits of developers. For example, they can be traditional machine learning models such as Hidden Markov Model (HMM) and Dynamic Time Warping (DTW); deep learning models such as DNN-HMM based on acoustic models or end-to-end speech recognition models, specifically models such as Attention model, Transformer model, RNN-Transducer (RNN-T), etc.; self-supervised learning models such as Wav2Vec model and a series of HuBERT models; or hybrid models such as large language models (such as GPT, T5), or multi-task learning models such as Speech2Vec. Of course, the above-mentioned auxiliary feature data can also be obtained through a multi-task learning model.

[0083] When the resume file is in video format, the method of extracting the resume text from the resume file is similar to that of the audio format resume file. The extracted auxiliary feature data includes not only the speech, intonation, high-frequency keywords, stress, weak stress, specific events, and emotion types of the person extracted from the audio format resume file, but also expressions and body movements. And the emotion type extracted is obtained by referring to expressions and body movements, and its accuracy is higher.

[0084] From the above processing steps for resume files of various formats, it can be seen that the present invention can not only extract resume text data but also other types of feature data. These feature data can reflect a large amount of key information that cannot be carried by document resumes and can deconstruct new information or associated information during the subsequent parsing process of resume data.

[0085] In one embodiment, when extracting the aforementioned resume text information, the resume text can also be marked according to the extracted auxiliary feature data. For example, for a document resume file, while extracting the text information, document auxiliary features such as text style, text position, image position, table position, and table attributes are also extracted, and these data are converted into marking information to mark the corresponding text blocks, thus facilitating the subsequent parsing of structured resume data. For example, when extracting a text content, mark the text style such as the font, font size, and paragraph spacing of the text content. When extracting text from a table, mark the text content as a table, and specifically mark the data formats (such as text, currency, number, etc.) of the row, column, and cell where it is located; for another example, for image auxiliary features, mark the relevant color features of a text block, such as background color, text color, or texture features, etc. These colors and textures are helpful for information classification during subsequent parsing. Similarly, for audio auxiliary features, convert speech intonation, speech rate, stress, weak stress, etc. into marking information and mark the corresponding text.

[0086] In addition, after extracting the text information, a text cleaning step is also included. For example, replace special characters, perform text cleaning based on rules, perform general text cleaning, perform text cleaning based on parsing configuration, correct easily confused characters after OCR recognition, and so on.

[0087] In an alternative embodiment, after extracting the resume text and auxiliary feature data from a resume file in a certain format, the content of the resume text can also be verified using the auxiliary feature data. For example, query the text segment corresponding to the auxiliary feature data from the resume text; when the text segment corresponding to the auxiliary feature data is found, verify the authenticity and / or consistency of the content of the text segment based on the auxiliary feature data. For example, when extracting a degree certificate image from a resume file in document format, generate a complete description content of the degree certificate image based on the degree certificate image and the text recognized in the image, such as "This is a degree certificate: Name: XXX, School: XXXXX...". Query the resume text based on the field keywords "School", "Major", "Degree", etc. in this content, extract the corresponding content therefrom, and then compare whether the content extracted from the degree certificate image is consistent with the corresponding content in the resume text. For another example, compare the image features of the certificate image extracted from the resume file in document format with the image features of the certificate images in the database to determine whether the certificate in the resume is authentic. In an embodiment of the present invention, the database stores images and their image features of various skill certificates, degree certificates, etc.

[0088] Since the professional vocabularies in different industries usually vary greatly, in order to improve the parsing accuracy, the model for parsing resumes in the present invention is a deep learning model trained based on industry categories and industry knowledge graphs. Therefore, in step S13, the industry category involved in the resume file content is identified according to the resume text.

[0089] In one embodiment, an industry keyword library is stored in the database. Corresponding industry keywords are stored in the industry keyword library by industry category. Therefore, industry keywords can be extracted from the resume text; the industry keyword library is queried based on the industry keywords to determine the matching industry category. In another embodiment, industry keywords are used as training data to train a model capable of performing industry classification according to the input text. Therefore, the resume text can be input to the industry classification model, and the industry category involved in the resume file content can be obtained through the industry classification model. The industry classification model analyzes the employing company, position, and industry terms in the resume, and reasons through the context relationship to obtain which sub-industry the current resume belongs to, specifically including one or more of the following processing processes:

[0090] Entity recognition and classification: Through natural language processing (NLP) technology, entities related to the working industry are recognized from the resume text, such as company names, position names, industry terms, etc. These entities can be classified according to existing industry classification standards, such as IT, finance, manufacturing, etc.

[0091] Industry keyword extraction: Since the industry classification model analyzes a large amount of text data related to industries, job descriptions, company introductions, etc. during the training process, keywords specific to a particular industry can be extracted. These keywords are closely related to industry-related technologies, functions, positions, etc., and can help the model identify the working industry.

[0092] Relationship extraction and reasoning: Mine the relationships between entities. For example, through the relationship between a position and the job content, or the relationship between a company and its affiliated industry, it can be inferred which industry a certain position or job belongs to.

[0093] Context analysis: Through context information, determine which industry a position or job task belongs to. For example, "software development" and "programming" are usually associated with the IT industry, while "financial analysis" and "budget" are common in the financial or accounting industry.

[0094] After the accurate industry category is determined through the processing of the industry classification model, the corresponding trained resume parsing model can be determined accordingly. The resume parsing model can be a combination of multiple models or a single model, and can output structured resume data based on the input data.

[0095] See Figure 2 ,Figure 2 It is a flowchart of a method for training a resume parsing model according to an embodiment of the present invention. In this embodiment, the method includes the following steps:

[0096] Step S101, collect resume files in various formats, where the formats include document format, audio format, and video format.

[0097] Step S102, respectively extract the original resume text and corresponding auxiliary feature data from each format of resume file.

[0098] Step S103, process the original resume text and corresponding auxiliary feature data to obtain samples. Specifically, fuse the auxiliary feature data extracted from the resume files of the same user (such as a job-seeking user or an applicant user) with the original resume text, and perform operations such as content deduplication and cleaning on the fused text to obtain a sample. In addition, the sample can be preprocessed according to the model to be used later, such as performing sentence splitting, word segmentation, and data augmentation in combination with knowledge graph entities.

[0099] Step S104, label the samples and determine the corresponding knowledge graph information. For example, label the industries involved in the resume content, and label the resume fields and their contents.

[0100] Step S105, construct a model module. In one embodiment, the model module includes a pre-trained language model BERT (Bidirectional Encoder Representations from Transformers) layer, a CRF layer (Conditional Random Field), and a knowledge graph layer.

[0101] Step S106, input the labeled samples into the model module for model training. The schematic diagram of the model module in this embodiment is as Figure 3 shown, Figure 3 It is a schematic diagram of the principle of a resume parsing model according to an embodiment of the present invention. Among them, Figure 3The dashed line represents the training process, and the solid line represents the inference process for actual applications. During the training process, the BERT layer extracts text features from the input samples. The knowledge graph layer uses the TransE, DistMult, or graph neural network (GNN) model to encode the knowledge graph information in the samples into vectors, which are then concatenated with the text features as knowledge graph features and passed to the CRF layer. The CRF layer learns the association between fields and domain knowledge to complete the annotation of the sequence (field types and field contents in the resume text). Field types are, for example, "name", "educational background", "work experience", and field contents are, for example, "Tsinghua University", "computer science", etc. By embedding the knowledge graph features, the input obtained by the CRF layer is more abundant, including both the context semantic information provided by the BERT layer and the prior knowledge and entity relationships provided by the knowledge graph. The CRF layer uses these features to globally optimize the entire sequence and generate more accurate annotation results. Among them, the model module further includes a post-processing layer for correcting field contents in the annotation results, complementing fields or field contents, verifying the logical consistency between field contents, etc. during the inference process of actual applications. Additionally, the model module can also adopt a single model capable of multi-task learning. The model can adopt a multi-head structure, with one head for text understanding and sequence annotation, and the other head for aligning field contents with the knowledge graph and generation, that is, combining the industry knowledge graph during the generation of structured data.

[0102] Step S107, calculate the losses (including annotation loss and knowledge graph matching loss), optimize the model, and store the model data.

[0103] Combining the industry knowledge graph during the training process of the present invention can help the model learn the commonalities and differences between entities in different fields, form a deeper understanding, better generalize to unseen entities, improve the accuracy of preliminary annotation, reduce the post-processing burden, and improve the overall processing efficiency of the model.

[0104] In an embodiment of the present invention, multiple resume parsing models for different industries and industry knowledge graphs are trained according to the above process. In another embodiment, multiple industry knowledge graphs can be used to train the resume parsing model, so that the resume parsing model can well understand the relevant professional vocabulary of multiple industries. To reduce the processing volume during application, according to the identified industry category, the industry knowledge graph of the same category is accessed during parsing.

[0105] In another embodiment, when constructing a knowledge graph, the entities in the knowledge graph include industry attributes, through which the industries of the entities that can have relationships with it can be determined. For example, for the triples (xxx Company - Industry - Technology), (xxx Company - Recruitment - Software Engineer) in the knowledge graph, it can be obtained that the current position "Software Engineer" is a position in the technology industry.

[0106] The industry knowledge graph in the present invention can be an existing industry knowledge graph or an industry knowledge graph constructed based on the needs of the application. Refer to Figure 4 , Figure 4 which is a flowchart of a method for constructing an industry knowledge graph according to an embodiment of the present invention. The method for constructing an industry knowledge graph mainly includes the following content:

[0107] Step S21, collect industry data. In one embodiment, collect information about companies, enterprises, institutions, and positions in the target industry. Specifically, data can be obtained through channels such as job recruitment websites, company official websites, social media (such as LinkedIn), and industry reports. The data types of the obtained data include structured data and unstructured data. Structured data such as job information on recruitment websites and data obtained from company databases, and unstructured data such as articles, news, job descriptions on social media, and company dynamics, etc.

[0108] Step S22, perform entity extraction (NER, named entity recognition). This includes company entities, position entities, keywords representing attributes, keywords related to responsibilities, etc. Among them, natural language processing (NLP) models, such as SpaCy, BERT, etc. models or industry-specific dictionaries are used to identify entities. The company entities are company names, such as "xxx Company", "xxx Research Institute", "xxx University", "Computer Science", etc.; the position entities are position names, such as "Software Engineer", "Product Manager", "Data Scientist", etc.; the attributes include industry, location, time, etc. For example, "Technology", "Internet" representing the industry, "Beijing", "Shenzhen" representing the location, "2000" representing the time, etc.; keywords related to responsibilities such as "Familiar with Java", "C++", "Have more than 5 years of work experience", "Good communication skills", etc.

[0109] Step S23: Perform relation extraction. The relations include relations between companies, relations between positions, relations between companies and positions, relations between companies and attributes, relations between positions and responsibilities / requirements, etc. Methods such as dependency syntactic analysis, relation extraction models, or template-based rules can be used for relation extraction. For example, from the recruitment information of software engineer posted by xxx company, the relation between "xxx company" and "software engineer" is extracted as "recruitment"; from the introduction of xxx company, the relation between "xxx company" and "2000" is extracted as "establishment time"; from the subject introduction of xx university, the relation between "xx university" and "computer science" is extracted as "undergraduate major", etc.

[0110] Step S24: Perform data annotation. For example, manually annotate the possible relations between companies, attributes, positions, responsibilities, etc. based on expert knowledge. For large-scale data, automated tools can be combined for annotation, that is, annotation is performed while extracting entities and relations as described above, and then the incorrect annotation results are manually verified and corrected. For example, the relations between the annotated company and attribute are such as "xxx company - industry - technology", "xxx company - location - Beijing", etc.; the relations between the annotated company and position are such as "xxx company – recruitment - software engineer"; the relations between the annotated position and responsibility are such as "software engineer - requirement - familiar with Java, C++, with more than 5 years of work experience", etc.

[0111] Step S25: Represent information using a triple data structure. Such as (xxx company, industry, technology), (xxx company, location, Beijing”), (software engineer, requirement, familiar with Java), etc.

[0112] Step S26: Build a knowledge graph model to store information. In one embodiment, the extracted entities are stored as nodes in a graph database (such as Neo4j) or in RDF in triple format, and the relationship between two entities is represented by a connection between the nodes. For example, "xxx company" and "Beijing" are nodes, and the connection between the two nodes is "location".

[0113] In step S15, in one embodiment, before parsing the resume text, auxiliary feature data is fused into the resume text to obtain a new resume text, and the new resume text is processed by the resume parsing model to obtain structured resume data. See Figure 5 , Figure 5 is a flowchart of the method for fusing auxiliary feature data into the resume text according to an embodiment of the present invention. In this embodiment, the method for fusing auxiliary feature data into the resume text specifically includes the following steps:

[0114] Step S511: Obtain an auxiliary feature data and determine its type.

[0115] Step S512: Convert the auxiliary feature data into a first description text according to the conversion method determined by the auxiliary feature type. In the present invention, some auxiliary feature data can be converted into description text to play a role in information supplementation. For example, the text extracted from an image and the text description generated for the image are used as the first description text. Or the specific events, emotion types, high-frequency keywords, etc. extracted from the audio are converted into description text.

[0116] Step S513: Query whether the corresponding second description text is included in the resume text. If the corresponding second description text is included in the resume text, execute Step S514. If the corresponding second description text is not included in the resume text, execute Step S518.

[0117] For example, after obtaining the description text generated for the degree certificate image and the text recognized from the image from the image features, query whether there is corresponding content in the resume text according to the education keywords, school keywords, major keywords, etc. in the description text.

[0118] Step S514: Compare the semantic scopes of the second description text and the first description text. When the semantic scope of the second description text is greater than or equal to that of the first description text, in Step S515, retain the second description text; when the semantic scopes of the first description text and the second description text intersect, in Step S516, merge the first description text and the second description text and remove the duplicates; when the corresponding second description text is not included in the resume text, in Step S517, merge the first description text into the resume text.

[0119] For example, after obtaining the description text generated for the degree certificate image and the text recognized from the image from the image features, corresponding school keywords and major keywords are queried in the resume text according to the education keywords, school keywords, major keywords in the description text and used as the second description text. Since there is no education keyword in the second description text, there is intersecting content between the first description text and the second description text, so the first description text and the second description text are merged and de-duplicated, thus adding new education information to the resume text. For the description text converted from specific events, emotion types, high-frequency keywords, etc. extracted from audio and video, usually the information is not in the resume text, that is, the corresponding content is not included in the resume text, so the description text converted from specific events, emotion types, high-frequency keywords, etc. is merged into the resume text.

[0120] Step S518: Determine whether there is still auxiliary feature data. If there is, return to step S511; if not, end the fusion process.

[0121] In one embodiment, the resume text obtained when extracting the resume file or the new resume text after the above fusion is a word list including annotation information, and the annotation information can be text style, position, stress, weak stress, intonation, tone, etc. Then, such a text is input to a sentence segmentation model for sentence segmentation to obtain segmented text. The sentence segmentation model is, for example, a model constructed by combining a bidirectional long short-term memory (LSTM) network structure with CRF. With the help of the annotation information, the sentence segmentation model can better understand the resume text, so as to reasonably and effectively segment the current resume text content. For example, it can accurately segment a resume text including "Self-introduction: xxxx; Education experience: xxxx; Work experience: xxxx" into three parts: "Self-introduction: xxxx / Education experience: xxxx / Work experience: xxxx". Further, the text in each part can be segmented in more detail. For example, the sentence "I graduated from xxx University" in the self-introduction part can be segmented into "I / graduated / from / xxx University".

[0122] Subsequently, the text with sentence segmentation is input into the trained resume parsing model. The BERT layer in the resume parsing model extracts text features from the segmented text and passes them to the CRF layer. The CRF layer completes sequence labeling based on the learned associations between fields and domain knowledge. After the sequence labeling is completed, the obtained fields and field contents are, for example: "Name: xxx", "Educational Background: 'School: xxxx'; Major: xxxx...", "Work Experience: ': xxxx'; ': xxxx'...", etc. Among them, since the resume parsing model uses a knowledge graph during training, it can well understand the relevant expressions in this industry and can accurately determine the corresponding field types. For example, when "xxx University" appears in the resume, it can determine whether the corresponding field type is "School" or "Work Unit" based on other entity words related to it, so the parsing accuracy is high. The post-processing layer in the resume parsing model queries the knowledge graph according to the fields and field contents obtained after labeling to verify, supplement, and correct the information. Specifically, it searches for matching entities in the knowledge graph based on the field content, and then verifies whether the field content obtained by labeling is correct according to the relationship between this entity and other entities in the knowledge graph. For example, when the model extracts "xxx University" as the school name, it searches the knowledge graph to confirm that it is indeed a university. On the other hand, it determines whether it is necessary to supplement the fields and field contents. For example, when entities matching "xxx University" and "Computer Science" are found in the knowledge graph, according to the relationship between the entities in the knowledge graph, it can be known that "Computer Science" is an undergraduate major. When this field is not in the fields extracted from the resume text, the undergraduate major obtained from the knowledge graph can be supplemented as the field content of the education level into the resume data. Further, the relationship in the knowledge graph can also be used for context relationship verification to verify the logical consistency between field contents. For example, when the content of the company name field extracted from the work experience field in the resume text is "xx Company" and the content of the position field is "President", when querying the knowledge graph, all positions of xx Company can be obtained through the knowledge graph, and then whether the field content "President" in the resume text exists among these positions is queried to verify whether the content of the position field extracted from the resume matches the knowledge graph.

[0123] By aligning the resume field content and knowledge graph information as described above, the extracted structured resume data can be verified and information can be supplemented, making the output resume data more accurate and comprehensive. In one embodiment, the finally obtained structured resume data in Json format is as follows:

[0124]

[0125]

[0126] In another embodiment, when the collected resume files of candidates include multiple formats, for example, including the document resume provided by the candidate and the interview video. In one embodiment, for the document resume and the interview video respectively, the corresponding structured resume data is obtained according to the foregoing Figure 1 and then the two structured resume data are fused based on the fusion strategy. The fusion strategy includes but is not limited to field completion, field content splicing, voting decision, etc. In a better embodiment, when there are resume files in multiple formats, after the resume parsing model obtains the initial structured resume data, the foregoing fusion is performed, and then information alignment is performed by querying the knowledge graph based on the fused fields and field content.

[0127] In another embodiment, when the collected resume files of candidates include multiple formats, after the processing of step S12, each format of file respectively obtains the corresponding resume text and the corresponding auxiliary feature data. Then, the multiple resume texts are merged and de-duplicated to obtain a text, and the corresponding auxiliary feature data is fused into the text. For example, as described in the foregoing method, the auxiliary feature data is converted into a description text and merged into the resume text, or converted into tag information and added to the text. Then, the text is input to the resume parsing model, and the final structured resume data is output after the processing of the resume parsing model.

[0128] In another embodiment, when users, such as recruitment users, enterprise human resources personnel (abbreviated as HR), etc. are screening the resumes of job applicants, they not only hope to obtain the structured resume data of the job applicants, but also hope to obtain the structured resume data of the job applicants that matches their recruitment requirements. Therefore, in this embodiment, when parsing the resume data, refer to Figure 1 , in step S11, when collecting the resume files of candidates, recruitment requirement information is also obtained; in step S12, it further includes extracting multiple recruitment feature data from the recruitment requirement information, where the recruitment feature data includes recruitment feature fields and field content. Compared with the model in the foregoing embodiment, the resume parsing model in this embodiment adds a matching processing layer, refer to Figure 6 , Figure 6 is a schematic diagram of the principle of the resume parsing model according to another embodiment of the present invention. The matching processing layer in this embodiment is used to match the obtained structured resume data with the recruitment feature data and output the corresponding structured resume data according to a specific output strategy. The output strategy is, for example, while outputting all the complete structured resume data, marking the fields and field content that match the recruitment feature data, or outputting comparison data, for example, outputting the fields and field content obtained from the resume and the recruitment fields and field content obtained based on the recruitment requirements in a table comparison manner and arranging them in a comparative manner.

[0129] On the other hand, the present invention also provides a resume data parsing device. Refer to Figure 7 , Figure 7 which is a schematic block diagram of a resume data parsing device according to an embodiment of the present invention. The resume data parsing device 10 in this embodiment includes a file collection module 11, a data extraction module 12, an industry category recognition module 13, and a parsing module 14. Among them, the file collection module 11 collects resume files of candidates, and the resume files include resume files in document format and / or non-document format. The non-document format is, for example, image format, audio format, or video format. The data extraction module 12 extracts resume text and auxiliary feature data from the resume files. The industry category recognition module 13 recognizes the industry category involved in the content of the resume file according to the resume text, and determines a corresponding resume parsing model according to the industry category. Among them, the resume parsing model is a deep learning model trained based on multi-modal resume data and industry knowledge graph of the industry category. The parsing module 14 obtains structured resume data based on the resume text and auxiliary feature data by using the resume parsing model. The structured resume data includes multiple resume fields and field contents that match the industry category. In addition, the present invention also provides a model training module 15 for training models as needed. For example, training the aforementioned resume parsing model, training an industry classification model for industry classification, or training various models required for extracting resume text and auxiliary feature data from resume files. After training is completed, they are stored in the database 20. The resume data parsing device 10 calls the corresponding models as needed when parsing resumes. For specific details, please refer to the relevant descriptions in the foregoing method, which will not be elaborated here.

[0130] In addition, in another embodiment, when the file collection module 11 collects resume files of candidates, it also collects recruitment requirement information of recruitment users. Correspondingly, the data extraction module 12 extracts multiple recruitment feature data from the recruitment requirement information, where the recruitment feature data includes recruitment feature fields and field contents. The parsing module 14 obtains structured resume data based on the resume text, recruitment feature data, and auxiliary feature data by using the resume parsing model.

[0131] Figure 8It is a schematic block diagram of a first application system according to an application embodiment of the present invention. In this embodiment, the first application system 100 includes a client and a server. The client is installed in the user terminal device 101 as a client, and it includes a user interaction module. The user terminal device 101 can be a smart phone, a desktop computer, a laptop computer, etc. The server includes a resume data parsing device 10 installed in the server 102. The user can send a resume parsing request to the server through the user interaction module. When sending the resume parsing request, the user can also input a candidate resume file or specify a candidate. In a specific application, when the user uses the first application system 100, the user logs in to the first application system 100 through the user name. Therefore, when the server receives the resume parsing request, it sends a notification to the resume data parsing device 10. The resume data parsing device 10 determines the user according to the user identifier in the request, receives the candidate resume file submitted by the user, or reads the candidate resume file from the database according to the specified candidate, and at the same time queries whether there is audio and video data of the user's interview with the candidate stored in the database. If so, it is used for parsing together with the resume file. After the resume data parsing device 10 of the server processes the resume file to obtain structured resume data, it sends it to the client 101, and the specific content can be directly displayed through the user interaction module, or returned to the client in the form of a file for the user to download to the local.

[0132] Figure 9 It is a schematic block diagram of a second application system according to another application embodiment of the present invention. In this embodiment, the second application system 200 communicates with the third-party system 300, and processes the resumes in the third-party database 301 according to the requirements of the third party to obtain structured resume data, and stores it in a public platform database 400 for use by systems with permissions. In a specific embodiment, the third-party system 300 is, for example, the management center of a recruitment platform, the third-party database 301 is, for example, the job-seeking user database of a recruitment platform, and the public platform database 400 is, for example, a public database of a recruitment platform. When each functional module in the recruitment platform needs to use the resume data of job-seeking users, it can read the structured resume data from the public platform database 400 as needed for convenient processing. The functional modules are, for example, the position / resume recommendation module, the advertisement recommendation module, and so on.

[0133] According to another aspect of the present invention, the present invention also provides an electronic device. Refer to Figure 10 , Figure 10It is a schematic diagram of the hardware structure principle of an electronic device according to an embodiment of the present invention. The electronic device can be implemented as the server in the foregoing application system, which includes a processor 601 and a memory 602. A set of program instructions is stored on the memory 602, and when the processor 601 executes the set of program instructions on the memory 602, the foregoing resume data parsing method is implemented.

[0134] Specifically, the foregoing processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0135] The memory 602 may include a mass storage for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 602 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 602 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid state memory.

[0136] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to execute the resume data parsing method provided by the present invention.

[0137] In one example, the electronic device may further include a communication interface 603 and a bus 604. The processor 601, the memory 602, and the communication interface 603 are connected through the bus 604 and complete communication with each other.

[0138] The communication interface 603 is mainly used to implement communication between various modules, devices, units, and / or devices in the embodiments of the present invention.

[0139] The bus 604 includes hardware, software, or both, and couples the components of the online data flow metering device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 604 may include one or more buses. Although embodiments of the present invention describe and illustrate specific buses, the present invention contemplates any suitable bus or interconnect.

[0140] The present invention also provides a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, any one of the resume data parsing methods in the foregoing embodiments can be implemented. The computer-readable storage medium may be any medium that tangibly contains or stores computer-executable instructions for use by or in connection with an instruction execution system, apparatus, and device. The storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium. Non-transient computer-readable storage media may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Corresponding embodiments of such storage devices include, for example, magnetic disks, optical discs based on CD, DVD, or Blu-ray technology, and persistent solid-state memories such as flash memory, solid-state drives, and the like.

[0141] The foregoing embodiments are only for illustrative purposes of the present invention and are not limitations on the present invention. Those of ordinary skill in the relevant art can make various changes and modifications without departing from the scope of the present invention. Therefore, all equivalent technical solutions should also fall within the scope of the disclosure of the present invention.

Claims

1. A resume data parsing method, characterized in that: include: Collecting resume files of candidates, wherein the resume files include resume files in document format and / or resume files in non-document format; Extracting resume text and auxiliary feature data from the resume file; Identify the industry category involved in the content of the resume file according to the resume text; Determine a corresponding resume parsing model according to the industry category, wherein the resume parsing model is a deep learning model trained based on multimodal resume data and industry knowledge graph of the industry category; as well as The resume parsing model is used to obtain structured resume data based on the resume text and the auxiliary feature data. The structured resume data includes a plurality of resume fields and field contents.

2. The resume data parsing method according to claim 1, characterized in that: When the resume file includes a resume file in a document format, the auxiliary feature data extracted from the resume file includes one or more of a text style, a text position, a table corresponding to the text, and an image.

3. The resume data parsing method according to claim 2, characterized in that: When the resume file in document format includes an image, the method further includes: extracting one or more of text information, image description in text form, and visual image feature data from the image.

4. The resume data parsing method according to claim 1, characterized in that: The non-document format is one or more of a picture format, an audio format and a video format.

5. The resume data parsing method according to claim 4, characterized in that: When the resume file is in picture format, the resume text extracted from the resume file is the resume text obtained by recognizing the text image in the picture and the table description and / or image description in text form; The auxiliary feature data includes visual image feature data.

6. The resume data parsing method according to claim 4, characterized in that: When the resume file is in audio format, the resume text extracted from the resume file includes the resume text obtained through speech recognition; the auxiliary feature data includes one or more of voice, intonation, high-frequency keywords, stress, weak sound, specific events and emotion types.

7. The resume data parsing method according to claim 4, characterized in that: When the resume file is in video format, the resume text extracted from the resume file includes the resume text obtained through speech recognition; the auxiliary feature data includes one or more of the voice, intonation, high-frequency keywords, stress, weak sound, specific events, expressions, body movements and emotional types of the characters in the video.

8. The resume data parsing method according to claim 1 is characterized in that after extracting the resume text and auxiliary feature data from the resume file, it further comprises: Based on the auxiliary feature type, converting the auxiliary feature data into marking information; as well as The marking information is used to mark the corresponding resume text block.

9. The resume data parsing method according to claim 1, characterized in that: The step of identifying the industry category involved in the resume file content according to the resume text includes: Extracting industry keywords from the resume text; and Based on the industry keywords, query the industry keyword library to determine matching industry categories; Wherein, the industry keyword library stores corresponding industry keywords by industry category; or Inputting the resume text into an industry classification model, wherein the industry classification model is a deep learning model trained using industry keywords as training data; and The industry category involved in the content of the resume file is obtained through the output of the industry classification model.

10. The resume data parsing method according to claim 1, characterized in that: Based on the resume text and the auxiliary feature data, the steps of obtaining structured resume data using the resume parsing model include: Fusion of auxiliary feature data into the resume text to obtain a new resume text; Inputting new resume text into the resume parsing model; and After being processed by the resume parsing model, structured resume data is output.

11. The resume data parsing method according to claim 10, characterized in that: Before fusing the auxiliary feature data into the resume text include: Querying the text segment corresponding to the auxiliary feature data from the resume text; and In response to finding a text segment corresponding to the auxiliary feature data from the resume text, the authenticity and / or consistency of the content of the text segment is verified based on the auxiliary feature data.

12. The resume data parsing method according to claim 10, characterized in that: The steps to fuse auxiliary feature data into resume text include: Based on the auxiliary feature type, converting the auxiliary feature data into a first description text; Check whether the resume text contains the corresponding second description text; In response to the resume text including the corresponding second description text, comparing the semantic scope of the first description text and the second description text; In response to the semantic scope of the second description text being greater than or equal to the first description text, retaining the second description text; in response to the semantic scope of the first description text and the semantic scope of the second description text intersecting, merging the first description text and the second description text and removing text with duplicate semantics; and In response to the resume text not containing the corresponding second description text, the first description text is merged into the resume text.

13. The resume data parsing method according to claim 1, characterized in that: When the resume files include resume files of multiple formats, the resume parsing model is used to obtain the structured resume data corresponding to each resume file based on the resume text and auxiliary feature data obtained from each resume file; as well as Based on the fusion strategy, the structured resume data obtained from each resume file is fused to obtain the final structured resume data.

14. The resume data parsing method according to claim 13, characterized in that: After obtaining the respective resume text and auxiliary feature data based on each resume file, the resume texts obtained from different resume files are cross-validated and the inconsistent content is marked; correspondingly, in the obtained structured resume data, the corresponding resume fields are marked; when the structured resume data obtained from each resume file are merged, the marked resume field content is retained.

15. The resume data parsing method according to claim 1, 10 or 13, characterized in that: Before inputting the resume text or new resume text into the resume parsing model, the sentence segmentation model is used to perform sentence segmentation to obtain segmented text; Inputting the segmented text into the resume parsing model; and After being processed by the resume parsing model, structured resume data is output.

16. The resume data analysis method according to claim 1, characterized in that Further including: Acquire recruitment demand information, and extract multiple recruitment feature data from the recruitment demand information, wherein the recruitment feature data includes recruitment feature fields and field contents; Correspondingly, based on the resume text, recruitment feature data and auxiliary feature data, a resume parsing model is used to obtain structured resume data; wherein, the resume parsing model is a deep learning model trained based on the multimodal resume data, recruitment feature data and industry knowledge graph of the industry category, and the deep learning model is used to extract fields and field contents that match the industry category and recruitment features from the multimodal resume data.

17. A resume data parsing device, characterized in that: include: A file collection module, configured to collect resume files of candidates, wherein the resume files include resume files in document format and / or resume files in non-document format; A data extraction module configured to extract resume text and auxiliary feature data from the resume file; An industry category identification module, configured to identify the industry category involved in the content of the resume file according to the resume text; as well as The parsing module is configured to determine a corresponding resume parsing model according to the industry category, and obtain structured resume data based on the resume text and auxiliary feature data using the resume parsing model. The structured resume data includes multiple resume fields and field contents, wherein the resume parsing model is a deep learning model obtained by training based on the multimodal resume data and industry knowledge graph of the industry category.

18. An electronic device comprising a processor and a memory, characterized in that: The memory stores a computer program instruction set, and when the processor executes the computer program instruction set on the memory, the resume data parsing method described in any one of claims 1-16 is implemented.

19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program instruction set, and when the computer program instruction set is executed by a processor, the resume data parsing method described in any one of claims 1-16 is implemented.

Citation Information

Patent Citations

  • A resume layout analysis method and device integrating multi-modality

    CN113743052B

  • Resume information extraction method and device, equipment and storage medium

    CN118314594A

Cited By

  • Resume analysis method and system based on dynamic semantic network

    CN120448469A

  • Dynamic resume evaluation method based on multi-modal large model

    CN121526543A