Large model-based resume key field extraction and analysis system and method
The large model-based resume analysis system addresses inefficiencies in traditional resume screening by converting formats, extracting key fields, and performing deep analysis to enhance accuracy and efficiency in resume screening.
Patent Information
- Application Number
- CN202510451251.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-15
AI Technical Summary
The existing resume screening system has inefficient and costly manual screening, and the inability to accurately identify and maintain the text structure of diverse resumes, resulting in the loss of key information or erroneous analysis.
The resume key field extraction and analysis system based on the big model is adopted, and the powerful learning and processing capabilities of the big model are used, combined with the propt technology, the resume format adaptation, keyword extraction and in-depth resume analysis are realized. Name, email, phone number and other information are accurately extracted through customized propt templates, and structured resume data is generated.
It has achieved rapid and accurate screening of resumes that meet the job requirements from a large number of resumes, improved recruitment efficiency and accuracy, avoided subjective factors and information loss caused by manual classification, and provided accurate resume analysis results.
Smart Images

Figure CN120317239A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human resource management, and specifically to a resume keyword field extraction and analysis system and method based on a large model. Background Art
[0002] With the development of Internet technology and the mature application of artificial intelligence technology, the recruitment industry is constantly exploring the use of new technologies to improve recruitment efficiency and accuracy. However, the existing resume screening systems have the problems of low efficiency and high cost of manual screening. Traditional resume screening methods are usually completed manually, collecting resume files manually and manually classifying resume texts, resulting in low accuracy of resume data information parsing, because manual classification processing often has subjective factors and is prone to information duplication or information loss.
[0003] Facing a large number of electronic resumes, the workload of manual screening is large and the efficiency is low. Existing resume parsing methods often cannot accurately identify and maintain the original structure of the text when dealing with diverse and complex-layout resumes, resulting in the loss or misparsing of key information. Therefore, how to quickly and accurately screen out electronic resumes that meet the job requirements from a large number of electronic resumes and efficiently extract key information from the resumes has become an urgent technical problem to be solved. Summary of the Invention
[0004] The purpose of the present invention is to provide a resume keyword field extraction and analysis system and method based on a large model, using the powerful learning and processing capabilities of the large model and combining prompt technology to accurately guide the model output, so as to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solutions: A resume keyword field extraction and analysis system based on a large model, including:
[0006] Resume upload module: providing a front-end batch upload interface, supporting the reception and format adaptation of resumes in PDF, Word, and TXT formats;
[0007] Keyword extraction module: taking the large model as the core, accurately extracting names, email addresses, and phone numbers through a customized prompt template;
[0008] Resume analysis module: deeply analyzing based on the large model, extracting educational information, work experience, and historical project information through complex prompt design, and generating structured resume data;
[0009] Front-end display layer: presenting the extraction and analysis results in a visual form in real time, supporting user interactive confirmation and export.
[0010] Preferably, the resume upload module includes:
[0011] Adopt a professional parsing engine for PDF files to convert complex layouts into text streams that preserve tables / charts;
[0012] When extracting text from Word files, mark special formats such as bold, underline, and title styles;
[0013] Perform whitespace character cleaning, garbled code filtering, and symbol standardization processing on TXT files;
[0014] All formatted resumes are finally uniformly converted into numbered plain text format for subsequent module calls.
[0015] Preferably, the keyword extraction module is used for:
[0016] Name extraction: Combine the surname library, common Chinese character library for personal names, and position features, and use the prompt to guide the model to identify rare surnames and special name combinations;
[0017] Email extraction: Build a regularization matching logic based on the "@" symbol and domain name rules, and use the prompt to enhance the model's semantic understanding of email formats;
[0018] Phone number extraction: Define the number of digits of the phone number, area code format, and international number identification rules through the prompt to support accurate positioning of multi-format phone numbers;
[0019] The extraction results are evaluated for confidence through a natural language processing model, and only high-confidence data is fed back to the front end.
[0020] Preferably, the resume analysis module includes:
[0021] Educational information extraction: Use the prompt to guide the model to identify school names, time intervals, educational levels, and major names, and build a chronological logic chain;
[0022] Work experience extraction: Based on the keyword library for duty descriptions and entity recognition of job titles, extract company names, positions, working time, and core work achievements;
[0023] Project information extraction: Define a four-element structure of project name, cycle, role, and achievement through the prompt to support parallel parsing of multiple projects;
[0024] Adopt a text structuring algorithm to logically arrange the extraction results and generate a visual resume graph containing a timeline and correlation analysis.
[0025] Preferably, the large model optimizes the domain adaptability to resume text through continuous fine-tuning, where:
[0026] Build a pre-training corpus containing millions of resume samples, covering Chinese-English bilingual and multi-industry terms;
[0027] Design a dynamic prompt generation mechanism to automatically adjust the focus of information extraction according to the resume type;
[0028] The system integrates a user feedback mechanism, provides an interface for manual correction of low-confidence extraction results, and iteratively optimizes model parameters based on the correction data.
[0029] A method for a resume keyword field extraction and analysis system based on a large model, comprising the following steps:
[0030] Receive resume files: Batch receive resumes in PDF, Word, and TXT formats uploaded by users through the front-end interface;
[0031] Format adaptation and cleaning: Parse multi-format resumes, convert them to a unified text format, retain key elements, and mark special formats;
[0032] Keyword field extraction: Use a large model combined with a customized prompt template to accurately extract name, email, and phone number from the text;
[0033] In-depth resume analysis: Process complex prompt instructions through a large model to extract education information, work experience, and historical project information;
[0034] Structured display: Visualize the extraction and analysis results through the front-end page, supporting user interaction verification and export.
[0035] Preferably, the format adaptation and cleaning step includes:
[0036] PDF processing: Use a professional parsing engine to convert the complex layout into an ordered text stream, retaining nested tables and charts;
[0037] Word processing: Mark special formats such as bold, underline, and title styles when extracting text, and generate metadata tags;
[0038] TXT processing: Execute a text cleaning process, including removing blank characters, correcting garbled characters, and standardizing symbols;
[0039] Unified output: Generate a numbered plain text file and attach format metadata for subsequent module calls.
[0040] Preferably, the keyword field extraction step includes:
[0041] Name extraction: Guide the model through prompts in combination with a surname library, a common Chinese character library for personal names, and position features to support the recognition of rare surnames;
[0042] Email extraction: Build a regularization matching logic based on the "@" symbol and domain name rules to enhance the model's semantic understanding of email formats;
[0043] Telephone extraction: Define the number of digits of the telephone number, the area code format, and the international number identification rule, and support accurate positioning of multi-format numbers;
[0044] Confidence evaluation: Perform a confidence score on the extraction result, and only output high-confidence data to the front end.
[0045] Preferably, the in-depth resume analysis step includes:
[0046] Educational information extraction: Guide the model to identify the school name, time interval, educational level, and major name through prompts, and construct a time-series logic chain;
[0047] Work experience extraction: Based on the keyword library of responsibility descriptions and entity recognition of job titles, extract the company name, position, working time, and core achievements;
[0048] Project information extraction: Define the four-element structure of project name, cycle, role, and achievement, and support parallel parsing of multiple projects;
[0049] Structured arrangement: Use text structuring algorithms to perform logical sorting on the extraction results, and generate a timeline and relevance analysis graph.
[0050] Preferably, it also includes an optimization mechanism:
[0051] Model fine-tuning: Based on the pre-trained corpus of millions of resume samples, continuously fine-tune to optimize the domain adaptability of the model to resume text;
[0052] Dynamic Prompt generation: Dynamically adjust the prompt instructions according to the resume type, and customize the focus of information extraction;
[0053] User feedback loop: Integrate an artificial correction interface to correct low-confidence results, and use the corrected data to iteratively optimize the model parameters.
[0054] Compared with the prior art, the beneficial effects of the present invention are:
[0055] The resume keyword field extraction and analysis system and method based on a large model proposed by the present invention, by using the form of large model prompts, uses the large model to extract key information such as educational experience, platform information, and project experience in each resume, and can quickly and accurately screen out resumes that meet the job requirements from a large number of electronic resumes, effectively solving the problems of low efficiency and high cost of manual screening.
[0056] Utilizing the powerful semantic understanding ability of the large model, it can deeply analyze and understand the content of the resume text, accurately extract the text containing valuable information and keyword fields, and avoid the subjective factors and situations of information duplication or information missing caused by manual classification processing.
[0057] Based on the powerful semantic understanding ability of large models, the present invention can effectively improve the accuracy and efficiency of analyzing user behavior, and provide more accurate resume analysis results for the company's talent recruitment specialists.
[0058] By using prompts and custom development methods, functions that meet business requirements are completed, effectively connecting functions such as resume keyword extraction and resume content analysis, forming a systematic process. Brief Description of the Drawings
[0059] Figure 1 This is the system architecture diagram of the present invention. Detailed Implementation Manner
[0060] In order to clearly and completely describe the purpose, technical solution of the present invention, and make the advantages more clear, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0061] Embodiment 1, please refer to Figure 1 The present invention provides a technical solution: a resume keyword field extraction and analysis system based on a large model, and its system structure diagram is as follows Figure 1 As shown, it mainly includes a resume upload module, a keyword extraction module, and a resume analysis module, etc. The resume upload module can facilitate users to batch upload resumes to be parsed at the front end and convert the uploaded resumes into an easily readable format. The keyword analysis module takes the large model as the core, and through prompt custom development of the keyword extraction function, can extract key information such as the name, email, and phone number contained in the resume and display it at the front end. The resume analysis module also takes the large model as the core, and through prompt custom development of the resume analysis function, automatically extracts information such as the educational information, work experience, and historical projects of the resume submitter and displays it on the front-end page. The detailed working principles of each module are as follows:
[0062] (1) Resume upload module:
[0063] This module provides users with a simple and intuitive front-end operation interface, supporting batch uploads of multiple common resume formats, such as PDF, Word, etc. After the user selects the resume files to be uploaded, the system will quickly initiate the file format adaptation process. For PDF-format resumes, a professional PDF parsing engine is used to convert their complex page layouts into an ordered text stream, while accurately identifying and retaining elements such as nested tables and charts that may contain key information. For Word-format resumes, it directly delves into the document's interior to efficiently extract the text content and marks and records special formats in the document, such as bold, underlined, and title styles, so that subsequent modules can better understand the importance and structure of the text. If a TXT-format resume is uploaded, it will perform preliminary text cleaning to remove redundant whitespace characters, garbled codes, and special symbol interferences, uniformly convert all uploaded resumes into a plain text format convenient for subsequent system processing, and number them in the order of upload to prepare for subsequent keyword extraction and resume analysis work.
[0064] (2) Keyword Extraction Module:
[0065] This module is driven by a large model at its core. First, for the extraction of basic key information such as name, email, and phone number, specially adapted prompt templates are developed. For example, for name extraction, the prompt will guide the large model to focus on words in the text that conform to the writing norms and position characteristics of personal names, and make a comprehensive judgment in combination with the surname library and the common Chinese character library for personal names. Even when encountering rare surnames or special name combinations, it can accurately identify them. When extracting an email, the prompt will let the large model quickly locate from the resume text based on the specific character combination rules of the email address, such as containing the "@" symbol and specific domain name formats. For phone number extraction, the prompt guides the large model to match common phone number digits, area code formats, and other characteristics. After receiving the resume text and the customized prompt, the large model uses its powerful natural language processing capabilities to perform in-depth semantic analysis on the text, accurately extracts the corresponding keywords, and immediately feeds back the extraction results to the front-end page, presenting them to the user in a clear and understandable manner for the user to intuitively confirm the accuracy of the key information.
[0066] (3) Resume Analysis Module:
[0067] Also built based on large models, it achieves in-depth information extraction through a series of carefully designed complex prompts. When extracting educational information, the prompts guide the large model to identify key elements such as school names, enrollment time, graduation time, major names, educational levels, etc., and utilize the model's understanding ability of education-related vocabulary and time series to sort out the complete educational context of the candidate. For work experience extraction, the prompts enable the large model to focus on information such as company names, job titles, working time, and job responsibility descriptions, and accurately extract and organize work resumes by analyzing the logical structure and language expression characteristics of work experience in the text. When mining historical project information, the prompts guide the large model to identify key content such as project names, project start and end times, project results, and personal roles and contributions in the project. After the large model completes information extraction, the system uses advanced text structuring algorithms to organize and arrange the extracted information in a reasonable logical order and display it in a clear section form on the front-end page. Recruitment specialists can quickly obtain comprehensive and well-organized core resume information of candidates without cumbersome manual screening and sorting, greatly improving the efficiency and decision-making accuracy of the recruitment work.
[0068] Example 2, based on Example 1, proposes a method for a resume keyword field extraction and analysis system based on a large model, including the following steps:
[0069] Receive resume files: Batch receive resumes in PDF, Word, and TXT formats uploaded by users through the front-end interface;
[0070] Format adaptation and cleaning: Parse multi-format resumes, convert them into a unified text format, retain key elements and mark special formats; including: PDF processing: Use a professional parsing engine to convert complex layouts into an ordered text stream, retaining nested tables and charts; Word processing: Mark special formats such as bold, underlined, and title styles when extracting text, and generate metadata tags; TXT processing: Execute a text cleaning process, including removing blank characters, correcting garbled characters, and standardizing symbols; Unified output: Generate a numbered plain text file and attach format metadata for subsequent modules to call.
[0071] Keyword field extraction: Use a large model combined with a customized prompt template to accurately extract names, email addresses, and phone numbers from the text; including: Name extraction: Guide the model to combine a surname library, a common Chinese character library for personal names, and position features through prompts to support the recognition of rare surnames; Email extraction: Build a regularization matching logic based on the "@" symbol and domain name rules to enhance the model's semantic understanding of email formats; Phone extraction: Define phone number digits, area code formats, and international number identification rules to support accurate positioning of multi-format numbers; Confidence evaluation: Score the confidence of the extraction results and only output high-confidence data to the front end.
[0072] In-depth resume analysis: Process complex prompt instructions through large models to extract educational information, work experience, and historical project information, including: Educational information extraction: Guide the model to identify school names, time intervals, educational levels, and major names through prompts, and construct a chronological logic chain; Work experience extraction: Based on the keyword library of responsibility descriptions and entity recognition of job names, extract company names, positions, working time, and core achievements; Project information extraction: Define a quadruple structure of project names, cycles, roles, and achievements to support parallel parsing of multiple projects; Structured arrangement: Use text structuring algorithms to logically sort the extraction results and generate a timeline and relevance analysis graph.
[0073] Structured display: Visualize the extraction and analysis results through a front-end page, supporting user interaction verification and export.
[0074] It also includes an optimization mechanism: Model fine-tuning: Based on a pre-trained corpus of millions of resume samples, continuously fine-tune to optimize the domain adaptability of the model to resume text; Dynamic Prompt generation: Dynamically adjust prompt instructions according to resume types to customize the focus of information extraction; User feedback loop: Integrate an artificial correction interface to correct low-confidence results and use the corrected data to iteratively optimize model parameters.
[0075] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A resume keyword field extraction and analysis system based on a large model, characterized in that: It includes: Resume Upload Module: Provides a front-end batch upload interface, supporting the reception and format adaptation of resumes in PDF, Word, and TXT formats; Keyword Extraction Module: With a large model at its core, it accurately extracts names, email addresses, and phone numbers by customizing the prompt template; Resume Analysis Module: Based on in-depth analysis by a large model, it extracts educational information, work experience, and historical project information through complex prompt design and generates structured resume data; Front-end Display Layer: Presents the extraction and analysis results in a visual form in real time, supporting user interactive confirmation and export.
2. The resume keyword field extraction and analysis system based on a large model according to claim 1, wherein: The Resume Upload Module includes: For PDF files, a professional parsing engine is used to convert the complex layout into a text stream that retains tables / charts; When extracting text from Word files, special formats such as bold, underlined, and title styles are marked; For TXT files, blank character cleaning, garbled character filtering, and symbol standardization processing are performed; All format resumes are finally uniformly converted into a numbered plain text format for subsequent module calls.
3. The resume keyword field extraction and analysis system based on a large model according to claim 2, wherein: The Keyword Extraction Module is used for: Name extraction: Combining the surname library, common Chinese character library for personal names, and position features, the model is guided by prompts to identify rare surnames and special name combinations; Email extraction: Based on the "@" symbol and domain name rules, a regularization matching logic is constructed, and the model's semantic understanding of email formats is enhanced through prompts; Phone number extraction: By defining the number of digits, area code format, and international number identification rules for phone numbers through prompts, accurate positioning of multi-format phone numbers is supported; The extraction results are evaluated for confidence by a natural language processing model, and only high-confidence data is fed back to the front end.
4. The resume keyword field extraction and analysis system based on a large model according to claim 3, characterized in that: The Resume Analysis Module includes: Educational information extraction: Guiding the model to identify school names, time intervals, educational levels, and major names through prompts and constructing a time sequence logic chain; Work experience extraction: Based on the keyword library for duty descriptions and entity recognition of job titles, it extracts company names, positions, working time, and core work achievements; Project information extraction: By defining the four-element structure of project name, cycle, role, and achievements through prompts, parallel parsing of multiple projects is supported; A text structuring algorithm is used to logically arrange the extraction results to generate a visual resume graph containing a timeline and correlation analysis.
5. The resume keyword field extraction and analysis system based on a large model according to claim 4, characterized in that: The large model continuously fine-tunes and optimizes the domain adaptability to resume text, where: A pre-training corpus containing millions of resume samples is established, covering Chinese-English bilingual and multi-industry terms; A dynamic prompt generation mechanism is designed to automatically adjust the focus of information extraction according to resume types; The system integrates a user feedback mechanism, provides an artificial correction interface for low-confidence extraction results, and iteratively optimizes model parameters based on the correction data.
6. A method for a resume keyword field extraction and analysis system based on a large model according to claim 5, characterized in that: It includes the following steps: Receive resume files: Batch receive resumes in PDF, Word, and TXT formats uploaded by users through the front-end interface; Format adaptation and cleaning: Parse multi-format resumes, convert them into a unified text format, retain key elements, and mark special formats; Keyword field extraction: Use a large model combined with a customized prompt template to accurately extract names, email addresses, and phone numbers from the text; Deep resume analysis: Process complex prompt instructions through large models to extract educational information, work experience, and historical project information; Structured display: Visualize the extraction and analysis results through the front-end page, supporting user interaction verification and export.
7. A method according to claim 6, characterized in that: Format adaptation and cleaning steps include: PDF processing: Use a professional parsing engine to convert complex layouts into ordered text streams, retaining nested tables and charts; Word processing: Mark special formats such as bold, underlined, and title styles when extracting text, and generate metadata tags; TXT processing: Execute the text cleaning process, including removing blank characters, correcting garbled characters, and standardizing symbols; Unified output: Generate a numbered plain text file and attach format metadata for subsequent module calls.
8. A method according to claim 7, characterized in that: Keyword extraction steps include: Name extraction: Guide the model through prompts to combine the surname library, common Chinese character library for personal names, and position features to support the recognition of rare surnames; Email extraction: Build a regularized matching logic based on the "@" symbol and domain name rules to enhance the model's semantic understanding of email formats; Phone number extraction: Define the number of digits of the phone number, area code format, and international number identification rules to support accurate positioning of multi-format numbers; Confidence evaluation: Score the confidence of the extraction results and only output high-confidence data to the front end.
9. A method according to claim 8, characterized in that: Deep resume analysis steps include: Educational information extraction: Guide the model through prompts to identify school names, time intervals, educational levels, and major names, and build a chronological logic chain; Work experience extraction: Based on the keyword library for responsibility descriptions and entity recognition of job titles, extract company names, positions, working times, and core achievements; Project information extraction: Define the four-tuple structure of project name, cycle, role, and achievement to support parallel parsing of multiple projects; Structured arrangement: Use text structuring algorithms to logically sort the extraction results and generate a timeline and relevance analysis graph.
10. A method according to claim 9, wherein: It also includes an optimization mechanism: Model fine-tuning: Based on the pre-trained corpus of millions of resume samples, continuously fine-tune to optimize the model's domain adaptability to resume text; Dynamic Prompt generation: Dynamically adjust prompt instructions according to resume types to customize the focus of information extraction; User feedback loop: Integrate an artificial correction interface to correct low-confidence results and use the corrected data to iteratively optimize model parameters.