An intelligent resume generation system based on natural language processing

The intelligent resume generation system based on natural language processing automatically retrieves information from recruitment websites and filters content from job seekers' work experience, solving the problem of time-consuming and labor-intensive traditional resume generation. It achieves efficient and targeted resume generation, improving the efficiency of job seekers' applications.

CN119558293BActive Publication Date: 2026-04-17HUBEI JINXIU TALENT TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUBEI JINXIU TALENT TECH GRP CO LTD
Filing Date
2024-11-13
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Traditional resume generation is time-consuming and labor-intensive. Job seekers need to manually modify a large number of resumes to adapt to the recruitment requirements of different companies, making it difficult to improve efficiency when submitting a large number of applications.

Method used

An intelligent resume generation system based on natural language processing is adopted. It crawls information from recruitment websites through a job information acquisition module, uses natural language processing technology to filter information from job seekers' work experience, and automatically fills it into the resume template to generate a highly targeted resume.

Benefits of technology

It enables the rapid generation of targeted resumes, improves the efficiency of job seekers submitting resumes in batches, ensures that each resume is suitable for the recruitment requirements of different companies, and enhances the effectiveness of resume submission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119558293B_ABST
    Figure CN119558293B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of automatic resume generation and discloses an intelligent resume generation system based on natural language processing. The system includes a job information acquisition module, a resume information storage module, a template storage module, and a resume generation module. The job information acquisition module retrieves job postings from recruitment websites. The resume information storage module stores the job seeker's work experience information. The template storage module stores resume templates, which include a basic information field and a work experience field. The basic information field contains the job seeker's basic information pre-filled. The resume generation module uses natural language processing technology to filter the work experience information based on the job postings, obtains filtered information, and fills this filtered information into the work experience field of the resume template, thereby generating a resume. This invention achieves intelligent resume generation, effectively improving the efficiency of job seekers submitting resumes in batches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic resume generation, and more particularly to an intelligent resume generation system based on natural language processing. Background Technology

[0002] The traditional resume creation process typically involves job seekers manually modifying resume templates to meet the specific requirements of different job postings. For the same type of position, different companies have different priorities, so job seekers need to manually revise their resumes to better suit the varying job requirements. For job seekers aiming to gain more interview opportunities by sending out a large number of resumes, revising numerous resumes is extremely time-consuming. Summary of the Invention

[0003] The purpose of this invention is to disclose an intelligent resume generation system based on natural language processing, which solves the technical problem of how to help job seekers generate resumes more quickly.

[0004] To achieve the above objectives, the present invention adopts the following technical solution:

[0005] This invention provides an intelligent resume generation system based on natural language processing, including a job information acquisition module, a resume information storage module, a template storage module, and a resume generation module;

[0006] The job information acquisition module is used to obtain job information from recruitment websites;

[0007] The resume information storage module is used to store job seekers' work experience information;

[0008] The template storage module is used to store resume templates. The resume templates include a basic information field and a work experience field. The basic information field contains the job seeker's basic information in advance.

[0009] The resume generation module uses natural language processing technology to filter work experience information based on job postings, obtain filtering information, and fill in the salary and experience fields in the resume template to generate a resume.

[0010] Preferably, the job information acquisition module includes a control unit and a crawler unit;

[0011] The control unit is used to adaptively calculate the access interval based on the crawled records;

[0012] The crawler unit is used to access recruitment websites based on access intervals and crawl the recruitment information published by the recruitment websites.

[0013] Preferably, the job information acquisition module also includes a comparison unit and a storage unit;

[0014] The comparison unit is used to compare all the recruitment information obtained in this crawl with all the recruitment information obtained in the previous crawl to determine whether there is any newly added recruitment information. If so, the newly added recruitment information is stored and sent to the storage unit for storage.

[0015] Preferably, the crawling records include the crawling time, the server response time, the received HTTP status codes, and the number of newly added job postings obtained during the crawl.

[0016] Preferably, the recruitment information includes basic information, job description, educational requirements, work experience requirements, skill requirements, and certification requirements;

[0017] The job posting website contains multiple pages, and each page contains multiple job postings.

[0018] Preferably, the basic information includes the job title, work location, and company name.

[0019] The job description includes the job duties;

[0020] Work experience requirements include project experience requirements and years of experience in the position;

[0021] Skill requirements are the skills that the personnel recruited for this position must possess;

[0022] Skill requirements refer to the certificates that the personnel recruited for this position must possess.

[0023] Preferably, the resume generation module includes a keyword acquisition unit and an extraction unit;

[0024] The keyword acquisition unit is used to obtain a set of keywords from recruitment information;

[0025] The extraction unit is used to obtain filtering information from job seekers' work experience information based on a set of keywords.

[0026] Preferably, the keyword set obtained from the recruitment information includes:

[0027] The recruitment information text is segmented into words to obtain a word set;

[0028] Calculate the ranking coefficient for each word in the word set;

[0029] Obtain a set of keywords based on ranking coefficients.

[0030] Preferably, the recruitment information text is segmented to obtain a word set, including:

[0031] Remove useless characters from the recruitment information text to obtain the text to be segmented.

[0032] The word segmentation algorithm is used to segment the text to be segmented, resulting in multiple words. The obtained words are then stored in a word set.

[0033] Remove stop words from the word set.

[0034] Preferably, the keyword set obtained based on the ranking coefficient includes:

[0035] The top S×N words with the highest ranking coefficients are stored as keywords in the keyword set, where N represents the total number of words in the keyword set and S represents the preset proportion.

[0036] Beneficial effects:

[0037] This invention automatically retrieves recruitment information from job websites and uses natural language processing technology to automatically extract filtering information from job seekers' work experience information to fill in the work experience section of the resume template, thereby achieving intelligent resume generation. This invention can quickly generate a single resume text when job seekers need to submit resumes in batches, effectively improving the efficiency of batch resume submissions. Furthermore, the generated resumes are targeted, not submitted to all companies with the same resume, thus achieving better resume submission results. Attached Figure Description

[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a schematic diagram of an intelligent resume generation system based on natural language processing according to the present invention.

[0040] Figure 2 This is a schematic diagram illustrating the process of obtaining a word set according to the present invention. Detailed Implementation

[0041] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0042] like Figure 1 As shown in one embodiment, the present invention provides an intelligent resume generation system based on natural language processing, including a job information acquisition module, a resume information storage module, a template storage module, and a resume generation module;

[0043] The job information acquisition module is used to obtain job information from recruitment websites;

[0044] The resume information storage module is used to store job seekers' work experience information;

[0045] The template storage module is used to store resume templates. The resume templates include a basic information field and a work experience field. The basic information field contains the job seeker's basic information in advance.

[0046] The resume generation module uses natural language processing technology to filter work experience information based on job postings, obtain filtering information, and fill in the salary and experience fields in the resume template to generate a resume.

[0047] The above process automatically retrieves recruitment information from job websites and uses natural language processing technology to automatically extract filtering information from job seekers' work experience information to fill in the work experience section of the resume template, thus achieving intelligent resume generation. It can quickly generate a single resume text when job seekers need to submit resumes in batches, effectively improving the efficiency of batch resume submissions. Furthermore, the generated resumes are targeted, not submitted to all companies with the same resume, resulting in better resume submission outcomes.

[0048] Specifically, the basic information field in the resume template is for filling in the job seeker's basic information, including age, contact information, name, ID number, education, and certificates.

[0049] The basic information field is located above the work experience field.

[0050] Preferably, the job information acquisition module includes a control unit and a crawler unit;

[0051] The control unit is used to adaptively calculate the access interval based on the crawled records;

[0052] The crawler unit is used to access recruitment websites based on access intervals and crawl the recruitment information published by the recruitment websites.

[0053] Using adaptive access intervals to retrieve job postings can effectively reduce the probability of problems arising from fixed access intervals being too long or too short. If the access interval is too long, the latest job postings cannot be retrieved in a timely manner, preventing job seekers from applying for those positions promptly. Conversely, if the access interval is set too short, it is easily identified as a web crawler by the server, triggering anti-crawler mechanisms and requiring operations such as changing IP addresses to continue crawling job postings.

[0054] Preferably, the crawling records include the crawling time, the server response time, the received HTTP status codes, and the number of newly added job postings obtained during the crawl.

[0055] Specifically, server response time refers to the time required from when the client sends an access request to when the recruitment website's server completes the request and returns the response data.

[0056] HTTP status codes are three-digit codes returned by a server in response to a client's request in the HTTP protocol. These status codes inform the client of the processing result of the request so that the client can take appropriate action based on the different status codes.

[0057] The number of newly added job postings obtained during this crawl refers to the number of new job postings published on the job website compared to the previous visit.

[0058] Preferably, the access interval is adaptively calculated based on the crawled records, including:

[0059] After completing the previous crawl, the interval for the next visit to the recruitment website is calculated:

[0060] Use S h Let S represent the access interval between the h-th and h-1-th visits to the recruitment website. h The calculation formula is:

[0061]

[0062] S h-1 This indicates the interval between the (h-1)th and (h-2)th visits to the recruitment website, bsh-1 and bs h-2 Let $\mathbf{h-1}$ and $\mathbf{h-2}$ represent the access characteristic values ​​of the h-th and h-th visits to the recruitment website, respectively, where $s$ represents the pre-set duration and $S$ represents the access characteristic values ​​of the h-th and h-th visits to the recruitment website. mi and S ma This indicates the minimum and maximum values ​​of the pre-set access interval;

[0063] The value of h is greater than 2, and the values ​​of S1 and S2 are S ma ;

[0064] For the i-th visit to the recruitment website, i∈{h-1,h-2}, the formula for calculating the visit feature value is:

[0065]

[0066] bs i Resp represents the access characteristic value of the i-th visit to the recruitment website. i Resp represents the server response time during the i-th access to the recruitment website. cmp wrognum represents the preset response time comparison value. i `wrgtypnum` represents the number of HTTP status codes of the preset type among all types of HTTP status codes returned by the server during the i-th access to the recruitment website, and `newnum` represents the total number of all types of HTTP status codes returned by the server during the i-th access to the recruitment website. i newnum represents the number of newly added job postings retrieved during the i-th visit to the job website. cmp d1 represents the preset quantity comparison value, d2 represents the response time calculation factor, d3 represents the status code calculation factor, and d3 represents the new quantity calculation factor.

[0067] The access interval of this invention varies based on the previous access interval, which reduces the impact of occasional high-frequency updates on the overall access interval, making the access interval more adaptable to different access results. Specifically, when the access feature value of the (h-1)th access is less than the access feature value of the (h-2)th access, the access interval of this invention will decrease accordingly; conversely, the access interval will increase accordingly. Furthermore, the magnitude of the increase or decrease is related to bs. h-1 and bs h-2 The absolute values ​​of the differences between them are related, bs h-1 and bs h-2The larger the absolute value of the difference between them, the greater the increase or decrease, which allows the access interval to follow the changes in the access results more closely, thus improving the effectiveness of the access interval. This can reduce the probability of being identified by anti-crawler mechanisms while avoiding excessively large access intervals that would prevent timely access to newly released job information on recruitment websites.

[0068] The access characteristic value of this invention is obtained through multi-faceted calculation based on response time, HTTP status codes, and the number of newly added job postings. This allows the access characteristic value to better represent the urgency of accessing the job recruitment website. Specifically, when the response time is longer, the number of HTTP status codes of the preset type among all types of HTTP status codes returned by the server is greater, and the number of newly added job postings is smaller, it indicates that the interval between the next access should be extended. Conversely, when the response time is shorter, the number of HTTP status codes of the preset type among all types of HTTP status codes returned by the server is smaller, and the number of newly added job postings is larger, it indicates that the interval between the next access should be shortened.

[0069] This feature value calculation method can not only effectively reduce the access pressure on recruitment websites caused by the crawling process of recruitment websites in this invention, but also obtain updated recruitment information more timely.

[0070] Preferably, the values ​​of d1, d2, and d3 are respectively and

[0071] Preferably, the preset response time comparison value is 5 seconds.

[0072] Preferably, the quantity comparison value is 1000.

[0073] Preferably, the default HTTP status codes include 429 and 430.

[0074] Preferably, the minimum and maximum access intervals are 1 minute and 12 minutes, respectively.

[0075] Preferably, the preset duration is 2 minutes.

[0076] Preferably, the recruitment website is accessed at intervals to crawl the recruitment information published on the website, including:

[0077] After the (h-1)th crawl is completed, calculate the access interval for the h-th visit to the recruitment website;

[0078] After calculating the interval of the h-th visit to the recruitment website, a process with a duration of S is performed. hThe countdown,

[0079] After the countdown ends, access the recruitment website. During the access, crawl the recruitment information from the recruitment website.

[0080] Preferably, the job information acquisition module also includes a comparison unit and a storage unit;

[0081] The comparison unit is used to compare all the recruitment information obtained in this crawl with all the recruitment information obtained in the previous crawl to determine whether there is any newly added recruitment information. If so, the newly added recruitment information is stored and sent to the storage unit for storage.

[0082] Specifically, in order to avoid missing some job postings due to the large amount of job postings published on job websites, this invention adopts a holistic crawling approach, that is, each crawl retrieves all job postings for the specified positions on the website.

[0083] All job postings obtained in this crawl and all job postings obtained in the previous crawl can be stored in two separate text files. The diff algorithm can then be used to extract the portion of the text containing the job postings obtained in this crawl that is subject to insertion, thereby obtaining the newly added job postings.

[0084] The Diff algorithm is a common text comparison algorithm that identifies differences between two texts. Many programming languages ​​have libraries that implement this algorithm, such as Python's difflib module.

[0085] The Diff algorithm primarily finds differences by comparing lines or characters between two texts and assigns an operation (insertion, deletion, or modification) to each difference. Classic Diff algorithms include the greedy algorithm and the Longest Common Subsequence (LCS) algorithm. The greedy algorithm compares lines or characters in two texts, matching as many common parts as possible and assigning deletion or insertion operations to the different parts. The LCS algorithm, on the other hand, finds the longest common subsequence between the two texts, identifying the most common parts and assigning deletion or insertion operations to the different parts.

[0086] Preferably, the recruitment information includes basic information, job description, educational requirements, work experience requirements, skill requirements, and certification requirements;

[0087] The job posting website contains multiple pages, and each page contains multiple job postings.

[0088] Specifically, when you enter a job title in the search box of a recruitment website, the search results are usually displayed in a list format. To avoid the data volume on a single page being too large, the search results are usually displayed in a paginated manner.

[0089] Preferably, the basic information includes the job title, work location, and company name.

[0090] The job description includes the job duties, such as what the main responsibilities are;

[0091] Work experience requirements include project experience requirements and years of service requirements. Project experience requirements are for recruiting people with relevant experience so that they can start working directly and reduce training time. Many positions have years of service requirements, such as more than 5 years of experience in human resources management.

[0092] Skill requirements refer to the skills that the person being recruited for this position must possess. Different positions require different skills. For example, for programmers, applicants are generally required to have proficiency in one or more programming algorithms.

[0093] Skill requirements refer to the certificates that the personnel recruited for the position must possess. For example, for lawyers, a lawyer's license is generally required.

[0094] Preferably, the resume generation module includes a keyword acquisition unit and an extraction unit;

[0095] The keyword acquisition unit is used to obtain a set of keywords from recruitment information;

[0096] The extraction unit is used to obtain filtering information from job seekers' work experience information based on a set of keywords.

[0097] Keyword acquisition algorithms can extract keywords from job postings, which can then be used to filter information and generate targeted text.

[0098] Preferably, the resume generation module also includes a fill-in unit, which is used to fill in the screening information into the salary and experience field in the resume template.

[0099] Preferably, the keyword set obtained from the recruitment information includes:

[0100] The recruitment information text is segmented into words to obtain a word set;

[0101] Calculate the ranking coefficient for each word in the word set;

[0102] Obtain a set of keywords based on ranking coefficients.

[0103] Preferably, the recruitment information text is segmented to obtain a word set, such as... Figure 2 As shown, it includes:

[0104] Remove the useless characters from the text of the recruitment information to obtain the text to be segmented;

[0105] Use a word segmentation algorithm to segment the text to be segmented, obtain multiple words, and store the obtained words in a word set;

[0106] Delete the stop words in the word set.

[0107] Specifically, the useless characters include punctuation marks.

[0108] Specifically, the word segmentation algorithm can be the forward maximum matching method, the backward maximum matching method, the bidirectional maximum matching method, etc.

[0109] The forward maximum matching method matches several consecutive characters in the text to be segmented with the word list from left to right. If the match is successful, a word is segmented. The core idea of the forward maximum matching method is to use the longest word as much as possible during the segmentation process to achieve the best segmentation effect.

[0110] The basic principle of the backward maximum matching method is similar to that of the forward maximum matching algorithm, but the segmentation direction is opposite. The backward maximum matching method starts from the end of the processed document for matching and scanning. Each time, the last i characters (the longest number of words in the dictionary) are used as the matching field. If the match fails, the first character of the matching field is removed and the matching continues. The word segmentation dictionary used by the backward maximum matching method is a reverse dictionary, and each entry in it will be stored in reverse order.

[0111] The bidirectional maximum matching method is an algorithm that combines the forward maximum matching method and the backward maximum matching method, aiming to improve the accuracy and efficiency of word segmentation. This method first segments the text using the forward maximum matching method and the backward maximum matching method respectively, and then compares the results of the two methods, and selects the one with fewer word segmentations as the final result.

[0112] Specifically, the stop words include "de", "le", "shi", etc.

[0113] The steps to delete the stop words in the word set are as follows:

[0114] Establish a stop word list before word segmentation: First, a stop word list needs to be prepared, which contains those words that are not very important in the analysis, frequently appear but contribute little to understanding the text (such as "de", "le", "shi", etc.).

[0115] Filter the word segmentation results: Compare the segmented words with the stop word list and remove the words in the list.

[0116] Preferably, obtain the keyword set based on the ranking coefficient, including:

[0117] The top S×N words with the highest ranking coefficients are stored as keywords in the keyword set, where N represents the total number of words in the keyword set and S represents the preset proportion.

[0118] For example, the preset ratio is 5%.

[0119] This invention does not obtain keywords by setting a fixed ranking coefficient threshold. This is because a higher frequency of a word usually indicates greater importance and a higher likelihood of it being selected as a keyword. If a fixed ranking coefficient threshold is set, then when the job posting text is too long, the calculated keyword ranking information will be too low, potentially resulting in the absence of words with a ranking coefficient higher than the threshold, leading to keyword acquisition failure.

[0120] The preferred formula for calculating the ranking coefficient is as follows:

[0121]

[0122] ordval v The rank coefficient of word v, numloca v Locarenum represents the total number of words (v) contained within a local scope. v α represents the total number of words v contained in the job posting text; ovenum represents the total number of words contained in the job posting text; and α represents the adaptive weight.

[0123] The ranking coefficient of this invention is calculated by comprehensively considering both the number of words "v" within a specific scope and the number of words "v" throughout the entire job posting text, thereby further improving the accuracy of the extracted keywords. If the ranking coefficient is calculated based solely on a local scope, a word with a low frequency within that local scope but a high frequency throughout the entire job posting text may not be selected as a keyword, which is clearly incorrect. Conversely, if the ranking coefficient is calculated based solely on a global scope, it is easy to select commonly used, frequently occurring industry-wide terms as the final keywords, failing to identify more targeted terms and resulting in poorly targeted resumes.

[0124] Specifically, the words counted by locarenum and ovenum are words from the word set.

[0125] The preferred formula for calculating adaptive weights is:

[0126]

[0127] Θ represents a preset value.

[0128] The weights in this invention are not fixed values, because fixed values ​​are not well-suited to different text distributions. When 'v' is a keyword related to job postings, if the word 'v', which is a keyword for job postings, only appears frequently in a local area and less frequently in other areas, setting the weight too high can easily lead to an overestimation of the calculated ranking coefficient. Conversely, setting the weight too low can lead to over-reliance on the local distribution characteristics of the words without considering the global distribution characteristics, thus reducing the accuracy of the final keyword. Therefore, this invention uses the distribution characteristics of words in the text to calculate adaptive weights, effectively improving the applicability of the calculated weights and, consequently, the applicability of the calculated ranking coefficient.

[0129] Preferably, the preset value is 1.

[0130] Preferably, the local area is determined in the following way:

[0131] When v first appears in the text of the job posting, position L v When the line is in the range [L], the text corresponding to the local scope is the line number in the recruitment information text. v -NL,L v [+NL] contains all text within the range, where NL represents the preset line number.

[0132] Preferably, the text of the recruitment information is formatted in the same way as the recruitment information is formatted at 1080p resolution in the original webpage. For example, the original webpage specifies that each line should contain M characters at 1080p resolution. Here, M can be determined based on the font size; the larger the font, the smaller M.

[0133] For example, if the width of the text is 10.8, then the value of M is 100.

[0134] Preferably, the preset number of rows is 5.

[0135] Preferably, filtering information is obtained from job seekers' work experience information based on keyword sets, including:

[0136] The job seeker's work experience information includes multiple project experiences;

[0137] The filter information is all project experience information whose text contains keywords belonging to the keyword set.

[0138] Project experience information includes the project name, the work responsibilities in the project, a brief introduction of the project, and the project results.

[0139] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An intelligent resume generation system based on natural language processing, characterized in that, It includes a job information acquisition module, a resume information storage module, a template storage module, and a resume generation module; The job information acquisition module is used to obtain job information from recruitment websites; The resume information storage module is used to store job seekers' work experience information; The template storage module is used to store resume templates. The resume templates include a basic information field and a work experience field. The basic information field contains the job seeker's basic information in advance. The resume generation module uses natural language processing technology to filter work experience information based on job postings, obtain filtering information, and fill in the salary and experience fields in the resume template to generate a resume. The job information acquisition module includes a control unit and a crawler unit; The control unit is used to adaptively calculate the access interval based on the crawled records; The crawler unit is used to access recruitment websites based on access intervals and crawl the recruitment information published by the recruitment websites; The access interval is adaptively calculated based on the crawled records, including: After completing the previous crawl, the interval for the next visit to the recruitment website is calculated: use Let h represent the interval between the h-th and h-1-th visits to the recruitment website. The calculation formula is: This indicates the interval between the (h-1)th and (h-2)th visits to the recruitment website. and Let represent the access feature values ​​for the (h-1)th and (h-2)th visits to the recruitment website, respectively. This indicates the pre-set duration. and This indicates the minimum and maximum values ​​of the pre-set access interval; The value of h is greater than 2. and The value is ; For the i-th visit to the recruitment website, i∈{h-1,h-2}, the formula for calculating the visit feature value is: This represents the access feature value of the i-th visit to the recruitment website. This represents the server response time during the i-th access to the recruitment website. This indicates the preset response time comparison value. This represents the number of HTTP status codes of the preset type among all types of HTTP status codes returned by the server during the i-th access to the recruitment website. This represents the total number of all types of HTTP status codes returned by the server during the i-th access to the recruitment website. This represents the number of newly added job postings obtained during the i-th visit to the job website. d1 represents the preset quantity comparison value, d2 represents the response time calculation factor, d3 represents the status code calculation factor, and d3 represents the new quantity calculation factor.

2. The intelligent resume generation system based on natural language processing according to claim 1, characterized in that, The job information acquisition module also includes a comparison unit and a storage unit; The comparison unit is used to compare all the recruitment information obtained in this crawl with all the recruitment information obtained in the previous crawl to determine whether there is any newly added recruitment information. If so, the newly added recruitment information is stored and sent to the storage unit for storage.

3. The intelligent resume generation system based on natural language processing according to claim 1, characterized in that, The crawling record includes the crawling time, server response time, received HTTP status codes, and the number of newly added job postings obtained during the crawl.

4. The intelligent resume generation system based on natural language processing according to claim 1, characterized in that, The recruitment information includes basic information, job description, educational requirements, work experience requirements, skill requirements, and certification requirements; The job posting website contains multiple pages, and each page contains multiple job postings.

5. The intelligent resume generation system based on natural language processing according to claim 4, characterized in that, Basic information includes the job title, work location, and company name. The job description includes the job duties; Work experience requirements include project experience requirements and years of experience in the position; Skill requirements are the skills that the personnel recruited for this position must possess; Skill requirements refer to the certificates that the personnel recruited for this position must possess.

6. The intelligent resume generation system based on natural language processing according to claim 1, characterized in that, The resume generation module includes a keyword acquisition unit and a keyword extraction unit; The keyword acquisition unit is used to obtain a set of keywords from recruitment information; The extraction unit is used to obtain filtering information from job seekers' work experience information based on a set of keywords.

7. The intelligent resume generation system based on natural language processing according to claim 6, characterized in that, Extract keyword sets from job postings, including: The recruitment information text is segmented into words to obtain a word set; Calculate the ranking coefficient for each word in the word set; Obtain a set of keywords based on ranking coefficients.

8. The intelligent resume generation system based on natural language processing according to claim 7, characterized in that, The recruitment information text is segmented to obtain a word set, including: Remove useless characters from the recruitment information text to obtain the text to be segmented. The word segmentation algorithm is used to segment the text to be segmented, resulting in multiple words. The obtained words are then stored in a word set. Remove stop words from the word set.

9. The intelligent resume generation system based on natural language processing according to claim 7, characterized in that, The keyword set is obtained based on the ranking coefficient, including: The top-ranked coefficient A number of words are stored as keywords in a keyword set, where N represents the total number of words in the keyword set and S represents the preset proportion.

Citation Information

Patent Citations

  • Resume generating method and resume generating system

    CN104050532A

  • Resume generation method based on AI

    CN117688920A