A resume search matching method and system for a resume database

By analyzing the relevance of keywords in resumes and the importance of column dimensions, and quantifying the depth of skill mastery indicators, this solves the problem of low resume matching accuracy in existing technologies and achieves more precise person-job matching.

CN121070974BActive Publication Date: 2026-04-28ZHANGJIAGANG HUMAN RESOURCES DEVELOPMENT CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHANGJIAGANG HUMAN RESOURCES DEVELOPMENT CO LTD
Filing Date
2025-08-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing semantic matching algorithms suffer from low matching accuracy in resume retrieval due to the varying depth and wording of skills expressed in resumes, making it impossible to accurately distinguish the level of candidates' skill mastery.

Method used

By acquiring job requirements from companies and skill keywords from resumes, we analyze the correlation and distribution of keywords, classify keywords using clustering algorithms, and quantify the degree to which each category reflects skills by combining the importance and correlation of each category dimension. We then construct a mastery depth index and match it with job requirements from companies for scoring.

Benefits of technology

It improves the accuracy of resume retrieval and the quality of job matching, and can better distinguish skill levels, thus achieving more accurate job matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121070974B_ABST
    Figure CN121070974B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, in particular to a resume retrieval matching method and system for a resume database, comprising: obtaining skill keywords of enterprise post requirements and keywords of each resume in the resume database; classifying all keywords according to the distribution of different keywords in the same resume in the resume database to obtain different keyword categories; obtaining a reflection degree according to the importance degree of each column dimension keyword in the corresponding resume, the distribution of the keyword and each skill keyword in the same resume and the correlation degree; obtaining a mastery depth index according to the correlation degree between each column dimension keyword and each skill keyword in each resume, combining the data distribution of the keyword category to which the skill keyword belongs and the reflection degree; and obtaining a resume matching result according to the mastery depth index and the matching condition. The present application can significantly improve the accuracy of resume retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to a resume retrieval and matching method and system for a resume database. Background Technology

[0002] With rapid economic and social development and an increasingly active job market, recruitment and job-seeking activities have become highly frequent and complex. To gain a competitive edge in the market, companies are increasingly eager to recruit suitable talent, making resume databases a crucial platform for screening and matching candidates. In recent years, with the continuous advancement of natural language processing and machine learning technologies, semantic matching algorithms have begun to be applied in resume retrieval. These algorithms can deeply understand the semantic information of text, improving the accuracy and intelligence of resume matching to a certain extent.

[0003] In existing technologies, semantic matching algorithms can semantically model resume text and search criteria, convert them into vector form, and evaluate the matching degree by calculating vector similarity. Finally, resumes can be recommended to different job requirements based on the matching degree from high to low. However, resumes often contain inconsistent descriptions of skills, which can lead to semantic misunderstandings by the algorithm. This can result in candidates lacking practical experience or with limited skills being misjudged as highly matched, affecting the accuracy and comprehensiveness of the matching process. Summary of the Invention

[0004] To address the technical problem that existing semantic matching algorithms suffer from low accuracy in resume retrieval and matching due to the influence of resume skill descriptions, the present invention aims to provide a resume retrieval and matching method and system for resume databases. The specific technical solution adopted is as follows:

[0005] In a first aspect, the present invention provides a resume retrieval and matching method for a resume database, comprising:

[0006] Obtain the skill keywords required for company positions and the keywords for each resume in the resume database; the keywords include skill keywords, and the resumes include different sections and dimensions;

[0007] Based on the distribution of different keywords in the same resume in the resume database, we analyze the degree of correlation between different keywords and classify all keywords to obtain different keyword categories;

[0008] Based on the importance of keywords in each section of each resume, the distribution of keywords and skill keywords in the same resume, and the degree of correlation, we can obtain the degree to which each section of each resume reflects each skill keyword.

[0009] Based on the degree of correlation between the keywords in each section of each resume and each skill keyword, combined with the data distribution and the degree of reflection of the keyword category to which the skill keyword belongs, the mastery depth index of each skill keyword in each resume is obtained;

[0010] Based on the aforementioned depth of understanding indicators, and combined with the matching results between the skill keywords required for the company's positions and the column dimensions of each resume, the resume matching results are obtained.

[0011] Preferably, the step of analyzing the correlation between different keywords based on their distribution in the same resume within the resume database, and classifying all keywords to obtain different keyword categories, specifically includes:

[0012] In the resume database, the degree of association between each pair of different keywords is obtained by analyzing the distribution of the frequency of occurrence of each pair of different keywords in the same resume and the distance distribution within the same resume.

[0013] Based on the degree of association, a clustering algorithm is used to classify all keywords in the resume database to obtain different keyword categories.

[0014] Preferably, the step of determining the correlation between two different keywords based on the distribution of the frequency of occurrence and the distance distribution within the same resume specifically includes:

[0015] The degree of association between each pair of different keywords is calculated by multiplying the percentage of times each pair of different keywords appears in the same resume by the mean negative correlation coefficient of the text distance between the two different keywords in the same resume.

[0016] Preferably, the step of determining the degree to which each section of each resume reflects each skill keyword based on the importance of the keywords in each section of each resume within the corresponding resume, the distribution of the keywords and each skill keyword in the same resume, and the degree of correlation, specifically includes:

[0017] In the resume database, the product of the percentage of times each keyword and each skill keyword appears in the same resume and the degree of association between each keyword and each other is used as the skill representation value of each keyword to each skill keyword.

[0018] Based on the importance of keywords in each section of each resume and the skill representation value, the basic performance indicators of each skill keyword in each section of each resume are obtained.

[0019] The degree of reflection of each skill keyword in each section of each resume is obtained by comparing the deviation between the basic performance indicators of each skill keyword in each section of each resume and the balance of the basic performance indicators of all sections in the same resume.

[0020] Preferably, the step of obtaining the basic performance index of each skill keyword in each section of each resume based on the importance of the keywords in each section of each resume and the skill representation value specifically includes:

[0021] For any column dimension in any resume, it is denoted as the selected dimension in the target resume; and for any skill keyword, it is denoted as the selected skill keyword.

[0022] The importance of each keyword contained in the selected dimensions of the target resume is used as the importance weight of each keyword contained in the selected dimensions of the target resume;

[0023] Using the aforementioned important weights, the skill representation values ​​of each keyword in the selected dimension of the target resume for the selected skill keyword are weighted and averaged to obtain the basic performance index of the selected dimension of the target resume for the selected skill keyword.

[0024] Preferably, the step of determining the degree to which each section of each resume reflects each skill keyword based on the deviation between the basic performance indicators of each section in each resume and the balance of the basic performance indicators of all sections in the same resume, specifically includes:

[0025] The average value of the basic performance indicators of all sections and dimensions in the target resume for the selected skill keywords is used as the balance indicator of the target resume for the selected skill keywords.

[0026] Based on the negative correlation coefficient of the relative difference between the basic performance indicators of the selected dimensions in the target resume and the equilibrium indicators, and the basic performance indicators of the selected dimensions in the target resume for the selected skill keywords, the degree of reflection of the selected dimensions in the target resume for the selected skill keywords is determined.

[0027] Preferably, the step of obtaining the mastery depth index of each skill keyword in each resume based on the correlation between the keywords of each section dimension and each skill keyword, combined with the data distribution of the keyword category to which the skill keyword belongs and the degree of reflection, specifically includes:

[0028] Based on the degree of correlation between the keywords in each section of each resume and each skill keyword, and combined with the data distribution of the keyword categories to which the skill keywords belong, the skill relevance between each section of each resume and each skill keyword is obtained;

[0029] Based on the skill relevance and the degree of responsiveness, a mastery depth index for each skill keyword in each resume is obtained. Both the skill relevance and the degree of responsiveness are positively correlated with the mastery depth index.

[0030] Preferably, the step of obtaining the skill relevance between each section dimension of each resume and each skill keyword based on the degree of correlation between the keywords and each skill keyword in each resume, combined with the data distribution of the keyword category to which the skill keyword belongs, specifically includes:

[0031] The maximum correlation between each keyword in the selected dimension of the target resume and the selected skill keyword is used as the first feature coefficient;

[0032] The second feature coefficient is obtained by normalizing the ratio between the number of overlaps between the keywords in the selected dimension of the target resume and the keywords in the keyword category of the selected skill keywords, and the number of keywords in the selected dimension of the target resume.

[0033] The product of the first and second feature coefficients is used as the skill relevance between the selected dimension and the selected skill keywords in the target resume.

[0034] Preferably, the step of obtaining resume matching results based on the mastery depth index, combined with the matching results between the skill keywords required for the enterprise's job and the column dimensions of each resume, specifically includes:

[0035] For any resume, a matching algorithm is used to determine the matching score between each skill keyword required by the company and the resume based on the similarity between each skill keyword required by the company and each section dimension of the resume.

[0036] By using the depth of mastery of each skill keyword required by the enterprise in the resume as an indicator, the matching score is weighted and summed to obtain a comprehensive evaluation index between the enterprise's job requirements and the resume.

[0037] Resumes will be recommended and displayed based on the company's job requirements, ranked from highest to lowest according to the comprehensive evaluation indicators.

[0038] In a second aspect, the present invention provides a resume retrieval and matching system for a resume database, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the computer program, when executed by the processor, implements the steps of a resume retrieval and matching method for a resume database.

[0039] The embodiments of the present invention have at least the following beneficial effects:

[0040] This invention first extracts skill keywords from job requirements and keywords from each section of a resume, focusing on the match between job requirements and resumes. Then, it clusters and categorizes keywords expressed differently in resumes, obtaining information on different categories of resume keywords to construct a more comprehensive skill performance cues. Secondly, it quantifies the representativeness of keywords in each section by utilizing the degree to which section-level information reflects skill keywords. This helps to identify how accurately each section of the resume reflects the characteristics of a particular skill, measuring the strength of each section's reflection of a skill and avoiding over-reliance on a single section to judge skill mastery. Furthermore, by integrating representative information from different sections of the resume and considering the correlation between data, it comprehensively assesses the depth of a candidate's mastery of the required skills. This not only differentiates skill mastery levels but also makes resume retrieval and job matching results more discriminative, accurate, and reliable, significantly improving the quality of person-job matching. Finally, by jointly retrieving skill mastery and basic semantic matching, it achieves more accurate and intelligent person-job matching, significantly improving the accuracy of resume retrieval and recruitment efficiency. Attached Figure Description

[0041] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart of the steps of a resume retrieval and matching method for a resume database provided by the present invention;

[0043] Figure 2 This is a flowchart of the steps of the method for obtaining the degree of reflection provided by the present invention;

[0044] Figure 3 This is a flowchart of the steps for obtaining depth indicators provided by the present invention;

[0045] Figure 4 This is a flowchart of the steps in the method for obtaining skill relevance provided by the present invention. Detailed Implementation

[0046] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a resume retrieval and matching method and system for a resume database proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0048] The following description, in conjunction with the accompanying drawings, details a specific scheme for a resume retrieval and matching method and system for a resume database provided by the present invention.

[0049] Please see Figure 1 The diagram illustrates a flowchart of a resume retrieval and matching method for a resume database according to an embodiment of the present invention. The method includes the following steps:

[0050] Step S100: Obtain the skill keywords required for the company's job positions and the keywords for each resume in the resume database; the keywords include skill keywords, and the resumes include different column dimensions.

[0051] First, we obtain the text data of each resume in the resume database and extract keywords from the resume text data to provide a data foundation for subsequent feature analysis processes such as resume matching. The keyword extraction technology is a well-known technology and will not be described in detail here.

[0052] Each resume contains different sections and dimensions, with keywords belonging to each section. For example, section dimensions include name, contact information, education background, work experience, project experience, skills list, certifications, and self-evaluation. Each section dimension represents a specific attribute within the resume. Essentially, a section dimension categorizes similar information modules within the resume.

[0053] As a concrete example, the "Project Experience" section in a resume might include project name, technologies used, job description, project goals, etc. The "Work Experience" section might include job title, daily work content, management / leadership / collaboration type, etc. The "Education Background" section might include major, course names, skill names, etc.

[0054] Next, the text data of the company's job requirements is obtained, and keywords are extracted from this text data to identify and extract skill-related information, resulting in skill keywords. For example, job requirement text generally includes responsibilities, requirements, and expected skills. This embodiment focuses on extracting skill keywords from job requirements to match different descriptions of the same skill in different resumes, improving the accuracy of matching between job requirements and resumes.

[0055] The key skills in job requirements are typically determined based on the job type, responsibilities, and industry characteristics, covering multiple dimensions such as professional skills, tool skills, and soft skills, to clearly define the core competencies a candidate must possess. For example, for a "Java Development Engineer" position, key skills might include the Java programming language, object-oriented programming (OOP), multithreading, JVM principles, distributed systems, microservice architecture, and the Spring Boot / Cloud framework. Similarly, for a "New Media Operations" position, key skills might include content planning, copywriting, event planning, user growth, user profiling, and conversion rate optimization.

[0056] It should be understood that the keywords contained in each section of each resume in the resume database are general terms. These keywords include skill keywords and keywords related to other attributes. Keywords related to other attributes can be generalized terms, conjunctions, or keywords with other category attributes, such as project names. The main purpose of this embodiment is to improve the accuracy of matching resumes with job requirements under different skill keyword descriptions; therefore, the focus is on extracting skill keywords from each text data item.

[0057] When companies search their resume databases for resumes that match job requirements, they use semantic matching algorithms to quickly filter suitable resumes. However, there is no unified standard for the depth and breadth of semantics; different resumes may describe the same skills differently, making it difficult for semantic matching algorithms to accurately grasp the semantic depth and determine which skills are more advanced or more basic, thus failing to achieve precise matching. Therefore, information from other dimensions of the resume, such as work experience and management projects, can demonstrate the application of skills and reflect the depth of mastery of the target skills in the applicant's resume. The specific quantification process is implemented in steps S200 to S400.

[0058] Step S200: Analyze the degree of correlation between different keywords based on the distribution of different keywords in the same resume in the resume database, and classify all keywords to obtain different keyword categories.

[0059] Because candidates may use different expressions when describing the same skills or experience, direct keyword matching can lead to semantic gaps or misjudgments. By statistically analyzing keyword co-occurrence relationships and contextual semantic distance metrics in a large-scale resume corpus, semantic associations between keywords can be constructed, identifying different words representing similar or related skills. Based on this, the relevance between all different keywords contained in all resumes in the resume database is first evaluated, and then similar or related keywords are grouped into the same category.

[0060] The first step is to determine the degree of association between two different keywords in the resume database by analyzing the distribution of the frequency of occurrence of each two different keywords in the same resume and their distance distribution within the same resume.

[0061] Specifically, the product of the percentage of times each two different keywords appear in the same resume and the mean negative correlation coefficient of the text distance between the two different keywords in the same resume is used as the degree of association between each two different keywords.

[0062] As a concrete example, let's take any two different keywords from all the keywords in all resumes in the resume database as an example. That is, the degree of relevance between keyword a and keyword b can be expressed by the formula:

[0063]

[0064] Where X(a,b) represents the degree of association between keyword a and keyword b, and N a,b N represents the number of times keyword a and keyword b appear in the same resume. ′ (a,b) represents the sum of the number of resumes containing keyword 'a' and the number of resumes containing keyword 'b'. This represents the mean text distance between keyword a and keyword b across all resumes.

[0065] If keyword A and keyword B appear together in the same resume, for example, Java (keyword A) and the Spring framework (keyword B) appearing in the same resume, it is counted as one occurrence. This indicates the percentage of times keyword a and keyword b appear in the same resume. The larger the value, the stronger the correlation between the two.

[0066] The mean negative correlation coefficient represents the text distance. The larger the value, the closer the positions of the keywords appearing in a resume are, and the greater the degree of semantic similarity. The process of obtaining it is to first screen out each resume that contains both keyword a and keyword b, calculate the text distance between keyword a and keyword b in each resume, and use methods such as Euclidean distance or Manhattan distance to calculate it. Then, calculate the mean of all text distances and take the inverse.

[0067] The second step is to use a clustering algorithm to classify all keywords in the resume database based on the degree of correlation, thereby obtaining different keyword categories.

[0068] To address the issues of diverse and semantically fragmented information presentation in resumes, clustering keywords based on their relationships can group semantically similar words into the same category, thus uniformly representing identical or similar information.

[0069] As a concrete example, this embodiment employs the DBSCAN clustering algorithm, using 1-Norm[X(a,b)] as the clustering distance metric between keyword a and keyword b to classify all keywords, automatically grouping semantically similar keywords into the same category. It should be noted that Norm is a normalization function, utilizing the negative correlation between any two distinct keywords as the difference distance between the two clustered objects, and employing existing clustering algorithms for classification. This is a well-known technique and will not be elaborated upon further here.

[0070] Step S300: Based on the importance of the keywords in each section of each resume in the corresponding resume, the distribution of the keywords and each skill keyword in the same resume, and the degree of correlation, obtain the degree of reflection of each section of each resume on each skill keyword.

[0071] When companies search and match resumes in their databases based on job requirements, considering the varying degrees to which different dimensions of information in resumes reflect skills, the system first filters resumes containing the same target skills from a large-scale resume sample. It then analyzes the representativeness of keywords and skills across each dimension, thereby constructing a mapping relationship between dimensional information and skills. Based on this, by examining the descriptions of keywords for each skill in a single resume, the system quantifies the degree to which textual information in a single section reflects a single skill.

[0072] In this embodiment, as Figure 2 As shown, the method for obtaining the degree of reflection can be implemented by steps S301 to S303.

[0073] Step S301: Based on the proportion of times each keyword and each skill keyword appears in the same resume in the resume database, and combined with the degree of association, obtain the skill representation value of each keyword to each skill keyword.

[0074] Specifically, in the resume database, the product of the percentage of occurrence of each keyword and each skill keyword in the same resume and the degree of relevance between each keyword is used as the skill representation value of each keyword to each skill keyword. It should be noted that this calculation is analogous to the calculation method for the percentage of occurrence of two different keywords in the same resume.

[0075] It should be understood that in the calculation of skill representation values, keywords are the different keywords contained in all resumes in the resume database, and skill keywords are the skill keywords present in the text data of all job requirements. The main purpose of measuring the skill representation value corresponding to each keyword in a resume is to measure whether the occurrence of keywords is concentrated on the semantic information of skills, to prevent "general terms" from being misjudged as representative terms, to determine whether a keyword exclusively represents a certain skill rather than being co-occurring with multiple skills, and to assess whether a keyword is a representative indicator of a certain skill.

[0076] More specifically, taking keyword a and skill keyword O as an example, the greater the proportion of times keyword a and skill keyword O appear in a resume, and the greater the correlation between keyword a and skill keyword O, the more similar or close the semantics of keyword a and skill keyword O are, the more representative the description of skill keyword O is by keyword a, and the greater the skill representation value of keyword a for skill keyword O.

[0077] By analyzing the frequency and co-occurrence relationships of keywords in a large-scale resume sample database, the skill representativeness of keywords was obtained, ensuring the objectivity and universality of keyword representativeness assessment and providing a reliable foundation for the subsequent construction of the mapping relationship between dimensional information and skills.

[0078] Step S302: Based on the importance of the keywords in each section of each resume in the corresponding resume and the skill representation value, obtain the basic representation index of each skill keyword in each section of each resume.

[0079] The representativeness of keywords in a specific category's information reflects the weight and importance of that category in expressing the skill. If certain keywords are frequently and significantly associated with skill keywords within a category, it indicates that that category is an important basis for assessing skill mastery. Based on this consideration, by analyzing the importance of keywords within a category and their description of the skill, a preliminary assessment of the representativeness of keywords for the skill is conducted.

[0080] Specifically, any section or dimension in any resume is designated as a selected dimension in the target resume, and any skill keyword is designated as a selected skill keyword. The importance of each keyword contained in the selected dimension of the target resume is used as the importance weight of each keyword contained in the selected dimension of the target resume. Using the importance weight, the skill representation value of each keyword in the selected dimension of the target resume in relation to the selected skill keyword is weighted and averaged to obtain the basic performance index of the selected dimension of the target resume in relation to the selected skill keyword.

[0081] More specifically, taking the nth resume as the target resume, the xth section dimension of the target resume as the selected dimension, and the mth skill keyword as the selected skill keyword, the basic performance index of the selected dimension of the target resume in relation to the selected skill keyword can be expressed by the formula: Among them, T n,x (m) represents the basic performance index of the selected dimension for the selected skill keyword in the target resume, where n represents the nth resume, x represents the xth section dimension, and m represents the mth skill keyword; N n,x ZY represents the total number of keywords contained in the x-th section dimension of the nth resume. n,x,i E represents the importance of the i-th keyword contained in the x-th section dimension of the n-th resume, i.e., the importance weight; n,x,i (m) represents the skill representation value of the i-th keyword in the x-th column dimension of the n-th resume for the m-th skill keyword.

[0082] The basic performance index characterizes the overall ability of the information contained in the selected dimensions of the target resume to represent the m-th skill keyword, reflecting whether this section dimension can effectively reflect the skill level. In other words, it comprehensively reflects the overall ability of important and representative keywords in the selected dimension to represent the skill, indicating the degree to which the information in the section dimension represents the skill in a basic way. The larger the value of the basic performance index, the more important the keyword is in the information expression of the resume's section dimension, and the more representative it is in the description of the skill, thus indicating a greater degree of basic representation of the skill keyword by the keyword.

[0083] It should be noted that in this embodiment, the TF-IDF value of each keyword included in each section of the resume is used as the importance of each keyword in each section of the resume. The calculation method of the TF-IDF value is a well-known technique and will not be described in detail here. A high TF-IDF value indicates that the word has a strong representation ability for the document topic. Therefore, the larger the TF-IDF value, the higher the importance of the corresponding keyword, and the greater the importance weight.

[0084] Step S303: Based on the degree of deviation between the basic performance indicators of each skill keyword in each column dimension of each resume and the balance of the basic performance indicators of all columns in the same resume, obtain the degree of reflection of each skill keyword in each column of each resume.

[0085] By comparing the correlation strength between keywords and skill keywords in different column dimensions, we can clarify the relative contribution and mapping relationship of each column dimension in reflecting the skill. Different column dimensions (such as project experience, job responsibilities, certifications, etc.) reflect skills in different ways and with varying degrees of importance. For example, project experience often reflects the practical application ability of a skill, while certifications reflect the level of skill certification. By comparing the balanced performance of a column dimension's information with all column dimensions in reflecting the basic level of skill keywords, we can quantify the mapping relationship between each column dimension's resume information and skill keywords, and evaluate the degree to which each column in each resume reflects each skill keyword.

[0086] Specifically, the average of the basic performance indicators of all columns and dimensions in the target resume for the selected skill keywords is used as the equilibrium indicator of the target resume for the selected skill keywords. Based on the negative correlation coefficient of the relative difference between the basic performance indicators of the selected dimensions in the target resume for the selected skill keywords and the equilibrium indicator, and the basic performance indicators of the selected dimensions in the target resume for the selected skill keywords, the degree of reflection of the selected dimensions in the target resume for the selected skill keywords is determined.

[0087] As a concrete example, the degree to which selected dimensions in a target resume reflect selected skill keywords can be expressed as follows:

[0088]

[0089] Among them, F n,x (m) represents the degree to which the selected dimension in the target resume reflects the selected skill keyword, where n represents the nth resume, x represents the xth section dimension, and m represents the mth skill keyword; T n,x (m) represents the basic performance index of the m-th skill keyword in the x-th section dimension of the n-th resume. represents the average of the basic performance indicators of all sections and dimensions in the nth resume for the mth skill keyword, which is also the equilibrium indicator; exp represents the exponential function with the natural constant e as the base.

[0090] This indicates the degree of difference between the basic performance indicators of skill keywords in a single column dimension and the equilibrium situation, reflecting the degree of deviation from the equilibrium situation. The larger the value, the greater the relative percentage of difference. The larger the value, the greater the difference between the column dimension and the semantic expression of the skill keywords, and the worse the semantic expression of the skill keywords by the column dimension, and the smaller the corresponding reflection value.

[0091] The degree of reflection is based on the basic degree of reflection of the selected skill keywords in the selected dimensions of the target resume. It also incorporates the deviation between the information contained in the selected dimensions of the target resume and the overall balance of the resume. This can more accurately represent the mapping relationship between a single column dimension and a single skill keyword in a single resume.

[0092] Step S400: Based on the degree of correlation between the keywords in each section dimension of each resume and each skill keyword, combined with the data distribution of the keyword category to which the skill keyword belongs and the degree of reflection, obtain the mastery depth index of each skill keyword in each resume.

[0093] Skills often appear in various ways on different resumes. It is difficult to judge the true extent of a job seeker's mastery based solely on skill keywords. Therefore, it is necessary to combine information from multiple dimensions such as project experience, job responsibilities, certificates, and outputs to comprehensively evaluate the depth of a job seeker's mastery of a particular skill. This not only improves the accuracy of skill identification but also better reflects the breadth and complexity of the candidate's skill application in actual work, which helps to achieve a more refined judgment on job suitability.

[0094] Based on this, firstly, the correlation between keywords and skill keywords appearing in a specific section of a resume is assessed, and the degree of correlation between them is evaluated. Secondly, the mapping relationship between a single section of a resume and a single skill keyword is analyzed. The combined results of these two analyses comprehensively reflect the depth of mastery of each skill keyword in each resume.

[0095] Specifically, such as Figure 3 As shown, the method for obtaining the mastery depth index of each skill keyword in each resume can be achieved by steps S401 and S402.

[0096] Step S401: Based on the degree of correlation between the keywords of each section dimension in each resume and each skill keyword, and combined with the data distribution of the keyword category to which the skill keyword belongs, obtain the skill relevance between each section dimension in each resume and each skill keyword.

[0097] If a resume contains a large number of keywords belonging to the same keyword category as the skill keywords in a certain section of information, it indicates that the semantics of that section of information and the skill keywords are similar. Furthermore, the greater the degree of correlation between the two, the more relevant that section of information is to the skills required for the job.

[0098] As a concrete example, such as Figure 4 As shown, the method for obtaining skill relevance can be implemented through steps S4011 to S4013.

[0099] Step S4011: The maximum value of the correlation between each keyword in the selected dimension of the target resume and the selected skill keyword is taken as the first feature coefficient.

[0100] Specifically, this explanation will be based on any section or dimension of any resume and any skill keyword from all job requirements. In other words, it will be based on the selected dimension and selected skill keyword in the target resume.

[0101] The maximum value of the correlation between each keyword in the selected dimension of the target resume and the selected skill keyword reflects the maximum relevance of the selected dimension to the selected skill keyword. The larger the value, the more similar or close the semantics of the selected dimension to the selected skill keyword in the target resume.

[0102] Thus, the first feature coefficient represents the semantic similarity between the selected dimensions and the selected skill keywords in the target resume.

[0103] Step S4012: Normalize the ratio between the number of overlaps between the keywords in the selected dimension of the target resume and the keywords in the keyword category of the selected skill keywords, and the number of keywords in the selected dimension of the target resume, to obtain the second feature coefficient.

[0104] It should be understood that the more overlap there is between the keywords in the selected dimension of the target resume and the keywords in the keyword category of the selected skill keywords, the more likely the information in the selected dimension of the target resume is a similar expression of the selected skill keywords.

[0105] The ratio of overlap to the number of keywords contained in the selected dimension of the target resume is used as the standard for the proportion of overlap. Therefore, the second feature coefficient reflects the relevance between the information contained in the selected dimension of the target resume and the selected skill keywords.

[0106] It should be noted that normalization methods are well-known techniques, and implementers can choose the appropriate method based on the specific implementation scenario, such as the minimax normalization method, etc.

[0107] Step S4013: The product of the first feature coefficient and the second feature coefficient is used as the skill relevance between the selected dimension and the selected skill keyword in the target resume.

[0108] By combining the results of the two correlation analyses, the basic correlation between the selected dimensions and selected skill keywords in the target resume, and the proportion of relevant keywords within the selected dimensions, the skill relevance is used to more comprehensively characterize the degree of correlation between the selected dimensions and selected skill keywords in the target resume.

[0109] Step S402: Based on the skill relevance and the degree of response, obtain the mastery depth index of each skill keyword in each resume. Both the skill relevance and the degree of response are positively correlated with the mastery depth index.

[0110] By combining the skill relevance of information in each section of the resume and the degree to which the information in each section reflects the skill keywords, we can obtain the depth of mastery of the skill keywords as shown in the resume.

[0111] Specifically, using the skill relevance as a weight, the weighted average of the degree to which the target resume reflects the selected skill keywords in each section is calculated to obtain the depth of mastery index of the selected skill keywords in the target resume. More specifically, taking the target resume and selected skill keywords as an example, the depth of mastery index can be expressed by the formula:

[0112]

[0113] Where Q(n,m) represents the mastery depth index of the m-th skill keyword in the n-th resume, which is also the mastery depth index of the selected skill keyword in the target resume. F n,x (m) represents the degree to which the x-th section dimension in the n-th resume reflects the m-th skill keyword, G n,x (m) represents the skill relevance between the x-th section dimension and the m-th skill keyword in the n-th resume, M n This represents the total number of column dimensions included in the nth resume.

[0114] In a single resume, the greater the degree to which each section dimension reflects the skill keywords, the more similar the meaning of the skill keywords in that section dimension is to the skill keywords. At the same time, the greater the skill relevance between each section dimension and the skill keywords, the greater the correlation between that section dimension and the skill keywords. By comprehensively and accurately representing the mastery depth of skill keywords in a single resume, the evaluation results of all section dimensions can be considered comprehensive. The higher the value of the mastery depth index, the more deeply the skill corresponding to the technical keyword is mastered in the target resume, rather than being mentioned superficially, or the stronger and closer the correlation between the two. This avoids problems such as discrepancies in skill descriptions and scattered dimensional information in resumes, and provides a data foundation for subsequent accurate assessment of skill mastery depth and job matching.

[0115] Step S500: Based on the mastery depth index, and combined with the matching between the skill keywords required by the enterprise and the column dimensions of each resume, the resume matching result is obtained.

[0116] Because resumes have a more complex information structure, including multiple sections and dimensions such as basic information, project experience, skills list, and certificates, directly extracting a single semantic feature vector from a single resume would lose dimensional details, hindering accurate matching. Therefore, this embodiment utilizes a semantic model to extract multiple vectors by dimension for semantic feature matching.

[0117] Specifically, the first step, for any given resume, uses a matching algorithm to determine the matching score between each skill keyword required by the company's job and the resume, based on the similarity between these keywords and each section / dimension of the resume. It should be understood that this explanation uses an arbitrary resume and a single company job requirement as an example, ultimately displaying the resume search and matching results for that specific job requirement.

[0118] As a concrete example, the BERT algorithm is used to extract the feature vector of each skill keyword in the job requirements of an enterprise. At the same time, the BERT algorithm is used to extract the feature vector of each section dimension in a resume. Furthermore, the cosine similarity between the feature vector of each skill keyword in the job requirements of an enterprise and the feature vector of each section dimension in the resume is used as the similarity between each skill keyword in the job requirements of an enterprise and each section dimension in the resume, reflecting the magnitude of semantic similarity between skill keywords and section dimensions.

[0119] It should be noted that the BERT model is a well-known technology and will not be discussed further here. In other embodiments, implementers can also use other large semantic models to extract vectors of skill keywords, as well as vectors of each section dimension in the resume, and then use them to evaluate the semantic similarity between the two.

[0120] Thus, each skill keyword corresponds to a similarity score with each section dimension of a resume. For any given skill keyword, the overall similarity across all sections dimension of the resume reflects the matching status between the skill keyword and the resume. Specifically, this embodiment uses the sum of the similarities between any given skill keyword required by the company and all sections dimension of the resume as the matching score between that skill keyword and the resume.

[0121] The second step involves using the depth of mastery of each skill keyword required by the company's job in the resume as an indicator to perform a weighted summation of the matching scores, thereby obtaining a comprehensive evaluation index between the company's job requirements and the resume.

[0122] It should be noted that, for the current job requirements of the company being searched, the mastery depth index of each skill keyword in the same resume is normalized to obtain the mastery weight corresponding to each skill keyword. The mastery weight of each skill keyword is then used to perform a weighted sum of the matching scores between each keyword and the resume to obtain the comprehensive evaluation index between the current job requirements of the company being searched and the resume.

[0123] The comprehensive evaluation index combines the semantic similarity between the company's job requirements and the resume, as well as the depth of the resume's understanding of the skills required for the job. The results of the feature analysis from these two aspects provide a more comprehensive and accurate representation of the matching between the company's job requirements and the resume. The higher the value, the better the match between the company's job requirements and the resume.

[0124] The third step is to recommend and display resumes for the company's job requirements in descending order of comprehensive evaluation indicators.

[0125] For the current job posting in the database, the higher the comprehensive evaluation index value for each resume, the better the resume matches the job requirements. In this case, the resume should be displayed higher in the search results, prioritizing candidates with a high degree of job fit. Conversely, the lower the comprehensive evaluation index value for each resume, the less suitable the resume is for the job requirements. In this case, the resume should be displayed lower in the search results.

[0126] In other embodiments, user behavior feedback learning can also be supported to dynamically adjust matching weights and ranking strategies, continuously optimize recommendation results, and thus help enterprises achieve efficient and intelligent talent screening.

[0127] In summary, during the process of resume retrieval and matching using semantic matching algorithms, the descriptions of work experience or skills in resumes may vary in depth, leading to inaccurate resume matching. Therefore, this invention collects resume data and performs standardized preprocessing, combined with systematic analysis of multi-dimensional information related to skills in the resume (such as certificates, project experience, job responsibilities, etc.), and constructs a mapping relationship between multi-dimensional information and skills in the resume database. This can compensate for the identification bias caused by differences in expression, thereby comprehensively assessing the depth and breadth of a job seeker's mastery of a particular skill, achieving more accurate and intelligent person-job matching, and significantly improving the accuracy of resume retrieval and recruitment efficiency.

[0128] By clustering and categorizing keywords expressed in different ways in resumes, and obtaining multi-dimensional information related to skills, a more comprehensive skill performance clue can be constructed. By establishing a database to quantify the representativeness of keywords in each dimension of information regarding skills, it helps to filter out information in other dimensions of the resume, such as certificates, job titles, or project descriptions, that most accurately reflects a particular skill. Establishing a correlation mapping model between skills and multi-dimensional information allows for a systematic measurement of the strength of each dimension's reflection of a particular skill, avoiding over-generalization or excessive reliance on a single dimension to judge skill mastery. By integrating highly representative information from different dimensions of the resume, a comprehensive assessment of the candidate's depth of mastery of required skills can be made. This not only distinguishes skill mastery levels but also makes resume retrieval and job matching results more differentiated, accurate, and credible, significantly improving the quality of person-job matching.

[0129] This invention also provides a resume retrieval and matching system for a resume database, including a memory, a processor, and a computer program stored in the memory and running on the processor. When executed by the processor, the computer program implements the steps of a resume retrieval and matching method for a resume database. Since an embodiment of a resume retrieval and matching method for a resume database has already been described in detail, further elaboration will not be repeated here.

[0130] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A resume retrieval and matching method for a resume database, characterized in that, The method includes the following steps: Obtain the skill keywords required for company positions and the keywords for each resume in the resume database; the keywords include skill keywords, and the resumes include different sections and dimensions; Based on the distribution of different keywords in the same resume in the resume database, we analyze the degree of correlation between different keywords and classify all keywords to obtain different keyword categories; Based on the importance of keywords in each section of each resume, the distribution of keywords and skill keywords within the same resume, and their correlation, the degree to which each section of each resume reflects each skill keyword is obtained, specifically including: In the resume database, the product of the percentage of times each keyword and each skill keyword appears in the same resume and the degree of association between each keyword and each other is used as the skill representation value of each keyword to each skill keyword. In the calculation of the skill representation value, the keywords are the different keywords contained in all resumes in the resume database, and the skill keywords are the skill keywords that exist in the text data of all job requirements. Based on the importance of keywords in each section of each resume and the skill representation value, the basic performance indicators of each skill keyword in each section of each resume are obtained. The degree of reflection of each skill keyword in each section of each resume is obtained by measuring the deviation between the basic performance indicators of each skill keyword in each section of each resume and the balance of the basic performance indicators of all sections of the same resume. Based on the degree of correlation between the keywords in each section of each resume and each skill keyword, combined with the data distribution and the degree of reflection of the keyword category to which the skill keyword belongs, the mastery depth index of each skill keyword in each resume is obtained; Based on the aforementioned depth of understanding indicators, and combined with the matching results between the skill keywords required for the company's positions and the column dimensions of each resume, the resume matching results are obtained.

2. The resume retrieval and matching method for a resume database according to claim 1, characterized in that, The method involves analyzing the distribution of different keywords within the same resume based on the resume database, determining the correlation between different keywords, and classifying all keywords into different keyword categories. Specifically, this includes: In the resume database, the degree of association between each pair of different keywords is obtained by analyzing the distribution of the frequency of occurrence of each pair of different keywords in the same resume and the distance distribution within the same resume. Based on the degree of association, a clustering algorithm is used to classify all keywords in the resume database to obtain different keyword categories.

3. The resume retrieval and matching method for a resume database according to claim 2, characterized in that, The degree of correlation between each pair of different keywords is determined by analyzing the distribution of their frequency of occurrence and their distance distribution within the same resume. Specifically, this includes: The degree of association between each pair of different keywords is calculated by multiplying the percentage of times each pair of different keywords appears in the same resume by the mean negative correlation coefficient of the text distance between the two different keywords in the same resume.

4. The resume retrieval and matching method for a resume database according to claim 1, characterized in that, The basic performance indicators for each skill keyword in each section of each resume are obtained based on the importance of the keywords in each section of each resume and the skill representation value. Specifically, these indicators include: For any column dimension in any resume, it is denoted as the selected dimension in the target resume; and for any skill keyword, it is denoted as the selected skill keyword. The importance of each keyword contained in the selected dimensions of the target resume is used as the importance weight of each keyword contained in the selected dimensions of the target resume; Using the aforementioned important weights, the skill representation values ​​of each keyword in the selected dimension of the target resume for the selected skill keyword are weighted and averaged to obtain the basic performance index of the selected dimension of the target resume for the selected skill keyword.

5. A resume retrieval and matching method for a resume database according to claim 4, characterized in that, The degree of reflection of each skill keyword in each section of each resume is obtained by calculating the deviation between the basic performance indicators of each skill keyword in each section dimension of each resume and the balance of the basic performance indicators of all sections dimensions in the same resume. Specifically, this includes: The average value of the basic performance indicators of all sections and dimensions in the target resume for the selected skill keywords is used as the balance indicator of the target resume for the selected skill keywords. Based on the negative correlation coefficient of the relative difference between the basic performance indicators of the selected dimensions in the target resume and the equilibrium indicators, and the basic performance indicators of the selected dimensions in the target resume for the selected skill keywords, the degree of reflection of the selected dimensions in the target resume for the selected skill keywords is determined.

6. A resume retrieval and matching method for a resume database according to claim 5, characterized in that, The method involves determining the depth of mastery of each skill keyword in each resume based on the correlation between keywords in each section and dimension of each resume, combined with the data distribution of the keyword category to which the skill keyword belongs and the degree of reflection. This specifically includes: Based on the degree of correlation between the keywords in each section of each resume and each skill keyword, and combined with the data distribution of the keyword categories to which the skill keywords belong, the skill relevance between each section of each resume and each skill keyword is obtained; Based on the skill relevance and the degree of responsiveness, a mastery depth index for each skill keyword in each resume is obtained. Both the skill relevance and the degree of responsiveness are positively correlated with the mastery depth index.

7. A resume retrieval and matching method for a resume database according to claim 6, characterized in that, The process involves determining the skill relevance between each section of each resume and each skill keyword based on the correlation between keywords in each section and each skill keyword, combined with the data distribution of the keyword categories to which the skill keywords belong. Specifically, this includes: The maximum correlation between each keyword in the selected dimension of the target resume and the selected skill keyword is used as the first feature coefficient; The second feature coefficient is obtained by normalizing the ratio between the number of overlaps between the keywords in the selected dimension of the target resume and the keywords in the keyword category of the selected skill keywords, and the number of keywords in the selected dimension of the target resume. The product of the first and second feature coefficients is used as the skill relevance between the selected dimension and the selected skill keywords in the target resume.

8. The resume retrieval and matching method for a resume database according to claim 1, characterized in that, The process involves using the mastery depth index, combined with the matching results between the skill keywords required for the company's job postings and the dimensions of each resume section, to obtain the resume matching results. Specifically, this includes: For any resume, a matching algorithm is used to determine the matching score between each skill keyword required by the company and the resume based on the similarity between each skill keyword required by the company and each section dimension of the resume. By using the depth of mastery of each skill keyword required by the enterprise in the resume as an indicator, the matching score is weighted and summed to obtain a comprehensive evaluation index between the enterprise's job requirements and the resume. Resumes will be recommended and displayed based on the company's job requirements, ranked from highest to lowest according to the comprehensive evaluation indicators.

9. A resume retrieval and matching system for a resume database, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of a resume retrieval and matching method for a resume database as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Method for extracting and storing multidimensional field key knowledge

    CN106446089A

  • Resume searching method, device and system, resume delivering method and device and system and electronic equipment

    CN110909120A