BERT model-based resume screening method
Through the resume screening method based on the BERT model, combined with the N-gram filtering technology of TFIDF-LDA and information entropy, the problem of inaccurate keyword matching in the existing resume screening methods is solved, and more accurate candidate ability assessment and efficient person-job matching are achieved.
Patent Information
- Application Number
- CN202510398882.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The existing resume screening methods are not accurate in keyword matching and cannot understand the context, making it difficult to accurately evaluate the candidate's abilities, unable to adapt to the needs of different positions, and have low accuracy in talent screening.
The resume screening method based on the BERT model is adopted, and the subject word extraction and category feature expansion are performed through the TFIDF-LDA model, combined with the N-gram word vocabulary filtering of information entropy, the comprehensive matching degree is calculated using the deep learning matching degree calculation module, and the resume text classification is used by the FE-BERT model.
It improves the accuracy and efficiency of resume screening, can more accurately evaluate whether candidates have the core abilities required for specific positions, and enhances the decision-making basis for job matching.
Smart Images

Figure CN120336524A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning and text classification. More specifically, it relates to a resume screening method based on the BERT model. Background Art
[0002] In the context of the current era of big data, various information data shows geometric growth. How to use information technology means to improve the current working mode and enhance the operation efficiency of enterprises is crucial. The essence of competition among enterprises is the competition for talents. The quality and quantity of talents determine whether a company can develop sustainably in the long run. How to quickly and accurately discover and recruit talents who meet the company's development needs has become the focus of common concern among many enterprises, experts, and scholars in recent years.
[0003] However, there are the following problems in the existing resume screening process: (1) Many enterprises use ATS (Applicant Tracking System) to automatically screen resumes. These systems often rely on keyword matching. If a candidate's resume does not contain specific keywords, even if they are very suitable, they may be automatically excluded by the system; (2) When screening resumes, it may be difficult for recruiters to judge whether a candidate possesses the core capabilities required for a specific position, especially for those positions with relatively special skills or high experience requirements; (3) Resume screening often focuses on a candidate's past experience and existing skills, while potentially ignoring an individual's potential and growth space.
[0004] After retrieval, Chinese Patent No. 202410613842.8 discloses a human resource resume screening method based on word classification. During the resume screening process of this application, the keyword matching method only focuses on whether a specific word is included in the resume, but cannot understand the actual meaning of the word, which may lead to inaccurate screening results; job seekers may increase their passing rate by stuffing a large number of keywords in their resumes, but this does not mean that they truly possess relevant capabilities, thus affecting the screening quality; at the same time, traditional methods are difficult to identify words with similar semantics but different expressions, and may miss suitable candidates. Summary of the Invention
[0005] The present invention provides a resume screening method based on the BERT model, which can solve the problems existing in the existing resume screening methods, such as inaccurate keyword matching, inability to understand the context, etc., resulting in difficulties in accurately evaluating the capabilities of candidates, inability to adapt to the requirements of different positions, or low accuracy of talent screening.
[0006] To achieve the above object, the technical solution provided by the present invention is as follows:
[0007] The first aspect of the present invention provides a resume screening method based on the BERT model, including:
[0008] Obtain the job requirement text data and the job seeker resume text data, and preprocess the text data respectively to obtain the corresponding preprocessed text vectors;
[0009] Based on the TFIDF-LDA model, extract the topic words and expand the category features of the preprocessed text vectors to obtain the corresponding topic word vectors after category feature expansion;
[0010] Perform word filtering on the topic word vectors after category feature expansion based on the N-gram of information entropy to obtain the filtered topic word vectors;
[0011] Combined with the deep learning matching degree calculation module, calculate the comprehensive matching degree between the job and the job seeker based on the filtered topic word vectors, and perform a preliminary screening of the job seeker resume in combination with the set comprehensive matching degree threshold;
[0012] Take the resume obtained after the preliminary screening, the preprocessed text vectors corresponding to the job requirement text, the topic word vectors after category feature expansion, and the filtered topic word vectors as the three-layer inputs of the FE-BERT model, and perform resume text classification through model processing to obtain the resume screening result.
[0013] According to any resume screening method described in the first aspect of the present invention, the preprocessing of the text data respectively includes:
[0014] Perform sentence segmentation and clause segmentation on the obtained job requirement text data and job seeker resume text data respectively to form short texts;
[0015] Perform preprocessing including data cleaning, Chinese word segmentation, and stop word removal on the obtained short texts respectively;
[0016] Convert the preprocessed short texts into numerical vectors respectively, that is, obtain the corresponding preprocessed text vectors.
[0017] According to any resume screening method described in the first aspect of the present invention, the method for extracting topic words and expanding category features of the preprocessed text vectors based on the TFIDF-LDA model to obtain the corresponding topic word vectors after category feature expansion includes:
[0018] Extract topic words from the preprocessed text vectors based on TF-IDF to obtain candidate topic word vectors;
[0019] Extract text topic words from the preprocessed text vectors based on the LDA model to obtain text topic word vectors;
[0020] Based on the set of the obtained candidate topic word vectors and text topic word vectors, construct a word co-occurrence matrix;
[0021] Count the frequency of each subject term within a certain window, and calculate the strength PMI of the co-occurrence relationship of words;
[0022] For phrases with relatively high PMI scores and reasonable combination patterns, consider them as potential new words and add them to the candidate subject terms and the set of text subject terms.
[0023] According to any of the resume screening methods described in the first aspect of the present invention, the calculation formula for the TF-IDF value is as follows:
[0024]
[0025] Where:
[0026] IDF is a way to measure the general importance of a specified vocabulary;
[0027] n i,j represents the number of times the vocabulary t i appears in the short text d j ;
[0028] |D| represents the total number of short texts;
[0029] |d j | is the length of the short text d j ;
[0030] is the average length of the short text;
[0031] |{ j :t i ∈d j} represents the number of short texts containing the vocabulary t i ;
[0032] λ ∈ [0,1]: used to control the smoothing degree of word frequency;
[0033] α ∈ [0,1]: used to control the normalization of text length;
[0034] μ ≥ 0: is a smoothing factor used to prevent outliers caused by too low occurrence times of some words;
[0035] P(t i |C) represents the generation probability of the vocabulary t i in the corpus C (preprocessed text vector), that is, the generation probability of the vocabulary t i in the corresponding job requirements or job application resume text.
[0036] According to any of the resume screening methods described in the first aspect of the present invention, in the LDA model, the generation probability formula for the entire document is:
[0037]
[0038] In the above formula:
[0039] θ mz is the distribution probability of topic z in document d m ;
[0040] is the probability of generating word w under topic z i ;
[0041] K is the total number of topics;
[0042] n is the total number of words in document d m ;
[0043] β > 0: Controls the sharpness of the topic distribution;
[0044] η ∈ [0, 1): Penalty coefficient, suppressing the strong association between high-frequency words and topics and enhancing the generalization ability of the model.
[0045] According to any resume screening method described in the first aspect of the present invention, the vocabulary filtering of the topic word vectors after category feature expansion based on information entropy of N-gram includes:
[0046] (1) Perform 2-gram and 3-gram word extraction on the topic word vectors after category feature expansion respectively, and form a sub-word list by taking the union of the two sub-words;
[0047] (2) Calculate the information entropy between sub-word categories and the information entropy within sub-word categories for the sub-word list respectively to obtain the sub-word information entropy IE(t); the calculation of the sub-word information entropy IE(t) is as follows:
[0048]
[0049] β w (H w (t)) = H w (t)
[0050] In the above formula:
[0051] β w (H w (t)) is the information entropy within the category of vocabulary t;
[0052] β b (H b (t)) is the information entropy between the categories of vocabulary t;
[0053] ε ∈ [0, 1): Used to control the attenuation rate of historical information;
[0054] v, ω > 0: Used to adjust the weight ratio of local and global entropy. v amplifies the current local entropy, and ω suppresses the inverse effect of global entropy;
[0055] δ ≥ 0: Used to balance the intensity of the interaction term H w (t)H b (t), enhancing the non - linear coupling effect;
[0056] (3) Take the quartiles of information entropy according to sub - word information entropy, and determine the preset threshold for filtering sub - words through multiple trials. Add the sub - words with information entropy greater than the preset threshold directly to the final sub - word list.
[0057] According to any resume screening method described in the first aspect of the present invention, the combined deep - learning matching degree calculation module calculates the comprehensive matching degree between the position and the job seeker based on the filtered subject - word vectors, and preliminarily screens the job seeker's resume in combination with the set comprehensive matching degree threshold, including:
[0058] Calculate the cosine similarity of the corresponding feature vectors based on the filtered job requirements and the subject - word vectors of the resume text
[0059] Calculate the weighted score of the interaction of each pair of features
[0060] Perform comprehensive matching degree calculation, and use the following scalar MatchScore(P, C) to represent the comprehensive matching degree between the resume and the position:
[0061]
[0062] Among them, is the dynamic weight, ranging in [0, 1], obtained through normalization or training and learning; ψ(h pi , h cj ) is the dynamic adjustment factor based on the feature vector; is the job stability weight; ContextAlign(P, C) is the context - based alignment factor, used to capture the global semantic relationship between the position and the job seeker; ξ is the coefficient of context alignment, ranging in [0, +∞), and needs to be adjusted to balance the weights of skill and non - skill factors.
[0063] According to any resume screening method described in the first aspect of the present invention, the calculation of the dynamic adjustment factor ψ(h pi , h cj ) based on the feature vector is as follows:
[0064]
[0065] The calculation formula of ContextAlign(P, C) is:
[0066]
[0067] X is an adjustable scale parameter used to control the decay rate of the reciprocal of the Euclidean distance between feature vectors; deg(p i ) and deg(c j ) represent the degrees of the job feature pi and the job seeker feature cj, respectively.
[0068] According to any of the resume screening methods described in the first aspect of the present invention, the FE-BERT model is obtained by improving on the basis of the BERT model, and it includes:
[0069] (1) Input layer
[0070] The input layer of the FE-BERT model converts the input into a vector representation of a fixed dimension through an embedding layer, which consists of three parts, namely, the word embedding representation of the preprocessing result, the word embedding representation of the topic feature after being augmented based on the category feature, and the word embedding representation after filtering by the take-word vocabulary. By stacking the word embedding representations of the three parts, the vector representation of the input short text sequence is obtained;
[0071] (2) Transformer Encoder layer
[0072] The Transformer Encoder is stacked by multiple self-attention layers and a feed-forward neural network, and it captures the connections between any words in the sentence by calculating Attention; the calculation process after improving the self-attention mechanism is as follows:
[0073]
[0074] In the above formula:
[0075] Q is the query vector; K is the key vector; V is the value vector; d k is the dimension of the key vector; ζ(Q,K) is the dynamic scaling term; is the context-aware term; P(Q,K) is the bias term; ρ is a learnable parameter or hyperparameter, and the initial value range is [0,1]; is the scalar weight parameter, and the range is (0,1].
[0076] (3) Output layer
[0077] The output layer adopts a fully connected Linear layer and a softmax layer.
[0078] For any resume screening method according to the first aspect of the present invention, the output layer takes the output vector at the [CLS] position for the classification task. For the sequence labeling task, the output vectors of each token are taken. Each token is passed to the softmax for calculation, and the result is mapped to the interval (0, 1).
[0079] Furthermore, by comparing the error between the predicted category output by the hierarchical softmax function and the actual category to which the resume text belongs, and using the stochastic gradient descent method as the optimizer, the value that minimizes the loss function in the gradient direction of the error with respect to the weight matrix is obtained, which is the optimal parameter of the FE-BERT resume text classification model.
[0080] The second aspect of the present invention also provides a resume screening system, including:
[0081] A data acquisition and preprocessing module, configured to acquire job requirement text data and job seeker resume text data, and respectively preprocess the text data to obtain corresponding preprocessed text vectors;
[0082] A subject word extraction and category feature expansion module, configured to perform subject word extraction and category feature expansion on the preprocessed text vectors based on the TFIDF-LDA model to obtain corresponding subject word vectors after category feature expansion;
[0083] A word-taking vocabulary filtering module, configured to perform word-taking vocabulary filtering on the subject word vectors after category feature expansion based on the N-gram of information entropy to obtain filtered subject word vectors;
[0084] A resume preliminary screening module, configured to calculate the comprehensive matching degree between the job and the job seeker based on the filtered subject word vectors in combination with a deep learning matching degree calculation module, and perform preliminary screening on the job seeker's resume in combination with a set comprehensive matching degree threshold;
[0085] A resume text classification module, configured to perform resume text classification based on the FE-BERT model to obtain a resume screening result; wherein, the preprocessed text vectors corresponding to the resumes and job requirement texts after preliminary screening, the subject word vectors after category feature expansion, and the filtered subject word vectors are used as the three-layer inputs of the FE-BERT model.
[0086] Adopting the technical solution provided by the present invention, compared with the prior art, the following beneficial effects can be achieved:
[0087] (1) The present invention proposes a resume screening method based on the BERT model. On the basis of the existing BERT model, by introducing a category feature expansion method based on TFIDF-LDA and an N-gram word selection vocabulary filtering method based on information entropy, the screening accuracy and efficiency of resumes can be effectively improved. Specifically, the category features obtained by TFIDF-LDA are more accurate and can better reflect the key features of the position; by using information entropy for N-gram word selection vocabulary filtering, it can not only alleviate the impact of meaningless words and low classification contribution words generated after word selection by N-gram features on the training of the classification model, enhance the feature representation ability of the text, but also reduce the training time of the model to a certain extent.
[0088] (2) The present invention further optimizes on the basis of the existing matching degree calculation formula, can comprehensively consider the local feature similarity and the global context relationship, generate a more accurate comprehensive matching degree score, and thus provide a high-quality decision-making basis for the person-job matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 is a flowchart of a resume screening method based on the BERT model provided by the present invention;
[0090] Figure 2 and Figure 3 is a comparison result graph of the classification effects of resume data of each position using different models on the test set;
[0091] Figure 4 is a comparison result graph of Macro-average of different models. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0092] The present invention provides a resume screening method based on the BERT model, especially involving feature mining, category feature expansion, vocabulary filtering and resume text classification of a self-built data set. This method can effectively solve the technical problem that recruiters may have difficulty quickly judging whether a candidate has the core capabilities required for a specific position, especially for those positions with special skills or high experience requirements. Therefore, it can be widely applied to career planning and analysis in the talent market, as well as for establishing database management and building a talent pool.
[0093] As Figure 1 shown, the resume screening method of the present invention specifically includes:
[0094] I. Obtain the text data of the position requirements and the text data of the job seeker's resume, and respectively preprocess the text data to obtain the corresponding preprocessed text vectors.
[0095] Among them, the recruitment position requirement text data includes keyword fields such as job description, essential skills, work location, and salary range; the personal resume data should cover core contents such as educational background, work experience, professional expertise, and project practice. It should be noted that the position requirement text data can be obtained from the recruitment platform or directly use the historical resume data of the position recruitment. Among them, for new positions or positions with large changes in recruitment requirements, the job description text can be used for experiments, and for positions with stable recruitment requirements, the historical resume of the position recruitment can be directly used for experiments.
[0096] Furthermore, the preprocessing of the text data respectively includes:
[0097] Segment and clause-process the obtained job requirement text data and job seeker resume text data into short texts respectively;
[0098] Perform preprocessing including data cleaning, Chinese word segmentation, and stop word removal on the obtained short texts respectively;
[0099] Since the text data is unstructured data, it is necessary to convert the text data into a numerical vector, that is, obtain the corresponding preprocessed text vector.
[0100] The present invention preferably adopts the distributed text representation method. Distributed representation is the now widely popular word vector. The word vector converts the segmented words into numerical vectors, which can express deep information such as the semantics of words, and its dimension can be customized, so as to avoid the high-dimensional situation.
[0101] Second, based on the TFIDF-LDA model, extract the topic words and expand the category features of the preprocessed text vector to obtain the corresponding topic word vector after category feature expansion, specifically including:
[0102] (1) Extract topic words from the preprocessed text vector based on TF-IDF to obtain the candidate topic word vector;
[0103] TF-IDF is a statistical-based feature processing method, which consists of two parts: TF (Term Frequency) and IDF (Inverse Document Frequency). The specific calculation formula for the TF-IDF value of a certain vocabulary is:
[0104] TF-IDF = TF × IDF
[0105] Among them, TF represents the number of times the specified vocabulary appears in all short texts.
[0106] For a certain short text d j the specified vocabulary t i, the TF value can be expressed as a formula:
[0107]
[0108] The numerator n i,j , represents the number of times the vocabulary t i appears in the short text d j ; the denominator ∑ k n k,j represents the sum of the number of times all vocabularies appear in the short text d j . The larger the TF value of a vocabulary, the more times the vocabulary appears in the short text, and the greater the probability that the vocabulary becomes a candidate topic word.
[0109] IDF is a way to measure the general importance of a specified vocabulary. It is obtained by dividing the total number of short texts in the corpus by the number of short texts containing the vocabulary and taking the logarithm of the quotient. The present invention has improved the calculation formula of the original TF-IDF method, so as to obtain a more accurate TF-IDF value:
[0110]
[0111] Among them, IDF is a way to measure the general importance of a specified vocabulary;
[0112] n i,j represents the number of times the vocabulary t i appears in the short text d j ;
[0113] |D| represents the total number of short texts;
[0114] |d j | is the length of the short text d j ;
[0115] is the average length of the short text;
[0116] |{j:t i ∈d j}| represents the number of short texts containing the vocabulary t i ;
[0117] λ ∈ [0, 1]: used to control the smoothing degree of word frequency to prevent large numerical values from affecting the calculation;
[0118] α ∈ [0, 1]: used to control the normalization of text length to prevent the word frequency in long documents from being overly amplified;
[0119] μ ≥ 0: is a smoothing factor used to prevent outliers caused by too low occurrence times of some words, taking 1 or determined by cross-validation;
[0120] P(t i |C) represents the generation probability of the vocabulary t i in the corpus C.
[0121] The present invention introduces terms in the above calculation formula to perform weighted adjustment on the document length, where is the ratio of the current document length to the average length, and the strength of the length influence is controlled by α, enabling the model to more flexibly adapt to documents of different lengths. At the same time, by adding a smoothing term μP(t i |C) to the denominator in the logarithmic term, the problem of weight explosion of rare words can be effectively alleviated, and the processing effect of low-frequency words can be improved.
[0122] (2) Based on the LDA model, extract text topic words from the preprocessed text vector to obtain a text topic word vector;
[0123] LDA (Latent Dirichlet Allocation) is a commonly used topic model for extracting topics from a document collection. In the LDA model, each document can be regarded as a mixture of multiple topics, and each topic can be regarded as a distribution of keywords. Its core idea is to regard each document as a probability mixture of multiple topics, and each topic is composed of a probability distribution of vocabulary. LDA estimates the parameters by establishing a document-topic distribution and a topic-word distribution and using probability inference methods (such as variational inference or Gibbs sampling), thereby automatically identifying the most representative topic set in the document.
[0124] The LDA view holds that a resume can be regarded as a set of ordered word sequence collections d = (w1, w2,..., w n ), each resume contains multiple topics, and each word is generated with a probability under a topic, that is, the generation of resume text is to first randomly select a topic from the topics of the document, then randomly select a topic word from that topic, and finally form a document consisting of N words. Under this rule, the documents are independent and unordered with each other, and the words under each document are also independent and unordered with each other. For the K words under a topic, denoted as ψ1, ψ2,..., ψ k , for each document d M in the corpus C = (d1, d2,..., d m ) containing M documents, there is a specific document-topic distribution θ m , and all the document-topic distributions are denoted as θ1, θ2,..., θ M .
[0125] Therefore, the generation probability formula for each word in the m-th document d m is:
[0126]
[0127] Therefore, the generation probability formula for the entire document is:
[0128]
[0129]
[0130] In the above formula:
[0131] is the probability of generating word w under topic z i ;
[0132] θ mz is the distribution probability of topic z in document d m ;
[0133] K is the total number of topics;
[0134] N is the total number of words in document d m ;
[0135] β > 0: Controls the sharpness of the topic distribution (similar to the concentration parameter of the Dirichlet prior);
[0136] η ∈ [0, 1): Penalty coefficient, which suppresses the strong association between high-frequency words and topics and improves the generalization ability of the model;
[0137] The Sigmoid function σ(·) maps to (0, 1), enhancing the robustness of the probability;
[0138] Topic weight adjustment Controls the sparsity of the topic distribution through the hyperparameter β (β > 1 strengthens the dominant topic, β < 1 smooths the distribution);
[0139] Penalty term Suppresses the excessive association between high-frequency words and topics through η to avoid overfitting.
[0140] (3) Based on the obtained set of candidate topic word vectors and text topic word vectors, construct a word co-occurrence matrix;
[0141] (4) Count the frequency of each topic word appearing within a certain window and calculate the strength of the word co-occurrence relationship PMI;
[0142] (5) For phrases with higher PMI scores and reasonable combination patterns, consider them as potential new words and add them to the topic word set.
[0143] The input of the LDA topic model is a list of words generated after preprocessing an article or sentence. There are a large number of words in the list that affect the mining of text topic-words. On the one hand, these words will increase the burden of document vector conversion and word distribution solving work, and increase the training time of the model. On the other hand, they are very likely to interfere with the solution of the text topic-word distribution, affecting the quality of the mined topic features. Therefore, the present invention proposes a method for expanding category features based on TFIDF-LDA, which can effectively reduce the training time of the LDA topic model while improving the accuracy of topic-word mining of the LDA topic model.
[0144] Third, N-gram based on information entropy is used to filter the words in the topic-word vector after category feature expansion to obtain the filtered topic-word vector.
[0145] N-gram features are added to the BERT model, which is likely to generate a large number of words that do not actually exist in Chinese texts. These words will lead to an overly large parameter space and make the vector sparsity more serious, resulting in excessive time consumption in the model training process. Therefore, the present invention further proposes a method for filtering words by N-gram based on information entropy, which can effectively alleviate the vector sparsity problem caused by N-gram in Chinese word extraction.
[0146] Specifically, the N-gram based on information entropy filters the words in the topic-word vector after category feature expansion, including:
[0147] (1) Perform 2-gram and 3-gram sliding word extraction on the topic-word vector after category feature expansion respectively, and take the union of the sub-words of the two to form a sub-word list;
[0148] (2) Calculate the information entropy between sub-word categories and the information entropy within sub-word categories for the sub-word list respectively to obtain the sub-word information entropy IE(t);
[0149] Information entropy β between vocabulary categories b (H b (t)) calculation formula design:
[0150] Assume that the short texts in a certain corpus are divided into W categories (c1,…,c w ,…,c W ), where c n contains K short texts (d1, d2,…,d K ), the frequencies of a certain vocabulary t in each category are (m1,…,m w ,…,m W ), and there is m1+…+m w +…+m W = M. Then the probability distribution of the vocabulary t between categories in the corpus can be expressed as
[0151] The information entropy H between the categories of vocabulary t b (t) can be defined by the formula:
[0152]
[0153] According to the definition of information entropy, when the occurrence frequencies of vocabulary t among various categories in the corpus are similar, that is, the more uniform the distribution is, the smaller its information entropy. Then, for the information entropy β between the categories of vocabulary t b (H b (t)) can be further defined by the formula:
[0154]
[0155] That is, the information entropy between the categories of vocabulary t should be inversely proportional to the degree of uniformity of its own distribution among the categories. The more uniform the distribution of t among the categories is, the smaller the contribution degree of t to category discrimination is, and the smaller the information entropy is.
[0156] The information entropy β within the category w (H w (t)) calculation formula design:
[0157] The number of times vocabulary t appears in K texts of category c n are respectively (s1, s2,..., s K ), and s1 + s2 +... + s K = S. From this, it can be known that the probability distribution of vocabulary t in category c n is Then the information entropy representation of vocabulary t within a certain category can be obtained and defined by the formula:
[0158]
[0159] According to the definition of information entropy, when the distribution of vocabulary t within the category is more uneven, its information entropy is larger. Then, for the information entropy β of vocabulary t within the category w (H w (t)) can be defined by the formula:
[0160] β w (H w (t)) = H w (t)
[0161] That is, the information entropy within the category of vocabulary t should be directly proportional to the degree of uniformity of its own distribution within the category. The more uneven the distribution of t within the category is, the greater the contribution degree of t to category discrimination is, and the greater the information is.
[0162] Therefore, the calculation of the sub - word information entropy IE(t) is as follows:
[0163]
[0164] In the above formulas:
[0165] β w (H w (t)) is the information entropy within the category of the term t;
[0166] β b (H b (t)) is the information entropy between the categories of the term t;
[0167] H w (t - 1) is the entropy value at the previous moment, making the calculation of entropy smoother in the time series;
[0168] ε ∈ [0, 1): used to control the decay rate of historical information. The larger the value, the more significant the historical influence;
[0169] v, ω > 0: used to adjust the weight ratio of local and global entropy. ν amplifies the current local entropy, and ω suppresses the inverse effect of the global entropy;
[0170] δ ≥ 0: used to balance the intensity of the interaction term H w (t)H b (t), enhancing the non - linear coupling effect.
[0171] By introducing a time - dependent term (such as εH w (t - 1)) for dynamic weight adjustment, the formula can respond more flexibly to data changes, improving the adaptability in dynamic scenarios; at the same time, enabling the model to have short - term memory, making it more suitable for time - series analysis; in addition, the present invention further combines local H w and global H b information entropy, and balances the contributions of the two through the parameters v, ω, δ, enhancing the comprehensive representation ability of the model and achieving multi - scale fusion.
[0172] (3) Take the information entropy quartiles of sub - word information entropy, determine the preset threshold for filtering sub - words through multiple experiments, and directly add the sub - words with information entropy greater than the preset threshold to the final sub - word list.
[0173] IV. Combine with the deep - learning matching degree calculation module, calculate the comprehensive matching degree between the position and the job seeker based on the filtered topic - word vectors, and conduct a preliminary screening of the job seeker's resume in combination with the set comprehensive matching degree threshold.
[0174] This step specifically includes:
[0175] Calculate the cosine similarity of the corresponding feature vectors based on the filtered job requirements and the resume text topic - word vectors
[0176] Calculate the weighted scores of the interactions of each pair of features
[0177] Perform a comprehensive matching degree calculation, and use the following scalar MatchScore(P, C) to represent the comprehensive matching degree between the resume and the position:
[0178]
[0179] Among them, P represents the set of position features, and C represents the set of job seeker features; h pi and h cj respectively represent the feature vectors of the position feature pi and the job seeker feature cj; is the dynamic weight, with a range of [0, 1], obtained through normalization or training and learning; is the job stability weight; ContextAlign(P, C) is the context-based alignment factor, used to capture the global semantic relationship between the position and the job seeker; ξ is the coefficient of context alignment, with a range of [0, +∞), and needs to be tuned to balance the weights of skill and non-skill factors; ψ(h pi , h cj ) is the dynamic adjustment factor based on the feature vector, and is calculated as follows:
[0180]
[0181] The calculation formula of ContextAlign(P, C) is:
[0182]
[0183]
[0184] represents the Euclidean distance between feature vectors; χ is an adjustable scale parameter, used to control the decay speed of the reciprocal of the Euclidean distance between feature vectors; deg(p i ) and deg(c j ) respectively represent the degrees of the position feature pi and the job seeker feature cj.
[0185] The present invention further improves on the existing matching degree calculation formula. Specifically, using ReLU(cos(h pi , h cj )) ensures that the matching degree is non-negative, avoids the interference of negative similarity, and at the same time retains the simplicity of linear calculation; the newly added ψ(h pi , h cj) It can capture high-order non-linear relationships (such as through neural networks or kernel functions) to make up for the deficiencies of cosine similarity in complex feature matching; ξ·ContextAlign(P,C) can be used to comprehensively evaluate non-skill factors (such as cultural fit, relevance of project experience) to enhance the comprehensiveness of matching. Among them, This item introduces the reciprocal decay of the Euclidean distance between feature vectors, making the contribution of feature pairs with closer distances to the alignment score larger and the contribution of feature pairs with farther distances to the alignment score smaller. This helps to improve the accuracy of alignment while avoiding the use of exponential functions.
[0186] V. Take the preprocessed text vector corresponding to the resume obtained after primary screening, the topic word vector after category feature expansion, and the filtered topic word vector as the three-layer input of the FE-BERT model. After processing by the model for resume text classification, the resume screening result can be obtained.
[0187] The classification model of the present invention is further improved on the basis of the existing BERT model. Compared with the original model, it adds filtering and expansion functions from a macroscopic perspective. Therefore, it is named the FE-BERT (Filter-Expand-BERT) resume screening model. Specifically, the FE-BERT model of the present invention includes:
[0188] (1) Input layer
[0189] The input layer of the FE-BERT model is composed of three parts, namely the word embedding representation V of the preprocessing result c =(w j,1 ,w j,2 ,w j,3 ,…,w j,n ), the word embedding representation V of the topic features after expansion based on category features feature =(w i,1 ,w i,2 ,w i,3 ,…,w i,k ), and the word embedding representation V of the words after vocabulary filtering gram =(w t,1 ,w t,2 ,w t,3 ,…,w t,m ). Superimpose the word embedding representations of the three parts to obtain the vector representation of the input short text sequence. Specifically, the input layer converts the three-layer input into the above-mentioned vector representation with a fixed dimension through the embedding layer.
[0190] The input layer of the FE-BERT short text classification model is constructed in the above manner, which can effectively alleviate the problem of sparse short text features and enhance the ability of features to vectorize short texts; at the same time, it alleviates the impact of meaningless words and words with low contribution to category discrimination in the sub-word list formed after N-gram word extraction on the vectorization representation of short texts and the training of classification models.
[0191] (2) Transformer Encoder layer
[0192] The Transformer Encoder is used to calculate Attention to capture the connection between any words in the sentence. It is stacked by multiple self-attention layers and feed-forward neural networks. The attention mechanism is the core of Transformer. When dealing with long sentences, the accumulation of information requires a more complex model to remember. In natural language processing, the attention mechanism is to explore the focus of attention of one sentence on another sentence, while the self-attention mechanism applies the attention mechanism to two adjacent sentences before and after, thereby calculating the association and semantic information between the words in the sentence. The word vector representation obtained by this calculation takes into account other words in the sentence and the relationship between the upper and lower sentences.
[0193] Furthermore, the improved calculation process of the self-attention mechanism in the present invention is as follows:
[0194]
[0195] In the above formula:
[0196] Q is the query vector; K is the key vector; V is the value vector; d k is the dimension of the key vector; ζ(Q, K) is the dynamic scaling term, which is used to solve the problem that the original attention is not sensitive to key features, such as highlighting important tokens in long texts; is the context-aware term. By combining the information of V, the attention mechanism can capture local semantics, not just the similarity between Q and K; P(Q, K) is the bias term. By explicitly adding position or syntactic structure information, the model's ability to model sequence order or hierarchical relationships is improved; ρ is a learnable parameter or hyperparameter, and the initial value range is [0, 1], which is used to control the contribution of the context-aware term to the attention distribution.
[0197] (3) Output layer
[0198] The output layer takes the output vector at the [CLS] position for the classification task. For the sequence labeling task, it takes the output vector of each token. Each token is passed to the softmax for calculation, and the result is mapped to the interval (0, 1) to represent the probability that the input sequence belongs to different categories, thereby performing multi-classification. The category with the highest probability is used as the predicted category and compared with the actual category. By calculating the error and selecting an appropriate optimizer, the parameters are updated. To reduce the training time of the FE-BERT resume text classification model, the output layer of the present invention preferably uses a fully connected Linear layer and a softmax layer, so that the multi-classification problem is converted into several binary classification problems by constructing a Huffman tree.
[0199] By comparing the error between the predicted category output by the hierarchical softmax function and the actual category of the resume text, and using the stochastic gradient descent method as the optimizer, the value that minimizes the loss function in the gradient direction of the error with respect to the weight matrix is obtained, which is the optimal parameter of the FE-BERT resume text classification model.
[0200] The present invention also provides a resume screening system, including:
[0201] A data acquisition and preprocessing module, which is used to acquire job requirement text data and job seeker resume text data, and preprocess the text data respectively to obtain corresponding preprocessed text vectors;
[0202] A keyword extraction and category feature expansion module, which is used to perform keyword extraction and category feature expansion on the preprocessed text vectors based on the TFIDF-LDA model to obtain corresponding keyword vectors after category feature expansion;
[0203] A word selection and vocabulary filtering module, which is used to perform word selection and vocabulary filtering on the keyword vectors after category feature expansion based on N-gram of information entropy to obtain filtered keyword vectors;
[0204] A resume preliminary screening module, which is used to combine with a deep learning matching degree calculation module to calculate the comprehensive matching degree between the job and the job seeker based on the filtered keyword vectors, and preliminarily screen the job seeker's resume in combination with a set comprehensive matching degree threshold;
[0205] A resume text classification module, which is used to perform resume text classification based on the FE-BERT model to obtain a resume screening result; among them, the preprocessed text vectors corresponding to the resumes and job requirement texts after preliminary screening, the keyword vectors after category feature expansion, and the filtered keyword vectors are used as the three-layer inputs of the FE-BERT model.
[0206] The specific working processes of the modules in the present invention are the same as those described above, and will not be elaborated here. By using the system of the present invention, the resume screening method of the present invention can be executed, improving the matching degree between the screened resumes and the positions, and enhancing the screening efficiency.
[0207] The present invention will be further described in detail below in conjunction with specific embodiments.
[0208] The resume screening method based on the BERT model in this embodiment includes:
[0209] Step 1: Obtain the text data of the job requirements and the text data of the job seeker's resume, and preprocess the text data respectively to obtain the corresponding preprocessed text vectors. The job requirements text and the job seeker's resume text in this embodiment are as follows.
[0210] Job requirements text:
[0211] Job title: Data Analyst
[0212] Job description:
[0213] (1) Be proficient in data analysis tools such as Python and R;
[0214] (2) Be familiar with SQL databases and be able to perform complex queries and data processing;
[0215] (3) Have a certain understanding of machine learning algorithms and be able to apply common algorithms to solve problems;
[0216] (4) Have strong logical thinking ability and data sensitivity, and be able to extract valuable information from data;
[0217] (5) Bachelor's degree or above, preferably in statistics, mathematics, computer-related majors.
[0218] Resume 1:
[0219] Educational background:
[0220] Bachelor's degree, Computer Science and Technology, XX University
[0221] Skills:
[0222] - Be proficient in Python and SQL;
[0223] - Have strong data analysis ability and have used R language for data visualization;
[0224] - Master the basic knowledge of machine learning and be able to apply Sklearn to implement simple prediction models;
[0225] Project experience:
[0226] - Internship in data analysis, completed a sales data analysis project, and optimized the inventory management strategy.
[0227] Resume 2:
[0228] Educational Background:
[0229] Master's degree in Statistics, XX University
[0230] Skills:
[0231] - Proficient in R language, good at statistical modeling and data processing;
[0232] - Master Python and SQL, able to complete the ETL process independently;
[0233] - Familiar with common machine learning algorithms such as decision trees and random forests.
[0234] Project Experience:
[0235] - Participated in a market analysis project, using regression models to predict sales.
[0236] Resume 3:
[0237] Educational Background:
[0238] Bachelor's degree in Information and Computational Science, XX University
[0239] Skills:
[0240] - Familiar with Java and C languages, with some data analysis experience;
[0241] - Can use Excel for data processing and simple statistical analysis;
[0242] - Have some understanding of Python, but lack practical experience.
[0243] Project Experience:
[0244] - Developed a student information management system based on Java.
[0245] Step 2: Based on the TFIDF-LDA model, perform topic word extraction and category feature augmentation on the preprocessed text vector to obtain the corresponding topic word vector after category feature augmentation;
[0246] (1) Based on TF-IDF, perform topic word extraction on the preprocessed text vector, and the obtained candidate topic word lists are as follows:
[0247] TF-IDF topic words of the job requirement text:
[0248] ['Data analysis', 'Python', 'R language', 'SQL', 'Machine learning', 'Logical thinking', 'Statistics', 'Mathematics', 'Computer']
[0249] TF-IDF topic words of resume text:
[0250] Resume 1: ['Python', 'SQL', 'Data analysis', 'R language', 'Sklearn', 'Prediction model', 'Sales data']
[0251] Resume 2: ['R language', 'Statistical modeling', 'Python', 'SQL', 'ETL', 'Machine learning', 'Random forest']
[0252] Resume 3: ['Java', 'C language', 'Data analysis', 'Excel', 'Statistical analysis', 'Python']
[0253] (2) Extract text topic words from the preprocessed text vectors based on the LDA model, and the obtained text topic word list is as follows:
[0254] LDA topic words of job requirement text:
[0255] ['Data analysis', 'Python', 'R language','Machine learning', 'SQL', 'Statistics']
[0256] LDA topic words of resume text:
[0257] Resume 1: ['Data analysis', 'Python', 'SQL', 'Prediction model']
[0258] Resume 2: ['Statistics', 'Python', 'R language', 'ETL']
[0259] Resume 3: ['Java', 'Data analysis', 'Statistical analysis', 'Python']
[0260] (3) The expanded topic word vectors after category feature expansion are as follows:
[0261] Expanded features of job requirements:
[0262] ['Data analysis', 'Python', 'R language', 'SQL', 'Machine learning', 'Logical thinking', 'Statistics', 'Mathematics', 'Computer']
[0263] Expanded features of resume text:
[0264] Resume 1: ['Python', 'SQL', 'Data Analysis', 'R Language', 'Sklearn', 'Prediction Model', 'Sales Data']
[0265] Resume 2: ['R Language', 'Statistical Modeling', 'Python', 'SQL', 'ETL', 'Machine Learning', 'Random Forest']
[0266] Resume 3: ['Java', 'C Language', 'Data Analysis', 'Excel', 'Statistical Analysis', 'Python']
[0267] Step 3: Based on information entropy N-gram word extraction and vocabulary filtering, the filtered keywords are as follows:
[0268] Job Requirements: ['Data Analysis', 'Python', 'SQL', 'R Language', 'Machine Learning']
[0269] Resume 1: ['Python', 'SQL', 'Data Analysis', 'R Language', 'Sklearn']
[0270] Resume 2: ['Python', 'SQL', 'Data Analysis', 'R Language', 'Sklearn']
[0271] Resume 3: ['Java', 'Data Analysis', 'Python']
[0272] Step 4: Combine the deep learning matching degree calculation module, calculate the comprehensive matching degree between the job and the job seekers based on the filtered topic word vectors, and perform a preliminary screening of the job seekers' resumes in combination with the set comprehensive matching degree threshold
[0273] Similarity calculation results:
[0274] Similarity between Resume 1 and job requirements: 0.85
[0275] Similarity between Resume 2 and job requirements: 0.88
[0276] Similarity between Resume 3 and job requirements: 0.65
[0277] According to the set threshold (such as 0.8), Resumes 1 and 2 are preliminarily screened out.
[0278] Step 5: Use the preprocessed text vectors corresponding to Resumes 1 and 2, the topic word vectors expanded by category features, and the filtered topic word vectors as the three-layer inputs of the FE-BERT model, and perform resume text classification through model processing to obtain the resume screening results.
[0279] Softmax mapping calculation results in:
[0280] Resume 1: P(inappropriate) ≈ 0.0832, P(appropriate) ≈ 0.8268
[0281] Resume 2: P(inappropriate) ≈ 0.0219, P(appropriate) ≈ 0.8781
[0282] This indicates that the matching degrees of Resume 1 and Resume 2 with the job requirements are 82.68% and 87.81% respectively. Therefore, Resume 2 is a more suitable candidate.
[0283] Verification and Result Analysis of FE-BERT Resume Text Classification Effect:
[0284] Since the number of resumes affects the performance of the method, the performance will be compared with other methods in multiple aspects under a certain number of resumes below. For the convenience of subsequent description, several groups of models involved in the experiment are named as follows:
[0285] NBC: Refers to the short text classification model based on Naive Bayes;
[0286] SVM: Refers to the short text classification model based on Support Vector Machine;
[0287] BERT: Refers to the basic BERT model, which does not perform category feature expansion and does not perform N-gram word extraction vocabulary filtering;
[0288] FE-BERT: Refers to the resume text classification model that uses the category feature expansion method based on TFIDF-LDA and the N-gram word extraction vocabulary filtering method based on information entropy to expand the category features of BERT and eliminate meaningless words and words with low contribution to category discrimination;
[0289] By performing model prediction on the job seeker resume text classification dataset for the four resume text classification models of NBC, SVM, BERT, and FE-BERT, the classification effect comparison is obtained, as shown in Figure 2 and Figure 3 shown. From the perspective of Macro-average, an overall comparison of each model is carried out, as shown in Figure 4 shown. It can be analyzed that the classification effect of FE-BERT has been significantly improved compared with BERT, SVM, and NBC, approximately increased by about 3% to 5%.
[0290] By respectively comparing and evaluating the classification effects of each category and the overall classification effect of the model, it is further verified that the present invention can effectively improve the resume screening result by introducing the category feature expansion method based on TFIDF-LDA and the N-gram word extraction vocabulary filtering method based on information entropy.
Claims
1. A resume screening method based on the BERT model, characterized in that, Including: Obtain the text data of job requirements and the text data of job seeker resumes, and preprocess the text data respectively to obtain the corresponding preprocessed text vectors; Based on the TFIDF-LDA model, extract the topic words and expand the category features of the preprocessed text vectors to obtain the corresponding topic word vectors after category feature expansion; Perform word filtering on the topic word vectors after category feature expansion based on the N-gram of information entropy to obtain the filtered topic word vectors; Combined with the deep learning matching degree calculation module, calculate the comprehensive matching degree between the job and the job seeker according to the filtered topic word vectors, and perform a preliminary screening of the job seeker resumes in combination with the set comprehensive matching degree threshold; Take the resumes obtained after the preliminary screening, the preprocessed text vectors corresponding to the job requirements text, the topic word vectors after category feature expansion, and the filtered topic word vectors as the three-layer inputs of the FE-BERT model, and perform resume text classification through model processing to obtain the resume screening results.
2. The resume screening method based on the BERT model according to claim 1, wherein The respective preprocessing of the text data includes: Perform paragraph and sentence segmentation on the obtained text data of job requirements and the text data of job seeker resumes into short texts; Perform preprocessing including data cleaning, Chinese word segmentation, and stop word removal on the obtained short texts respectively; Convert the preprocessed short texts into numerical vectors respectively, that is, obtain the corresponding preprocessed text vectors.
3. The resume screening method based on the BERT model according to claim 1, wherein The extracting of the topic words and expanding the category features of the preprocessed text vectors based on the TFIDF-LDA model to obtain the corresponding topic word vectors after category feature expansion includes: Extract the topic words from the preprocessed text vectors based on TF-IDF to obtain the candidate topic word vectors; Extract the text topic words from the preprocessed text vectors based on the LDA model to obtain the text topic word vectors; Based on the set of the obtained candidate topic word vectors and text topic word vectors, construct a word co-occurrence matrix; Count the frequencies of each topic word appearing within a certain window, and calculate the strength PMI of the word co-occurrence relationship; For the phrases with higher PMI scores and reasonable combination patterns, consider them as possible new words and add them to the set of candidate topic words and text topic words.
4. The resume screening method according to claim 3, wherein The calculation formula of the TF-IDF value is as follows: Where: IDF is a way to measure the general importance of a specified vocabulary; n i,j represents the number of occurrences of the vocabulary t i in the short text d j ; |D| represents the total number of short texts; |d j | is the length of the short text d j ; is the average length of the short text; |{j:t i ∈d j}| indicates the number of short texts containing the vocabulary t i ; λ∈[0,1]: used to control the smoothing degree of the word frequency; α∈[0,1]: used to control the normalization of the text length; μ≥0: is a smoothing factor used to prevent outliers caused by too low occurrence times of some words; P(t i |C) represents the generation probability of the vocabulary t i in the corpus C.
5. The resume screening method according to claim 3, wherein In the LDA model, the generation probability formula of the whole document is: In the above formula: θ mz is the distribution probability of the topic z in m document d; is the probability of generating word w under topic z i ; K is the total number of topics; N is the total number of words in document d m in the middle; β>0: controls the sharpness of the topic distribution; η∈[0,1): penalty coefficient, suppressing the strong association between high-frequency words and topics, and improving the generalization ability of the model.
6. The resume screening method according to any one of claims 1-5, characterized in that The performing of word filtering on the topic word vectors after category feature expansion based on the N-gram of information entropy includes: (1) Perform 2-gram and 3-gram word extraction on the topic word vectors after category feature expansion respectively, and take the union of the sub-words of the two to form a sub-word list; (2) Calculate the information entropy between sub - word categories and the information entropy within sub - word categories for the sub - word list respectively to obtain the sub - word information entropy IE(t); the calculation of the sub - word information entropy IE(t) is as follows: β w (H w (t)) = H w (t) In the above formula: β w (H w (t)) is the information entropy within the category of the vocabulary t; β b (H b (t)) is the information entropy between the categories of the vocabulary t; ε ∈ [0, 1): used to control the decay rate of historical information; ν, ω > 0: used to adjust the weight ratio of local and global entropy, ν amplifies the current local entropy, and ω suppresses the reverse effect of global entropy; δ≥0: used to balance the interaction term H w (t)H b (t) intensity to enhance the nonlinear coupling effect; (3) Take the information entropy quartiles according to the sub - word information entropy, determine the preset threshold for filtering sub - words through multiple experiments, and directly add the sub - words with information entropy greater than the preset threshold to the final sub - word list.
7. The resume screening method according to any one of claims 1-5, characterized in that The combined deep - learning matching degree calculation module calculates the comprehensive matching degree between the position and the job seeker based on the filtered topic - word vectors, and preliminarily screens the job seeker's resume in combination with the set comprehensive matching degree threshold, including: Calculate the cosine similarity of the corresponding feature vectors based on the filtered job requirements and the resume text topic word vectors Calculate the weighted score θ for the interaction of each pair of features ij ·ReLU(cos(h pi ,h cj ))·ψ(h pi ,h cj )·S j ; Perform comprehensive matching degree calculation, and use the following scalar MatchScore(P, C) to represent the comprehensive matching degree between the resume and the position: Among them, θ ij is the dynamic weight, with a range of [0, 1], obtained through normalization or training and learning; ψ(h pi , h cj ) is the dynamic adjustment factor based on the feature vector; is the job stability weight; ContextAlign(P, C) is the context-based alignment factor, used to capture the global semantic relationship between the position and the job seeker; ξ is the coefficient of context alignment, with a range of [0, +∞), and parameter tuning is required to balance the weights of skill and non-skill factors.
8. The resume screening method according to claim 7, wherein The dynamic adjustment factor ψ(h pi ,h cj ) based on the eigenvector is calculated as follows: The calculation formula of ContextAlign(P, C) is: ContextAlign(P, C) χ is an adjustable scale parameter used to control the decay rate of the reciprocal of the Euclidean distance between feature vectors; deg(p i ) and deg(c j ) represent the degrees of the job feature pi and the candidate feature cj, respectively.
9. The resume screening method according to any one of claims 1-5, characterized in that The FE - BERT model is improved based on the BERT model, and it includes: (1) Input layer The input layer of the FE - BERT model consists of three parts, namely the word - embedding representation of the pre - processing result, the word - embedding representation of the topic features expanded based on category features, and the word - embedding representation after filtering by the word - taking vocabulary. Stack the word - embedding representations of the three parts to obtain the vector representation of the input short - text sequence; (2) Transformer Encoder layer The Transformer Encoder is stacked by multiple self - attention layers and feed - forward neural networks. The calculation process after the improvement of the self - attention mechanism is as follows: In the above formula: Q is the query vector; K is the key vector; V is the value vector; d k is the dimension of the key vector; is the dynamic scaling term; is the context-aware term; P(Q, K) is the bias term; ρ is a learnable parameter or hyperparameter, with an initial value range in [0, 1]; is the scalar weight parameter, with a range of (0, 1]; (3) Output layer The output layer uses a fully - connected Linear layer and a softmax layer.
10. The resume screening method according to claim 9, wherein, By comparing the error between the predicted category output by the hierarchical softmax function and the actual category to which the resume text belongs, and using the stochastic gradient descent method as the optimizer, find the value that minimizes the loss function in the gradient direction of the error with respect to the weight matrix, which is the optimal parameter of the FE - BERT resume text classification model.
Citation Information
Patent Citations
A human resources resume screening method based on word classification
CN118427213B
Cited By
Intelligent employment data processing method and system based on cloud computing
CN120822930A
A cloud computing-based intelligent employment data processing method and system
CN120822930B
Intelligent generation system for resume file talent portrait report based on AI large model
CN120873179A
Job application data processing method and system for overseas recruitment
CN121352752A
Job Search Data Processing Methods and Systems for Overseas Recruitment
CN121352752B