Resume label generation method and device and medium
By preprocessing, word frequency weighting and nonlinear classification of resume text, combined with a hierarchical tag library, the problem of insufficient accuracy in resume tag generation in existing technologies is solved, and efficient and accurate resume tag generation and enterprise-level real-time processing are achieved.
Patent Information
- Application Number
- CN202510810951.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-30
Smart Images

Figure CN120724993A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a resume tag generation method, device, and medium. Background Art
[0002] With the widespread adoption of digital recruitment, companies are faced with the challenge of processing massive volumes of resumes daily. Manual resume screening is not only inefficient but also susceptible to subjective factors, making it difficult for companies to quickly and accurately match talent. Therefore, extracting key information from large volumes of resumes and performing classification and screening has become crucial for improving recruitment efficiency and quality.
[0003] Resume tagging technology, a core component of intelligent recruitment, aims to extract key information from resumes through automated means, transforming unstructured resume text into a structured tagging system, thereby enabling efficient resume classification, retrieval, and recommendation. However, existing resume tagging methods often focus on optimizing a single model, relying solely on a single model to extract and classify resume text. This makes it difficult to accurately capture the semantic connections between words in resume text, and can lead to information loss for long texts, resulting in inaccurate tagging and poor classification results. Summary of the Invention
[0004] In order to solve the above problems, this application proposes a resume tag generation method, including:
[0005] Preprocessing the original resume text to segment the original resume data into multiple text words;
[0006] Counting the word frequencies corresponding to the text words, and assigning corresponding importance weights to the text words according to the word frequencies;
[0007] Vectorizing the text words to obtain corresponding word vectors, and aggregating the word vectors according to the importance weights to obtain a resume text vector corresponding to the original resume text;
[0008] Based on a preset hierarchical label library, a nonlinear classification algorithm is used to perform multi-label classification on the text semantic vector, and a predicted label of the original resume text is output;
[0009] The tag prediction values are filtered based on a preset threshold to obtain a final resume tag set.
[0010] On the other hand, the present application also proposes a resume tag generation method and device, comprising:
[0011] at least one processor; and,
[0012] a memory communicatively connected to the at least one processor; wherein,
[0013] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a resume tag generation method as described in the above example.
[0014] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured as: a resume label generation method as described in the above example.
[0015] This application proposes a resume tag generation method that can bring the following beneficial effects:
[0016] By integrating frequency-weighted semantic representation with a nonlinear classification algorithm, the accuracy and adaptability of resume tag generation have been significantly improved. The frequency-weighted algorithm effectively suppresses high-frequency, general-purpose words and highlights low-frequency keywords. Combined with the semantic representation capabilities of subword embedding technology for unregistered words, the semantic vectors of resume text accurately capture core information. Furthermore, the kernel-based nonlinear classification algorithm can handle complex semantic boundaries and resolve the challenge of distinguishing similar tags. It also supports a dynamically updated hierarchical tag library, enabling rapid adaptation to changes in industry terminology and flexible response to recruitment needs in diverse fields.
[0017] In terms of model efficiency and interpretability, this paper constructs a lightweight hybrid architecture. By averaging pooled word vectors and optimizing the SVM classification decision process, it significantly improves processing speed, meeting the needs of enterprise-level real-time resume processing while significantly reducing computing resource consumption. Furthermore, the SVM's classification logic, based on the optimal hyperplane, is inherently interpretable. Combined with the hierarchical association and rule merging mechanisms of the label optimization module, a structured labeling system is formed, making it easier for recruiters to understand the basis for label generation and providing transparent and reliable support for intelligent recruitment decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0019] Figure 1 This is a flow chart of a resume tag generation method according to an embodiment of the present application;
[0020] Figure 2 This is a schematic diagram of a resume label generation method and device in an embodiment of the present application. DETAILED DESCRIPTION
[0021] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0022] The technical solutions provided by the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0023] like Figure 1 As shown, the embodiment of the present application provides a resume tag generation method, including:
[0024] S101: Preprocess the original resume text and segment the original resume data into multiple text words.
[0025] Specifically, it reads resume data from multiple channels, including PDF, Word, HTML and other formats, cleans the original resume data, removes HTML tags and special characters, obtains standardized resume text, uses word segmentation tools to segment the standardized resume text, obtains multiple text words, and filters out stop words in the text words.
[0026] In an embodiment of the present application, irrelevant content such as headers, footers, and image annotations are removed through regular expressions, the text is case-normalized, all English letters are converted to lowercase, Chinese word segmentation is performed using the Jieba word segmentation tool, a word list is generated, a custom stop word library is loaded, and stop words in the word segmentation results are filtered.
[0027] S102: Counting the word frequencies corresponding to the text words, and assigning corresponding importance weights to the text words according to the word frequencies.
[0028] Specifically, the word frequency corresponding to the text word is counted, the word probability is calculated, and the weight value is calculated based on the word occurrence probability. It is determined whether the word frequency is higher than a preset frequency threshold. If so, the text word corresponding to the word frequency is determined to be a word, and the high-frequency word is assigned a corresponding importance weight based on the high-frequency weight allocation standard; if not, the text word corresponding to the word frequency is determined to be a low-frequency word, and the low-frequency word is assigned a corresponding importance weight based on the low-frequency weight allocation standard.
[0029] It should be noted that high-frequency words are assigned lower weight values, and low-frequency words are assigned higher weight values. The weight calculation requires adjusting the degree of suppression of high-frequency words through a preset adjustable parameter, and the parameter value range is 0.001 to 0.1.
[0030] In the embodiment of the present application, the SIF algorithm is used to weight words according to their frequency of occurrence in the corpus. The calculation formula is: Among them, w i represents the SIF weight of the i-th word, α is a hyperparameter, usually ranging from 0.001 to 0.1, which is used to adjust the degree of suppression of high-frequency words, p(w i ) represents the word w i The probability of appearing in the corpus can be obtained by counting the words w in the corpus i The formula is obtained by dividing the number of occurrences of by the total number of words in the corpus. This formula reduces the influence of common words and highlights important words by assigning lower weights to high-frequency words and higher weights to low-frequency words.
[0031] S103: Vectorize the text words to obtain corresponding word vectors, aggregate the word vectors according to the importance weights, and obtain a resume text vector corresponding to the original resume text.
[0032] Specifically, the text words are split into subword units, which are vectorized to obtain corresponding subword vectors. The subword vectors are then superimposed to obtain the word vector corresponding to the text word. For unregistered words, vector representations are generated directly based on their subword combinations.
[0033] Furthermore, based on the importance weights corresponding to the text words, the average value of all word vectors in the original resume text is calculated, and the average value is aggregated to generate a resume text vector corresponding to the original resume text.
[0034] In the embodiment of the present application, the word vector is generated based on the subword information of the word through the fasttext model. For a word w, it will be split into multiple subwords g1, g2, ..., g n , each subword corresponds to a vector x g1 ,x g2 ,…,x gn Using the fasttext model, we can get the word vector v corresponding to each word w w In this way, reasonable word vector representations can be generated even for unregistered words, providing rich semantic information for the resume text.
[0035] Aggregate the word vectors of each word obtained, calculate the average value of all word vectors by average pooling, and obtain the vector representation x of the resume text doc , the formula is as follows: Where n is the number of words in the resume text, x i is the word vector of the i-th word. This vector can comprehensively reflect the semantic information of the resume text.
[0036] It should be noted that based on industry characteristics and recruitment needs, a tag library is pre-built, which contains various tags that may be used to describe the content of a resume, such as the first-level tags "skills" and "experience", and the second-level tags "programming language" and "project experience".
[0037] S104: Based on a preset hierarchical label library, multi-label classification is performed on the text semantic vector using a nonlinear classification algorithm to output a predicted label of the original resume text.
[0038] Specifically, the resume text vector is input into the pre-trained support vector machine model, and the resume text vector is mapped to a high-dimensional space through the kernel function in the support vector machine model. A high-dimensional linear optimization objective is constructed, and based on the high-dimensional linear optimization objective, a classification decision function is constructed. Through the classification decision function, the label scores of each predicted label in the resume text vector are output.
[0039] Among them, based on Lagrangian duality, the classification optimization objective corresponding to the resume text vector is converted into a dual optimization objective, the dual optimization objective is solved, the optimal Lagrangian multiplier and classification hyperplane parameters are determined, and the sample inner product in the high-dimensional space is calculated through the kernel function in the support vector machine model. Based on the optimal Lagrangian multiplier, classification hyperplane parameters and sample inner product, a high-dimensional linear optimization objective is constructed.
[0040] In the embodiment of the present application, the support vector machine (SVM) algorithm is used to classify the resume text vector. For the linearly separable case, the goal of SVM is to find an optimal hyperplane w T x+b=0, which maximizes the distance between different categories of data points and the hyperplane. The optimization objective function is: sty i (w T x i +b)≥1,i=1,2,...,nIn the actual resume label classification task, the data is usually linearly inseparable, that is, it is impossible to find a linear hyperplane to completely separate the resume text vectors of different label categories. Therefore, SVM introduces the kernel function k(x i ,x j ), mapping the original data from low-dimensional space to high-dimensional space, so that the resume text vector becomes linearly separable in the high-dimensional space. After introducing the kernel function, the dual problem optimization objective function of SVM becomes: Among them, C is the penalty parameter used to control the model's tolerance to misclassification. By solving this optimization problem, the optimal Lagrange multiplier α is obtained * , and then determine the classification decision function:
[0041] S105: Filter the tag prediction values based on a preset threshold to obtain a final resume tag set.
[0042] Specifically, the predicted tags whose tag scores are lower than a preset threshold are deleted, and the deleted predicted tags are supplemented or merged based on a preset hierarchical tag library to obtain a final tag set of the original resume text.
[0043] In the embodiment of the present application, for the input text vector x, it is input into the trained SVM model, and the classification decision function The output value is calculated and the label category of the resume is determined based on whether the output value is positive or negative. The labels obtained by SVM classification are filtered and optimized, removing some labels with low relevance to the resume content. At the same time, labels are merged and adjusted according to preset rules to obtain the final accurate and concise resume label.
[0044] By integrating frequency-weighted semantic representation with a nonlinear classification algorithm, the accuracy and adaptability of resume tag generation have been significantly improved. The frequency-weighted algorithm effectively suppresses high-frequency, general-purpose words and highlights low-frequency keywords. Combined with the semantic representation capabilities of subword embedding technology for unregistered words, the semantic vectors of resume text accurately capture core information. Furthermore, the kernel-based nonlinear classification algorithm can handle complex semantic boundaries and resolve the challenge of distinguishing similar tags. It also supports a dynamically updated hierarchical tag library, enabling rapid adaptation to changes in industry terminology and flexible response to recruitment needs in diverse fields.
[0045] In terms of model efficiency and interpretability, this paper constructs a lightweight hybrid architecture, which greatly improves the processing speed by averaging the pooled word vectors and optimizing the SVM classification decision process, meeting the enterprise-level real-time resume processing needs, and significantly reducing computing resource consumption. In addition, The classification logic based on the optimal hyperplane is naturally interpretable. Combined with the hierarchical association and rule merging mechanism of the label optimization module, a structured label system is formed, which makes it easier for recruiters to understand the basis for label generation and provides transparent and reliable support for intelligent recruitment decisions.
[0046] like Figure 2 As shown, the embodiment of the present application also proposes a resume tag generation method and device, including:
[0047] at least one processor; and,
[0048] a memory communicatively connected to the at least one processor; wherein,
[0049] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a resume tag generation method as described in any of the above embodiments.
[0050] An embodiment of the present application also provides a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured to be: a resume label generation method as described in any of the above embodiments.
[0051] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0052] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0053] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0054] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0055] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0056] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0057] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0058] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0059] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0060] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0061] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A resume tag generation method, characterized in that: include: Preprocessing the original resume text to segment the original resume data into multiple text words; Counting the word frequencies corresponding to the text words, and assigning corresponding importance weights to the text words according to the word frequencies; Vectorizing the text words to obtain corresponding word vectors, and aggregating the word vectors according to the importance weights to obtain a resume text vector corresponding to the original resume text; Based on a preset hierarchical label library, a nonlinear classification algorithm is used to perform multi-label classification on the text semantic vector, and a predicted label of the original resume text is output; The tag prediction values are filtered based on a preset threshold to obtain a final resume tag set.
2. A resume tag generation method according to claim 1, characterized in that: Assigning corresponding importance weights to the text words according to the word frequencies specifically includes: Determining whether the frequency of the word is higher than a preset frequency threshold; If so, determining that the text word corresponding to the word frequency is a high-frequency word, and assigning a corresponding importance weight to the high-frequency word based on a high-frequency weight allocation standard; If not, it is determined that the text word corresponding to the word frequency is a low-frequency word, and the low-frequency word is assigned a corresponding importance weight based on a low-frequency weight allocation standard.
3. A resume tag generation method according to claim 1, characterized in that: The vectorizing of the text words to obtain corresponding word vectors specifically includes: Splitting the text words into subword units, vectorizing the subword units to obtain corresponding subword vectors; The subword vectors are superimposed to obtain the word vector corresponding to the text word.
4. A resume tag generation method according to claim 3, characterized in that: The step of aggregating the word vectors according to the importance weights to obtain a resume text vector corresponding to the original resume text specifically includes: Calculate the average value of all word vectors in the original resume text based on the importance weights corresponding to the text words; Aggregate the average values to generate a resume text vector corresponding to the original resume text.
5. The resume tag generation method according to claim 1, characterized in that: The multi-label classification of the text semantic vector by a nonlinear classification algorithm and the output of multiple label prediction values of the original resume text specifically include: Inputting the resume text vector into a pre-trained support vector machine model, mapping the resume text vector into a high-dimensional space through a kernel function in the support vector machine model, and constructing a high-dimensional linear optimization objective; Based on the high-dimensional linear optimization objective, a classification decision function is constructed, and through the classification decision function, the label score of each predicted label in the resume text vector is output.
6. A resume tag generation method according to claim 5, characterized in that: The kernel function in the support vector machine model is used to map the resume text vector to a high-dimensional space to construct a high-dimensional linear optimization objective, specifically including: Based on Lagrangian duality, the classification optimization objective corresponding to the resume text vector is converted into a dual optimization objective, the dual optimization objective is solved, and the optimal Lagrangian multiplier and classification hyperplane parameters are determined; Calculate the inner product of samples in high-dimensional space through the kernel function in the support vector machine model; A high-dimensional linear optimization objective is constructed based on the optimal Lagrange multiplier, the classification hyperplane parameters and the sample inner product.
7. A resume tag generation method according to claim 5, characterized in that: The tag prediction values are filtered based on a preset threshold to obtain a final resume tag set, specifically including: Deleting the predicted labels whose label scores are lower than a preset threshold; The deleted predicted tags are supplemented based on a preset hierarchical tag library to obtain a final tag set for the original resume text.
8. The resume tag generation method according to claim 1, characterized in that: The pre-processing of the original resume text to segment the original resume data into multiple text words specifically includes: Obtaining original resume data, cleaning the original resume data, removing HTML tags and special characters, and obtaining standardized resume text; The standardized resume text is segmented using a word segmentation tool to obtain multiple text words, and stop words in the text words are filtered out.
9. A resume tag generation method and device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a resume label generation method as described in any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are set to: a resume label generation method as described in any one of claims 1 to 8.
Citation Information
Cited By
Demand processing method and device for fusing keywords and semantics
CN121189512A
A demand processing method and device fusing keywords and semantics
CN121189512B