Resume data processing method and device, equipment and storage medium
By using word segmentation and classification models to identify information entities in resume data, and combining similarity matching to achieve automated standardization of resume data, the problem of non-standard resume descriptions is solved, and processing efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202511402693.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2025-12-02
AI Technical Summary
Inaccurate and non-standard data presentation in resumes leads to low processing efficiency and susceptibility to subjective errors, affecting information screening and talent matching.
By using word segmentation and pre-trained classification models to identify information entities, and using similarity matching to determine target standard information entities, the automated and standardized processing of resume data is achieved.
Significantly improve resume processing efficiency, provide high-quality standardized data, support subsequent management and analysis, reduce manual intervention, and improve accuracy and consistency.
Smart Images

Figure CN121052250A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and more particularly to resume data processing methods, apparatus, devices, and storage media. Background Technology
[0002] Resume data is filled out by job seekers themselves, and there may be instances of non-standard or inaccurate expression. This directly affects the efficiency and accuracy of resume data processing, and impacts the effectiveness of information screening and talent matching.
[0003] In related technologies, resumes are often checked and processed manually, but this method is inefficient and susceptible to subjective errors. Summary of the Invention
[0004] This application provides a resume data processing method, apparatus, device, and storage medium, which can automatically solve the problems of inaccurate and non-standard information description in resumes and greatly improve resume processing efficiency.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0006] Firstly, a resume data processing method is provided, which includes:
[0007] Obtain the target text information from the resume data, and perform word segmentation on the target text information to obtain the text to be processed;
[0008] The text to be processed is input into a pre-trained classification model to obtain the predicted labels of each word segment in the text output by the classification model; wherein, the predicted labels represent whether the word segment belongs to an information entity and the position of the word segment in the information entity, including the information entity start label, the information entity internal label and the non-information entity label;
[0009] The word segment whose predicted label is the starting label of the information entity or the internal label of the entity is used as the target word segment, and the original position of each target word segment in the text to be processed is determined.
[0010] Starting with the target word corresponding to the initial tag of the information entity as the starting point of the concatenation, the target words corresponding to the tags inside the entity are concatenated sequentially according to the original position order in the text to be processed to obtain the information entity to be processed.
[0011] Based on the similarity between the information entity to be processed and the preset standard information entity, the target standard information entity corresponding to the information entity to be processed is determined;
[0012] Based on the target standard information entity, the target text information is standardized to obtain standardized resume data.
[0013] This embodiment first performs word segmentation on the target text, breaking down continuous text into independent word units to lay the foundation for entity recognition. Then, a classification model is used to output prediction results containing three categories of labels: start, interior, and non-entity, identifying the boundaries and structure of information entities and eliminating irrelevant word segmentation interference. Next, target words are filtered and concatenated according to their original positions to obtain complete information entities to be processed. Then, through similarity matching between the information entities to be processed and preset standard information entities, a mapping between non-standard and standard expressions is established to obtain the target standard information entities corresponding to the information entities to be processed. Finally, based on the target standard information entities, the target text information in the resume data is standardized to obtain standardized resume data. This entire process replaces manual operation, significantly improving resume processing efficiency and standardizing non-standard expressions in resumes, providing high-quality standardized data for subsequent resume management and analysis.
[0014] In another possible implementation of the first aspect, determining the target standard information entity corresponding to the information entity to be processed based on the similarity between the information entity to be processed and a preset standard information entity includes:
[0015] Calculate the similarity between the information entity to be processed and each preset standard information entity in the preset database;
[0016] Based on the similarity between the information entity to be processed and each standard information entity, a set of candidate information entities corresponding to the information entity to be processed is determined; wherein, the set of candidate information entities includes a preset number of standard information entities, and each standard information entity is arranged in descending order of similarity with the information entity to be processed;
[0017] The candidate information entity set is sent to the user-side device, and based on the user-side device, the user's instruction information is obtained; the instruction information is used to indicate the target standard information entity corresponding to the information entity to be processed.
[0018] The standard information entity in the candidate information entity set indicated by the instruction information is determined as the target standard information entity, or the custom information entity indicated by the instruction information is determined as the target standard information entity.
[0019] This embodiment first calculates similarity to generate a sorted set of candidate information entities, which helps users quickly focus on high-matching standard entities and reduce screening costs. Then, it supports users to select target standard information entities from the candidate entity set through instruction information, or to manually input custom entities as target standard information entities. This retains the efficiency of automatic matching and improves the accuracy of target standard entity determination by correcting possible misjudgments through manual intervention. At the same time, it supports custom entities as target standard entities, which can flexibly adapt to new information not included in the preset database and enhance the scalability of standardization processing.
[0020] In another possible implementation of the first aspect, calculating the similarity between the information entity to be processed and each standard information entity in a preset database includes:
[0021] Calculate the edit distance between the information entity to be processed and each standard information entity, and use the obtained edit distance as the corresponding similarity; wherein, the magnitude of the edit distance is inversely proportional to the magnitude of the similarity.
[0022] Here, the edit distance quantifies the character differences between the entity to be processed and the standard entity. The calculation logic is simple and intuitive. In this embodiment, the edit distance between characters is used as a similarity index to provide a reliable basis for the ranking of candidate sets and improve the operability and accuracy of subsequent entity matching.
[0023] In another possible implementation of the first aspect, after determining the target standard information entity corresponding to the information entity to be processed, the method further includes:
[0024] Traverse the resume data in the resume database and identify non-standard resume data in the resume database that includes the target text information;
[0025] Based on the target standard information entity, the target text information in the non-standard resume data is batch standardized and corresponding log information is generated.
[0026] This embodiment traverses the resume database to identify non-standard resume data containing the target text information, and standardizes it in batches. The same mapping rule is applied to all corresponding resumes at once, avoiding repetitive operations and significantly improving resume processing efficiency. At the same time, it ensures that resumes with the same non-standard information are updated uniformly to avoid data inconsistency. In addition, the generated log records also provide a basis for subsequent data traceability and verification.
[0027] In another possible implementation of the first aspect, the classification model includes a preprocessing module and a fully connected module;
[0028] The step of inputting the text to be processed into a pre-trained classification model to obtain the predicted labels of each word segment in the text to be processed output by the classification model includes:
[0029] The text to be processed is input into a pre-trained classification model, and the semantic feature vector of each word in the text to be processed is extracted using the preprocessing module.
[0030] Based on the fully connected module, the semantic feature vectors of each word segment are mapped to obtain the predicted label corresponding to each word in the text to be processed.
[0031] In this embodiment, the semantic feature vectors of word segmentation are extracted through the preprocessing module of the classification model to capture the deeper meaning of the context; then, the fully connected module is used to map the semantic features to the label space dimension, achieving a precise conversion from semantic features to predicted labels. These two methods work together to improve the accuracy of label prediction, laying the foundation for accurate extraction of information entities in the subsequent process.
[0032] In another possible implementation of the first aspect, the classification model is trained through the following process:
[0033] The text to be processed from multiple resume data is input into the initial classification model to obtain the predicted label corresponding to each word in each text to be processed;
[0034] The cross-entropy loss between the predicted and true labels corresponding to each word segment in each text to be processed is summed to obtain the loss function value.
[0035] If the loss function value is greater than a preset threshold, the model parameters of the initial classification model are adjusted along the gradient descent direction of the loss function value according to the preset learning rate to obtain the parameter-adjusted classification model.
[0036] Using the parameter-adjusted classification model as a new initial classification model, the process of inputting the text to be processed from multiple resume data into the initial classification model to obtain the predicted label corresponding to each word in each text to be processed is repeated until the loss function value is less than the preset threshold, thus obtaining the trained classification model.
[0037] In training the classification model, the cross-entropy loss function is used to measure the difference between the predicted and true labels. The parameters are adjusted using a gradient descent algorithm at a preset learning rate until the loss function value is less than a threshold, ensuring the model converges to its optimal state. This training process allows the classification model to fully learn the word segmentation-label correspondence rules of resume text, improving its adaptability to resume data and ensuring the accuracy of predicted labels.
[0038] In another possible implementation of the first aspect, after obtaining the standardized resume data, the method further includes:
[0039] Based on job requirements, matching rules are generated for standardized target text information; wherein, the target text information includes one or more of the following: educational background information, previous employer information, work skills information, and qualification information; according to the matching rules, the target text information in multiple standardized resume data is matched in batches to obtain target resumes.
[0040] In this embodiment, based on the standardized resumes, matching rules are generated according to job requirements. By batch matching, target resumes are screened from a large number of resumes, which can avoid omissions and errors caused by non-standard expressions in the resumes, and improve the efficiency and accuracy of recruitment screening.
[0041] Secondly, a resume data processing apparatus is provided, the apparatus comprising:
[0042] The preprocessing unit is used to acquire target text information from resume data and perform word segmentation on the target text information to obtain the text to be processed.
[0043] The prediction unit is used to input the text to be processed into a pre-trained classification model to obtain the predicted labels of each word segment in the text to be processed output by the classification model; wherein, the predicted labels are used to characterize whether the word segment belongs to an information entity and the position of the word segment in the information entity, and the predicted labels include information entity start label, information entity internal label and non-information entity label.
[0044] The determining unit is used to take the word segment whose predicted label is the starting label of the information entity or the internal label of the entity as the target word segment, and determine the original position of each target word segment in the text to be processed;
[0045] The first processing unit is used to take the target word corresponding to the starting tag of the information entity as the starting point of the splicing, and splice the target words corresponding to the tags inside the entity in sequence according to the original position order in the text to be processed, so as to obtain the information entity to be processed.
[0046] The second processing unit is used to determine the target standard information entity corresponding to the information entity to be processed based on the similarity between the information entity to be processed and the preset standard information entity.
[0047] The standardization processing unit is used to standardize the target text information based on the target standard information entity to obtain standardized resume data.
[0048] Thirdly, an electronic device is provided, the method comprising: a memory and at least one processor. The memory is communicatively connected to the processor. The memory is used to store computer program code, the computer program code including computer instructions. When the processor executes the computer instructions, it causes the electronic device to perform the method as described in the first aspect and any possible implementation thereof.
[0049] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions. When executed by a processor, these computer instructions are used to implement the method described in the first aspect and any possible implementation thereof.
[0050] Fifthly, embodiments of this application provide a computer program product that, when run on a computer or executed by a computer's processor, implements the method described in the first aspect and any possible design thereof. The computer may be the electronic device described in the second aspect and any possible implementation thereof.
[0051] It is understood that the beneficial effects achieved by the resume data processing device described in the second aspect, the electronic device described in the third aspect, the computer-readable storage medium described in the fourth aspect, and the computer program product described in the fifth aspect can be referred to as the beneficial effects in the first aspect and any possible implementation thereof, which will not be repeated here. Attached Figure Description
[0052] Figure 1 This is a flowchart illustrating a resume data processing method provided in an embodiment of this application;
[0053] Figure 2 This is a flowchart illustrating another resume data processing method provided in an embodiment of this application;
[0054] Figure 3 This is a flowchart illustrating another resume data processing method provided in an embodiment of this application;
[0055] Figure 4 This is a schematic diagram of the structure of a resume data processing device provided in an embodiment of this application;
[0056] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0057] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.
[0058] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0059] The technical solutions provided in this application, including the collection, storage, use, processing, transmission, provision, and disclosure of financial data or user data, comply with relevant laws and regulations and do not violate public order and good morals.
[0060] The personal information used in this application's technical solution is limited to information for which individual consent has been obtained, including but not limited to notifying and reminding users to read and sign the relevant user agreement before they use the function, and authorizing the relevant user information to be signed.
[0061] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0062] Resume data is filled out by job seekers themselves, and there are common problems with non-standard and inaccurate descriptions, directly affecting the efficiency and accuracy of resume data processing, and thus interfering with information screening and talent matching. Current technologies largely rely on manual review and processing of resumes, which is not only inefficient but also susceptible to subjective errors, failing to effectively address information expression issues and unable to meet the demands for efficient processing.
[0063] To address the issues of inaccurate and non-standard information in resumes and improve resume processing efficiency, the implementation method of this application first segments the target text information in the resume into words, then obtains predicted labels representing whether each segmented word is an information entity and its location through a pre-trained classification model, and uses these labels to filter and splice out the information entities to be processed; subsequently, the similarity between the information entities to be processed and preset standard information entities is combined to determine the target standard entities, and finally the standardization of the target text information is completed. This automated process achieves the standardization of resume data, while replacing manual work and improving resume processing efficiency.
[0064] This application provides a resume data processing method that can be applied to electronic devices. The electronic device can be a single server, a server cluster consisting of multiple servers, a cloud computing platform with data processing capabilities, an edge computing device, a chip, or a device with computing capabilities. This application does not limit the specific form of the electronic device.
[0065] Figure 1 This is a flowchart illustrating a resume data processing method provided in an embodiment of this application. Figure 1 As shown, the method in the embodiments of this application may include:
[0066] S101. Obtain the target text information from the resume data and perform word segmentation on the target text information to obtain the text to be processed.
[0067] The target text information can be one or more of the following: educational background information, previous employer information, work skills information, and qualification information from the resume.
[0068] This embodiment can directly read resume data stored in the resume database of an enterprise recruitment system; it can also receive resume files submitted by job seekers through recruitment platforms, extract the text content in the resume through a text parsing tool, and filter to obtain the fields corresponding to the target text information, i.e., the target text information, such as work experience, employer, education background, degree, years of work experience, etc. in the resume.
[0069] After obtaining the target text information, this embodiment decomposes the target text information into word segments according to semantic logic. For example, the target text information "formerly employed by XX Co., Ltd. in City A" is processed into a word segmentation set {"formerly", "employed by", "City A", "XX", "Co., Ltd."} after word segmentation, and this word segmentation set is the text to be processed.
[0070] S102. Input the text to be processed into a pre-trained classification model to obtain the predicted labels of each word in the text to be processed output by the classification model; wherein, the predicted labels are used to characterize whether the word belongs to an information entity and the position of the word in the information entity. The predicted labels include the information entity start label, the information entity internal label, and the non-information entity label.
[0071] Among them, information entities refer to information in the target text that has complete semantics and can independently represent a specific meaning, such as "xx company", "5 years of work experience", "bachelor's degree", etc.
[0072] The predicted tags specifically include the information entity start tag, denoted as B; the information entity internal tag, denoted as I; and the non-information entity tag, denoted as "O", indicating that the segment does not belong to any information entity.
[0073] This embodiment utilizes a pre-trained classification model to output the predicted label corresponding to each word segment, clarifying its attributes: if the word segment is the first component of a certain type of information entity, the information entity start label is output (such as "B-ORG" or "B-EDU", where suffixes such as "ORG" and "EDU" indicate the entity type, ORG represents an institutional entity, and EDU represents an educational background entity); if the word segment is not a component of a certain type of information entity, the information entity internal label is output (such as "I-ORG" or "I-EDU"); if the word segment does not belong to any information entity, the non-information entity label "O" is output.
[0074] S103. Take the word segment whose predicted label is the starting label of the information entity or the internal label of the entity as the target word segment, and determine the original position of each target word segment in the text to be processed.
[0075] For example, in this embodiment, all word segments and their corresponding predicted tags in the text to be processed are traversed, word segments with the tags "information entity start tag" and "information entity internal tag" are retained, and word segments with the tag "non-information entity tag" ("O") are removed to obtain the target word segment set.
[0076] Moreover, determine the original positions of each target token in the text to be processed. For example, assign a unique index value to each token in the text to be processed, representing the original position of the token in the text to be processed (the index starts from 0 or 1 and increases sequentially according to the position order of the tokens in the text to be processed), and establish a "token-index" correspondence; then, based on the set of target tokens, extract the index value of each target token from the "token-index" correspondence as its original position. If the "token-index" of the text to be processed is {"曾":0, "任职于":1, "A市":2, "XX":3, "互联网":4, "科技":5, "有限公司":6}, then the original position of the target token "A市" is 2, "XX" is 3, "互联网" is 4, "科技" is 5, and "有限公司" is 6.
[0077] S104. Take the target token corresponding to the starting tag of the information entity as the splicing starting point, and sequentially splice the target tokens corresponding to the internal tags of the entity according to the order of the original positions in the text to be processed, to obtain the information entity to be processed.
[0078] Exemplarily, in this embodiment, traverse the set of target tokens and their predicted tags, and filter out the target tokens with the tag of "starting tag of information entity" as the splicing starting point.
[0079] If there are starting tags in the set of target tokens, which respectively correspond to multiple information entities, then assign independent splicing tasks to each starting tag. For example, when the text to be processed simultaneously contains two entities "A City XX Technology" and "B City YY Company", it is necessary to start splicing with "A City" and "B City" respectively.
[0080] In this embodiment, take the target token corresponding to the starting tag of the information entity as the splicing starting point, and sequentially traverse the original positions of the target tokens according to the position order in the text to be processed: If the original position of a certain target token is after the starting token and its tag is "internal tag of information entity", then splice it to the text that has been spliced; If the original position of a certain target token is after the starting token, but its tag is "starting tag of information entity", then after the current information entity to be processed is constructed, use the target token corresponding to the tag of this information entity as the splicing starting point to construct another information entity to be processed; If all tokens after the starting token in the set of target tokens have been traversed, then automatically stop splicing.
[0081] Repeat the above process until all splicing tasks corresponding to the starting tags are completed.
[0082] S105. Determine the target standard information entity corresponding to the information entity to be processed according to the similarity between the information entity to be processed and the preset standard information entity.
[0083] The preset standard information entities include multiple types, and the preset standard information entities correspond to the types of target text information. For example, if the target text information is the employer information in the resume, then the preset standard information entities are the standardized and normalized descriptions of the employer.
[0084] For example, this embodiment can calculate the similarity between the information entity to be processed and a preset standard information entity using algorithms such as edit distance, cosine similarity, and Jaccard similarity. Then, based on the magnitude of the similarity, the target standard information entity corresponding to the information entity to be processed is determined. For example, the standard information entity with the highest similarity is determined as the target standard information entity. Alternatively, after obtaining the similarity between the information entity to be processed and each preset standard information entity, the user further confirms and determines the target standard information entity corresponding to the information entity to be processed.
[0085] S106. Based on the target standard information entity, standardize the target text information to obtain standardized resume data.
[0086] For example, this embodiment can locate the specific position of the information entity to be processed in the target text information through string matching; for example, if the original target text information is "employed at XX company from 2020 to 2023", after entity extraction by the classification model, the information entity to be processed is "XX company", then the position of "XX company" in the text is located after "employed at".
[0087] Replace the located information entity to be processed with the corresponding target standard information entity: replace "XX Company" with "XX Group Holdings Co., Ltd." After the replacement, generate standardized target text information: "Worked at XX Group Holdings Co., Ltd. from 2020 to 2023." Update the corresponding fields in the original resume data with the standardized target text information to obtain the standardized resume data.
[0088] In one feasible implementation, this embodiment can also perform format verification on the updated resume data to ensure field integrity and consistency of expression. Once the verification is passed, it becomes standardized resume data, which can be stored in a standardized resume database for subsequent resume screening, talent matching and other scenarios.
[0089] In summary, this embodiment first performs word segmentation on the target text, breaking down continuous text into independent word units to lay the foundation for entity recognition. Then, a classification model is used to output prediction results containing three categories of labels: start, internal, and non-entity, identifying the boundaries and structure of information entities and eliminating irrelevant word segmentation interference. Next, target words are selected and concatenated according to their original positions to obtain complete information entities to be processed. Then, through similarity matching between the information entities to be processed and preset standard information entities, a mapping between non-standard and standard expressions is established to obtain the target standard information entities corresponding to the information entities to be processed. Finally, based on the target standard information entities, the target text information in the resume data is standardized to obtain standardized resume data. This entire process replaces manual operation, significantly improving resume processing efficiency and standardizing non-standard expressions in resumes, providing high-quality standardized data for subsequent resume management and analysis.
[0090] Figure 2 This is a flowchart illustrating another resume data processing method provided in an embodiment of this application. Figure 2 As shown, the method in the embodiments of this application may include:
[0091] S201. Obtain the target text information from the resume data and perform word segmentation on the target text information to obtain the text to be processed.
[0092] The target text information includes one or more of the following: educational background information, previous employer information, work skills information, and qualification information.
[0093] For example, this step is the same as step 101, and will not be repeated here.
[0094] S202. Input the text to be processed into the pre-trained classification model and use the preprocessing module to extract the semantic feature vector of each word in the text to be processed.
[0095] This embodiment uses RoBERTa as the base model, whose output is represented as H. A fully connected layer, represented as W, is added on top of the RoBERTa output layer to form a classification model for classification tasks.
[0096] The output of the fully connected layer is represented as Z, and its calculation expression is: Z = H·W + b, where b is the bias term.
[0097] In this embodiment, the RoBERTa model is used as a preprocessing module to extract the semantic feature vector of each word in the text to be processed.
[0098] For example, in this embodiment, the text to be processed is input into the classification model in its original order. The preprocessing module first performs tokenization on each word segment, mapping the word-level segment to a tag ID that the model can recognize. Then, the internal Transformer encoder performs bidirectional semantic modeling on the tag ID sequence, outputting a fixed-dimensional semantic feature vector, denoted as H, for each tag. For example, when the text to be processed contains 5 word-level segments, the H dimension output by RoBERTa is "5×768", and each 768-dimensional vector corresponds to the contextual semantics of a word-level segment.
[0099] S203. The fully connected module based on the classification model maps the semantic feature vectors of each word segment to obtain the predicted label corresponding to each word in the text to be processed.
[0100] Among them, the prediction label is used to characterize whether the word segment belongs to the information entity and the position of the word segment in the information entity. The prediction label includes the information entity start label, the information entity internal label and the non-information entity label.
[0101] For example, this embodiment performs "dimensional transformation and scoring" through a fully connected module: using the trained weight matrix, multiplying it by a 768-dimensional feature vector, and adding a bias, the semantic feature vector is mapped and transformed to a dimension corresponding to the number of labels. For example, if there are 3 types of labels, a 3-dimensional vector is output.
[0102] Each value in these 3-dimensional vectors can be understood as "the probability score of the word segment belonging to the corresponding label". For example, the first dimension corresponds to a score of 0.8 for "starting label", the second dimension to a score of 0.1 for "internal label", and the third dimension to a score of 0.1 for "non-entity label". Finally, the label corresponding to the highest probability score is taken as the predicted label of the word segment.
[0103] The classification model is trained through the following process:
[0104] The text to be processed from multiple resume data is input into the initial classification model to obtain the predicted label corresponding to each word in each text to be processed.
[0105] The cross-entropy loss between the predicted and true labels corresponding to each word segment in each text to be processed is summed to obtain the loss function value.
[0106] If the loss function value is greater than the preset threshold, the model parameters of the initial classification model are adjusted along the gradient descent direction of the loss function value according to the preset learning rate, so as to obtain the classification model after parameter adjustment.
[0107] Using the parameter-adjusted classification model as the new initial classification model, the steps of "inputting the text to be processed from multiple resume data into the initial classification model and obtaining the predicted label corresponding to each word in each text to be processed" are repeated until the loss function value is less than the preset threshold, and the trained classification model is obtained.
[0108] This embodiment acquires multiple resume data, segments the target text information in each resume data into words, and then annotates the segmented words using the Begin-Inside-Outside (BIO) annotation method to obtain training data. Here, B represents the beginning of an information entity, I represents the internal part of an information entity, and O represents a non-information entity.
[0109] The classification model is then trained using the training data. This includes training the fully connected modules within the classification model.
[0110] In this process, the weights w and biases b of the fully connected modules are initialized in the initial stage using a random method, such as a uniform distribution or a normal distribution.
[0111] This embodiment uses the cross-entropy loss function to optimize the model parameters. The expression for the cross-entropy loss function is shown in equation (1) below:
[0112]
[0113] Where n is the number of samples, y i For real labels, To predict probabilities.
[0114] Furthermore, the gradient descent algorithm is used to minimize the loss function and update the model parameters θ (including W and b). The parameter update expression is shown in equation (2) below:
[0115]
[0116] Where η is the learning rate. This is the gradient of the loss function.
[0117] Furthermore, this embodiment employs multi-round iterative model training, continuously adjusting model parameters to gradually improve model performance. During training, the model's performance on the validation set is periodically evaluated, and the best-performing model parameters are saved to ensure efficient completion of the classification task.
[0118] S204. Take the word segment whose predicted label is the starting label of the information entity or the internal label of the entity as the target word segment, and determine the original position of each target word segment in the text to be processed.
[0119] For example, this step is the same as step 103, and will not be repeated here.
[0120] S205. Using the target word corresponding to the starting label of the information entity as the starting point of the splicing, according to the original position order in the text to be processed, the target words corresponding to the labels inside the entity are spliced in sequence to obtain the information entity to be processed.
[0121] For example, this step is the same as step 104, and will not be repeated here.
[0122] S206. Based on the similarity between the information entity to be processed and the preset standard information entity, determine the target standard information entity corresponding to the information entity to be processed.
[0123] In one feasible implementation, step 206 includes the following steps:
[0124] Calculate the similarity between the information entity to be processed and each preset standard information entity in the preset database.
[0125] Based on the similarity between the information entity to be processed and each standard information entity, a set of candidate information entities corresponding to the information entity to be processed is determined; wherein, the set of candidate information entities includes a preset number of standard information entities, and each standard information entity is arranged in descending order of similarity with the information entity to be processed.
[0126] The candidate information entity set is sent to the user-side device, and the user's instruction information is obtained based on the user-side device; the instruction information is used to indicate the target standard information entity corresponding to the information entity to be processed.
[0127] The standard information entity in the candidate information entity set indicated by the instruction information is determined as the target standard information entity, or the custom information entity indicated by the instruction information is determined as the target standard information entity.
[0128] In one feasible implementation, this embodiment pre-constructs a preset database, including storing standard information entities categorized according to common types of target information in resume data, such as employer, education level, years of work experience, qualifications, and work skills, covering standardized and normative expressions of common information entities. For example, the employer category stores the full name of the company, and the education level category stores the standardized names of academic qualifications recognized by the education department; and indexes are created for the standard information entities: retrieval indexes are created according to entity type, first letter of characters, and other dimensions to improve the matching efficiency during subsequent similarity calculations and avoid the time-consuming problem caused by traversing the entire database.
[0129] For example, this embodiment uses a single information entity to be processed as a benchmark and performs similarity calculations one by one with all preset standard information entities of the same type in the preset database. For example, if the information entity to be processed is "employer type", then only the "employer type" standard entities in the database are matched to avoid invalid cross-type matching.
[0130] When calculating similarity, the following algorithms can be used: Edit distance algorithm: This calculates the minimum number of character editing operations (including insertion, deletion, and replacement) required to convert the information entity to be processed into a standard information entity. The fewer the number of operations, the higher the similarity. For example, the edit distance between "XX Technology" and "XX Internet Technology Co., Ltd." is 8, while the edit distance between "XX Technology Company" and "YY Technology Company" is 12, indicating a higher similarity. Alternatively, the cosine similarity algorithm can be used: This converts the information entity to be processed and the standard information entity into vector form and calculates the cosine of the angle between the two vectors. The closer the cosine is to 1, the higher the semantic similarity. Finally, the Jaccard similarity algorithm can be used: This calculates the ratio of the intersection to the union of the character sets of the information entity to be processed and the standard information entity. The larger the ratio, the higher the character overlap.
[0131] In one feasible implementation, this embodiment flexibly sets matching thresholds based on the type of information entity (e.g., edit distance threshold ≤ 15 for entities of employers, cosine similarity threshold ≥ 0.8 for entities of educational background), filters out standard information entities whose similarity meets the threshold, and forms a candidate information entity set; or, selects the top 10 standard information entities in similarity ranking to form a candidate information entity set.
[0132] From the candidate information entity set, select the standard information entity with the highest similarity as the target standard information entity corresponding to the information entity to be processed. Alternatively, send it to the user for confirmation.
[0133] In one example, Figure 3 This is a flowchart illustrating another resume data processing method provided in an embodiment of this application, as shown below. Figure 3 As shown, if the target text information is the historical employer information in a resume, after word segmentation, classification model prediction, and entity information extraction, the entity to be processed is an employer-type entity, i.e., the aforementioned organization-type entity ORZ. If the entity to be processed is a non-standard organization, such as xx Technology, then this embodiment determines the corresponding candidate entity set by calculating the similarity between the entity to be processed and various standard information entities in a preset database. Figure 3 As shown, the system includes multiple candidate standard organizations. This set of candidate information entities is sent to the user's device so that HR can select and confirm it. The target standard organization among the candidate standard organizations is determined. Based on the mapping relationship between the target standard organization and non-standard organizations, the target text information is replaced to obtain the updated resume data.
[0134] If the candidate standard entity set is empty or the user confirms that there is no correct standard information entity in the candidate information entity set, then the information entity to be processed is recorded, a custom standard information entity is manually entered as the target standard information entity, and the target standard information entity is added to the preset database, or the custom standard information entity entry process is triggered.
[0135] In one example, this embodiment calculates the edit distance between the information entity to be processed and each standard information entity, and uses the obtained edit distance as the corresponding similarity; wherein, the magnitude of the edit distance is inversely proportional to the magnitude of the similarity.
[0136] For example, in this embodiment, the Levenshtein distance between the information entity to be processed and each standard information entity is calculated as the corresponding similarity. The smaller the calculated distance, the more similar the entities are. The edit distance calculation expression is shown in the following formula (3):
[0137]
[0138] In the formula, lev(a, b) i,j Let L be the Levinstein distance between the first i characters of string a and the first j characters of string b.
[0139] In one feasible implementation, after determining the target standard information entity corresponding to the information entity to be processed, this embodiment may further:
[0140] Iterate through the resume data in the resume database and identify non-standard resume data in the resume database that includes target text information;
[0141] Based on the target standard information entity, the target text information in non-standard resume data is batch standardized and corresponding log information is generated.
[0142] For example, in this embodiment, all resume data in the resume database are traversed to filter out non-standard resume data containing target text information.
[0143] For example, if the target standard entity corresponding to "XX Technology" was previously determined to be "XX Internet Technology Co., Ltd.", then the database is traversed to find all resumes containing the non-standard expression "XX Technology".
[0144] Based on the established "target standard information entity" (such as "XX Internet Technology Co., Ltd."), the "target text information" in all identified non-standard resume data is uniformly replaced, transforming non-standard expressions into standard expressions. For example, "XX Technology" in all resumes is uniformly replaced with "XX Internet Technology Co., Ltd." to achieve consistency in the expression of similar information.
[0145] During the processing, key information is recorded and stored to form a log, including: the resume ID processed, the non-standard text before replacement, the standard text after replacement, the processing time, and whether it was successful. The log is used for subsequent tracking of processing results, verification of standardization effectiveness, or to locate problems when errors occur.
[0146] S207. Based on the target standard information entity, standardize the target text information to obtain standardized resume data.
[0147] In one feasible implementation, this embodiment can also generate matching rules for standardized target text information based on job requirements, and then perform batch matching of target text information in multiple standardized resume data according to the matching rules to filter and obtain target resumes.
[0148] For example, in this embodiment, vague requirements are transformed into explicit matching rules according to job requirements. For instance, priority is given to matching those who have worked at xxx company. Based on the generated matching rules, standardized resumes are automatically matched in batches, and all resumes that meet the rules are filtered out to obtain the target resume.
[0149] In summary, this embodiment first extracts key target text information such as educational background, previous employers, work skills, and qualifications from resume data and performs word segmentation. Then, a classification model is constructed using the RoBERTa model combined with fully connected layers. Simultaneously, the model parameters are optimized using the cross-entropy loss function, gradient descent algorithm, and multiple rounds of iterative training to ensure accurate mapping of word segmentation to labels representing entity attributes and locations, providing reliable support for subsequent information entity extraction. Next, target word segments with entity labels are selected and concatenated based on their original positions in the text to be processed, effectively eliminating interference from non-entity word segmentation and ensuring the integrity of the information entities to be processed. Finally, algorithms such as edit distance are used to calculate the relationship between the entity to be processed and a preset... The standard entity similarity algorithm also incorporates a user confirmation mechanism to supplement custom standard entities, which avoids invalid cross-type matching, improves the preset database, and enhances matching accuracy. After determining the target standard entity, the resume database is traversed to identify resume data containing corresponding non-standard text. Non-standard expressions are replaced in batches based on the standard entity, and a log containing information such as resume ID, text before and after replacement, and processing time is generated. This significantly improves the standardization of resume information in the entire database and makes the processing process traceable. Finally, based on the standardized resume data, matching rules are generated according to job requirements, and target resumes are filtered in batches. This effectively solves the pain points of inefficiency and subjectivity in manual screening in traditional recruitment and significantly optimizes the quality of resume processing.
[0150] Figure 4 This is a schematic diagram of a resume data processing device provided in an embodiment of this application. Figure 4 As shown, the resume data processing device includes: a preprocessing unit 401, a prediction unit 402, a determination unit 403, a first processing unit 404, a second processing unit 405, and a standardization processing unit 406.
[0151] The preprocessing unit 401 is used to obtain the target text information in the resume data and perform word segmentation on the target text information to obtain the text to be processed.
[0152] The prediction unit 402 is used to input the text to be processed into a pre-trained classification model to obtain the predicted labels of each word in the text to be processed output by the classification model. The predicted labels are used to characterize whether the word belongs to an information entity and the position of the word in the information entity. The predicted labels include the information entity start label, the information entity internal label, and the non-information entity label.
[0153] The determination unit 403 is used to take the word segment whose predicted label is the starting label of the information entity or the internal label of the entity as the target word segment, and determine the original position of each target word segment in the text to be processed.
[0154] The first processing unit 404 is used to take the target word corresponding to the starting tag of the information entity as the starting point of the splicing, and splice the target words corresponding to the tags inside the entity in sequence according to the original position order in the text to be processed, so as to obtain the information entity to be processed.
[0155] The second processing unit 405 is used to determine the target standard information entity corresponding to the information entity to be processed based on the similarity between the information entity to be processed and the preset standard information entity.
[0156] The standardization processing unit 406 is used to standardize the target text information based on the target standard information entity to obtain standardized resume data.
[0157] In other embodiments, the second processing unit 405 described above is further configured to:
[0158] Calculate the similarity between the information entity to be processed and each preset standard information entity in the preset database.
[0159] Based on the similarity between the information entity to be processed and each standard information entity, a set of candidate information entities corresponding to the information entity to be processed is determined; wherein, the set of candidate information entities includes a preset number of standard information entities, and each standard information entity is arranged in descending order of similarity with the information entity to be processed.
[0160] The candidate information entity set is sent to the user-side device, and the user's instruction information is obtained based on the user-side device; the instruction information is used to indicate the target standard information entity corresponding to the information entity to be processed.
[0161] The standard information entity in the candidate information entity set indicated by the instruction information is determined as the target standard information entity, or the custom information entity indicated by the instruction information is determined as the target standard information entity.
[0162] In other embodiments, the second processing unit 405 is further configured to calculate the edit distance between the information entity to be processed and each standard information entity, and use the obtained edit distance as the corresponding similarity; wherein the magnitude of the edit distance is inversely proportional to the magnitude of the similarity.
[0163] In other embodiments, after the second processing unit 405, the above-mentioned resume data processing device may further include a first execution unit, which is used to: traverse the resume data in the resume database, identify non-standard resume data in the resume database that includes target text information; perform batch standardization processing on the target text information in the non-standard resume data based on the target standard information entity, and generate corresponding log information.
[0164] In other embodiments, the prediction unit 402 is further configured to: input the text to be processed into a pre-trained classification model, extract the semantic feature vector of each word in the text to be processed using a preprocessing module, and map the semantic feature vector of each word based on a fully connected module to obtain the predicted label corresponding to each word in the text to be processed.
[0165] In other embodiments, the above-described resume data processing apparatus further includes a training unit for training a classification model, including:
[0166] The text to be processed from multiple resume data is input into the initial classification model to obtain the predicted label corresponding to each word in each text to be processed.
[0167] The cross-entropy loss between the predicted and true labels corresponding to each word segment in each text to be processed is summed to obtain the loss function value.
[0168] If the loss function value is greater than the preset threshold, the model parameters of the initial classification model are adjusted along the gradient descent direction of the loss function value according to the preset learning rate, so as to obtain the classification model after parameter adjustment.
[0169] Using the parameter-adjusted classification model as the new initial classification model, the steps of "inputting the text to be processed from multiple resume data into the initial classification model and obtaining the predicted label corresponding to each word in each text to be processed" are repeated until the loss function value is less than the preset threshold, and the trained classification model is obtained.
[0170] In other embodiments, the above-mentioned resume data processing device further includes a second execution unit, configured to: generate matching rules for standardized target text information based on job requirements; wherein the target text information includes one or more of educational background information, previous employer information, work skills information, and qualification information; and perform batch matching of the target text information in multiple standardized resume data according to the matching rules to obtain target resumes.
[0171] The resume data processing device provided in this application embodiment can execute the method shown in the above method embodiment. Its implementation principle and beneficial effects can be referred to the relevant description in the method embodiment, and will not be repeated here.
[0172] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device includes a memory 501 and at least one processor 502.
[0173] The memory 501 is used to store computer program code, which includes computer instructions. These computer instructions run in the aforementioned electronic device to implement the method shown in the above-described method embodiments. For example, the memory may include high-speed random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk storage device, or a USB flash drive, portable hard drive, read-only memory, magnetic disk, or optical disk, etc.
[0174] Processor 502 can be a general-purpose processor, including a Central Processing Unit (CPU), a network processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Processor 502 can also be other general-purpose processors. The general-purpose processor can be a microprocessor or any conventional processor.
[0175] The memory 501 and processor 502 are connected for communication. For example, the memory 501 can be connected to the processor 502 via a system bus to complete communication between them. The system bus can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, an industry standard architecture (ISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the figure, but this does not mean that there is only one bus or one type of bus.
[0176] Optionally, the memory 501 can be either standalone or integrated with the processor 502. When the memory 501 is set up independently, it is connected to the processor 502 via a system bus.
[0177] This application also provides a chip for executing instructions, which is used to execute the technical solution of the data processing method described in the above embodiments.
[0178] This application also provides a computer-readable storage medium storing computer instructions. When these computer instructions are executed by a processor, they are used to implement the technical solution of the resume data processing method described in the above embodiments. Specifically, when the computer instructions are executed by a processor, the electronic device can perform the technical solution of the resume data processing method described in the above embodiments.
[0179] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, it can implement the technical solution of the data processing method described in the above embodiments.
[0180] The aforementioned computer-readable storage media can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The computer-readable storage media can be any available medium accessible to a general-purpose or special-purpose computer.
[0181] An exemplary computer-readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the computer-readable storage medium can also be a component of the processor. The processor and the computer-readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the computer-readable storage medium can exist as discrete components in an electronic control unit or main control device; this application does not limit this.
[0182] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0183] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.
[0184] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit composed of the above modules can be implemented in hardware or in the form of hardware plus software functional units.
[0185] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0186] It should be understood that the steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.
[0187] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for processing resume data, characterized in that, include: Obtain the target text information from the resume data, and perform word segmentation on the target text information to obtain the text to be processed; The text to be processed is input into a pre-trained classification model to obtain the predicted labels of each word segment in the text to be processed output by the classification model; wherein, the predicted labels are used to characterize whether the word segment belongs to an information entity and the position of the word segment in the information entity, and the predicted labels include information entity start label, information entity internal label and non-information entity label; The word segment whose predicted label is the starting label of the information entity or the internal label of the entity is used as the target word segment, and the original position of each target word segment in the text to be processed is determined. Starting with the target word corresponding to the initial tag of the information entity as the starting point of the concatenation, the target words corresponding to the tags inside the entity are concatenated sequentially according to the original position order in the text to be processed to obtain the information entity to be processed. Based on the similarity between the information entity to be processed and the preset standard information entity, the target standard information entity corresponding to the information entity to be processed is determined; Based on the target standard information entity, the target text information is standardized to obtain standardized resume data.
2. The resume data processing method according to claim 1, characterized in that, The step of determining the target standard information entity corresponding to the information entity to be processed based on the similarity between the information entity to be processed and the preset standard information entity includes: Calculate the similarity between the information entity to be processed and each preset standard information entity in the preset database; Based on the similarity between the information entity to be processed and each standard information entity, a set of candidate information entities corresponding to the information entity to be processed is determined; wherein, the set of candidate information entities includes a preset number of standard information entities, and each standard information entity is arranged in descending order of similarity with the information entity to be processed; The candidate information entity set is sent to the user-side device, and based on the user-side device, the user's instruction information is obtained; the instruction information is used to indicate the target standard information entity corresponding to the information entity to be processed. The standard information entity in the candidate information entity set indicated by the instruction information is determined as the target standard information entity, or the custom information entity indicated by the instruction information is determined as the target standard information entity.
3. The resume data processing method according to claim 2, characterized in that, The calculation of the similarity between the information entity to be processed and each standard information entity in the preset database includes: Calculate the edit distance between the information entity to be processed and each standard information entity, and use the obtained edit distance as the corresponding similarity; wherein, the magnitude of the edit distance is inversely proportional to the magnitude of the similarity.
4. The resume data processing method according to claim 1, characterized in that, After determining the target standard information entity corresponding to the information entity to be processed, the method further includes: Traverse the resume data in the resume database and identify non-standard resume data in the resume database that includes the target text information; Based on the target standard information entity, the target text information in the non-standard resume data is batch standardized and corresponding log information is generated.
5. The resume data processing method according to claim 1, characterized in that, The classification model includes a preprocessing module and a fully connected module; The step of inputting the text to be processed into a pre-trained classification model to obtain the predicted labels of each word segment in the text to be processed output by the classification model includes: The text to be processed is input into a pre-trained classification model, and the semantic feature vector of each word in the text to be processed is extracted using the preprocessing module. Based on the fully connected module, the semantic feature vectors of each word segment are mapped to obtain the predicted label corresponding to each word in the text to be processed.
6. The resume data processing method according to claim 5, characterized in that, The classification model was trained through the following process: The text to be processed from multiple resume data is input into the initial classification model to obtain the predicted label corresponding to each word in each text to be processed; The cross-entropy loss between the predicted and true labels corresponding to each word segment in each text to be processed is summed to obtain the loss function value. If the loss function value is greater than a preset threshold, the model parameters of the initial classification model are adjusted along the gradient descent direction of the loss function value according to the preset learning rate to obtain the parameter-adjusted classification model. Using the parameter-adjusted classification model as a new initial classification model, the process of inputting the text to be processed from multiple resume data into the initial classification model to obtain the predicted label corresponding to each word in each text to be processed is repeated until the loss function value is less than the preset threshold, thus obtaining the trained classification model.
7. The resume data processing method according to any one of claims 1 to 6, characterized in that, After obtaining the standardized resume data, the method further includes: Based on job requirements, matching rules are generated for standardized target text information; wherein, the target text information includes one or more of the following: educational background information, previous employer information, work skills information, and qualification information; According to the matching rules, the target text information in multiple standardized resume data is matched in batches to obtain the target resume.
8. A resume data processing device, characterized in that, The device includes: The preprocessing unit is used to acquire target text information from resume data and perform word segmentation on the target text information to obtain the text to be processed. The prediction unit is used to input the text to be processed into a pre-trained classification model to obtain the predicted labels of each word segment in the text to be processed output by the classification model; wherein, the predicted labels are used to characterize whether the word segment belongs to an information entity and the position of the word segment in the information entity, and the predicted labels include information entity start label, information entity internal label and non-information entity label. The determining unit is used to take the word segment whose predicted label is the starting label of the information entity or the internal label of the entity as the target word segment, and determine the original position of each target word segment in the text to be processed; The first processing unit is used to take the target word corresponding to the starting tag of the information entity as the starting point of the splicing, and splice the target words corresponding to the tags inside the entity in sequence according to the original position order in the text to be processed, so as to obtain the information entity to be processed. The second processing unit is used to determine the target standard information entity corresponding to the information entity to be processed based on the similarity between the information entity to be processed and the preset standard information entity. The standardization processing unit is used to standardize the target text information based on the target standard information entity to obtain standardized resume data.
9. An electronic device, characterized in that, include: The electronic device includes a memory and at least one processor; the memory is communicatively connected to the processor; the memory is used to store computer program code, the computer program code including computer instructions; when the processor executes the computer instructions, the electronic device performs the resume data processing method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, which, when executed by a processor, are used to implement the resume data processing method as described in any one of claims 1-7.
11. A computer program product, characterized in that, When the computer program product is run on a computer / executed by the computer's processor, it implements the resume data processing method as described in any one of claims 1-7.