Intelligent talent information analysis method based on deep learning
Through the intelligent talent information analysis method based on deep learning, the resume information is automatically processed, and the problems of low efficiency and low accuracy of traditional manual screening are solved, and efficient and accurate resume information extraction and classification are achieved.
Patent Information
- Application Number
- CN202510097681.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
AI Technical Summary
The traditional resume screening process relies on manual review, which is time-consuming, labor-intensive and subjective judgments, resulting in low efficiency and unstable results.
The intelligent talent information analysis method based on deep learning is adopted, including converting resume files into image formats, correcting processing, image text analysis, naming entity recognition and structured data cleaning, to realize automated document classification and key information extraction.
It significantly improves resume processing efficiency, reduces the probability of errors, reduces labor and time costs, and provides efficient, accurate and reliable information processing solutions for talent recruitment.
Smart Images

Figure CN120069823A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis in the field of human resources, and particularly to an intelligent talent information parsing method based on deep learning. Background Art
[0002] The traditional resume screening process mainly relies on manual review. Specifically, recruiters read resumes one by one according to the job requirements and select the candidates who meet the conditions. This method obviously has several main problems. First, due to the large number of resumes to be processed, manual review is an extremely time-consuming and energy-consuming process. Second, during the manual review process, the subjective judgment of the recruiter will directly affect the final result. In this regard, it also has a large degree of randomness, and the accuracy and stability of the result are often difficult to guarantee. Finally, due to the need to process a large number of resumes, the efficiency of manual review is extremely low, and it is difficult to complete the screening in a short time.
[0003] In today's industrial society, the field of artificial intelligence has developed vigorously. Especially in the process of talent recruitment, the application of artificial intelligence technology is particularly important. In the above process, how to effectively process and analyze a large amount of resume information has always been an important challenge faced by researchers and enterprises.
[0004] To achieve "intelligent information parsing" for the human resources department, the first problem to be solved is the text parsing of resumes and candidate information.
[0005] Text parsing technologies can be roughly divided into the following two categories: 1. Template matching method, which uses predefined character templates to match the text in the image; 2. Optical Character Recognition (OCR): using a neural network to extract text features and perform recognition.
[0006] In the article "Template matching technique for searching words in document images", the template matching technique is studied for searching and locating a template image (a part of an image) in a large image. First, a template image is selected and prepared as the object of target recognition, then the template image is compared with the input image, and a performance metric method is used to evaluate the matching result to identify single characters or multiple characters (such as words, sentences or paragraphs). Finally, the accuracy and efficiency of template matching are evaluated by analyzing the comparison results, and the method is adjusted and optimized as needed. However, a major challenge that the disclosed template matching technique encounters in the field of intelligent resume parsing is that the technique requires a large number of pre-set templates to match different text styles and variations. This not only limits the adaptability of the technique in dealing with complex resume scenarios and large-scale data sets, but also shows deficiencies in the context recognition and semantic analysis of resume information, which directly affects the accuracy and efficiency of resume information recognition and information extraction.
[0007] The article "Focusing Attention: Towards Accurate Text Recognition in Natural Images" proposes an innovative method called FAN (Focusing Attention Network). The method aims to improve the attention shift problem faced by traditional attention mechanism-based text recognition models when dealing with scenes with mixed elements or poor image quality. By combining the Attention Network (AN) and the Focusing Network (FN), FAN implements an automatic attention adjustment mechanism in the model to ensure the precise alignment between the character target and the image feature region during the recognition process, thus improving the recognition accuracy. In this literature, although aiming to improve the recognition accuracy, the introduced additional complexity and computational cost result in the consumption of a large amount of computing resources in practical applications. For the intelligent resume parsing task that requires rapid response, this method may affect the overall parsing efficiency due to slow execution speed, thus reducing its practical value in a high-speed environment.
[0008] Obtaining the required structured, standardized and valuable information from a large amount of chaotic text information is a core issue that needs to be solved urgently. There are four main types of research on text information extraction: 1. Rule-based methods, which determine the pattern and structure of information through preset logical rules; 2. Statistical methods, which rely on the statistical properties of data to infer the key elements of the text; 3. Machine learning-based methods, which automatically identify and extract information by learning from data; 4. Deep learning-based methods, which use complex neural network structures to understand text and extract information. Rule-based recognition methods rely on carefully designed rule sets that need to cover many possible resume formats and structures. However, this method is difficult to exhaust all changes, especially when encountering non-standard or innovative resume formats, which may lead to insufficient recognition accuracy.
[0009] In the article "Statistical-based Chinese Text Keyword Extraction Method for Industries", the study focused on the extraction technology of Chinese text keywords in vertical search engines. The researchers expanded and innovated the traditional statistical methods and took into account the position information and word span information of word occurrence, which is a beneficial supplement to the TF-IDF method. This method not only considers the conventional word frequency (TF) and inverse document frequency (IDF), but also innovatively introduces position information and word span information to calculate the criticality of words. This comprehensive indicator fully reflects the importance of a word, thereby greatly improving the accuracy and effectiveness of keyword extraction. The keyword extraction mechanism relies too much on word frequency statistics, and may mistakenly regard some high-frequency stop words that do not contribute significantly to the resume content as key information. This not only reduces the accuracy of information extraction, but also increases the burden on the system to process irrelevant information.
[0010] The article "Text Information Extraction Method Based on Clustered Hidden Markov Model" proposes an innovative method using hidden Markov model (HMM) to effectively extract text information by introducing clustering (Clustered-HMM, C-HMM for short). This method first divides a large amount of text into several subgroups through cluster analysis as a preparatory stage for information extraction; then, the Markov chain model is trained using the specific text data obtained from each cluster, and all the training data is used to optimize the unified probability output matrix. This method can effectively improve the performance of the model in processing text information with different structures, thereby greatly improving the performance of information extraction. However, when parsing structured resume information (such as personal information), the efficiency and accuracy of this document are highly dependent on good quality training data and appropriate clustering algorithms. Improper data processing and algorithm selection may weaken the accuracy and effectiveness of the resume information extraction process. Summary of the invention
[0011] The technical problem to be solved by the embodiments of the present invention is to provide an intelligent talent information parsing method based on deep learning, which can overcome the deficiencies existing in the prior art in talent resume parsing.
[0012] To solve the above technical problem, the embodiments of the present invention provide an intelligent talent information parsing method based on deep learning, including the following steps: S110: Convert the resume file into a picture format and perform correction processing on the picture; S120: Parse the image text; S130: Extract information from the obtained text data through a named entity recognition method to obtain semi-structured information; S140: Clean the structured data of the information to obtain standardized data.
[0013] Further, the S3 specifically includes the steps: S131: Determine entity tags according to the recognition requirements; S132: Establish a training set and label the corresponding tags in the entity part; S133: Segment the labeled training set text to obtain a word sequence; S134: Input the word sequence into the BERT-CRF model for training to obtain an entity recognition model; S135: Input the text to be recognized into the entity model to obtain a recognition result, and combine regular expressions to extract the part with a longer number of characters.
[0014] Further, the BERT-CRF model uses BERT as a feature extractor to obtain the context-related representation of each word in the input sentence, and the representation will be used as a feature to input into the CRF model to learn the label transition probability in the sequence labeling task, and obtain a predicted annotation sequence.
[0015] Further, absolute position encoding is added to each word in the input sequence of the BERT-CRF model, and the process includes the following steps: Calculate an Attention matrix once with the standard RoPE:
[0016] where is the rotation matrix of RoPE, and are two input vectors, representing the query and key vectors respectively; Let be the position of the word in the sentence, be its corresponding position encoding vector, is the dimension of the positional encoding, , then the function is the positional encoding function, and its definition is as follows:
[0017] Among them, , for each word, three vectors are calculated, and these vectors are obtained by multiplying the sum of the word embedding and the positional encoding by three different matrices; In the above,
[0018] Among them, are the parameters of the model; Calculate the Attention matrix of RoPE with a calculation interval of :
[0019] Among them, is a constant that adjusts the interval of the relative positional encoding. Finally, according to the condition , the two matrices are combined to obtain:
[0020] Through the above formula, a non-linear relative position is obtained. Next, calculate the attention score of each word to other words, and use to represent the attention score of the th word to the th word;
[0021] Among them is an activation function with a non-negative value range, is the length of the input sequence.
[0022] Furthermore, the S130 further includes the steps of: Pass the output of the self-attention mechanism through a feed-forward neural network; The feed-forward neural network includes two fully connected layers and a ReLu activation function
[0023] Among them, are the parameters of the model, is the activation function.
[0024] Send the vector output in BERT into the CRF model, and the objective function of the CRF model can be defined as:
[0025] Among them, is the feature function weight, and are the current and previous labels respectively, is the input sequence, is the weighted sum of all feature functions; it is measured by the ratio of the score of the predicted optimal path to the scores of all paths. The larger the ratio, the more reasonable it is. The loss function is , that is, negative log-likelihood, The larger it is, the smaller the loss. Therefore, the model is optimized by optimizing the loss function for training.
[0026] Furthermore, the S130 also includes a method using label confusion learning to learn the relationships between labels, including the steps of: First, use the BERT model to learn the context information of the text to be recognized and generate word vector representations, and then feed them into the CRF to obtain the predicted label distribution:
[0027] Where is BERT, for the input sequence , with length , the output sequence representation with length , and dimension , is the predicted label distribution Finally, add the original BIOES distribution with the control parameter α to the label confusion distribution, and then normalize it through the softmax function to generate the simulated label distribution. The control parameter α determines how much the label confusion distribution will change the BIOES distribution. The above process is expressed as:
[0028] Where, is the label encoder function, which is an LSTM network for transmitting labels represents the label representation matrix , C is the number of entity categories, is the label confusion distribution, is the simulated label distribution.
[0029] Furthermore, the S135 also includes the steps of: using the method of dynamically and precisely dividing sliding windows to correct incorrect paragraph segmentation according to the calculated similarity between paragraphs, including the steps of: Divide the entire text context into sentences: Among them, Represent the th sentence, calculate the embedding of each sentence , and split the text context into blocks of size , obtaining:
[0030] Next, perform a block max-pooling operation on each block to extract the main semantic information of the entire block, obtaining , and then calculate the cosine similarity between adjacent blocks and to represent the semantic similarity degree between them:
[0031] In the adjacent cosine similarity sequence, identify the local minimum points :
[0032] These points represent the possible segmentation positions in the text. Finally, compare the identified segmentation points with the original segmentation boundaries . If a certain segmentation point is not in , it is regarded as an incorrect segmentation, and the original segmentation method is modified, and the above operations are repeated until the most perfect segmentation method is achieved.
[0033] Furthermore, the S140 specifically includes the steps of: S41: Missing value processing: Delete missing values or fill in missing values; S42: Outlier processing: Delete outliers or transform outliers; S43: Data format conversion: Unify different date formats into a standard format; S44: Data deduplication: Compare all data and delete completely duplicate data; S45: Data verification and consistency check: Compare the consistency between different data to ensure the accuracy and integrity of the data.
[0034] Implementing the embodiments of the present invention has the following beneficial effects: Traditional resume parsing and information extraction often require manual intervention and multi-step processing, while the present invention greatly improves the processing efficiency and reduces the error probability through automated document classification and key information extraction, thereby significantly reducing the labor cost and time cost, bringing an efficient, accurate and reliable information processing solution to the talent recruitment process. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is the overall process schematic diagram of the present invention; Figure 2 It is the main process schematic diagram of the present invention; Figure 3 It is the process schematic diagram of the specific embodiment of academic qualification information parsing; Figure 4 It is the process schematic diagram of the specific embodiment of ID card information parsing; Figure 5 It is the process schematic diagram of the specific embodiment of bank card information parsing; Figure 6 It is the process schematic diagram of the specific embodiment of image text parsing; Figure 7 It is the schematic diagram of the optimized residual structure and attention structure; Figure 8 It is the process schematic diagram of the data filtering training scheme; Figure 9 It is the process schematic diagram of the specific embodiment of semi-structured information extraction; Figure 10 It is the process schematic diagram of the specific embodiment of card image correction; Figure 11 It is the process schematic diagram of the specific embodiment of bank card information extraction. Specific Embodiment
[0035] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0036] Refer to Figure 1 , Figure 2 the schematic diagrams shown.
[0037] An intelligent talent information parsing method based on deep learning according to an embodiment of the present invention is implemented through the following steps.
[0038] S110: Convert the resume file into a picture format and perform correction processing on the picture; S120: Parse the image text; S130: Extract information from the obtained text data through a named entity recognition method to obtain semi-structured information; S140: Clean the structured data of the information to obtain standardized data.
[0039] In S110, the resume file data saved in the talent pool is converted into PNG image format data through the pdf2image library in Python.
[0040] In S120, OCR technology is mainly used to recognize text in the acquired image format data. Among them, the DB model is used for text detection, and the SVTR model is used for text recognition. The algorithm steps are as Figure 5 shown: The specific implementation example of image text analysis should include the following steps: S1201. Image preprocessing: Process the input image to better extract and recognize text, including the steps: S12011 Grayscale conversion: Convert the color image to a grayscale image to simplify subsequent processing steps. S12012 Denoising: Remove noise and unnecessary details in the image to reduce interference with subsequent processing steps. S12013 Edge enhancement: Highlight the text edges in the image to make the text area clearer and more prominent.
[0041] Adding the operation of image preprocessing before S1202 helps to improve the robustness of the model, remove noise and highlight text features, while accelerating the convergence process of the model and optimizing its performance. Therefore, the image optimized through the above preprocessing can improve the accuracy and efficiency of the text detection model and make it more suitable for complex scene text recognition tasks.
[0042] For the above steps, in S12011, the weighted average method is used to perform weighted summation on the RGB channels of the color image to convert the color image into a single-channel grayscale image to extract the brightness information in the image, so as to facilitate subsequent character recognition and analysis. The calculation method of its grayscale value is as follows:
[0043] This formula represents the process of converting the RGB values of each pixel of the color image into grayscale values where respectively represent the values of the red, green, and blue channels of the pixel. These weighting coefficients 0.299, 0.587, 0.114 are determined according to the perception weights of the human eye for different colors, and they represent the relative importance of the three colors of red, green, and blue for the brightness perception of the human eye. This weighted method can better reflect the perception characteristics of the human eye for image brightness, so the obtained grayscale image will be more in line with the perception of human vision.
[0044] In S12012, the median filtering method is used. Its basic idea is to use the median of the grayscale values in the neighborhood of the pixel point to replace the grayscale value of the pixel point, so as to achieve the purpose of removing noise. The median filtering algorithm steps are as follows: a) For each pixel point of the image, take the pixel values in its neighborhood to form a set.
[0045] b) Sort this set and select the middle value as the new value of this pixel point.
[0046] c) Repeat the above steps until all pixel points of the entire image are processed.
[0047] Let the original image be , and apply a median filter of size at . Then the output image of the median filter can be expressed as:
[0048] where represents the operation of taking the median. This formula means that at each pixel point , the pixel values within the neighborhood centered on it are taken, then the median of these pixel values is calculated, and the median is used as the new value of this pixel point.
[0049] In S12013, the Sobel operator is used to highlight the edge features of characters, separate the characters from the background, and thus extract the important features in the image.
[0050]
[0051] where represents the gray value of the pixel at position in the image, and are the template coefficients of the Sobel operator, represents the horizontal gradient intensity of the image at the current position , which is used to calculate the horizontal gradient of the current position and its neighborhood; represents the vertical gradient intensity of the image at the current position ; and represents the comprehensive intensity of the gradient at the current position, which can reflect the edge intensity of the image and is used for edge detection.
[0052] S1202, Text detection, uses the DB model to locate the text regions in the image.
[0053] DB is a text detection algorithm based on segmentation. Its proposed DifferentiableBinarization module (DB module) uses a dynamic threshold to distinguish text regions from the background. The input image extracts features through the network Backbone and FPN, and the extracted features are cascaded together to obtain the original Figure 4Features of one - tenth size are then obtained, and the convolution layer is used to obtain the text region prediction probability map and the threshold map respectively. Then, the text bounding curve is obtained through the post - processing of DB. Its model can be divided into three parts: 1) Bankbone network: responsible for extracting the features of the image; 2) FPN structure: the feature pyramid structure enhances the features; 3) Head network, calculating the text region probability map.
[0054] The Bankbone network is the basic part of the DB model and is responsible for extracting the features of the image. It usually uses pre - trained deep - learning models, such as ResNet or VGG, to extract low - level and high - level features in the image, and these features are used as the input for subsequent steps. Convolution operations are usually used to extract features, and its formula is:
[0055] Among them, and respectively represent the ordinate (row number) and abscissa (column number) of the current image pixel, and they refer to the position being processed in the convolution result. and are the coordinate offsets of the convolution kernel (Kernel), used to represent the positions of the elements in the kernel. is the input image, is the convolution kernel, represents the convolution operation.
[0056] FPN is a structure used to enhance features. It extracts features at different scales and then fuses these features together to obtain richer and more comprehensive image features. The fusion of feature maps is usually carried out by means of upsampling and element - wise addition, and its formula is as follows:
[0057] Among them, is the feature map, and are the interpolation positions.
[0058] As Figure 7 shown, in this embodiment, a residual structure and a channel attention structure are introduced into the FPN structure to optimize the FPN. The modified FPN structure is shown in the following figure. The residual structure and the channel attention structure are used to replace the conventional convolution layer in the FPN to enhance the representation ability of the feature map. Considering that the number of channels in the FPN in the DB model is small, only 64, directly using the squeeze - and - excitation module to replace the convolution will suppress the features of some channels, thus reducing the accuracy. Therefore, the residual attention convolution introduces a residual structure to alleviate this problem and improves the effect of text detection.
[0059] In the squeezing stage, the global average pooling layer is applied to the input feature map , obtaining a -dimensional vector. This process can be expressed by the following formula:
[0060] where is the value of the input feature map at the th channel and position, and are the height and width of the feature map respectively.
[0061] In the excitation stage, first, through a fully connected (FC) layer and a ReLU activation function, the -dimensional vector obtained in the squeezing stage is mapped to a new -dimensional vector , where is a hyperparameter, usually set to 16. Then, through another FC layer and a Hard-Sigmoid activation function, is mapped back to the -dimensional vector . This process is expressed as:
[0062] where and are the weights of the two fully connected layers, is the sigmoid activation function, and ReLU is the rectified linear unit function.
[0063] Finally, the resulting -dimensional vector is added to the result of the channel-wise multiplication of the input feature map and the original input feature map to form a residual connection, obtaining the output feature .
[0064]
[0065] where is the value of the vector at the th channel, is the value of the input feature map at the th channel and position. Thus, the output feature map is the input feature map and the excitation vector The sum of the products with the input feature map can help the network better learn features.
[0066] The Head network is responsible for calculating the text region probability map. In this embodiment, the initial Head network structure is modified, and the calculation process of the modified Head structure is described as follows.
[0067] First, the output feature map of the FPN is processed through a convolutional layer to obtain the output .
[0068]
[0069] Then the data is split into two different paths, one is the transposed convolution , and the other is the convolution after upsampling , and the operation is as follows. convolution .
[0070]
[0071] Next, the outputs and of these two paths are concatenated to obtain , and then processed through a convolutional layer to obtain the output .
[0072]
[0073] Finally, the output of this convolution is added to the output of the transposed convolution to obtain the final probability map . This process can be represented by the following formula:
[0074] This probability map represents the probability that each pixel point in the image is a text region. After obtaining the probability map, the differentiable binarization method will be used for processing to obtain the text detection result, and the formula is as follows.
[0075]
[0076] Differentiable binarization is essentially a sigmoid function with a coefficient , and the value range is (0, 1). is the dilation factor, refers to the probability image pixel point, refers to the threshold image pixel point.
[0077] In this embodiment, a distillation strategy is introduced during model training. The core idea of the distillation strategy is to transfer the knowledge of a complex model (teacher model) to a simplified model (student model), thereby improving the performance and generalization ability of the student model while reducing the computational and storage costs.
[0078] During the training process of the DB model, the distillation strategy can be summarized as providing a large DB model (teacher model) , and a small DB model (student model) . The training dataset of the present invention is , where is the input data, is the original label. In the distillation strategy, first use the classroom model to make predictions on the training data to generate soft labels . Then, the present invention trains the student model . During the training process, the student model not only learns from the original label , but also can learn from the predictions (soft labels) of the classroom model . Compared with traditional one-hot labels, soft labels are a probability distribution, which is more rich and continuous. By learning these soft labels, the student model can better capture the data distribution and the uncertainty of the model, and improve the generalization ability.
[0079] This process can be represented as the following optimization problem:
[0080] where, is the loss function, is a weight parameter used to balance the importance of the original label and the soft label. In this way, the student model can achieve performance similar to that of the teacher model, while having higher running efficiency and lower resource requirements, and is particularly suitable for running on resource-constrained devices.
[0081] S1203, Character recognition, use the SVTR model to extract the text area and perform character recognition.
[0082] SVTR is a character recognition algorithm based on a single vision model. By introducing local and global hybrid blocks, stroke features and inter-character correlations are extracted respectively, and combined with a multi-scale backbone to form a multi-granularity feature description. The specific structure of SVTR is as follows: Progressive Overlapping Patch Embedding: Embed the input image into character components; MixBlock: Used to extract global and local features; Hierarchical stage: Obtain the character sequence through parallel linear prediction.
[0083] Progressive Overlapping Patch Embedding: This step embeds the input image into character components. The SVTR model uses two cascaded convolutional modules to implement Progressive Overlapping Patch Embedding, embedding the input image into character components, which are used to represent the stroke features of characters.
[0084] MixBlock is used to extract global and local features. The SVTR model designs two MixBlocks with different receptive fields to complete the perception and extraction of these two granularity features. Among them, Global Mixing perceives the dependencies between all character components and extracts global features. Local Mixing perceives the correlations between components within a predefined receptive field (7×11) and extracts local features.
[0085] The SVTR model performs feature extraction through three stages. Each stage consists of a series of MixBlock and Merging or Combing operations. Feature extraction is performed at different scales to generate a representation called C, and through parallel linear prediction, the character sequence is obtained.
[0086] Since the SVTR model itself has good performance and is also a relatively new open-source model, it is difficult to achieve a significant performance improvement by optimizing parts such as the model network structure. Therefore, in this embodiment, starting from the perspective of model training, an excellent data filtering scheme is designed and implemented.
[0087] Data filtering is a simple and effective data mining scheme. The core idea is to use an existing model to predict training data and screen the full amount of training data through information such as confidence and prediction results. Specifically: Suppose there is a set of training datasets , which contains samples, and each sample is represented by , where is the sample feature vector and is the corresponding label. First, a low-precision model can be obtained through a quick training, which can be expressed as . Use this model to predict all training data to obtain the confidence of each sample, and then remove the samples with confidence greater than the threshold Samples in this part are considered redundant samples that are ineffective in improving the model accuracy, i.e.:
[0088] Next, use a high-precision model (such as the CRNN model) to predict the remaining samples, obtaining the confidence , and then remove the samples with a confidence less than the threshold . Samples in this part are considered samples that are difficult to recognize or of very poor quality, i.e.:
[0089] In this way, after two rounds of screening, a refined training dataset is obtained, which contains far fewer samples than the initial samples. Using this strategy, tens of millions of training data are refined to the million level, which can greatly reduce the model training time and improve the accuracy at the same time. The process schematic diagram is as shown in Figure 8 .
[0090] When introducing the distillation strategy during model training, its meaning and operation are the same as S12022.
[0091] S1204. Post-processing: The recognized text results may need to be post-processed, such as character-level correction, dictionary matching, or error correction by a language model, to improve the recognition accuracy and fluency.
[0092] For the DB model in S1202, in this embodiment, it is based on the ResNet50_vd backbone network and has been trained and adjusted a lot on two public text detection benchmark datasets, ICDAR2025 and TD_TR.
[0093] For the SVTR model in S1203, in this embodiment, it is trained on the English scene recognition public dataset IC15 and the Chinese dataset Chinese_scene_test respectively, so that the model can accurately recognize both Chinese and English at the same time.
[0094] After the above operations, an OCR joint model that can be applied to multiple scenarios and has a high accuracy is obtained for recognizing text in images.
[0095] In S130, mainly use named entity recognition technology to extract information from the text data obtained in S120, aiming to extract the semi-structured information required by the present invention from a large amount of text. The specific embodiments of semi-structured information extraction should include steps as shown in Figure 9 .
[0096] S1301. Determine entity tags according to recognition requirements. The entity tags used for resume information extraction can be roughly divided into the following parts according to the text content and the number of characters. Each entity category corresponds to a specific tag: For resume information extraction, the entity categories used are as follows: 1) Personal information: name, gender, mobile phone number, date of birth, email, ethnicity, education level, graduated school, graduation time, type of education, major, expected salary, work experience, job hunting status, city of employment, arrival time, job hunting intention (position); 2) Work experience: company name, position name, job responsibilities, start time, end time, project name, project description, project responsibilities, project start time, project end time; 3) Personal skills: professional skills, self-evaluation.
[0097] For academic information extraction, the entity categories used are as follows: name, gender, date of birth, enrollment date, graduation (completion) date, school name, major, type of academic degree, length of schooling, form of study, level, graduation (completion), name of school (college) president, certificate number.
[0098] For ID card information extraction, the entity categories used are as follows: name, ID card number, gender, ethnicity, date of birth, address, issuing authority, start date of validity, end date of validity.
[0099] In this embodiment, a highly adaptable and targeted model is trained for each part to achieve industrial recognition accuracy.
[0100] S1302. Obtain a training set, and the training set is labeled according to entity tags. Use the steps of S110 and S120 for the desensitized resume data. The recognized text of each resume data is stored in the corresponding txt file to complete the collection of the data set. Next, use the Label-Studio annotation tool to manually annotate each file, and label the entity parts in the text with the corresponding tags for the model to learn.
[0101] In this embodiment, the BIOES format is used to label entity tags, aiming to achieve more accurate and refined named entity recognition. The IOES format takes into account the start and end positions of entities, making the annotation more structured and information-rich.
[0102] BIOES - Five-digit sequence annotation method (B - begin, I - inside, O - outside, E - end, S - single): Add the end identifier of the E entity.
[0103] · B indicates the start · I indicates the inside ·O represents non-entity ·E represents the end of an entity ·S represents that the word itself is an entity Since information such as resumes, ID cards, and educational backgrounds involves personal privacy, the initial amount of data is not large. For each task in the present invention, about 1500 pieces are labeled for each task. In order to make the model training more effective, data augmentation operations need to be carried out.
[0104] Data Augmentation refers to a method of expanding the training dataset by performing a series of transformations or processes on the original data to generate new data samples. The purpose of data augmentation is to improve the generalization ability of the model, reduce overfitting, and improve the performance of the model on unseen data. The data augmentation method used in this embodiment is as follows: 1) Random word replacement: Randomly select certain words and replace them with other words. A word vector model can be used to select similar words for replacement.
[0105] 2) Sentence restructuring: Adjust the order or structure of words in a sentence.
[0106] 3) Synonym replacement: Use a thesaurus to replace some words in a sentence while keeping the sentence semantics unchanged.
[0107] 4) Adding noise: Introduce random noise into the text, such as randomly inserting, deleting, or changing characters.
[0108] S1303. Segment the labeled training set text to obtain a word sequence. The purpose of segmentation is to convert a continuous text sequence into a discrete word sequence so that the subsequent model can understand and process it. Here, the BERT model is used for segmentation. The steps are as follows: The labeled dataset, which contains multiple text samples , each sample is a string, denoted as , where is the th word or sub-word in the sample . First, each text sample is basically segmented according to certain rules into a sequence of words or sub-words. This can be expressed as: .
[0109] Next, each word or sub-word is further decomposed into word fragments, denoted as , where It is a word fragment. "playing" may be decomposed into "play" and "##ing", where "##" indicates that it is part of the previous fragment. During the word fragmentation process, some special tokens such as [CLS], [SEP], etc. are added to form the final model input sequence. Usually, the [CLS] token is added at the beginning of the entire input sequence, and the [SEP] token is added between different sentences. For the entire text sample After word fragmentation, we get .
[0110] The tokenized text sequence is the input of the model. Each word fragment will be mapped to the corresponding word embedding, and then input into the BERT model for processing and learning.
[0111] S1304. Input the word sequence into the BERT-CRF model for training to obtain an entity recognition model. The BERT-CRF model uses BERT as a feature extractor to obtain the context-related representations of each word in the input sentence. These representations will be used as features and input into the CRF model to learn the label transition probabilities in the sequence labeling task, obtaining a predicted annotation sequence, and then extracting and classifying each entity in the sequence to complete the task of Chinese entity recognition. Therefore, the entire model can be divided into two parts: 1) The BERT part: It is used to generate the context encoding of each word in the input sentence. For an input sentence, BERT will output the word vector representation of each word, which contains context information.
[0112] 2) The CRF part: It is used to learn the transition probabilities between labels. Based on the word vectors output by BERT, the CRF model will learn how to optimally assign labels to each word to maximize the overall sentence annotation probability.
[0113] For each word in the input sequence, there is a word embedding vector. This vector is obtained by looking up the word embedding matrix, where each row corresponds to the embedding of a word. If the input sequence is , then the input embedding of the present invention is , where is Word embeddings. To enable the model to consider the order of words, a positional encoding needs to be added to each word. The relative positional encoding is used in BERT, and in this embodiment, an absolute positional encoding is used. Generally speaking, the absolute positional encoding has advantages such as simple implementation and fast calculation speed, while the relative positional encoding directly reflects the relative position signals. The original relative positional encoding is used in BERT. If the relative positional encoding can be achieved through the absolute positional encoding method, good results can be obtained.
[0114] Rotary Position Embedding (RoPE) is a design that can achieve "relative positional encoding in the way of absolute positional encoding" in conjunction with the Attention mechanism. In this embodiment, some improvements are made based on RoPE. Specifically, first calculate the Attention matrix (before Softmax) using the standard RoPE.
[0115]
[0116] The first equal sign here is the implementation method, and the second equal sign is the equivalent result, where is the rotation matrix of RoPE, and are two input vectors, representing the query and key vectors respectively, obtained through the initial positional encoding calculation method of BERT. Let be the position of the word in the sentence, be its corresponding positional encoding vector, then the function is the positional encoding function, and its definition is as follows:
[0117] Among them, , is the dimension of the positional encoding, . For each word, three vectors are calculated. These vectors are obtained by multiplying the sum of the word embedding and the positional encoding by three different matrices.
[0118]
[0119] Among them, is a parameter of the model. For simplicity, the scale factor of Attention is ignored in the present invention. Then, it is necessary to calculate the Attention matrix of RoPE with an interval of :
[0120] Among them, is a constant that adjusts the interval of the relative positional encoding. Finally, according to This condition combines two matrices to obtain:
[0121] In this way, a non-linear relative position is obtained , and then calculate the attention score of each word to other words, using to represent the attention score of the -th word to the -th word.
[0122]
[0123] Among them is an activation function with a non-negative value range, is the length of the input sequence.
[0124] Finally, the output of the self-attention mechanism passes through a feed-forward neural network, which contains two fully connected layers and a ReLu activation function.
[0125]
[0126] Among them, are the parameters of the model, is the activation function.
[0127] In the Chinese word segmentation recognition task, BERT is good at processing long-distance text information, but it cannot handle the dependencies between adjacent labels. BERT can only use the context information of the sequence to give the most likely prediction, but it cannot use the predicted information to help with the prediction. Specifically, the labels for Chinese word segmentation are "BIOES" in total. If a word is predicted as I, then the previous word must be B and cannot be E or S. This is the constraint relationship between labels, and BERT cannot utilize this relationship. However, CRF can obtain an optimal prediction sequence through the relationship between adjacent labels, which can make up for the shortcomings of BERT. The CRF layer performs sentence-level sequence annotation.
[0128] Send the vector output from BERT into the CRF model. The objective function of the CRF model can be defined as:
[0129] Among them, is the -th feature function, which depends on the -th element of the input sequence , the current output label and the previous output label , is the number of feature functions, is the length of the sequence. is the feature function weight of, is the output sequence score function of, defined as the weighted sum of all feature functions, is the given input sequence when, the output sequence conditional probability of. In the above formula, the numerator is the score of the current label sequence, and the denominator is the sum of the scores of all possible label sequences, also known as the partition function, which is used to ensure that the sum of probabilities is 1. It is measured by the ratio of the score of the predicted optimal path to the scores of all paths. The larger the ratio, the more reasonable it is. The loss function is , that is, negative log-likelihood, the larger, the smaller the loss, so the model is optimized by optimizing the loss function for training.
[0130] Named entity recognition is usually regarded as a classification task. This task requires annotating the input text sequence, and the output of the annotation is a sequence, and each position has a label indicating which category of entity the word or phrase at that position belongs to. During training, the BIOES annotation method is usually used, but this method often ignores the impact of noisy data on model recognition. Using the BIOES annotation method directly for training is particularly prone to mislabeled samples. For this reason, this embodiment proposes and uses the method of label confusion learning to learn the relationship between labels and guide the model to be able to recognize noisy data.
[0131] This method first uses the BERT model to learn the context information of the text to be recognized and generate word vector representations, and then feeds them into the CRF to obtain the predicted label distribution:
[0132] where represents using BERT to map the input into the feature space; is the input vector, where represents the th element in the input sequence, with a length of . is the output vector of the function , where represents the th feature. is the conditional random field function that converts the input vector into the output vector . With a length of , and a dimension of , is the predicted label distribution.
[0133] In this example, a label encoder module and a simulated label distribution calculation block are additionally designed. The label encoder uses a deep neural network to generate a label representation matrix. The simulated label distribution calculation block consists of a similarity layer and a calculation layer. The similarity layer takes the label representation and the current instance representation as inputs, calculates their similarity values through dot product, and then uses a neural network activated by softmax to obtain the label confusion distribution. The label confusion distribution captures the dependencies between labels by calculating the similarity between instances and labels. Therefore, the label confusion distribution is a dynamic, instance-dependent distribution, which is better than the distribution that only considers the similarity between labels and also better than the simple uniform noise distribution in label smoothing.
[0134] Finally, the original BIOES distribution is added to the label confusion distribution with a control parameter α, and then normalized through the softmax function to generate the simulated label distribution. The control parameter α determines how much the label confusion distribution will change the BIOES distribution. The above process can be expressed as:
[0135] where, is a label encoder function, specifically implemented as an LSTM network, which receives the label sequence as input and encodes it into a representation matrix ; is the specific form of the representation matrix output by the label encoder. Each represents the encoding vector of the -th label, and is the number of entity categories. is the label confusion distribution, representing the prediction distribution of the model for labels. This distribution is generated through the function to ensure that the sum of the output probabilities is 1. is the transpose of the input vector , is a weight matrix for linear transformation, is a bias vector for adjusting the result of linear transformation, is the simulated label distribution, is a hyperparameter for controlling the probability distribution.
[0136] To enhance the recognition accuracy of the model and make it suitable for the current task, this embodiment performs partial optimization on the BERT-CRF model. The optimization aspects are as follows: S1305. Input the text to be recognized into the entity model to obtain the recognition result, and combine regular expressions to extract the part with a longer number of characters. Using regular expressions to extract specific information from long texts, such as the personal skills part, is an effective auxiliary method, especially when the recognition accuracy using NER technology is insufficient. Regular expressions can help capture text fragments with specific patterns, thereby extracting the required information.
[0137] For resume texts, their character lengths often far exceed 512. If directly using the model for recognition, it will exceed the input length limit of the BERT model. For this problem, most researchers use the sliding window strategy to solve it. The core of the sliding window strategy is: divide the long text into subsequences (windows) of a fixed length. For example, divide the entire text into several subsequences of length k, where k is usually the maximum input length limit of the BERT model. However, there is a fatal flaw in this problem. When dividing, it may destroy the semantic coherence between paragraphs, resulting in incorrect model recognition results. Therefore, this embodiment designs and implements a strategy for dynamically and accurately dividing the sliding window, and corrects the incorrect segmentation by calculating the similarity between paragraphs. The specific method is as follows: First, divide the entire text context into sentences: , where represents the -th sentence. Calculate the embedding of each sentence , and divide the text context into blocks of size , obtaining:
[0138] Next, perform a block maximum pooling operation on each block to extract the main semantic information of the entire block, obtaining . Then calculate the cosine similarity between adjacent blocks and to represent their semantic similarity degree:
[0139] where, and are the row and column indices of the output feature map, is the size of the pooling window, is the stride, is the input feature map at the position value. is the dot product of two vectors, and are the norms of these two vectors respectively Identify local minimum points in adjacent cosine similarity sequences :
[0140] These points represent possible segmentation positions in the text. Finally, the identified segmentation points are compared with the original segmentation boundaries If a certain segmentation point is not within , it is regarded as an incorrect segmentation, and the original segmentation method is modified. Repeat the above operations until the most perfect segmentation method is achieved. Through the above strategy of dynamically and precisely dividing the sliding window, the problem of the input length limit of BERT can be effectively solved, and the coherence between the divided paragraphs can be ensured without being split, thereby further improving the accuracy of the BERT-CRF model.
[0141] In S140, for structured data cleaning, since there is no fixed template for resumes and they are written entirely according to the preferences of job seekers, how to process the recognition results in S130 into standardized structured data is also a problem that needs to be solved. For this problem, in this embodiment, regular expressions are mainly used to clean the data to obtain standardized data. The processing contents include: Missing value processing: Delete missing values or fill in missing values. For example, if the expected salary recognition is missing, "Salary negotiable" can be filled in to ensure the integrity of the data; Outlier processing: Delete outliers or convert outliers. For example, if the age is misrecognized as "99" years old, etc., the outliers can be directly deleted; Data format conversion: Unify different date formats into a standard format. For example, if the enrollment time is recognized as "September 1, 2022", it is converted to "2022-09-01"; Data deduplication: Compare all the data and delete completely duplicate data. For example, if two "Bachelor's degree" are recognized for the education level, only one piece of data is retained; Data verification and consistency check: Compare the consistency between different data to ensure the accuracy and integrity of the data. For example, if the enrollment time is "2018", the graduation time is "2019", and the education level is "Bachelor's degree", it can be judged through the education level that the graduation time and the enrollment time are not consistent, and then it is modified.
[0142] The embodiment of the present invention further includes a step of correcting the card certificate image, such as Figure 5As shown in the figure, when photographing an ID card or a bank card, due to factors such as the angle and distance between the camera and the card, the card image may be distorted or tilted. Through image correction, the card image can be converted into a front view to improve the accuracy when extracting card information. The specific embodiments of card image correction include the following steps: S2101. Key point data annotation: Four corner points need to be marked for the card part in the image data, namely: the upper left corner point, the upper right corner point, the lower right corner point, and the lower left corner point.
[0143] S2102. Key point detection: Use the dataset marked in S2101 to train the Yolov8-Pose model, and use the trained model to detect key points in the image. Yolov8-Pose is an object detection algorithm for human pose estimation. It is an improved version based on the Yolo series and requires the annotation of a main body detection and 17 key points. In this embodiment, based on the application of this model to card key point detection, only 4 corner points of the card area need to be marked. The principle of the model is as follows: Feature extraction: Yolov8-Pose uses Darknet53 as the backbone network to extract the features of the input image through multiple convolutional layers and residual blocks; Feature fusion: By introducing a feature fusion module, feature maps of different scales are fused to obtain a more global and detailed feature representation.
[0144] Object detection: The model uses the Anchor-Based method for object detection, divides the image into multiple grids, and predicts a fixed number of bounding boxes and class probabilities for each grid. These bounding boxes are used to detect key points, thereby obtaining the coordinates of each key point.
[0145] S2103 Key point matching: During the key point matching process, the Euclidean distance between the key points and the reference points is calculated to find the key point coordinates corresponding to each corner point of the card. Assume that the reference point is the upper left corner point with coordinates and the coordinates of the key point are , then the Euclidean distance between the two points can be calculated by the following formula:
[0146] S2104 Perspective transformation calculation: First, after the calculation in S2103, the corner points are further sorted in a clockwise direction, and the four corner points of the perspective image card and are used. The planar homography method is adopted. Let the corner point coordinates of the corrected standard card image be , where , , , , are the standard width and standard height of the card image respectively. The planar homography method is a necessary step for perspective transformation. Assume that and are two points on the planes and respectively. The operation of converting point to point is defined as follows:
[0147] is a matrix. This homography matrix can be used to correct the perspective distortion of the input image. Calculate the matrix and from two given planes . First, find the corresponding transformation from a single point and then extend it to four points. Considering any pair of corresponding points and , the above formula can be transformed into:
[0148] Solve for and :
[0149] From this, the following two linear equations can be obtained:
[0150] At this time, apply the same method to four pairs of two corresponding points and to obtain the matrix for each corresponding point.
[0151] S2105. Image correction. After obtaining the perspective transformation matrix, it is necessary to use this matrix to convert the two-dimensional coordinate vector into the two-dimensional coordinates on the screen after perspective projection. The specific method is as follows: Step1: For each pixel in the original image, apply the perspective transformation matrix to obtain the transformed pixel coordinates.
[0152] Step2: Since the transformed pixel coordinates may no longer be integers, interpolation methods need to be used to determine the color of the new pixel.
[0153] Step3: Generate a new image based on the transformed pixel coordinates and the color obtained by interpolation. This new image is the original image after perspective transformation.
[0154] The embodiments of the present invention further include the step of extracting bank card information, such as Figure 11 shown. In the bank card image, it is necessary to extract the bank card number and the expiration date. If the OCR technology is directly used, it may be affected by factors such as font, size, color, and layout, resulting in poor recognition effect. In addition, the bank card number is more important than the information in the resume. For the security of talent information parsing, in this embodiment, a new method different from the OCR technology is designed for bank card information extraction, which can further ensure the recognition accuracy. The specific implementation steps of bank card information extraction are as follows: S3201. Obtain the regions of the card number and the expiration date in the bank card image, and use the target detection model Yolov5 to locate the regions where the card number and the expiration date are located; Use Mosaic data augmentation to randomly scale, randomly crop, and randomly arrange 4 pictures and then splice them; Adaptive anchor box calculation. For different data sets, there are anchor boxes with initially set length and width. During network training, the network outputs prediction boxes based on the initial anchor boxes, and then compares them with the real boxes, calculates the difference between the two, and updates them in reverse to iterate the network parameters.
[0155] Use the Letterbox method to adaptively add the least amount of black edges to the scaled image; Use the Backbone structure for feature extraction, and extract the object information in the image through the convolutional network for subsequent object detection.
[0156] Use the Neck structure to mix and combine features, enhance the robustness of the network, strengthen the object detection ability, and transfer these features to the Head layer for prediction.
[0157] To measure the difference between the predicted bounding box and the real bounding box, use the CloU loss for optimization, and its formula is:
[0158] where is the square of the diagonal length of the smallest rectangle containing the predicted box and the real box, is the sum of the squares of the distances between the centers of the predicted box and the real box.
[0159] In object detection, a region may be detected by multiple prediction boxes. To avoid detecting the same region multiple times, it is necessary to filter out the duplicate prediction boxes. This process is non-maximum suppression (NMS), and the steps are as follows: Sorting: First, sort all the candidate boxes in descending order of confidence score.
[0160] Select and Remove: Select the candidate box with the highest confidence and remove it from the candidate box list.
[0161] Calculate Intersection over Union (IoU): Calculate the Intersection over Union (IoU) between the remaining candidate boxes and the currently selected candidate box.
[0162] Remove overlapping boxes: If the IoU of a candidate box with the currently selected candidate box is greater than a preset threshold, remove that candidate box from the list.
[0163] Repeat: Repeat steps 2 - 4 until the candidate box list is empty.
[0164] Among them, the Intersection over Union is an index to measure the overlapping degree of two bounding boxes, and its mathematical formula can be expressed as:
[0165] Among them, and are two bounding boxes, is and is the area of the intersection region of is and is the area of the union region of
[0166] S3202. Crop the card number region and the expiration date region in the bank card image. After S3201, the coordinate values of the four corner points of the two candidate boxes will be obtained, and the initial image will be cropped according to the coordinates to obtain two effective regions for recognition.
[0167] S3203. Regarding the uniqueness of the bank card number, in this embodiment, a bank card number dataset is constructed and trained using the SVTR model used in S1203, so as to identify the bank card number and the effective region according to the candidate region, and judge whether the result is the bank card number or the effective region according to the character length of the recognition result.
[0168] The advantages of the embodiments of the present invention are concentrated in four aspects: intelligent parsing and document classification, complete recognition and extraction of long text information, image processing and card correction functions, and data correction and structured output. Around these four aspects, the present invention is applied in the field of talent recruitment.
[0169] Through intelligent parsing and document classification, this invention uses deep learning technology and an improved OCR model to achieve automated extraction and classification of resume, academic degree, ID card, and bank card information. Secondly, by combining a sliding window with a BERT-CRF model, it breaks through the length limitation of the traditional BERT model for long texts and can completely identify and extract information from long documents. Next, a card and certificate correction model based on key point detection is introduced. Through Yolov8-Pose and perspective transformation operations, it realizes the automatic correction of the image areas of ID cards and bank cards. Finally, a data correction and cleaning method is proposed. Using the BERT-CRF model to filter and uniformly format the recognized text information greatly improves the efficiency and structuring degree of information processing.
[0170] Traditional document recognition technologies are often limited to the extraction of single-type information, while the method of this invention integrates the extraction requirements of multiple document types. Based on the optimized extraction technology of deep learning models, it achieves highly intelligent document classification and information extraction, becoming an important technological breakthrough in the field of talent recruitment.
[0171] The above-disclosed is only a preferred embodiment of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. An intelligent talent information analysis method based on deep learning, characterized in that: The following steps are involved: S110: converting the resume file into an image format and performing correction processing on the image; S120: parsing the image text; S130: extracting information from the obtained text data by using a named entity recognition method to obtain semi-structured information; S140: Perform structured data cleaning on the information to obtain standardized data.
2. The method for analyzing intelligent talent information based on deep learning according to claim 1 is characterized in that: The S130 specifically includes the following steps: S131: Determine an entity tag according to identification requirements; S132: Establish a training set and annotate corresponding labels of the entity parts therein; S133: Segment the labeled training set text to obtain a word sequence; S134: inputting the word sequence into a BERT-CRF model for training to obtain an entity recognition model; S135: Input the text to be recognized into the entity model to obtain a recognition result, and use a regular expression to extract the part with longer characters.
3. The method for analyzing intelligent talent information based on deep learning according to claim 2 is characterized in that: The BERT-CRF model uses BERT as a feature extractor to obtain context-dependent representations of each word in an input sentence. The representations are input as features into a CRF model to learn label transition probabilities in a sequence labeling task and obtain a predicted labeling sequence.
4. The method for analyzing intelligent talent information based on deep learning according to claim 3 is characterized in that: The BERT-CRF model adds absolute position encoding to each word in the input sequence, and the process includes the following steps: Use the standard RoPE to calculate the Attention matrix once: in is the rotation matrix of RoPE, and are two input vectors, representing the query and key vectors respectively; set up is the position of the word in the sentence, is the corresponding position encoding vector, is the dimension of the position encoding, , then the function It is the position encoding function, which is defined as follows: in, , for each word, three vectors are calculated , these vectors are obtained by multiplying the sum of word embeddings and positional encodings with three different matrices; In the above, in, are the parameters of the model; The calculation interval is The Attention matrix of RoPE: in, is a constant that adjusts the interval of relative position encoding. Finally, according to the conditions , merging the two matrices together, we get: Through the above formula, we can get the relative position of nonlinearity , then calculate the attention score of each word to other words, using Indicates Word pair The attention score of each word; in is an activation function with a non-negative range, is the length of the input sequence.
5. The method for analyzing intelligent talent information based on deep learning according to claim 4 is characterized in that: The S130 further includes the steps of: Pass the output of the self-attention mechanism through a feed-forward neural network; The feedforward neural network includes two fully connected layers and a ReLu activation function in, are the parameters of the model, is the activation function. The vector output from BERT is fed into the CRF model. The objective function of the CRF model can be defined as: in, is the characteristic function The weight of and are the current and previous tags, is the input sequence, is the weighted sum of all feature functions; it is measured by the ratio of the predicted optimal path score to the scores of all paths. The larger the ratio, the more reasonable it is. The loss function is , that is, the negative log-likelihood, The larger it is, the smaller the loss is, thus optimizing the model for training by optimizing the loss function.
6. The method for analyzing intelligent talent information based on deep learning according to claim 5 is characterized in that: The S130 also includes a label confusion learning method for learning the relationship between labels, including the steps of: First, use the BERT model to learn contextual information for the text to be recognized and generate word vector representations, which are then fed into the CRF to obtain the predicted label distribution: in For BERT, for the input sequence , the length is , the output sequence is represented by Length is , the dimension is , To predict the label distribution Finally, the original BIOES distribution is added to the label confusion distribution with the control parameter α, and then normalized by the softmax function to generate a simulated label distribution. The control parameter α determines how much the label confusion distribution will change the BIOES distribution. The above process can be expressed as: in, is the label encoder function, which is an LSTM network used to transfer labels Represents the label representation matrix , C is the number of entity categories, is the label confusion distribution, To simulate label distribution.
7. The method for analyzing intelligent talent information based on deep learning according to claim 6 is characterized in that: The S135 further includes the steps of: using a method of dynamically and accurately dividing a sliding window to correct erroneous segmentation according to the similarity between the calculated divided paragraphs, including the steps of: Divide the entire text context into sentences: ,in Indicates sentences, calculate each sentence Embed , and split the text context into The block is: Next, for each block Perform block-wise maximum pooling to extract the main semantic information of the entire block, and obtain , then calculate the adjacent blocks and The cosine similarity between them is used to indicate the degree of semantic similarity between them: Identify local minima in a sequence of adjacent cosine similarities : These points represent possible segmentation locations in the text, and finally the identified segmentation points are With the original segment boundary Compare, if a segment point , not in , it is considered an erroneous segment and the original segment Modify the method and repeat the above steps until the most perfect segmentation method is achieved.
8. The method for analyzing intelligent talent information based on deep learning according to claim 1, characterized in that: The S140 specifically includes the following steps: S41: Missing value processing: delete missing values or fill in missing values; S42: outlier processing: deleting outliers or converting outliers; S43: Data format conversion: unify different date formats into a standard format; S44: Data deduplication: compare all data and delete completely duplicate data; S45: Data verification and consistency check: Compare the consistency between different data to ensure the accuracy and completeness of the data.
Citation Information
Cited By
Artificial intelligence-based talent evaluation management method and system
CN121190025A
A talent evaluation and management method and system based on artificial intelligence
CN121190025B