Chinese electronic resume named entity identification method and device, equipment and storage medium
By using a multi-layer attention mechanism recognition recognition model in Chinese electronic resume naming entity recognition, multi-knowledge features of characters, words and glyphs are extracted and fused, the problem of low accuracy of unlogged words and complex context recognition in the prior art is solved, and a higher accuracy of naming entity recognition is achieved.
Patent Information
- Application Number
- CN202510483810.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing Chinese electronic resume naming entity recognition methods mainly rely on regular expressions and domain dictionaries, making it difficult to deal with unlogged words and complex contexts, resulting in low recognition accuracy.
The recognition model constructed by a multi-layer attention mechanism is used to preprocess the text of the electronic resume, and the named entity recognition is performed through the multi-knowledge feature extraction layer, the multi-knowledge attention layer and the label prediction layer. The model is able to extract multi-knowledge features of characters, words, and glyphs and fuse these features through attention mechanisms to improve the accuracy of recognition.
By capturing the correlation between entity semantics and context in the resume, the accuracy of electronic resume naming entity recognition is improved, and complex Chinese resume text can be processed more effectively.
Smart Images

Figure CN120068873A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of recognition technology, and particularly relates to a method, device, equipment and storage medium for named entity recognition of Chinese electronic resumes. Background Art
[0002] With the rapid development of digital recruitment and talent management, electronic resumes, as the core information carrier between job seekers and employers, have an increasingly urgent need for structured parsing. Named Entity Recognition (NER) technology aims to extract predefined key entities (such as person names, organizations, time, etc.) from unstructured text, and the NER task of Chinese electronic resumes faces multiple challenges due to its unique language characteristics and application scenarios.
[0003] Chinese resume texts lack explicit delimiters (such as spaces) and contain a large number of domain-specific vocabulary (such as "Java development engineer", "National Encouragement Scholarship"), and need to handle word segmentation ambiguity, nested entities, and non-standard expressions (such as the mixed use of "2020.09 - 2023.06" and "from September 2020 to June 2023").
[0004] Moreover, Chinese resume NER mainly adopts rule- and dictionary-driven methods, which quickly locate fixed-pattern entities based on regular expressions (such as matching phone numbers and email addresses) and domain dictionaries (such as college directories and job title libraries), but are sensitive to out-of-vocabulary words and complex contexts, resulting in low recognition accuracy. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, device, equipment and storage medium for named entity recognition of Chinese electronic resumes, so as to solve the problem that the existing method quickly locates fixed-pattern entities based on regular expressions and domain dictionaries, but is sensitive to out-of-vocabulary words and complex contexts, resulting in low recognition accuracy.
[0006] To achieve the above-mentioned invention purpose, the technical solutions adopted by the present invention are as follows: In the first aspect, the present invention provides a method for named entity recognition of Chinese electronic resumes, and the method includes: Obtain the electronic resume to be recognized, perform text extraction on the electronic resume to obtain the text to be recognized; Preprocess the text to be recognized to obtain the target text; Based on a pre-constructed recognition model, recognize the named entities of the target text to obtain a recognition result, wherein the recognition model is constructed based on a multi-layer attention mechanism; Visualize and display the recognition result.
[0007] Preferably, the recognition model includes: a multi-knowledge feature extraction layer, a multi-knowledge attention layer, and a label prediction layer; The multi-knowledge feature extraction layer is used to extract features from the target text to obtain a multi-knowledge feature vector; The multi-knowledge attention layer is used to fuse the multi-knowledge feature vectors to obtain an interactive feature representation; The label prediction layer is used to perform label prediction on the interactive feature representation to obtain the labels corresponding to each character in the target text, and use the labels corresponding to each character in the target text as the recognition result.
[0008] Preferably, the multi-knowledge feature vector includes: a character vector, a word vector, and a glyph vector.
[0009] Preferably, the multi-knowledge feature extraction layer includes: A character feature extraction module, which is used to extract characters from the target text to obtain a character vector; A word feature extraction module, which is used to extract words from the target text to obtain a word vector; A glyph feature extraction module, which is used to extract glyphs from the target text to obtain a glyph vector.
[0010] Preferably, the multi-knowledge attention layer includes: a first attention module, a second attention module, and a third attention module, and both the first attention module and the second attention module are connected to the third attention module.
[0011] Preferably, fusing the multi-knowledge feature vectors to obtain an interactive feature representation includes: Construct a ternary matrix of the first attention module based on the character vector and the word vector, and construct a ternary matrix of the second attention module based on the character vector and the glyph vector; Determine the attention output of the first attention module based on the ternary matrix of the first attention module; Determine the attention output of the second attention module based on the ternary matrix of the second attention module; Construct a ternary matrix of the third attention module based on the attention output of the first attention module and the attention output of the second attention module; Determine the attention output of the third attention module based on the ternary matrix of the third attention module, and use the attention output of the third attention module as the interactive feature representation.
[0012] Preferably, the ternary matrix includes: a query matrix, a key matrix, and a value matrix, and the calculation steps of the attention output of the first attention module, the attention output of the second attention module, and the attention output of the third attention module all include: Taking the query matrix, key matrix, and value matrix as the inputs of the attention function, the attention function outputs the first matrix; Based on a preset similarity function, calculate the similarity between the first matrix and the query matrix to obtain the second matrix; Based on a preset linear transformation function, perform transformations on the first matrix and the second matrix respectively to obtain a third matrix corresponding to the first matrix and a fourth matrix corresponding to the second matrix; Based on a preset non - linear transformation function, perform a transformation on the fourth matrix to obtain the fifth matrix; Based on the fifth matrix and the third matrix, obtain the attention output.
[0013] In a second aspect, the present invention provides a Chinese electronic resume named entity recognition device for implementing the above - mentioned Chinese electronic resume named entity recognition method. The device includes: A text extraction module for obtaining an electronic resume to be recognized, extracting text from the electronic resume to obtain the text to be recognized; A text processing module for pre - processing the text to be recognized to obtain the target text; An entity recognition module for recognizing the named entities of the target text based on a pre - constructed recognition model, obtaining a recognition result, wherein the recognition model is constructed based on a multi - layer attention mechanism; A recognition display module for visually displaying the recognition result.
[0014] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above - mentioned Chinese electronic resume named entity recognition method is implemented.
[0015] In a fourth aspect, the present invention provides a computer - readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above - mentioned Chinese electronic resume named entity recognition method is implemented.
[0016] The beneficial effects of the present invention are mainly reflected in: The present invention extracts text from the electronic resume to obtain the text to be recognized, pre - processes the text to be recognized to obtain the target text, and then uses a recognition model constructed by a multi - layer attention mechanism to recognize the named entities of the target text, which can capture the association between the entity semantics of the resume and the context, and improve the accuracy of named entity recognition of the electronic resume. Description of the Drawings
[0017] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not limit the embodiments of the present invention. In the accompanying drawings: Figure 1 is a flowchart of a method for named entity recognition in a Chinese electronic resume provided by an embodiment of the present invention; Figure 2 is a block diagram of a device for named entity recognition in a Chinese electronic resume provided by an embodiment of the present invention. Detailed Description of the Embodiments
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the present invention will be briefly introduced below in combination with the accompanying drawings and the description of the embodiments or the prior art. Obviously, the following description of the structures of the accompanying drawings is only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts. It should be noted here that the description of these embodiments is used to help understand the present invention, but does not limit the present invention.
[0019] Embodiment 1 Figure 1 is a flowchart of a method for named entity recognition in a Chinese electronic resume provided by an embodiment of the present invention. As Figure 1 shown, this embodiment provides a method for named entity recognition in a Chinese electronic resume, and the method includes: Step S10: Obtain the electronic resume to be recognized, perform text extraction on the electronic resume, and obtain the text to be recognized.
[0020] In this embodiment, the electronic resume is usually in PDF format or word format, and tools such as pdfminer, python-docx, or OCR tools (such as PaddleOCR) can be used to extract text, and the extracted text is used as the text to be recognized.
[0021] Among them, pdfminer is an open-source tool, mainly a Python library for extracting various information from PDF documents; it can parse PDF content into text, pictures, and other metadata, even if the PDF contains complex layouts and formats.
[0022] Among them, python-docx is a Python library for creating and updating Microsoft Word (.docx) files. It allows users to create new Word documents from scratch or modify existing.docx files, including adding elements such as text, pictures, headers, footers, and tables.
[0023] Among them, OCR (Optical Character Recognition) is a recognition technology that can convert the text in different types of documents (such as scanned paper documents, PDF files or images) into a machine-editable text format.
[0024] Step S20: Preprocess the text to be recognized to obtain the target text.
[0025] In this embodiment, the preprocessing is mainly used to remove the garbled characters and irrelevant characters in the text to be recognized, merge the broken lines and process the table content.
[0026] The types of garbled characters and irrelevant characters are mainly the following types: Encoding error characters: such as the garbled characters caused by the failure of UTF-8 decoding; Special symbols: irrelevant HTML tags ( ), LaTeX control characters ( ), invisible characters; Redundant characters: consecutive repeated punctuation (------), advertising text.
[0027] Step S30: Recognize the named entities of the target text based on the pre-constructed recognition model to obtain the recognition result, where the recognition model is constructed based on a multi-layer attention mechanism.
[0028] In this embodiment, the recognition model includes: a multi-knowledge feature extraction layer, a multi-knowledge attention layer, and a label prediction layer.
[0029] Among them, the multi-knowledge feature extraction layer is used to extract features from the target text to obtain a multi-knowledge feature vector; the multi-knowledge attention layer is used to fuse the multi-knowledge feature vectors to obtain an interactive feature representation; the label prediction layer is used to perform label prediction on the interactive feature representation to obtain the labels corresponding to each character in the target text, and use the labels corresponding to each character in the target text as the recognition result.
[0030] In this embodiment, due to the limited context information of the electronic resume, the semantic information learned by the attention mechanism is limited, which will lead to poor performance of the final recognition result. Therefore, this application constructs a multi-knowledge feature extraction layer and a multi-knowledge attention layer; the multi-knowledge feature extraction layer can extract multi-knowledge features such as character vectors, word vectors, and glyph vectors, and then the multi-knowledge attention layer performs interactive fusion on the character vectors, word vectors, and glyph vectors to extract more semantic knowledge, enrich the context, and thus improve the recognition accuracy.
[0031] In this embodiment, the recognition model is trained using a sample data set and deployed after training; during the training process, data collection and labeling are performed on the electronic resume: Annotated entity types: name, contact information (phone / email), educational background (school, major, degree, time), work experience (company, position, time), skills, projects, certificates, etc.
[0032] Tool annotation: Use annotation tools (such as BRAT, Label Studio) for manual annotation, or use semi-automatic tools (such as pre-annotation with regular expressions).
[0033] Public datasets: If there is no annotated data, Chinese resume datasets such as MSRA-NER and ResumeNER can be reused.
[0034] In this embodiment, the label prediction layer adopts the CRF (conditional random field algorithm) algorithm. By using the CRF algorithm, the transition probabilities between different labels are predicted, thereby reducing the occurrence of unreasonable label sequence combinations and improving the prediction accuracy of the recognition results.
[0035] As a further optimization of this embodiment, the multi-knowledge feature extraction layer includes: a character feature extraction module, a word feature extraction module, and a glyph feature extraction module.
[0036] Among them, the character feature extraction module is used to extract characters from the target text to obtain character vectors.
[0037] In this embodiment, the character feature extraction module uses a pre-trained BERT model to extract character knowledge. Since it is trained on a large amount of pre-trained corpus, the BERT model has learned a large amount of semantic knowledge, can extract richer character semantic information, and effectively avoids the noise caused by character ambiguity in the case of joint context.
[0038] Among them, the word feature extraction module is used to extract words from the target text to obtain word vectors.
[0039] In this embodiment, the word feature extraction module can adopt the Lattice-LSTM model. The Lattice-LSTM model is based on a vocabulary enhancement method and combines character and vocabulary information; when processing a sentence, it will consider all potential vocabulary in the sentence to form a structure similar to a lattice (Lattice). This structure can avoid entity recognition errors caused by word segmentation errors, thereby improving the accuracy of NER.
[0040] Among them, the glyph feature extraction module is used to extract glyphs from the target text to obtain glyph vectors.
[0041] In this embodiment, since the Chinese of the target text contains glyph information, the semantics contained in the capturer can be captured by converting the characters in the target text into corresponding Wubi codes. Therefore, the glyph feature extraction module uses an existing Wubi code conversion table to convert the characters of the target text into corresponding Wubi codes, and then uses a convolutional neural network to extract features from the Wubi codes of the target text to extract the semantic information contained in the glyphs and obtain a glyph vector.
[0042] As a further optimization of this embodiment, since there is a certain heterogeneity among characters, words, and glyphs, directly fusing the three is not conducive to exploring the complementarity between multiple knowledge; therefore, in this embodiment, the multi-knowledge attention layer uses the attention mechanism to fuse the word vector, glyph vector, and character vector to strengthen the representation of the text and fully integrate the context information.
[0043] In this embodiment, the multi-knowledge attention layer includes: a first attention module, a second attention module, and a third attention module, and both the first attention module and the second attention module are connected to the third attention module.
[0044] Among them, fusing the multi-knowledge feature vectors to obtain an interactive feature representation includes: Step A1: Construct a ternary matrix of the first attention module based on the character vector and the word vector, and construct a ternary matrix of the second attention module based on the character vector and the glyph vector.
[0045] In this embodiment, the character vector and the word vector are concatenated to obtain a character-word vector, and then the character-word vector is used as the input of the first attention module. Similarly, the character vector and the glyph vector are concatenated to obtain a character-glyph vector, and then the character-glyph vector is used as the input of the second attention module.
[0046] Step A2: Determine the attention output of the first attention module based on the ternary matrix of the first attention module.
[0047] In this embodiment, the ternary matrix includes: a query matrix, a key matrix, and a value matrix, and the function expression of the ternary matrix is: (1); In formula (1), Q is the query matrix, K is the key matrix, V is the value matrix, X is the input vector, that is, the input of the first attention module, the input of the second attention module, and the input of the third attention module in the following text; is the weight matrix of the query, is the weight matrix of the key, is the weight matrix of the value; where, , and It can be obtained through learning.
[0048] In this embodiment, the word vector and the character vector are fused through the first attention module. The fused character vector contains the semantic information of the words in the given sentence as additional context information, enriching the semantic knowledge.
[0049] Step A3: Based on the triple matrix of the second attention module, determine the attention output of the second attention module.
[0050] In this embodiment, the second attention module calculates the correlation between characters and glyphs to obtain a glyph-enhanced character feature representation.
[0051] Step A4: Based on the attention output of the first attention module and the attention output of the second attention module, construct the triple matrix of the third attention module.
[0052] In this embodiment, the attention output of the first attention module is concatenated with the attention output of the second attention module to obtain a concatenated vector. The concatenated vector is used as the input of the third attention module, and the triple matrix of the third attention module is also calculated using formula (1).
[0053] Step A5: Based on the triple matrix of the third attention module, determine the attention output of the third attention module, and use the attention output of the third attention module as the interaction feature representation.
[0054] In this embodiment, the third attention module fuses the attention output of the first attention module and the attention output of the second attention module. At this time, the attention output of the third attention module contains the knowledge of text, words, and glyphs, which can effectively fuse multiple information, realize the interaction of semantic information of multiple knowledge, and can better help the model understand entity semantics and locate entity boundaries, improving the accuracy of entity recognition.
[0055] Step S40: Visualize the recognition result.
[0056] In this embodiment, the recognition result is converted into a graph model or a table, the graph model or the table is displayed, or an entity report is generated, and the entity report is sent to the user terminal of the user (such as devices such as mobile phones and computers), and the entity report is visually displayed through the user terminal.
[0057] As a further optimization of this embodiment, since the query matrix Q of the three attention modules outputs a weighted average value, and whether the query matrix is relevant to the key matrix or the value matrix, the three attention modules will still generate a weighted average vector; therefore, when there is no correlation between the query matrix and the key matrix or the value matrix, the output of the three attention modules will mislead the result, thereby affecting the quality of the recognition result.
[0058] Therefore, the calculation steps of the attention output of the first attention module, the attention output of the second attention module, and the attention output of the third attention module all include: Step B10: Use the query matrix, the key matrix, and the value matrix as the inputs of the attention function, and the attention function outputs the first matrix.
[0059] Step B20: Calculate the similarity between the first matrix and the query matrix based on a preset similarity function to obtain the second matrix; among them, the similarity function uses the cosine similarity function.
[0060] Step B30: Perform transformations on the first matrix and the second matrix respectively based on a preset linear transformation function to obtain a third matrix corresponding to the first matrix and a fourth matrix corresponding to the second matrix.
[0061] Among them, the function expressions of the third matrix and the fourth matrix are: (2); (3); In formula (2), X 1 is the first matrix, X 2 is the second matrix, X 3 is the third matrix, X 4 is the fourth matrix, linear () is the linear transformation function.
[0062] Step B40: Perform a transformation on the fourth matrix based on a preset non-linear transformation function to obtain the fifth matrix.
[0063] Among them, the function expression of the fifth matrix is: (4); In formula (4), X5 is the fifth matrix, and sigmoid() is the non-linear transformation function.
[0064] Step B50: Obtain the attention output based on the fifth matrix and the third matrix.
[0065] Among them, the function expression of the attention output is: (5); In formula (6), X 6 is the attention output, and W is the fusion weight matrix.
[0066] In this embodiment, in each attention calculation process, the cosine similarity function can be used to calculate the similarity between the first matrix and the query matrix. When the similarity is close to 1, it indicates that the first matrix and the query matrix are more similar. When the similarity is close to 0, it indicates that there is no association between the first matrix and the query matrix. After the fourth matrix undergoes a non-linear transformation, if the similarity is close to 0 and there is no correlation between the two features, the corresponding weight can be cleared, which can better enhance the reasoning ability and further improve the quality of feature fusion.
[0067] The word features and glyph features of the present invention provide effective assistance in the entity recognition task of electronic Chinese resumes. Especially when there are more words involved in the entity, the contribution of the word features is greater; when there are fewer words involved in the entity, the Chinese character features can provide greater help to improve the recognition accuracy.
[0068] Therefore, the present invention extracts text from the electronic resume to obtain the text to be recognized, preprocesses the text to be recognized to obtain the target text, and then uses the recognition model constructed by the multi-layer attention mechanism to recognize the named entities of the target text, which can capture the association between the entity semantics and context of the resume and improve the accuracy of named entity recognition of the electronic resume.
[0069] Embodiment 2 Figure 2 This is a named entity recognition device for Chinese electronic resumes provided by an embodiment of the present invention. As Figure 2 shown, this embodiment provides a named entity recognition device for Chinese electronic resumes. The device is used to implement the named entity recognition method for Chinese electronic resumes in Embodiment 1. The device includes: A text extraction module, configured to obtain the electronic resume to be recognized and extract text from the electronic resume to obtain the text to be recognized; A text processing module, configured to preprocess the text to be recognized to obtain the target text; An entity recognition module, configured to recognize the named entities of the target text based on a pre-constructed recognition model to obtain a recognition result, where the recognition model is constructed based on a multi-layer attention mechanism; A recognition display module, configured to visually display the recognition result.
[0070] This embodiment also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the named entity recognition method for Chinese electronic resumes in Embodiment 1.
[0071] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the Chinese electronic resume named entity recognition method in Embodiment 1.
[0072] The present invention extracts text from an electronic resume to obtain text to be recognized, preprocesses the text to be recognized to obtain target text, and then uses a recognition model constructed by a multi-layer attention mechanism to recognize the named entities in the target text, which can capture the association between the entity semantics of the resume and the context, and improve the accuracy of named entity recognition in electronic resumes.
[0073] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0074] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one Figure 1 one process or multiple processes and / or blocks Figure 1 a block or multiple blocks.
[0075] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for identifying named entities in Chinese electronic resumes, characterized in that: The method comprises: Obtaining an electronic resume to be identified, performing text extraction on the electronic resume, and obtaining the text to be identified; Preprocess the text to be recognized to obtain the target text; Recognize the named entities of the target text based on a pre-built recognition model to obtain a recognition result, wherein the recognition model is built based on a multi-layer attention mechanism; Visualize the recognition results.
2. The method for identifying named entities in Chinese electronic resumes according to claim 1, characterized in that: The recognition model includes: a multi-knowledge feature extraction layer, a multi-knowledge attention layer and a label prediction layer; The multi-knowledge feature extraction layer is used to extract features from the target text to obtain a multi-knowledge feature vector; The multi-knowledge attention layer is used to fuse the multi-knowledge feature vectors to obtain interactive feature representation; The label prediction layer is used to perform label prediction on the interactive feature representation to obtain the label corresponding to each character in the target text, and use the label corresponding to each character in the target text as the recognition result.
3. The method for identifying named entities in Chinese electronic resumes according to claim 2, characterized in that: The multi-knowledge feature vector includes: a character vector, a word vector and a glyph vector.
4. The method for identifying named entities in Chinese electronic resumes according to claim 3 is characterized in that: The multi-knowledge feature extraction layer includes: Character feature extraction module, used to extract characters from target text and obtain character vectors; The word feature extraction module is used to extract words from the target text and obtain word vectors; The glyph feature extraction module is used to extract glyphs from the target text and obtain glyph vectors.
5. The method for identifying named entities in Chinese electronic resumes according to claim 3, characterized in that: The multi-knowledge attention layer includes: a first attention module, a second attention module and a third attention module, and the first attention module and the second attention module are both connected to the third attention module.
6. The method for identifying named entities in Chinese electronic resumes according to claim 5, characterized in that: Multiple knowledge feature vectors are fused to obtain interactive feature representations, including: A ternary matrix of the first attention module is constructed based on the character vector and the word vector, and a ternary matrix of the second attention module is constructed based on the character vector and the glyph vector; Determining an attention output of the first attention module based on the ternary matrix of the first attention module; Determining an attention output of the second attention module based on the ternary matrix of the second attention module; Construct a ternary matrix of a third attention module based on the attention output of the first attention module and the attention output of the second attention module; Based on the ternary matrix of the third attention module, the attention output of the third attention module is determined, and the attention output of the third attention module is used as the interaction feature representation.
7. The method for identifying named entities in Chinese electronic resumes according to claim 6, characterized in that: The ternary matrix includes: a query matrix, a key matrix and a value matrix, and the calculation steps of the attention output of the first attention module, the attention output of the second attention module and the attention output of the third attention module all include: Taking the query matrix, the key matrix and the value matrix as inputs of an attention function, the attention function outputs a first matrix; Based on a preset similarity function, the similarity between the first matrix and the query matrix is calculated to obtain a second matrix; Based on a preset linear transformation function, the first matrix and the second matrix are transformed respectively to obtain a third matrix corresponding to the first matrix and a fourth matrix corresponding to the second matrix; transforming the fourth matrix based on a preset nonlinear transformation function to obtain a fifth matrix; Based on the fifth matrix and the third matrix, the attention output is obtained.
8. A device for identifying named entities in Chinese electronic resumes, used to implement the method for identifying named entities in Chinese electronic resumes as described in any one of claims 1 to 7, characterized in that: The device comprises: The text extraction module is used to obtain the electronic resume to be identified, extract the text from the electronic resume, and obtain the text to be identified; A text processing module is used to pre-process the text to be recognized to obtain the target text; An entity recognition module is used to recognize named entities in the target text based on a pre-built recognition model to obtain a recognition result, wherein the recognition model is built based on a multi-layer attention mechanism; The recognition display module is used to visualize the recognition results.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for recognizing named entities in Chinese electronic resumes described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for recognizing named entities in Chinese electronic resumes described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Improved Chinese named entity identification method based on Lattice-LSTM
CN111476031A
Chinese named entity recognition method and device based on multi-view semantic feature fusion
CN114580416A
Named entity identification method and terminal
CN115221880A
Chinese electronic medical record named entity identification method and system
CN117057350A