Entity recognition method and device, electronic device and storage medium
By iteratively training and adjusting parameters of multiple models to generate a target model, the problems of high manual labeling costs and data omissions and errors in entity recognition are solved, and automated and efficient entity recognition is achieved.
Patent Information
- Application Number
- CN202210307561.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2042-03-25
AI Technical Summary
The entity recognition methods in the existing technology have the problems of high manual labeling cost, low efficiency, data omission and incorrect labeling.
By iteratively training and adjusting parameters of multiple models, a target model is generated to identify entity categories in text data. This model includes a pre-trained second model and a self-learned third model, combined with fragment sequence labeling and latent vector processing to achieve automated recognition.
It realizes automatic entity recognition, saves manpower, improves recognition efficiency, and is more accurate in identifying entity categories, reducing data omissions and errors.
Smart Images

Figure CN114626380B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a method and device for entity identification, an electronic device, and a storage medium. Background Art
[0002] Named Entity Recognition (NER) is the task of detecting real entities from text and classifying them into predetermined types (e.g., places, people, objects, and organizations). It is a core task in knowledge extraction and a fundamental and important foundation for various downstream applications, such as search engines, question-answering systems, and dialogue systems.
[0003] Traditional NER methods primarily train sequence annotation models, such as hidden Markov models and conditional random fields, based on manually designed features. With the development of deep neural networks, deep learning models can leverage deep neural networks to automatically extract effective features, alleviating the burden of manually designed features. Therefore, deep learning models for NER tasks have been proposed and demonstrated strong performance. However, most deep learning methods rely on a large amount of annotated training data. Since NER tasks require character (token)-level labels, in sequence annotation-based NER models, a token can often only be annotated as one entity, failing to address the problem of entity nesting. Deep learning-based NER often requires a large amount of annotated data. This presents challenges for NER tasks in domains lacking a large amount of annotated data, such as the high cost, low timeliness, and potential for errors in manual annotation. Some NER models employing remote supervision utilize existing knowledge bases or domain dictionaries for data annotation, which can lead to missed data labels due to the limited coverage of the knowledge base.
[0004] Therefore, the related technologies have problems such as high cost, low efficiency, missed data and incorrect labeling of manual labeling. Summary of the Invention
[0005] The present application provides a method and apparatus for entity recognition, an electronic device, and a storage medium to at least solve the problems of high cost, low efficiency, missed data, and incorrect data labeling in the related art.
[0006] According to one aspect of an embodiment of the present application, a method for entity recognition is provided, the method comprising:
[0007] Obtain target text data to be recognized;
[0008] The target text data is input into a target model to obtain a target entity category to which the target text data belongs, wherein the target model is used to obtain annotation information of the text data and identify the target entity category based on the annotation information. The target model is a final model obtained by adjusting the third model parameters of the third model. The third model is a model that uses the second model parameters in the second model to pre-train the training set. The second model is a model obtained after iterative training of the first model for a preset number of times. The preset number of times is obtained by processing the training set using the fourth model.
[0009] According to another aspect of an embodiment of the present application, a device for entity identification is further provided, the device comprising:
[0010] A first acquiring unit, configured to acquire target text data to be recognized;
[0011] The first input unit is used to input the target text data into the target model to obtain the target entity category to which the target text data belongs, wherein the target model is used to obtain the annotation information of the text data and identify the target entity category based on the annotation information. The target model is a final model obtained by adjusting the third model parameters of the third model. The third model is a model that uses the second model parameters in the second model to pre-train the training set. The second model is a model obtained after the first model is iteratively trained for a preset number of times, and the preset number of times is obtained by processing the training set using the fourth model.
[0012] Optionally, the device further comprises:
[0013] A second acquisition unit, configured to acquire training text data before acquiring the target text data to be recognized;
[0014] a splicing unit, configured to perform segment-wise splicing of characters in the training text data according to a preset scheme to generate a plurality of segment sequences;
[0015] a matching unit, configured to perform text matching on each character in the segment sequence with a preset entity name to determine the entity type to which the training text data belongs;
[0016] A setting unit is used to use the training text data and the entity type as the training set.
[0017] Optionally, the splicing unit includes:
[0018] A division module is used to divide the training text data into single characters and perform character labeling on each divided character;
[0019] The splicing module is used to perform segment-wise splicing on the character annotations to generate a plurality of segment sequences.
[0020] Optionally, the splicing module includes:
[0021] a determination subunit, configured to determine a preset window length, wherein the preset window length is a maximum value of the total number of characters allowed to be contained in each of the segment sequences;
[0022] The splicing subunit is used to splice the head character and the tail character contained in each segment within the range of the preset window length to obtain a plurality of segment sequences, wherein each segment contains at least one character.
[0023] Optionally, the device further comprises:
[0024] A second input unit is configured to generate a plurality of latent vectors corresponding to each of the segment sequences according to the training text data and the first model after the training text data and the entity type are used as the training set;
[0025] a third input unit, configured to input the plurality of latent vectors into the feedforward neural network of the first model to obtain a first probability value of each latent vector belonging to the entity type;
[0026] a first adjusting unit, configured to adjust first model parameters of the first model according to the first probability value and after the preset number of iterations to obtain the second model;
[0027] The second adjustment unit is used to adjust the third model parameters of the third model based on the second model and the multiple fragment sequences to obtain the target model.
[0028] Optionally, the second adjustment unit includes:
[0029] an initialization module, configured to initialize the third model using the second model parameters of the second model, wherein the third model parameters in the third model are currently equal to the second model parameters;
[0030] an input module, configured to input the plurality of latent vectors into the third model to obtain a reference probability value of each of the segment sequences belonging to the entity type;
[0031] The first adjustment module is used to train the third model using the mean square error loss function, adjust the third model parameters of the third model until the reference probability value is greater than or equal to a preset threshold, and obtain the target model, wherein the preset threshold is the minimum value for stopping adjusting the third model parameters.
[0032] Optionally, the first adjustment module includes:
[0033] A first input subunit is configured to input the plurality of latent vectors into the first submodel of the third model to obtain a second probability value;
[0034] a training subunit, configured to train the second submodel of the third model based on the second probability value using the mean square error loss function until the preset number of iterations are completed, thereby obtaining second submodel parameters of the trained second submodel;
[0035] an updating sub-unit, configured to update first sub-model parameters in the first sub-model using the second sub-model parameters to obtain an updated first sub-model;
[0036] A second input subunit is configured to input the plurality of latent vectors into the updated first submodel to obtain a third probability value;
[0037] The second adjustment module is used to adjust the second sub-model parameters based on the third probability value until the reference probability value output by the second sub-model is greater than or equal to the preset threshold, stop adjusting the second sub-model parameters, and obtain the target model.
[0038] According to another aspect of the embodiments of the present application, an electronic device is also provided, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; wherein the memory is used to store computer programs; and the processor is used to execute the method steps in any of the above embodiments by running the computer program stored on the memory.
[0039] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the method steps in any of the above embodiments when run.
[0040] The embodiment of the present application constructs a graph in the field of knowledge graph technology. In the embodiment of the present application, by obtaining target text data to be identified; inputting the target text data into a target model, obtaining the target entity category to which the target text data belongs, wherein the target model is used to obtain the annotation information of the text data, and identify the target entity category based on the annotation information, the target model is a final model obtained by adjusting the third model parameters of the third model, the third model is a model that pre-trains the training set using the second model parameters in the second model, and the second model is a model obtained after iterative training of the first model for a preset number of times, and the preset number of times is obtained by processing the training set using the fourth model. Since the embodiment of the present application uses the trained target model as the final model for processing the entity category to which the target text data to be identified belongs, the effect of automatic recognition is achieved, manpower is saved, and recognition efficiency is improved, and the target model is completed by continuous training and parameter adjustment through self-learning of the first model, the second model, the third model, and the fourth model, which is more accurate in identifying entity categories and solves the problems of missed and incorrect data labels. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0043] Figure 1 is a schematic diagram of a hardware environment of an optional entity recognition method according to an embodiment of the present invention;
[0044] Figure 2 is a flowchart of an optional entity identification method according to an embodiment of the present application;
[0045] Figure 3 1 is a schematic diagram of an optional domain BERT_NER model structure according to an embodiment of the present invention;
[0046] Figure 4 1 is a schematic diagram of the overall training process of an optional entity recognition method according to an embodiment of the present application;
[0047] Figure 5 This is a structural block diagram of an optional entity identification device according to an embodiment of the present application;
[0048] Figure 6This is a structural block diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0049] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0051] According to one aspect of the embodiment of the present application, a method for entity recognition is provided. Optionally, in this embodiment, the above-mentioned method for entity recognition can be applied to Figure 1 In the hardware environment shown in the figure. Figure 1 As shown, the terminal 102 may include a memory 104, a processor 106, and a display 108 (optional component). The terminal 102 may be communicatively connected to a server 112 via a network 110. The server 112 may be configured to provide services (such as application services) for the terminal or a client installed on the terminal. A database 114 may be provided on or independent of the server 112 to provide data storage services for the server 112. In addition, a processing engine 116 may be running on the server 112, which may be configured to execute the steps performed by the server 112.
[0052] Optionally, the terminal 102 may be, but is not limited to, a terminal capable of computing data, such as a mobile terminal (e.g., a mobile phone, tablet computer), a laptop computer, or a PC (Personal Computer). The aforementioned network may include, but is not limited to, a wireless network or a wired network. Wireless networks include Bluetooth, Wi-Fi (Wireless Fidelity), and other networks that enable wireless communication. Wired networks may include, but are not limited to, wide area networks, metropolitan area networks, and local area networks. The aforementioned server 112 may include, but is not limited to, any hardware device capable of computing.
[0053] Furthermore, in this embodiment, the entity identification method described above can also be applied, but is not limited to, to an independent processing device with relatively powerful processing capabilities, without requiring data exchange. For example, the processing device can be, but is not limited to, a terminal device with relatively powerful processing capabilities. That is, the various operations in the entity identification method described above can be integrated into a single independent processing device. The above is merely an example, and this embodiment does not impose any limitation thereto.
[0054] Optionally, in this embodiment, the above-mentioned entity identification method can be executed by the server 112, or by the terminal 102, or can be executed jointly by the server 112 and the terminal 102. The entity identification method of the embodiment of the present application performed by the terminal 102 can also be executed by a client installed thereon.
[0055] Taking running on the server as an example, Figure 2 is a flow chart of an optional entity recognition method according to an embodiment of the present application, such as Figure 2 As shown, the process of the method may include the following steps:
[0056] Step S201, obtaining target text data to be recognized;
[0057] In step S202, the target text data is input into the target model to obtain the target entity category to which the target text data belongs, wherein the target model is used to obtain the annotation information of the text data and identify the target entity category based on the annotation information. The target model is a final model obtained by adjusting the third model parameters of the third model. The third model is a model that uses the second model parameters in the second model to pre-train the training set. The second model is a model obtained after iterative training of the first model for a preset number of times. The preset number of times is obtained by processing the training set using the fourth model.
[0058] Optionally, in this embodiment of the present application, the server will obtain the target text data to be recognized, and then input the target text data into the target model of this embodiment of the present application. The target model will output the target entity category to which the target text data belongs. It should be noted that an entity often refers to a collection of certain types of things, and the individual data objects of each type are called entities. Therefore, the entity types that can be included in this embodiment of the present application include: brand, product, etc.
[0059] In addition, in addition to the target model, the embodiment of the present application also includes multiple other models, namely: a first model (which can be a BERT_NER model), a second model (which can be a BERT_NER model after multiple iterations), a third model (which can be a self-learning teacher-student model), and a fourth model (which can be an early stopping method (early_stop) model). The relationship between the multiple models is: the target model is the final model determined by adjusting the third model parameters of the third model, and the third model is associated with the second model, and the third model parameters of the third model are the same as the second model parameters of the second model, that is, the third model is a model that uses the second model parameters in the second model to pre-train the training set, and then the first model constructs an initial domain BERT_NER model for BERT, which can be the cosmetics field, etc., and uses the BERT_NER model to identify which entity category in the cosmetics field the data to be labeled belongs to, and then the second model is the model generated after multiple iterative training of the first model.
[0060] It should be noted that, in the above-mentioned multiple iterative training, when determining the specific value of "multiple times", the fourth model can be used to process the training set, and the obtained number of training iterations, for example, the obtained number of iterations is a preset number: 5 times, etc.
[0061] In an embodiment of the present application, by obtaining target text data to be identified; inputting the target text data into a target model, obtaining the target entity category to which the target text data belongs, wherein the target model is used to obtain the annotation information of the text data, and identify the target entity category based on the annotation information, the target model is a final model obtained by adjusting the third model parameters of the third model, the third model is a model that pre-trains the training set using the second model parameters in the second model, and the second model is a model obtained after iterative training of the first model for a preset number of times, the preset number of times being obtained by processing the training set using the fourth model. Since the embodiment of the present application uses the trained target model as the final model for processing the entity category to which the target text data to be identified belongs, the effect of automated recognition is achieved, manpower is saved, and recognition efficiency is improved, and the target model is completed by continuous training and parameter adjustment through self-learning of the first model, the second model, the third model, and the fourth model, and is more accurate in identifying entity categories, solving the problems of missed and incorrect data labels.
[0062] As an optional embodiment, before obtaining the target text data to be recognized, the method further includes:
[0063] Get training text data;
[0064] Perform segment-wise splicing of characters in the training text data according to a preset scheme to generate multiple segment sequences;
[0065] Perform text matching on each character in the segment sequence with the preset entity name to determine the entity type to which the training text data belongs;
[0066] The training text data and entity types are used as the training set.
[0067] Optionally, in an embodiment of the present application, a training set needs to be generated in advance to train the model. First, training text data is obtained. These training text data can be domain data based on a multi-source knowledge base. Then, the characters in the training text data are segmented. For example, the training text data is divided into single characters, and each divided character is labeled. For example: training text data: L'Oreal Moisturizing Eye Cream, character labeling: [[O], [L], [Ya], [Bao], [Wet], [Yan], [Bao], [Wet], [Yan]], then [O], [L], [Ya], [Bao], [Wet], [Yan], [Cream] are segmented to obtain multiple segment sequences as shown in Table 1. Then, each character in the segment sequence is text-matched with the preset entity name. In the case of a complete match, the corresponding entity type (entity category) is determined. For example, if the preset entity name is "L'Oreal", its corresponding entity category is "brand", and if the preset entity name is "L'Oreal Moisturizing Eye Cream", its corresponding entity category is "product", as long as it matches "L'Oreal" and "L'Oreal Moisturizing Eye Cream", the entity type can be obtained. For details, please refer to the content of Table 1.
[0068] Table 1
[0069]
[0070]
[0071] The training text data and corresponding entity types obtained above are used as training sample sets for subsequent model training.
[0072] In the embodiment of the present application, a method of remotely supervised segment sequence labeling of labeled data is adopted to solve the problem of entity type nesting.
[0073] As an optional embodiment, performing segment-wise splicing on character annotations to generate multiple segment sequences includes:
[0074] Determine a preset window length, wherein the preset window length is the maximum value of the total number of characters allowed to be contained in each segment sequence;
[0075] Within the range of a preset window length, the head character and the tail character contained in each segment are concatenated to obtain multiple segment sequences, wherein each segment contains at least one character.
[0076] Optionally, when segmenting the training text data, a preset window length can be set, which is the maximum value of the total number of characters allowed in each segment sequence, usually less than 10, such as set to 9, etc. Then, when segmenting, the head and tail characters included in each segment need to be concatenated within the preset window length. For example, segment sequence 1 in Table 1: "Ou", segment sequence 2: "Oulai", and so on, to obtain multiple segment sequences. It can be known that the number of characters included in each segment sequence is at least 1.
[0077] In the embodiment of the present application, since too long entities account for a relatively small proportion in the training samples, the length of the segment sequence is controlled by setting the maximum length of the segment sequence to reduce resource waste.
[0078] As an optional embodiment, after using the training text data and entity types as the training set, the method further includes:
[0079] Generating multiple hidden vectors corresponding to each segment sequence according to the training text data and the first model;
[0080] Inputting the multiple hidden vectors into the feed-forward neural network of the first model to obtain the first probability value of each hidden vector belonging to the entity type;
[0081] According to the first probability value, after a preset number of iterations, adjusting the first model parameters of the first model to obtain a second model;
[0082] Based on the second model and multiple segment sequences, adjusting the third model parameters of the third model to obtain a target model.
[0083] Optionally, use the pre-trained model BERT to construct an initial domain BERT_NER model, that is, the first model. Use the first model as the encoder, input the training text data into the first model, and obtain the hidden vector corresponding to each character, as Figure 3 shown. Then each segment sequence is composed of the combination (vector addition, subtraction, dot product) of the hidden vectors of the head and tail characters in the segment to generate multiple hidden vectors corresponding to each segment sequence; finally, use a feed-forward neural network as the classifier for the segment entity type, input the multiple hidden vectors corresponding to each segment sequence into the feed-forward neural network of the first model, and obtain the first probability value of each hidden vector belonging to the entity type.
[0084] The above content can be seen in Figure 3 , where E in the figure represents the hidden vector of the i-th character, H i represents the hidden vector of the i-th segment, and L i represents the label of the i-th segment sequence.
[0085] Based on the comparison between the first probability value and the entity type in the training set, if the label indicated by the first probability value is inconsistent with the entity type in the training set, the first model parameter of the first model is adjusted. At this time, the first model can be iterated for a preset number of times through the preset number of times obtained by the fourth model, thereby obtaining the second model (such as Figure 4 ).
[0086] Then, based on the obtained second model and multiple fragment sequences generated by training text data in the training set, the third model parameters of the third model are adjusted to obtain the target model.
[0087] As an optional embodiment, adjusting a third model parameter of a third model based on the second model and the plurality of segment sequences to obtain a target model includes:
[0088] Initializing a third model using the second model parameters of the second model, wherein the third model parameters in the current third model are equal to the second model parameters;
[0089] Input multiple latent vectors into the third model to obtain a reference probability value of each segment sequence belonging to the entity type;
[0090] The third model is trained using the mean square error loss function, and the third model parameters of the third model are adjusted until the reference probability value is greater than or equal to a preset threshold, so as to obtain a target model, wherein the preset threshold is the minimum value for stopping adjusting the third model parameters.
[0091] Optionally, the second model parameters in the second model are applied to the third model, that is, the BERT_NER model parameters obtained after a preset number of training iterations are initialized to the third model, and the third model parameters in the third model are set as the second model parameters.
[0092] Here, the third model is a teacher-student model. The teacher model can be used as the first sub-model of the third model, and the student model can be used as the second sub-model of the third model. The model structure of the teacher model and the student model is the same as that of the BERT_NER model.
[0093] Input multiple latent vectors into the first sub-model of the third model to obtain a second probability value; based on the second probability value, train the second sub-model of the third model using the mean square error loss function until a preset number of iterations are completed to obtain the second sub-model parameters of the trained second sub-model; use the second sub-model parameters to update the first sub-model parameters in the first sub-model to obtain an updated first sub-model; input multiple latent vectors into the updated first sub-model to obtain a third probability value; based on the third probability value, adjust the second sub-model parameters until the reference probability value output by the second sub-model is greater than or equal to the preset threshold, stop adjusting the second sub-model parameters, and obtain the target model.
[0094] The above is the self-learning training process of the third model, which can be found in Figure 4 : Use the second model parameters to fix the model parameters of the teacher model, and use the teacher model to predict the probability that each fragment sequence in the training set belongs to each entity type. Then use the teacher model to get the results, use the mean square error loss function to train the student model and update the student model parameters. After completing the preset number of iterations, use the model parameters of the student model to update the model parameters of the teacher model. Repeat the above steps N times until the last N-1 iterations obtain the reference probability value output by the student model that is greater than or equal to the preset threshold, then stop the iteration loop and use the current student model as the final NER model Finanl_NER.
[0095] In the embodiment of the present application, the NER model is trained in a self-learning manner to solve the problem of data missing labels.
[0096] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0097] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM (Read-Only Memory, Read-Only Memory) / RAM (Random Access Memory, Random Access Memory), a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of each embodiment of the present application.
[0098] According to another aspect of the embodiments of the present application, a device for entity identification for implementing the above-mentioned entity identification method is also provided. Figure 5 This is a structural block diagram of an optional entity recognition device according to an embodiment of the present application, such as Figure 5 As shown, the device may include:
[0099] The first acquisition unit 501 is used to acquire target text data to be recognized;
[0100] The first input unit 502 is connected to the first acquisition unit 501, and is used to input the target text data into the target model to obtain the target entity category to which the target text data belongs, wherein the target model is used to obtain the annotation information of the text data and identify the target entity category based on the annotation information. The target model is a final model obtained by adjusting the third model parameters of the third model. The third model is a model that uses the second model parameters in the second model to pre-train the training set. The second model is a model obtained after iterative training of the first model for a preset number of times. The preset number of times is obtained by processing the training set using the fourth model.
[0101] It should be noted that the first acquisition unit 501 in this embodiment can be used to execute the above step S201, and the first input unit 502 in this embodiment can be used to execute the above step S202.
[0102] Through the above modules, the trained target model is used as the final model for processing the entity category to which the target text data to be identified belongs, achieving the effect of automatic recognition, saving manpower, and improving recognition efficiency. The target model is completed through continuous training and parameter adjustment through self-learning of the first model, the second model, the third model, and the fourth model. It is more accurate in identifying entity categories and solves the problems of missed and incorrect data labels.
[0103] As an optional embodiment, the device further includes:
[0104] A second acquisition unit is used to acquire training text data before acquiring target text data to be recognized;
[0105] A splicing unit, configured to perform segment-wise splicing of characters in the training text data according to a preset scheme to generate multiple segment sequences;
[0106] A matching unit, configured to perform text matching between each character in the segment sequence and a preset entity name to determine the entity type to which the training text data belongs;
[0107] Set up the unit for training text data and entity types as training sets.
[0108] As an optional embodiment, the splicing unit includes:
[0109] The segmentation module is used to segment the training text data into single characters and perform character labeling on each segmented character;
[0110] The splicing module is used to perform fragment splicing on character annotations to generate multiple fragment sequences.
[0111] As an optional embodiment, the splicing module includes:
[0112] A determination subunit, configured to determine a preset window length, wherein the preset window length is the maximum value of the total number of characters allowed to be contained in each segment sequence;
[0113] The splicing subunit is used to splice the head character and the tail character contained in each segment within the range of a preset window length to obtain multiple segment sequences, wherein each segment contains at least one character.
[0114] As an optional embodiment, the device further includes:
[0115] A second input unit is configured to generate a plurality of latent vectors corresponding to each segment sequence according to the training text data and the first model after taking the training text data and the entity type as a training set;
[0116] A third input unit is configured to input the plurality of latent vectors into the feedforward neural network of the first model to obtain a first probability value of each latent vector belonging to an entity type;
[0117] A first adjustment unit is configured to adjust a first model parameter of the first model according to the first probability value through a preset number of iterations to obtain a second model;
[0118] The second adjustment unit is used to adjust the third model parameters of the third model based on the second model and the multiple segment sequences to obtain the target model.
[0119] As an optional embodiment, the second adjustment unit includes:
[0120] an initialization module, configured to initialize a third model using second model parameters of the second model, wherein the third model parameters in the current third model are equal to the second model parameters;
[0121] An input module, configured to input multiple latent vectors into a third model to obtain a reference probability value of each segment sequence belonging to an entity type;
[0122] The first adjustment module is used to train the third model using the mean square error loss function, adjust the third model parameters of the third model until the reference probability value is greater than or equal to a preset threshold, and obtain the target model, wherein the preset threshold is the minimum value for stopping adjusting the third model parameters.
[0123] As an optional embodiment, the first adjustment module includes:
[0124] A first input subunit is configured to input the plurality of latent vectors into the first submodel of the third model to obtain a second probability value;
[0125] a training subunit, configured to train the second sub-model of the third model based on the second probability value using a mean square error loss function until a preset number of iterations are completed, thereby obtaining second sub-model parameters of the trained second sub-model;
[0126] an updating sub-unit, configured to update the first sub-model parameters in the first sub-model using the second sub-model parameters to obtain an updated first sub-model;
[0127] A second input subunit is used to input the multiple latent vectors into the updated first sub-model to obtain a third probability value;
[0128] The second adjustment module is used to adjust the second sub-model parameters based on the third probability value until the reference probability value output by the second sub-model is greater than or equal to the preset threshold, stop adjusting the second sub-model parameters, and obtain the target model.
[0129] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments. Figure 1 The hardware environment shown can be implemented through software or hardware, wherein the hardware environment includes a network environment.
[0130] According to another aspect of the embodiments of the present application, an electronic device for implementing the above-mentioned entity recognition method is also provided. The electronic device may be a server, a terminal, or a combination thereof.
[0131] Figure 6 is a structural block diagram of an optional electronic device according to an embodiment of the present application, such as Figure 6 As shown, it includes a processor 601, a communication interface 602, a memory 603 and a communication bus 604, wherein the processor 601, the communication interface 602 and the memory 603 communicate with each other through the communication bus 604, wherein,
[0132] Memory 603, used to store computer programs;
[0133] The processor 601 is configured to execute the computer program stored in the memory 603 to implement the following steps:
[0134] Obtain target text data to be recognized;
[0135] The target text data is input into the target model to obtain the target entity category to which the target text data belongs, wherein the target model is used to obtain the annotation information of the text data and identify the target entity category based on the annotation information. The target model is a final model obtained by adjusting the third model parameters of the third model. The third model is a model that uses the second model parameters in the second model to pre-train the training set. The second model is a model obtained after iterative training of the first model for a preset number of times. The preset number of times is obtained by processing the training set using the fourth model.
[0136] Optionally, in this embodiment, the communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0137] The communication interface is used for communication between the above electronic device and other devices.
[0138] The memory may include RAM, or may include non-volatile memory, such as at least one disk memory. Alternatively, the memory may also be at least one storage device located away from the aforementioned processor.
[0139] As an example, Figure 6 As shown, the memory 603 may include, but is not limited to, the first acquisition unit 501 and the first input unit 502 in the entity recognition device. In addition, it may also include, but is not limited to, other module units in the entity recognition device, which will not be repeated in this example.
[0140] The above-mentioned processor can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0141] In addition, the electronic device further includes: a display for displaying the result of entity recognition.
[0142] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.
[0143] It can be understood by those skilled in the art that Figure 6 The structure shown is for illustration only. The device for implementing the above-mentioned entity recognition method may be a terminal device, which may be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, and other terminal devices. Figure 6 It does not limit the structure of the above electronic equipment. For example, the terminal device may also include Figure 6 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 6 Different configurations shown.
[0144] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which can include: a flash drive, ROM, RAM, a magnetic disk or an optical disk, etc.
[0145] According to another aspect of the embodiment of the present application, a storage medium is further provided. Optionally, in this embodiment, the storage medium can be used to execute program code of the entity recognition method.
[0146] Optionally, in this embodiment, the above-mentioned storage medium may be located on at least one network device among the multiple network devices in the network shown in the above-mentioned embodiment.
[0147] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps:
[0148] Obtain target text data to be recognized;
[0149] The target text data is input into the target model to obtain the target entity category to which the target text data belongs, wherein the target model is used to obtain the annotation information of the text data and identify the target entity category based on the annotation information. The target model is a final model obtained by adjusting the third model parameters of the third model. The third model is a model that uses the second model parameters in the second model to pre-train the training set. The second model is a model obtained after iterative training of the first model for a preset number of times. The preset number of times is obtained by processing the training set using the fourth model.
[0150] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, which will not be described in detail in this embodiment.
[0151] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a ROM, a RAM, a mobile hard disk, a magnetic disk, or an optical disk.
[0152] According to another aspect of the embodiments of the present application, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the method steps of entity recognition in any of the above embodiments.
[0153] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0154] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to perform all or part of the steps of the entity recognition method of each embodiment of the present application.
[0155] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, there may be other division methods, such as combining or integrating multiple units or components into another system, or ignoring or not implementing some features. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interface, indirect coupling or communication connection of units or modules, and may be electrical or other forms.
[0157] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the purpose of the solution provided in this embodiment.
[0158] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0159] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for entity recognition, characterized in that: The method comprises: Obtain target text data to be recognized; Inputting the target text data into a target model to obtain a target entity category to which the target text data belongs, wherein the target model is used to obtain annotation information of the text data and identify the target entity category according to the annotation information, the target model is a final model obtained by adjusting the third model parameters of the third model, the third model is a model that pre-trains the training set using the second model parameters in the second model, the second model is a model obtained by iteratively training the first model for a preset number of times, the preset number of times being obtained by processing the training set using the fourth model, the first model is a BERT_NER model, the second model is a BERT_NER model after multiple iterations, the third model is a teacher-student model, and the fourth model is an early_stop model; Wherein, when obtaining the target model: multiple fragment sequences are input into the first sub-model of the third model to obtain a second probability value, and the fragment sequence is obtained by fragment-wise splicing of character annotations of characters obtained by dividing the training text data. Based on the second probability value, the second sub-model of the third model is trained using the mean square error loss function until a preset number of iterations are completed to obtain the second sub-model parameters of the trained second sub-model; the first sub-model parameters in the first sub-model are updated using the second sub-model parameters to obtain the updated first sub-model; multiple latent vectors of the fragment sequence are input into the updated first sub-model to obtain a third probability value; based on the third probability value, the second sub-model parameters are adjusted until the reference probability value output by the second sub-model is greater than or equal to a preset threshold, and the adjustment of the second sub-model parameters is stopped to obtain the target model, the first sub-model of the third model is a teacher model, and the second sub-model of the third model is a student model.
2. The method according to claim 1, characterized in that Before obtaining the target text data to be recognized, the method further includes: Get training text data; Performing segment-wise splicing of characters in the training text data according to a preset scheme to generate multiple segment sequences; Performing text matching on each character in the segment sequence with a preset entity name to determine the entity type to which the training text data belongs; The training text data and the entity type are used as the training set.
3. The method according to claim 2, characterized in that The segment-wise splicing of characters in the training text data according to a preset scheme to generate a plurality of segment sequences comprises: Dividing the training text data into single-character forms and labeling each divided character; The character annotations are spliced in segments to generate a plurality of segment sequences.
4. The method according to claim 3, characterized in that The segment-wise splicing of the character annotations to generate a plurality of segment sequences includes: Determining a preset window length, wherein the preset window length is the maximum value of the total number of characters allowed to be contained in each of the segment sequences; Within the range of the preset window length, the head character and the tail character contained in each segment are spliced to obtain a plurality of segment sequences, wherein each segment contains at least one character.
5. The method according to claim 2, characterized in that After taking the training text data and the entity type as the training set, the method further includes: Generating a plurality of latent vectors corresponding to each of the segment sequences according to the training text data and the first model; Inputting the plurality of latent vectors into the feedforward neural network of the first model to obtain a first probability value of each latent vector belonging to the entity type; According to the first probability value, after the preset number of iterations, adjusting the first model parameters of the first model to obtain the second model; Based on the second model and the plurality of segment sequences, third model parameters of the third model are adjusted to obtain the target model.
6. The method according to claim 5, characterized in that The adjusting the third model parameters of the third model based on the second model and the plurality of fragment sequences to obtain the target model comprises: Initializing the third model using the second model parameters of the second model, wherein the third model parameters in the third model are currently equal to the second model parameters; Inputting the plurality of latent vectors into the third model to obtain a reference probability value of each of the segment sequences belonging to the entity type; The third model is trained using a mean square error loss function, and the third model parameters of the third model are adjusted until the reference probability value is greater than or equal to a preset threshold, so as to obtain the target model, wherein the preset threshold is the minimum value for stopping adjusting the third model parameters.
7. A device for entity recognition, characterized in that: The device comprises: A first acquiring unit, configured to acquire target text data to be recognized; a first input unit, configured to input the target text data into a target model to obtain a target entity category to which the target text data belongs, wherein the target model is configured to obtain annotation information of the text data and identify the target entity category based on the annotation information; the target model is a final model obtained by adjusting a third model parameter of a third model; the third model is a model that pre-trains a training set using the second model parameter in the second model; the second model is a model obtained by iteratively training the first model for a preset number of times; the preset number of times is obtained by processing the training set using a fourth model; the first model is a BERT_NER model; the second model is a BERT_NER model after multiple iterations; the third model is a teacher-student model; and the fourth model is an early_stop model; Wherein, when obtaining the target model: multiple fragment sequences are input into the first sub-model of the third model to obtain a second probability value, and the fragment sequence is obtained by fragment-wise splicing of character annotations of characters obtained by dividing the training text data. Based on the second probability value, the second sub-model of the third model is trained using the mean square error loss function until a preset number of iterations are completed to obtain the second sub-model parameters of the trained second sub-model; the first sub-model parameters in the first sub-model are updated using the second sub-model parameters to obtain the updated first sub-model; multiple latent vectors of the fragment sequence are input into the updated first sub-model to obtain a third probability value; based on the third probability value, the second sub-model parameters are adjusted until the reference probability value output by the second sub-model is greater than or equal to a preset threshold, and the adjustment of the second sub-model parameters is stopped to obtain the target model, the first sub-model of the third model is a teacher model, and the second sub-model of the third model is a student model.
8. An electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein: The processor, the communication interface and the memory communicate with each other via the communication bus, wherein: The memory is used to store computer programs; The processor is configured to execute the method steps according to any one of claims 1 to 6 by running the computer program stored in the memory.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program implements the method steps described in any one of claims 1 to 6 when executed by a processor.
Citation Information
Patent Citations
Training method and device of named entity recognition model
CN111523324A
Named entity recognition model training method, named entity recognition method and medium
CN111738003A
Named entity recognition model training method and device, equipment and storage medium
CN113836927A