Named entity recognition method and device and readable storage medium

By introducing span detection and distribution calibration modules to optimize the pre-trained language model, the problems of overfitting and class boundary deviation in small sample named entity recognition are solved, and efficient named entity recognition in the industrial chain and supply chain are achieved.

CN120509407AActive Publication Date: 2025-08-19SHENZHEN ACAD OF INSPECTION & QUARANTINE +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510405327.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-08-19
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In the industrial chain and supply chain, existing small sample named entity recognition methods are prone to overfitting or class boundary deviations, resulting in poor generalization capabilities of the model and the inability to accurately identify and extract named entities.

Method used

Parallel span detection and multi-task methods are introduced to separate entities from non-entities, combine the distribution calibration module, optimize the small sample naming and recognition model, and optimize the pre-trained language model, fine-tuning and feature calibration using data in the industrial chain and supply chain field to improve the generalization ability of the model.

Benefits of technology

The accuracy and generalization ability of naming entity recognition have been improved, and the naming entities related to the industrial chain and supply chain can be better identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509407A_ABST
    Figure CN120509407A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of natural language processing in the field of computer science and technology, and provides a named entity recognition method and device and a readable storage medium, and the method comprises the steps: obtaining first text information and a first category label; inputting the first text information and the first category label into an optimized pre-trained language model to obtain a plurality of first feature vectors and a plurality of second feature vectors; each second feature vector is calibrated, and a third feature vector corresponding to each second feature vector is obtained; and identifying the first text information according to the plurality of first feature vectors and the plurality of third feature vectors so as to identify a plurality of named entities in the first text information. According to the method, the named entities related to the industrial chain and the supply chain can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of natural language processing technology in computer science and technology, and in particular relates to a method, device and readable storage medium for named entity recognition. Background Art

[0002] The security of the "dual chain" of industrial and supply chains is fundamental to my country's high-quality economic development. Faced with multiple risks and challenges amid global industrial development trends, countries around the world are conducting relevant technical research. Knowledge graphs have been successfully applied in various fields, but research on industrial and supply chain graphs is relatively limited. "Dual chain" knowledge graph technology faces the challenge of diverse and complex data sources. Named entity recognition is key to knowledge graph construction. Within industrial and supply chains, the amount of labeled named entity recognition sample data is small and the cost of obtaining annotations is high. Therefore, there is an urgent need for small-sample named entity recognition systems.

[0003] Metric learning has become a common technique for small-sample named entity recognition. However, existing metric learning-based methods for small-sample named entity recognition have largely ignored the limitations imposed by the small sample size problem. When the number of support set samples is extremely small, overfitting or class boundary deviation are prone to occur, resulting in poor generalization and inability to accurately identify and extract named entities. Summary of the Invention

[0004] The embodiments of the present application provide a method, apparatus, terminal device, and computer-readable storage medium for named entity recognition, which can accurately identify named entities related to the industrial chain and supply chain.

[0005] In a first aspect, an embodiment of the present application provides a method for named entity recognition, comprising:

[0006] Acquire first text information and a first category label; wherein the first text information includes a plurality of first words; and the first category label represents a label corresponding to a plurality of named entities of different entity types in a target domain;

[0007] Inputting the first text information and the first category label into an optimized pre-trained language model to obtain a plurality of first feature vectors and a plurality of second feature vectors; wherein each first feature vector represents feature information corresponding to one of the first words in the first text information; and each second feature vector represents feature information corresponding to one of the entity types in the first category label;

[0008] calibrating each of the second eigenvectors to obtain a third eigenvector corresponding to each of the second eigenvectors;

[0009] The first text information is recognized based on the plurality of first feature vectors and the plurality of third feature vectors to recognize a plurality of named entities in the first text information.

[0010] In an embodiment of the present application, the first text information and the first category label are input into a pre-trained language model to obtain a plurality of first feature vectors (corresponding to the feature information of each first word in the first text information) and a plurality of second feature vectors (corresponding to the feature information of each entity type in the first category label). The pre-trained language model can extract rich semantic and grammatical information from the text and label and convert it into a feature vector representation. Each second feature vector is calibrated to obtain a plurality of third feature vectors. The calibration process is to adjust the category feature vector, improve the generalization ability of the model so that it can more accurately represent the corresponding entity type, so the above method can improve the accuracy of named entity recognition.

[0011] In a possible implementation of the first aspect, the training process of the pre-trained language model includes:

[0012] Obtain a first training set, the first training set including a plurality of second text information and a plurality of second category labels; wherein the second text information includes a plurality of second words; and the second category labels represent labels corresponding to named entities of a plurality of different entity types in N fields; wherein N>1;

[0013] For each second text information in the first training set, inputting the second text information into a language recognition model to be trained to obtain a fourth feature vector corresponding to each second word in the second text information;

[0014] Mapping the second category label through a mapping function to obtain a first semantic label corresponding to each entity type in the second category label;

[0015] Mapping the non-named entity and the named entity corresponding to the second category label respectively through a mapping function to obtain a second semantic label corresponding to each of the named entities and a third semantic label corresponding to each of the non-named entities;

[0016] The language recognition model to be trained is trained according to the third feature vector, the first semantic label, the second semantic label, and the third semantic label to obtain the pre-trained language model.

[0017] In a possible implementation of the first aspect, the training the language recognition model to be trained based on the third feature vector, the first semantic label, the second semantic label, and the third semantic label to obtain the pre-trained language model includes:

[0018] Inputting the first semantic label, the second semantic label, and the third semantic label into a language recognition model to be trained, obtaining a fifth feature vector corresponding to the first semantic label, a sixth feature vector corresponding to the second semantic label, and a seventh feature vector corresponding to the third semantic label;

[0019] Obtaining a first loss function according to the fourth eigenvector and the fifth eigenvector;

[0020] Obtaining a second loss function according to the fourth eigenvector, the sixth eigenvector, and the seventh eigenvector;

[0021] The language recognition model to be trained is trained according to the first loss function and the second loss function to obtain the pre-trained language model.

[0022] In a possible implementation of the first aspect, the training the to-be-trained language recognition model according to the first loss function and the second loss function to obtain the pre-trained language model includes:

[0023] Adding the first loss and the second loss to obtain a total loss;

[0024] If the total loss reaches the convergence condition, the current language recognition model is recorded as the pre-trained language model;

[0025] If the total loss does not reach the convergence condition, the parameters of the language recognition model to be trained are adjusted according to the total loss until the total loss function reaches the convergence condition, thereby obtaining the pre-trained language model.

[0026] In a possible implementation of the first aspect, the method further includes:

[0027] Obtaining a second training set, including third text information and a third category label in a target domain of the second training set; the third category label represents labels corresponding to respective named entities of a plurality of different entity types in the target domain;

[0028] The trained language recognition model is optimized according to the second training set to obtain the optimized pre-trained language model.

[0029] In a possible implementation manner of the first aspect, calibrating each second eigenvector to obtain multiple third eigenvectors includes:

[0030] Obtaining a label whose distance from the second feature vector is within a preset distance from the first label to obtain a third label;

[0031] generating an eighth feature vector according to the third label;

[0032] An average of the second eigenvector and the eighth eigenvector is calculated to obtain the third eigenvector.

[0033] In a possible implementation of the first aspect, identifying the first text information based on the multiple first feature vectors and the multiple third feature vectors to identify multiple named entities in the first text information includes:

[0034] calculating a first distance between each of the first eigenvectors and each of the third eigenvectors;

[0035] The entity category corresponding to the calculated minimum first distance is determined as the named entity.

[0036] In a second aspect, an embodiment of the present application provides a named entity recognition device, comprising:

[0037] An information acquisition module, configured to acquire first text information and a first category label, wherein the first text information includes a plurality of first words; and the first category label represents a plurality of different first entity labels in a target domain;

[0038] a feature generation module, configured to input the first text information and the first category label into an optimized pre-trained language model to obtain a plurality of first feature vectors and a plurality of second feature vectors; wherein each first feature vector represents feature information corresponding to one of the first words in the first text information; and each second feature vector represents feature information corresponding to one of the entity types in the first category label;

[0039] a feature calibration module, configured to calibrate each of the second feature vectors to obtain a third feature vector corresponding to each of the second feature vectors;

[0040] An entity recognition module is used to recognize the first text information based on the multiple first feature vectors and the multiple third feature vectors to recognize the multiple first named entities in the first text information.

[0041] In a third aspect, an embodiment of the present application provides a terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the named entity recognition method as described in any one of the first aspects above is implemented.

[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the named entity recognition method as described in any one of the above-mentioned first aspects is implemented.

[0043] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, enables the terminal device to execute the named entity recognition method described in any one of the above-mentioned first aspects.

[0044] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0046] Figure 1 Schematic diagram of the process of the named entity recognition method provided in the embodiment of the present application;

[0047] Figure 2 This is a schematic diagram of the model training process provided in the embodiment of this application Figure 1 ;

[0048] Figure 3 This is a schematic diagram of the model training process provided in the embodiment of this application Figure 2 ;

[0049] Figure 4 This is a schematic diagram of the model training process provided in the embodiment of this application Figure 3 ;

[0050] Figure 5 This is a schematic diagram of the optimization model provided in the embodiment of this application. Figure 1 ;

[0051] Figure 6 1 is a flow chart of feature calibration provided in an embodiment of the present application;

[0052] Figure 7 is a schematic diagram of the structure of the feature calibration provided in an embodiment of the present application;

[0053] Figure 8 This is a flowchart of named entity recognition provided by an embodiment of the present application;

[0054] Figure 9This is a schematic diagram of the overall structure of named entity recognition provided by an embodiment of the present application;

[0055] Figure 10 Schematic diagram of the training set for model training provided in the embodiment of the present application;

[0056] Figure 11 It is the experimental result of the model training provided by the implementation of this application;

[0057] Figure 12 This is a schematic diagram of the structure of a named entity recognition device provided by an embodiment of the present application;

[0058] Figure 13 This is a structural diagram of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0059] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0060] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0061] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0062] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0063] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0064] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized.

[0065] The security of the "dual chain" of industrial and supply chains is fundamental to my country's high-quality economic development. Faced with multiple risks and challenges amid global industrial development trends, countries around the world are conducting relevant technical research. Knowledge graphs have been successfully applied in various fields, but research on industrial and supply chain graphs is relatively limited. "Dual chain" knowledge graph technology faces the challenge of diverse and complex data sources. Named entity recognition is key to knowledge graph construction. Within industrial and supply chains, the amount of labeled named entity recognition sample data is small and the cost of obtaining annotations is high. Therefore, there is an urgent need for small-sample named entity recognition systems.

[0066] Metric learning has become a common technique for small-sample named entity recognition. However, existing metric learning-based methods for small-sample named entity recognition have largely ignored the limitations imposed by the small sample size problem. When the number of support set samples is extremely small, overfitting or class boundary deviation are prone to occur, resulting in poor generalization and inability to accurately identify and extract named entities.

[0067] In order to solve the problems in the above-mentioned related technologies, the embodiments of the present application provide a named entity recognition method, apparatus, terminal device and computer-readable storage medium. In this application, parallel span detection is introduced on the basis of metric learning to fully utilize the information contained in the sample, and entities and non-entities are separated in a multi-task manner without affecting the entity classification of the main part. At the same time, in order to deal with the overfitting and class boundary deviation problems of classifiers learned from very few support set samples, this article applies distribution accuracy to small sample naming recognition, filling the gap in optimizing small sample named entity recognition models in the prediction stage.

[0068] See also Figure 1 , is a flowchart of a method for named entity recognition provided in an embodiment of the present application. As an example and not a limitation, the method may include the following steps:

[0069] S101, obtaining first text information and a first category label; wherein the first text information includes a plurality of first words; and the first category label represents a label corresponding to each of a plurality of named entities of different entity types in a target domain.

[0070] In the embodiments of this application, the first text information is the original text content to be processed, which is usually from relevant text in the target field (such as the industrial chain and supply chain field). The first text information is composed of multiple first words, which are the basic units of the text. The subsequent named entity recognition work is to analyze these words and determine which entity type they belong to.

[0071] First-category labels represent the labels corresponding to named entities of different entity types in the target domain. There are many different types of entities in the target domain, and to classify and identify them, we assign a specific label to each entity type. These first-category labels clearly identify the identifiers corresponding to each possible entity type. During the subsequent recognition process, we match the words in the first text message with these labels to determine the entity type to which each word belongs.

[0072] As an example, assuming that the target field is the smartphone industry chain and supply chain, the first text information can be a news report about the production and sales of smartphones, such as: "Apple recently released a new iPhone 16 mobile phone, which uses Qualcomm's chip and sells well in the Chinese market." In this text, "Apple", "recently", "released", "new iPhone 16 mobile phone", "Qualcomm", "chip", "Chinese market", etc. are all first words, which together constitute the first text information.

[0073] In the target field of the smartphone industry chain supply chain, common entity types and their corresponding first-category labels are as follows:

[0074] Company: The label is "COM", such as "Apple" and "Qualcomm" belong to this entity type;

[0075] Products: Products labeled "PRO", "New iPhone 16 Phone", and "Chip" belong to this category;

[0076] Location: The label is "LOC", and the entity type corresponding to "China Market" is location.

[0077] By obtaining such first text information and first category labels, we can use subsequent models and algorithms to analyze the words in the first text information and match them with the first category labels, thereby identifying the entity type corresponding to each word, for example, determining that the entity type label of "Apple Inc." is "COM".

[0078] S102, inputting the first text information and the first category label into the optimized pre-trained language model to obtain multiple first feature vectors and multiple second feature vectors; wherein, each first feature vector represents feature information corresponding to one of the first words in the first text information; and each second feature vector represents feature information corresponding to one of the entity types in the first category label.

[0079] In an embodiment of the present application, the optimized pre-trained language model (hereinafter referred to as the "optimized recognition model") is obtained by combining the existing pre-trained model with the data in the field of industrial chain and supply chain and a specific training method. Taking the selected COPNER model as an example, pre-training is first performed on large-scale general text data to learn the general features and patterns of the language. Afterwards, the training data in the field of industrial chain and supply chain is used to fine-tune it through relevant methods. In this way, the model can better capture the characteristics of the industrial chain and supply chain text, understand the relationship between entities and vocabulary in this field, and lay the foundation for subsequent feature vector generation.

[0080] The first text information is text from the field of industrial chain and supply chain, containing multiple first words. After being input into the optimized recognition model, the optimized recognition model analyzes and processes each word. For example, for the text "Huawei produces 5G base station equipment in Dongguan", the optimization model will extract features from the semantics, grammar, and position of words in the text for words such as "Huawei", "Dongguan", and "5G base station" based on its internal neural network structure, word vector representation and other technologies, and integrate these features into a latent vector form, namely the first feature vector. Each first word has a corresponding first feature vector. The first feature vector of "Huawei" contains information such as its relevant semantics as an enterprise in the industrial chain and its role in the text; the first feature vector of "Dongguan" contains information such as its association with production activities as a location. These vectors provide a basis for subsequent judgment of the entity type to which the word belongs.

[0081] Similarly, the first category labels represent labels corresponding to different entity types in the industrial chain and supply chain field, such as "company", "place", "product", etc. The optimized recognition model takes these labels as input and also undergoes internal processing. The optimized recognition model will generate the corresponding second feature vector based on the semantics represented by the label and the characteristics in the field. For example, the second feature vector corresponding to the "place" label will contain feature information related to the geographical location and industrial distribution. These second feature vectors are a feature representation of different entity types by the model, which are used for subsequent comparison and matching with the first feature vector.

[0082] In one embodiment, see Figure 2 , is a schematic diagram of the model training process provided in the embodiment of this application Figure 1 ,like Figure 2As shown, the training process of the pre-trained language model in step S102 includes:

[0083] S1021, obtaining a first training set, wherein the first training set includes multiple second text information and multiple second category labels; wherein the second text information includes multiple second words; the second category labels represent labels corresponding to named entities of multiple different entity types in N fields; wherein N>1.

[0084] In this embodiment of the present application, the first training set is a dataset used to train a pre-trained speech recognition model. It consists of multiple second text messages and multiple corresponding second category labels. This data provides the model with learning samples, allowing it to gradually grasp the entity characteristics and patterns in different text messages during training, thereby improving its ability to recognize named entities.

[0085] Among them, the second text information in the first training set is text content from multiple fields (N fields, N>1), including text information from various fields such as finance, medical care, and education. These texts contain rich and diverse information. Each second text information is composed of multiple second words, which are the basic components of the text. And the second category label is used to mark the type of named entity that appears in the second text information. Since the second text information comes from N fields, the second category label covers the labels corresponding to named entities of different entity types in multiple fields.

[0086] As an example, the second text information can be information collected about people in a certain field, such as the collected information "Li Ming was born in 1970" (second text information), where "Li Ming", "birth", and "1970" all belong to the second vocabulary, and "Li Ming" corresponds to a named entity, whose corresponding named entity is "person", and its corresponding named entity label is "PER (person)", and "1970" corresponds to another named entity, whose corresponding named entity label is "DATE", etc. Different fields include multiple different text information and labels corresponding to multiple different categories of named entities.

[0087] S1022: For each second text information in the first training set, input the second text information into a language recognition model to be trained to obtain a fourth feature vector corresponding to each second word in the second text information.

[0088] In an embodiment of the present application, a non-entity separation module is provided. The module includes an entity recognition part. The part is based on COPNER (a language recognition model to be trained). After each second text information in the first training set is input into COPNER, a prototype of each second vocabulary word in each second text information is obtained, i.e., a fourth feature vector. The formula is as follows:

[0089] H=[h1,h2,…,h t ]=PLM([x1,x2,…,x t ]) (1)

[0090] Among them, t is the length of the sentence, x i is the i-th word in the sentence, and h1 is its corresponding fourth eigenvector.

[0091] S1023: Map the second category label through a mapping function to obtain a first semantic label corresponding to each entity type in the second category label.

[0092] In the embodiments of the present application, category labels are typically concise symbols or abbreviations that represent specific entity types. In named entity recognition tasks, to facilitate computer processing and recognition, people use simple codes to represent different entity categories, such as "LOC" representing a location entity. However, these concise labels lack clear semantic information, making it difficult for computers to directly understand the specific meaning of the entity type they represent.

[0093] Specifically, the labels of each entity type need to be mapped using a mapping function M. The purpose of the mapping function M is to convert these concise category labels into labels with clear semantics. For example, "LOC" is mapped to "location" (the first semantic label). This conversion makes the labels more readable and understandable, allowing the computer to better grasp the essential characteristics of each entity type based on this semantic information. Through the mapping function, the model can more accurately match labels with actual semantic concepts, laying the foundation for subsequent learning and recognition.

[0094] S1024: Map the non-named entities and the named entities corresponding to the second category labels respectively through a mapping function to obtain a second semantic label corresponding to each of the named entities and a third semantic label corresponding to each of the non-named entities.

[0095] In the embodiments of the present application, considering that the difference between entity and non-entity words should be greater than the difference between entities, a binary classification is performed in the Span Detection task in the non-entity separation module to determine whether the input word is an entity. In this way, while considering the noise of non-entity words, non-entity words and entity words are effectively separated.

[0096] Specifically, the second category label represents the labels corresponding to named entities of different entity types in N domains. For each named entity, the mapping function M is used to convert it into the corresponding second semantic label. For example, if entity types such as "Person (PER)", "Location (LOC)", "Date (DATE)" in the second category label are uniformly called "entity", these named entities can be uniformly mapped to the unified entity semantic label (second semantic label) through the mapping function.

[0097] In addition, non-named entities are words in the text that do not represent specific entities, such as some function words ("of", "in", "and", etc.), auxiliary words ("le", "ah", "ne", etc.) and some common notional words ("carry out", "occur", etc.). The mapping function M is also used to process them to obtain the third semantic label corresponding to each non-named entity. Usually, the third semantic label is represented by a unified identifier to indicate the non-entity attribute, such as "outside" commonly used.

[0098] S1025, training the to-be-trained language recognition model according to the fourth feature vector, the first semantic label, the second semantic label and the third semantic label to obtain the pre-trained language model.

[0099] In the embodiments of the present application, the fourth feature vector is the prototype corresponding to each second word in the second text information, and the first semantic label is obtained by mapping the second category label through the mapping function. Each first semantic label corresponds to an entity type. For example, mapping "LOC" to "location", these semantic labels can enable the model to more intuitively understand the meaning of each entity type, which helps the model learn the features and patterns of different entity types.

[0100] The second semantic label is the semantic label obtained by the entity word through the mapping function, which clarifies the entity type to which the entity word belongs and is represented by "entitiy"; the third semantic label is the semantic label corresponding to the non-entity word, usually represented by "OUTSIDE". Through these two semantic labels, the model can clearly distinguish entity words and non-entity words in the text, and then learn the feature differences between entities and non-entities. The to-be-trained model is trained according to the above-mentioned fourth feature vector and multiple semantic labels.

[0101] In one embodiment, see Figure 3 , is a schematic diagram of the model training process provided in the embodiment of this application Figure 2 ,like Figure 3 As shown, step S1025 includes:

[0102] S10251: Input the first semantic label, the second semantic label, and the third semantic label into a language recognition model to be trained to obtain a fifth feature vector corresponding to the first semantic label, a sixth feature vector corresponding to the second semantic label, and a seventh feature vector corresponding to the third semantic label.

[0103] In an embodiment of the present application, the mapped first semantic label, second semantic label, and third semantic label are input into the language recognition model to be trained, and a feature vector corresponding to each semantic label can be obtained.

[0104] Specifically, each category label is mapped to a semantic label, i.e., a second semantic label, through a mapping function M, such as Inputting it into COPNER will generate the prototype of the class (fifth eigenvector), the formula of which is as follows:

[0105] P=[p1,p2,…p N+1 ]=PLM([l1,l2,…,l N+1 ])(2)

[0106] Among them, N is the number of categories, plus the non-entity Outside class, a total of N+1 prototypes, l i Represents the label word corresponding to the i-th category after the mapping function, p i Represents the prototype (fifth eigenvector) corresponding to the i-th class.

[0107] Similarly, if we input the entity and non-entity labels as "entity" and "outside" respectively into COPNER, we will generate the prototypes of named entities and non-named entities (the sixth eigenvector and the seventh eigenvector), and the formula is as follows:

[0108] E=[e1,e2]=PLM([entity,outside])(3)

[0109] Among them, e1 and e2 are the prototypes of entity class and non-entity class respectively.

[0110] S10252: Obtain a first loss function according to the fourth eigenvector and the fifth eigenvector.

[0111] In the embodiment of the present application, during the training process, COPNER calculates the loss function based on the input data. The training goal of COPNER is to make the distance between the latent vector (fourth eigenvector) generated by the vocabulary and the correct prototype (fifth eigenvector) close and the distance between the incorrect prototype is far away. p The loss calculation formula is as follows:

[0112]

[0113] where p′ p (name) represents x p The corresponding correct prototype, d represents the Euclidean distance, and the calculation formula is Finally, the loss of the EntityTyping part is the sum of the losses of all words (the first loss function) and the formula is as follows:

[0114]

[0115] in, represents all second words in the first training set.

[0116] S10253: Obtain a second loss function according to the fourth eigenvector, the sixth eigenvector, and the seventh eigenvector.

[0117] In the embodiment of the present application, during the training process, COPNER calculates the loss function based on the input data. The training of COPNER is to make the hidden vector (fourth eigenvector) generated by the entity vocabulary close to the prototype of the named entity (sixth eigenvector) and to make it farther away from the prototype of the non-named entity (seventh eigenvector). p The loss calculation formula is as follows:

[0118]

[0119] Among them, e′ p When h p If it is an entity, it is e1, otherwise it is e2.

[0120] Similarly, the total loss function (second loss function) of the SpanDetection part is as shown in formula (7).

[0121]

[0122] S10254: Train the language recognition model to be trained according to the first loss function and the second loss function to obtain the pre-trained language model.

[0123] In one embodiment, see Figure 4, is a schematic diagram of the model training process provided in the embodiment of this application Figure 3 ,like Figure 4 As shown, the implementation of step S10254 is as follows:

[0124] S102541: Add the first loss and the second loss to obtain a total loss.

[0125] In the embodiment of the present application, the first loss L1 and the second loss L2 are added together to obtain the total loss, which is expressed as follows:

[0126]

[0127] The total loss comprehensively considers the errors in both entity type classification and entity and non-entity distinction, and can more comprehensively evaluate the performance of the model in the named entity recognition task.

[0128] S102542: If the total loss reaches a convergence condition, the current language recognition model is recorded as the pre-trained language model.

[0129] In the embodiments of the present application, the convergence condition is a criterion for determining whether the COPNER is sufficiently trained. It can typically be set to the point where the loss function value no longer decreases significantly, or reaches a preset threshold. When the total loss reaches the convergence condition, it means that the model has reached a relatively stable and optimal state under the current training data. At this point, the current speech recognition model can be recorded as a pre-trained language model.

[0130] S102543: If the total loss does not reach the convergence condition, the parameters of the language recognition model to be trained are adjusted according to the total loss until the total loss function reaches the convergence condition, thereby obtaining the pre-trained language model.

[0131] In an embodiment of the present application, if the total loss does not reach the convergence condition, it means that the model needs further learning and optimization. At this time, the parameters of the model will be adjusted using an optimization algorithm (such as stochastic gradient descent, etc.) based on the value of the total loss. The optimization algorithm will determine the direction and step size of the parameter adjustment based on the gradient information of the loss function, so that the model will continuously reduce the error in subsequent training and gradually approach the convergence condition. By continuously iterating this process, the model parameters are continuously adjusted until the total loss function finally reaches the convergence condition, thereby obtaining a pre-trained language model with good performance.

[0132] In one embodiment, see Figure 5 , is a flow chart of the optimization model provided in the embodiment of the present application, such as Figure 5 As shown, after obtaining the pre-trained language model in S102543, the method further includes:

[0133] S201, obtaining a second training set, third text information and a third category label in a target domain of the second training set; the third category label represents labels corresponding to a plurality of named entities of different entity types in the target domain.

[0134] In the embodiment of the present application, the second training set consists of third text information and third category labels in the target field. Among them, the target field specifically refers to the field of industrial chain and supply chain. The third text information is selected from the actual text data in this field. These texts contain various information related to the industrial chain and supply chain, and are composed of multiple words. They are text materials for further learning of the model. The third category labels clearly indicate the labels corresponding to the multiple named entities of different entity types that appear in these texts.

[0135] The purpose of obtaining a second training set is to optimize the initially trained language recognition model. Previously, the model may have been pre-trained on data from a general domain or a mix of multiple domains, but this training may not fully adapt to the specific needs of the industrial chain and supply chain. By using a second training set for the target domain, the model can learn the domain-specific language expressions, entity type characteristics, and the relationships between them.

[0136] S202: Optimize the trained language recognition model according to the second training set to obtain the optimized pre-trained language model.

[0137] In the embodiment of the present application, based on the above formula (8), a small number of samples from the target domain are used to fine-tune the model. Generally speaking, only K samples per category in the target domain can be used for fine-tuning. After fine-tuning, our model can be more suitable for the target domain, that is, the industrial chain supply chain field. After fine-tuning the pre-trained language model, an optimized pre-trained language recognition model is obtained.

[0138] S103: Calibrate each of the second eigenvectors to obtain a third eigenvector corresponding to each of the second eigenvectors.

[0139] In the embodiment of the present application, due to the small number of training sets or support sets in the target domain, i.e., the industrial chain, which may lead to overfitting and deviation from class boundaries, the embodiment of the present application provides a distribution calibration module to address the above problems. Distribution calibration attempts to calibrate the distribution of unknown categories in the visual domain by transferring the statistical data of known categories with sufficient samples to those with only a small number of samples.

[0140] After inputting the second category labels in the target domain into the optimized recognition model and obtaining the second feature vector corresponding to each named entity type, each second feature vector is further calibrated using the distribution calibration module to accurately identify the named entities in the first text information.

[0141] In one embodiment, see Figure 6 , is a flow chart of feature calibration provided in an embodiment of the present application, such as Figure 6 As shown, step S103 includes:

[0142] S1031: Acquire a label whose distance from the second feature vector is within a preset distance from the first labels to obtain a third label.

[0143] In the embodiment of the present application, for each second feature vector, that is, the prototype corresponding to a certain category of the second category label, see Figure 7 , is a schematic diagram of the structure of the feature calibration provided in the embodiment of the present application, such as Figure 7 As shown, the second eigenvector (solid star) has been generated. In order to calibrate the second eigenvector, it is necessary to obtain multiple category labels (third labels) of ungenerated prototypes near the second eigenvector (within a preset distance). Figure 7 As shown, the adjacent ones are multiple category labels (third labels) (solid triangles) for which prototypes are not generated. Figure 7 As shown, the second eigenvector (solid star) has been generated, and the multiple category labels (third labels) (solid triangle) nearby are not generated prototypes.

[0144] S1032: Generate an eighth feature vector according to the third label.

[0145] In the application embodiment, it is first assumed that the accuracy of the multiple third labels of the query samples gathered near the prototype generated by the model, that is, the second eigenvector, is high enough, then a pseudo label can be generated for the points near these prototypes, and the label value is the category label corresponding to these prototypes.

[0146] Specifically, we can use the above multiple point labels to generate a new prototype, namely the eighth eigenvector (hollow star), and its formula is as follows:

[0147]

[0148] Among them, n represents the nth class, Represents the new prototype corresponding to the nth class, i.e., the eighth eigenvector. R represents each class selecting R nearest neighbor samples (third label) to generate pseudo labels. Represents the hidden vector (feature vector corresponding to the pseudo label) of the query set sample that is closest to the i-th old prototype of the n-th old prototype.

[0149] S1033: Calculate an average value of the second eigenvector and the eighth eigenvector to obtain the third eigenvector.

[0150] In the embodiment of the present application, after generating a new prototype (eighth eigenvector), the final prototype P (third eigenvector after calibration) is the midpoint between the new and old prototypes, completing the feature calibration, and the formula is as follows.

[0151]

[0152] It should be noted that the distribution accuracy module is only used in the inference phase and does not participate in model training and fine-tuning, thereby reducing the complexity of the model.

[0153] S104: Identify the first text information according to the plurality of first feature vectors and the plurality of third feature vectors to identify a plurality of named entities in the first text information.

[0154] In an embodiment of the present application, named entities are recognized in the inference stage based on the calibrated feature vector and the first feature vector corresponding to each word in the input text information.

[0155] In one embodiment, see Figure 8 , is a flowchart of named entity recognition provided by an embodiment of the present application, such as Figure 8 As shown, step S104 includes:

[0156] S1041: Calculate a first distance between each of the first eigenvectors and each of the third eigenvectors.

[0157] In the embodiment of the present application, the distance (first distance) between the prototype (first eigenvector) corresponding to each word in each first text message and the prototype (third eigenvector) corresponding to all category labels can be calculated, and the formula is as follows:

[0158]

[0159] where h i represents the i-th word in the input sentence, p j represents the measure of the jth entity category after accurate distribution. Through the inference formula, we can find the entity category corresponding to the i-th word.

[0160] S1042: Determine the entity category corresponding to the calculated minimum first distance as the named entity.

[0161] In an embodiment of the present application, after calculating the distance between the vocabulary feature vector (first feature vector) and all entity type feature vectors (third feature vectors), these distance values are compared to find the smallest distance, which is called the "smallest first distance".

[0162] The entity type feature vector corresponding to the smallest first distance represents the entity type most similar to the features of the vocabulary to be recognized. Therefore, the entity category corresponding to this smallest distance is determined as the named entity category represented by the vocabulary to be recognized.

[0163] See also Figure 9 , is a schematic diagram of the overall structure of named entity recognition provided by the embodiment of the present application, such as Figure 9 As shown, the steps are as follows;

[0164] 1) Source domain training

[0165] The pre-trained language model is trained using text information from multiple fields, such as "Li Ming was born in 1970". The non-entity separation module is used to separate named entities and non-named entities and classify them. For example, the named entity "Li Ming" has a corresponding classification label of "PER (person)" and the named entity "1970" has a corresponding classification label of "DATE (date)". After multiple iterative training, the trained pre-trained language model is obtained.

[0166] 2) Fine-tuning using the support set

[0167] The pre-trained language model is optimized using the target domain, i.e., product chain and supply chain, and is adjusted and optimized using text information in the target domain, such as "Microsoft was founded in Redmond," to obtain an optimized pre-trained language model.

[0168] 3) Reasoning in the target domain

[0169] In the inference stage, the text information that needs to be inferred in the target domain is input into the optimized pre-trained language model, and the prototype of a certain category entity generated by the optimized pre-trained language model, i.e., the feature vector, is calibrated using the distribution calibration module to obtain the calibrated entity category feature vector. The distance between the calibrated entity category feature vector and the feature vector of the entity vocabulary in the text is measured to accurately identify the named entities in the text information to be inferred.

[0170] This method of determining named entity categories by calculating distance is based on the principle of similarity between feature vectors. In named entity recognition tasks, the model effectively leverages the learned feature information to more accurately determine the entity type to which words in a text belong. This allows for the identification and classification of named entities in the text, providing a foundation for subsequent natural language processing tasks such as information extraction and knowledge graph construction.

[0171] The following are the experimental results of this application:

[0172] a. Dataset and Experimental Settings

[0173] See also Figure 10 , is a schematic diagram of the training set for the model training provided in the embodiment of the present application, such as Figure 10 As shown, this application selected Few-NERD, OntoNotes, CoNLL, WNUT and I2B2 as data sets to verify the effectiveness of the model.

[0174] Currently, two experimental settings are generally used in the field of small-sample named entity recognition. One is the cross-label setting. For this setting, this paper uses the Few-NERD dataset to evaluate the model, which contains two different tasks, Few-NERD (INTER) and Few-NERD (INTRA). For INTER, all fine-grained entity categories are disjoint in the training set, development set, and test set, while the coarse-grained categories are shared. For INTRA, fine-grained entity categories in different sets belong to different coarse-grained categories. Due to the limitation of shared coarse-grained types, Few-NERD (INTRA) is more challenging. The other setting is domain transfer. The focus of this setting is to transfer a NER model to a new domain. Specifically, this paper trains the model on the OntoNotes 5.0 dataset from the general domain and evaluates it on the test sets of CoNLL'03, WNUT'17, and I2B2'14 from the news, social, and medical fields. The support set of the target domain is provided by Yang and Katiyar (2020). The characteristics of the datasets mentioned above are as follows: Figure 11 The following is a description of the data set provided in the embodiment of the present application.

[0175] b. Evaluation indicators

[0176] In the field of named entity recognition, F1-Score is generally used as an evaluation indicator.

[0177]

[0178] in Indicates the percentage of correct entities output by the model, Indicates the percentage of correct entities found by the model. In named entity recognition tasks, both precision and recall are important metrics. Considering only one of them is not feasible, so we use the F1-Score, which takes both metrics into consideration, as the evaluation metric.

[0179] Experimental results

[0180] Crossover trial results

[0181] See also Figure 11 , is the experimental result of the model training provided by this application, from Figure 11 As can be seen from a in Figure 1, under the Few-NERD (INTER) setting, the model proposed in this paper achieves the best results in three of the four specific small-sample settings, surpassing all one-stage methods. For the two-stage method, it fails to surpass ESD in the 5-way 5-shot setting, but still achieves SOTA results in the other three settings.

[0182] from Figure 11 As can be seen from Figure 2, under the Few-NERD (INTRA) setting, the proposed model achieves the best results in three of the four specific small-sample settings. Like the Inter setting, it surpasses all one-stage methods. For the two-stage method, it fails to surpass DecomMETA in the 10-way 5-shot setting, but achieves the best results in the other three settings.

[0183] for Figure 11 c in and Figure 11 As can be seen from the figure, the model proposed in this paper has achieved the best results in the domain transfer problem, especially in the 5-shot setting, surpassing the most SOTA algorithm by an average of 8.1%.

[0184] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0185] Corresponding to the named entity recognition method described in the above embodiment, Figure 12 This is a structural block diagram of the named entity recognition device provided in an embodiment of the present application. For the sake of convenience, only the parts related to the embodiment of the present application are shown.

[0186] Reference Figure 12 , the device comprises:

[0187] The model training module 121 is used to:

[0188] Obtain a first training set, the first training set including a plurality of second text information and a plurality of second category labels; wherein the second text information includes a plurality of second words; and the second category labels represent labels corresponding to named entities of a plurality of different entity types in N fields; wherein N>1;

[0189] For each second text information in the first training set, inputting the second text information into a language recognition model to be trained to obtain a fourth feature vector corresponding to each second word in the second text information;

[0190] Mapping the second category label through a mapping function to obtain a first semantic label corresponding to each entity type in the second category label;

[0191] Mapping the non-named entity and the named entity corresponding to the second category label respectively through a mapping function to obtain a second semantic label corresponding to each of the named entities and a third semantic label corresponding to each of the non-named entities;

[0192] The language recognition model to be trained is trained according to the fourth feature vector, the first semantic label, the second semantic label, and the third semantic label to obtain the pre-trained language model.

[0193] The model optimization module 122 is used to:

[0194] Obtaining a second training set, including third text information and a third category label in a target domain of the second training set; the third category label represents labels corresponding to respective named entities of a plurality of different entity types in the target domain;

[0195] The trained language recognition model is optimized according to the second training set to obtain the optimized pre-trained language model.

[0196] The information acquisition module 123 is configured to acquire first text information and a first category label, wherein the first text information includes a plurality of first words; and the first category label represents a plurality of different first entity labels in the target domain.

[0197] a feature generation module 124 configured to input the first text information and the first category label into an optimized pre-trained language model to obtain a plurality of first feature vectors and a plurality of second feature vectors; wherein each first feature vector represents feature information corresponding to one of the first words in the first text information; and each second feature vector represents feature information corresponding to one of the entity types in the first category label;

[0198] A feature calibration module 125 is configured to calibrate each of the second feature vectors to obtain a third feature vector corresponding to each of the second feature vectors;

[0199] The entity recognition module 126 is configured to recognize the first text information based on the plurality of first feature vectors and the plurality of third feature vectors, so as to recognize a plurality of first named entities in the first text information.

[0200] Optionally, the model training module 121 is further configured to:

[0201] Inputting the first semantic label, the second semantic label, and the third semantic label into a language recognition model to be trained, obtaining a fifth feature vector corresponding to the first semantic label, a sixth feature vector corresponding to the second semantic label, and a seventh feature vector corresponding to the third semantic label;

[0202] Obtaining a first loss function according to the fourth eigenvector and the fifth eigenvector;

[0203] Obtaining a second loss function according to the fourth eigenvector, the sixth eigenvector, and the seventh eigenvector;

[0204] The language recognition model to be trained is trained according to the first loss function and the second loss function to obtain the pre-trained language model.

[0205] Optionally, the model training module 121 is further configured to:

[0206] Adding the first loss and the second loss to obtain a total loss;

[0207] If the total loss reaches the convergence condition, the current language recognition model is recorded as the pre-trained language model;

[0208] If the total loss does not reach the convergence condition, the parameters of the language recognition model to be trained are adjusted according to the total loss until the total loss function reaches the convergence condition, thereby obtaining the pre-trained language model.

[0209] Optionally, the feature calibration module 125 is further configured to:

[0210] Obtaining a label whose distance from the second feature vector is within a preset distance from the first label to obtain a third label;

[0211] generating an eighth feature vector according to the third label;

[0212] An average of the second eigenvector and the eighth eigenvector is calculated to obtain the third eigenvector.

[0213] Optionally, the entity identification module 126 is further configured to:

[0214] calculating a first distance between each of the first eigenvectors and each of the third eigenvectors;

[0215] The entity category corresponding to the calculated minimum first distance is determined as the named entity.

[0216] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of this application. Their specific functions and technical effects can be found in the method embodiment section and will not be repeated here.

[0217] in addition, Figure 12 The named entity recognition device shown can be a software unit, a hardware unit, or a combination of software and hardware units built into an existing terminal device, or can be integrated into the terminal device as an independent accessory, or can exist as an independent terminal device.

[0218] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0219] Figure 13 This is a schematic diagram of the structure of the terminal device provided in the embodiment of the present application. Figure 13 As shown, the terminal device 13 of this embodiment includes: at least one processor 130 ( Figure 13 Only one is shown in the figure) a processor, a memory 131, and a computer program 132 stored in the memory 131 and capable of running on the at least one processor 130, and when the processor 130 executes the computer program 132, the steps in any of the above-mentioned named entity recognition method embodiments are implemented.

[0220] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that Figure 13 This is merely an example of the terminal device 13 and does not constitute a limitation on the terminal device 13 . The terminal device 13 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device 13 may also include input and output devices, network access devices, etc.

[0221] The processor 130 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.

[0222] In some embodiments, the memory 131 may be an internal storage unit of the terminal device 13, such as a hard disk or memory of the terminal device 13. In other embodiments, the memory 131 may also be an external storage device of the terminal device 13, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 13. Furthermore, the memory 131 may also include both an internal storage unit of the terminal device 13 and an external storage device. The memory 131 is used to store an operating system, an application program, a boot loader, data, and other programs, such as the program code of the computer program. The memory 131 may also be used to temporarily store data that has been output or is about to be output.

[0223] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0224] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0225] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device capable of carrying the computer program code to the device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, a computer-readable medium cannot be an electric carrier signal or a telecommunication signal.

[0226] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0227] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0228] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0229] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0230] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for named entity recognition, characterized in that: The method comprises: Acquire first text information and a first category label; wherein the first text information includes a plurality of first words; and the first category label represents a label corresponding to a plurality of named entities of different entity types in a target domain; Inputting the first text information and the first category label into an optimized pre-trained language model to obtain a plurality of first feature vectors and a plurality of second feature vectors; wherein each first feature vector represents feature information corresponding to one of the first words in the first text information; and each second feature vector represents feature information corresponding to one of the entity types in the first category label; calibrating each of the second eigenvectors to obtain a third eigenvector corresponding to each of the second eigenvectors; The first text information is recognized based on the plurality of first feature vectors and the plurality of third feature vectors to recognize a plurality of named entities in the first text information.

2. The named entity recognition method according to claim 1, wherein The training process of the pre-trained language model includes: Obtain a first training set, the first training set including a plurality of second text information and a plurality of second category labels; wherein the second text information includes a plurality of second words; and the second category labels represent labels corresponding to named entities of a plurality of different entity types in N fields; wherein N>1; For each second text information in the first training set, inputting the second text information into a language recognition model to be trained to obtain a fourth feature vector corresponding to each second word in the second text information; Mapping the second category label through a mapping function to obtain a first semantic label corresponding to each entity type in the second category label; Mapping the non-named entity and the named entity corresponding to the second category label respectively through a mapping function to obtain a second semantic label corresponding to each of the named entities and a third semantic label corresponding to each of the non-named entities; The language recognition model to be trained is trained according to the fourth feature vector, the first semantic label, the second semantic label, and the third semantic label to obtain the pre-trained language model.

3. The named entity recognition method according to claim 2, wherein: The training of the language recognition model to be trained according to the third feature vector, the first semantic label, the second semantic label, and the third semantic label to obtain the pre-trained language model includes: Inputting the first semantic label, the second semantic label, and the third semantic label into a language recognition model to be trained, obtaining a fifth feature vector corresponding to the first semantic label, a sixth feature vector corresponding to the second semantic label, and a seventh feature vector corresponding to the third semantic label; Obtaining a first loss function according to the fourth eigenvector and the fifth eigenvector; Obtaining a second loss function according to the fourth eigenvector, the sixth eigenvector, and the seventh eigenvector; The language recognition model to be trained is trained according to the first loss function and the second loss function to obtain the pre-trained language model.

4. The named entity recognition method according to claim 3, wherein: The training of the language recognition model to be trained according to the first loss function and the second loss function to obtain the pre-trained language model includes: Adding the first loss and the second loss to obtain a total loss; If the total loss reaches the convergence condition, the current language recognition model is recorded as the pre-trained language model; If the total loss does not reach the convergence condition, the parameters of the language recognition model to be trained are adjusted according to the total loss until the total loss function reaches the convergence condition, thereby obtaining the pre-trained language model.

5. The named entity recognition method according to claim 4, wherein: The method further comprises: Obtaining a second training set, including third text information and a third category label in a target domain of the second training set; the third category label represents labels corresponding to respective named entities of a plurality of different entity types in the target domain; The trained language recognition model is optimized according to the second training set to obtain the optimized pre-trained language model.

6. The method for named entity recognition according to claim 4, wherein: The step of calibrating each of the second eigenvectors to obtain a plurality of third eigenvectors includes: Obtaining a label whose distance from the second feature vector is within a preset distance from the first label to obtain a third label; generating an eighth feature vector according to the third label; An average of the second eigenvector and the eighth eigenvector is calculated to obtain the third eigenvector.

7. The named entity recognition method according to claim 1, wherein: The identifying the first text information according to the plurality of first feature vectors and the plurality of third feature vectors to identify the plurality of named entities in the first text information includes: calculating a first distance between each of the first eigenvectors and each of the third eigenvectors; The entity category corresponding to the calculated minimum first distance is determined as the named entity.

8. A named entity recognition device, characterized in that: include: An information acquisition module, configured to acquire first text information and a first category label, wherein the first text information includes a plurality of first words; and the first category label represents a plurality of different first entity labels in a target domain; a feature generation module, configured to input the first text information and the first category label into an optimized pre-trained language model to obtain a plurality of first feature vectors and a plurality of second feature vectors; wherein each first feature vector represents feature information corresponding to one of the first words in the first text information; and each second feature vector represents feature information corresponding to one of the entity types in the first category label; a feature calibration module, configured to calibrate each of the second feature vectors to obtain a third feature vector corresponding to each of the second feature vectors; An entity recognition module is used to recognize the first text information based on the multiple first feature vectors and the multiple third feature vectors to recognize the multiple first named entities in the first text information.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Named entity identification method and device

    CN110162772A

  • Survey learning-based small sample named entity identification method and equipment in surveying and mapping field

    CN118395984A

  • Small sample named entity recognition method based on metric learning

    CN118536509A

  • Model training method and apparatus

    US20250021888A1