Zero-sample named entity recognition method and device based on large model feedback optimization

By iterating the entity extraction, filtering and type matching strategies to optimize the big model, the problem of insufficient recall and accuracy in zero-sample named entity recognition is solved, and efficient named entity recognition in complex text environments is achieved.

CN120354855AActive Publication Date: 2025-07-22SHANDONG UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510855236.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-07-22
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

The prior art has insufficient recall and accuracy in naming entity recognition tasks in complex or professional texts, making it difficult to achieve accurate identification and classification, especially in zero-sample scenarios.

Method used

It adopts iterative entity extraction strategy, entity filtering strategy and entity type matching strategy to gradually improve recall and accuracy through large-scale feedback optimization.

Benefits of technology

It significantly improves the coverage and accuracy of named entity recognition, adapts to complex text environments, has good scalability and model universality, and adapts to open field tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354855A_ABST
    Figure CN120354855A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of electric digital data processing, and particularly relates to a zero-sample named entity recognition method and device based on large model feedback optimization. According to the method, an entity extraction strategy is iterated to force a large model to pay attention to other parts, except known entities, in a to-be-extracted text, so that the recall rate of entity recognition is increased; in order to give consideration to the recall rate and the accuracy rate, the method further provides an entity filtering strategy based on a large model, and screening of candidate entities is achieved. After screening, a large model is guided to focus on an entity type classification task, and classification errors possibly generated in the extraction process are effectively corrected; meanwhile, a text-based entity filtering strategy is introduced in the entity extraction process, so that invalid or wrong entities caused by a large model illusion problem are inhibited, and the accuracy and stability of overall recognition are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electronic digital data processing, and specifically relates to a zero-shot named entity recognition method and device based on large model feedback optimization. Background Art

[0002] A named entity refers to one or more entity names that always represent a certain entity with a specific meaning. Named entity recognition originally referred to the recognition of person names, place names, organization names, etc. in text. However, with the in-depth integration of technology research and industry needs, the definition scope of named entities has gradually extended bidirectionally. Vertically, there is a differentiation at the basic concept level. For example, from the general "person name" it is refined into occupational attribute classifications such as "researcher" and "biologist". At the same time, new entity types are horizontally extended, such as domain-specific entity types like "compound" and "protein". By accurately extracting entity information in text, it provides structured data support for knowledge graph construction, question answering, information retrieval, etc. It has important practical value in vertical fields such as biomedical literature mining, financial risk prediction, and judicial document structuring. However, such datasets often have high annotation costs, great difficulties, and the annotation process is cumbersome and time-consuming and relies on domain knowledge. These problems have prompted researchers to start studying the problem of named entity recognition in zero-shot scenarios.

[0003] The zero-shot named entity recognition task depends on the development of the large language model (LLM). Since ChatGPT was introduced in 2023, its demonstrated rich knowledge reserve, strong generalization ability, and excellent context understanding ability have prompted scholars to study its learning and reasoning abilities in zero-shot scenarios. Currently, the main perspectives for solving the zero-shot named entity recognition task based on large models include two types. One is to mine the internal knowledge of the large model through prompts or system design, and the other is to introduce external knowledge to focus on or supplement the internal knowledge of the large model. The former can adapt to zero-shot named entity recognition tasks in different fields by reserving category options in the prompt template, but it relies on a large number of repeated prompt engineering attempts. The latter, by introducing external knowledge, makes up for the poor timeliness of the large model's knowledge and enhances its performance in professional fields. However, its effect is limited by the quality of external knowledge and the efficiency of the retrieval algorithm.

[0004] The relevant cutting-edge work on the problem of professional term extraction mainly includes: Chinese Patent CN116245104A proposes a zero-shot named entity recognition method enhanced by external knowledge. This method obtains sentences containing the target entity category name from an external knowledge base, extracts the category semantic representation using a preprocessing model, and calculates the semantic similarity with the entity to be recognized to determine the entity category. Chinese Patent CN118114675A discloses a medical named entity recognition method based on a large language model. This method identifies candidate entity categories under various prompts, extracts arguments in combination with knowledge texts, and evaluates the correctness of various viewpoints to determine the entity category.

[0005] Overall, the above prompt design and task process are relatively concise. When dealing with complex, ambiguous, or highly professional texts, the above methods all face challenges. The semantic ambiguity, polysemy, and domain-specific knowledge and expressions in complex texts make it difficult to determine the entity boundaries and category judgments, making it difficult to achieve accurate recognition and classification, and restricting their performance in deep applications in professional fields and complex scenarios. Summary of the Invention

[0006] In view of the deficiencies of the prior art, the present invention discloses a zero-shot named entity recognition method optimized based on large model feedback. The method improves the recall rate of entity recognition by introducing an iterative entity extraction strategy to force the large model to focus on the remaining part of the text to be extracted except for the known entities. To balance the recall rate and precision, the method further proposes an entity filtering strategy based on the large model to screen candidate entities. After screening, by guiding the large model to focus on the entity type classification task, it can effectively correct possible classification errors in the extraction process.

[0007] The present invention also discloses a device for implementing the above method.

[0008] The present invention also discloses an electronic device for implementing the above method.

[0009] The present invention also discloses a machine-readable storage medium for implementing the above method.

[0010] To achieve the above object, the present invention adopts the following technical solutions: A zero-shot named entity recognition method optimized based on large model feedback, the method comprising: S1: Adopt an iterative entity extraction strategy based on a large model and prompt learning, and use entity extraction prompts to extract entities in the input text X contained therein to obtain an entity set ; S2: Adopt an entity filtering strategy based on a large model, and use entity screening prompts to the entity set obtained in step S1 Perform filtering to obtain the set of true entities E that meet the definition of named entity recognition, where , (1) represents the -th entity in the entity set , represents the size of the entity set, is the sequential index of the entities in the entity set E; S3: Adopt an entity type matching strategy based on a large model, and use the entity type matching hint to perform entity type matching on each named entity in the set of true entities E that meet the definition of named entity recognition obtained in step S2, and form a structured classification output.

[0011] Preferably, the entity extraction hint described in step S1 further includes a category set C module and a category explanation module. The category set C contains a predefined category set, and the category explanation is used to explain and describe each category; More preferably, the entity extraction hint further includes a task description module, which is used to introduce the modules included in the entity extraction hint and the meaning of each module, aiming to enable the large model to understand the specific utility and task requirements of each part; an explanation module, which is used to specifically describe the tasks that the large model should perform, aiming to enable the large model to extract as many entities as possible; an output module, which requires the large model to output in a specified format.

[0012] Preferably, step S1 is specifically: S11. Input the entity extraction hint into the large model to obtain an extraction result , and perform text-based entity screening on the obtained extraction result to obtain a screened extraction result ; S12. Determine whether there are new entities in the screened extraction result obtained in step S11. If there are no new entities, that is , terminate the iteration and output the entity set , where is the extraction result after the -th round of entity screening, is the entity set composed of all entities in the previous i - 1 rounds. For the first round, is initialized as an empty set; if there are new entities, then execute the masking strategy based on to obtain the input text , execute step S13, and continue the iteration; S13. Replace the input text obtained in step S12 in the entity extraction prompt with the corresponding content of the input text to obtain the entity extraction prompt after replacement, and repeat steps S11 and S12

[0013] More preferably, the entity screening based on the text for the obtained extraction result in step S11 is specifically as follows Filter according to whether the extraction result exists in the input text X. The formula is as follows (2) where is the extraction result after the th round of entity screening represents the elements in the set and represents the input text. The whole formula (2) means that for each element in , only if it satisfies existing in the input text will it be used as an element of the set

[0014] More preferably, the output entity set in step S12 , (3) where is the entity collection of the first i rounds is the entity collection of the first i - 1 rounds is the extraction result after the th round of entity screening The execution of the masking strategy based on in step S12 to obtain the input text is specifically as follows In the th iteration, according to the entity extraction results of the previous rounds mask the input text : Traverse the entities in and replace the entities in the input text X corresponding to those in with "[MASK]" to obtain the input text , as shown in formula (4) (4) where is the input text obtained after masking, and e represents the set​ The elements in " indicate replacement. The meaning of the whole formula is to traverse the elements in, and replace the entities existing in the original input text with the identifier "[MASK]".

[0015] Preferably, the entity screening prompt described in step S2 includes a named entity definition module for normalizing the description of named entities; an output format module that contains the entry "explanation" for explaining the reason for determining whether the extracted entity is a named entity.

[0016] Preferably, step S3 is specifically: combining each entity in the set of real entities E that conform to the named entity recognition definition obtained in step S2 with the existing category set C to form an entity type matching prompt , and providing as input to the large model to obtain the extraction result of each entity, that is, the matching category , and judging whether the extraction result exists in the category set C, as shown in formula (5) (5) Where is the output result after category screening, represents the entity category corresponding to a certain entity, is the extraction result, and C is the existing category set.

[0017] If the output result is an empty set, it indicates that the entity does not correspond to any category in the category set, and it is discarded; otherwise, select the model output category as the final category of the entity; When the constraint condition for ending the traversal is not satisfied: (6), Where is the sequential index of the entity in the entity set E, is the number of entities in the entity set E; Continue to loop through the entity set E and replace the entity , until the constraint condition is satisfied. After looping through all entities, each identified entity will finally be assigned a clear category to form a structured classification output.

[0018] Further preferably, the entity type matching prompt also includes a category explanation module for judging which category in the "category set" the entity extracted by the model is closest to, so as to verify the result of the entity extraction module.

[0019] In another aspect of the present invention, there is provided an apparatus for a zero-shot named entity recognition method optimized based on large model feedback. The apparatus includes: Prompt module: including entity extraction prompts , entity screening prompts , entity type matching prompts , respectively used for entity extraction, screening, and type matching; where aims to extract as many entities existing in the text as possible, aims to filter out the incorrect entities extracted by the entity extraction module based on coverage, aims to match the entity and type as accurately as possible; Entity extraction module based on coverage: built on the basis of large models and prompt learning, used to design entity extraction prompts according to expert knowledge and supplemented by a category set and category explanations to guide the large model to extract as many words or phrases that may be named entities in the input text as possible; Entity type matching module based on accuracy: through entity type matching prompts performs secondary verification on the extracted entities, performs type matching one by one, and allows the large model to select the most suitable type for the current entity from a large number of category sets, so as to ensure that the recognized entities meet the expected types and can be flexibly adapted to named entity recognition tasks in different fields, thereby enhancing the scalability of the method.

[0020] In another aspect of the present invention, there is also provided an electronic device, including: At least one processor; and, A memory that stores instructions, and when the instructions are executed by the at least one processor, the at least one processor executes the zero-shot named entity recognition method optimized based on large model feedback as described above.

[0021] In another aspect of the present invention, there is also provided a machine-readable storage medium that stores executable instructions, and when the instructions are executed, the machine executes the zero-shot named entity recognition method optimized based on large model feedback as described above.

[0022] Compared with the prior art, the beneficial effects of the present invention are: (1) From the perspective of entity evaluation, the present invention constructs an overall zero-shot named entity recognition method optimized based on large model feedback. By iteratively linking the two modules of entity extraction and type matching, a closed-loop entity recognition and correction mechanism is formed. Compared with the extraction process in the prior art that lacks a feedback channel and has a fixed structure, the present invention takes into account both semantic completion ability and output stability, and can effectively adapt to unannotated resources and complex context environments. While improving the recognition coverage and accuracy, this method has good scalability and model generality, demonstrating the generalization ability and application practicability for the named entity recognition task in the open domain.

[0023] (2) The present invention introduces a mask-based entity extraction strategy driven by large model feedback. Aiming at the problem of insufficient entity coverage in zero-shot named entity recognition, combined with the large model's understanding ability of language semantic structure, a round-by-round feedback completion mechanism is constructed. Compared with the prior methods that rely on static prompts or one-time extraction, based on the first extraction, the present invention uses entity omission prompts to generate supplementary masks and triggers the recognition process again, effectively expanding the boundary recognition ability and semantic extension ability of entities. This technical solution can significantly improve the extraction coverage rate of complex sentence patterns and implicit entities in the zero-shot scenario, and conforms to the regular characteristics of language redundant expression and implicit semantic transmission.

[0024] (3) Due to the hallucination phenomenon in the generation mechanism of the large model, it may output words or phrases outside the input text, thus affecting the final recognition quality. Therefore, the present invention designs a text-based entity filtering strategy in combination with the characteristics of the named entity recognition task to eliminate the wrong entities generated by the model. Specifically, the present invention establishes a strict matching rule between the input text and the output entities to ensure that all the finally retained entities actually exist in the input text, significantly improving the overall recognition accuracy and stability.

[0025] (4) The present invention proposes an entity type matching technology based on the large model. Compared with the traditional scheme based on single template comparison, the present invention comprehensively considers the similarity discrimination between the extraction results and type semantics, constructs a dynamic adaptation relationship between entities and types, and further eliminates low-confidence interference items and type-inconsistent items. This module can achieve refined entity judgment under multi-source prompts, strengthen the certainty and semantic consistency of the recognition results, and reveal the structural alignment law and context semantic coordination mechanism in the named entity type matching process. Description of the Drawings

[0026] Figure 1 is the overall structure diagram of the zero-shot named entity recognition method and device optimized based on large model feedback according to the present invention; Figure 2 is the design flow chart of the zero-shot named entity recognition method and device optimized based on large model feedback according to the present invention. Detailed implementation manners

[0027] The present invention will be described in detail below in conjunction with embodiments and the accompanying drawings of the specification, but is not limited thereto.

[0028] Explanation of technical terms: 1. Named entity: In the present invention, it refers to one or more entity names that always represent a certain entity with a specific meaning.

[0029] 2. Input text: In the present invention, it refers to the text content to be processed and analyzed to identify the named entities therein.

[0030] 3. Coverage: In the present invention, it corresponds to the recall rate in the evaluation index of named entity recognition, which mainly measures the proportion of all true entities correctly identified by the model.

[0031] 4. Accuracy: In the present invention, it corresponds to the precision rate in the evaluation index of named entity recognition, which mainly measures the proportion of actually correct samples predicted as entities.

[0032] 5. Large model: In the present invention, it refers to a large language model, the number of its parameters is usually in units of B, such as ChatGPT, Llama, Vicuna, etc.

[0033] 6. Prompt learning: In the present invention, it refers to a learning method that guides the model to complete a specific task without parameter update by inputting examples or instructions.

[0034] Embodiment 1 The present invention provides a zero-shot named entity recognition method based on large model feedback optimization, as Figure 1 shown, the method includes: S1: Adopt an iterative entity extraction strategy based on a large model and prompt learning, and use entity extraction prompts to extract entities in the input text X it contains, and obtain an entity set ; Step S1 of the present invention is based on a large model and prompt learning. Considering that directly using a large model for entity extraction will be affected by factors such as attention distribution and prompt word interference, it is difficult to extract all entities in one go. Therefore, an iterative entity extraction strategy is designed in this paper, so that the model can focus on the remaining text areas during the extraction process.

[0035] S11: Input the entity extraction prompt into the large model to obtain an extraction result , and perform text-based entity screening on the obtained extraction result to obtain a screened extraction result ; Since the generation results of large models are difficult to strictly constrain, the output content may contain irrelevant information beyond the text scope, and there may even be cases of fabricating entities and types out of thin air. This not only affects the final recognition effect but also reduces the performance of downstream tasks. Therefore, targeted strategies need to be introduced to avoid such interference. In the method of this paper, this problem is solved through the entity filtering strategy of the text. Specifically, filtering is performed based on whether the extraction result exists in the input text X.

[0036] (1) Among them, is the extraction result after the round of entity screening, represents the element in the set , represents the input text; the whole formula (1) means that for each element in , only if it exists in the input text will it be used as an element of the set.

[0037] S12. Determine whether there are new entities in the filtered extraction result obtained in step S11. If there are no new entities, that is, , terminate the iteration and output the entity set , where is the extraction result after the round of entity screening, is the entity set composed of all entities in the previous i - 1 rounds. For the first round, is initialized as an empty set; if there are new entities, then execute the masking strategy based on to obtain the input text , execute step S13, and continue the iteration; Specifically, the output entity set refers to the entity set obtained by accumulating the extraction result after the i - th round of entity screening into the output result of the previous i - 1 rounds. It is prepared for the next round of extraction, as shown in formula (2): (2) is the entity set of the previous i rounds, is the entity set of the previous i - 1 rounds, is the extraction result after the i - th round of entity screening; The execution of the masking strategy based on to obtain the input text is specifically: Specifically, in In the iteration, according to the previous Entity extraction results of the wheel For input text Mask and traverse The corresponding entity in the input text X is replaced with "[MASK]" to obtain the input text , as shown in formula (3): (3) in, is the input text after masking, and e represents the set The elements in the " means replacement, the meaning of the whole formula is to traverse The element will exist with the original input text Entities in are replaced with the identifier "[MASK]".

[0038] S13, the input text obtained in step S12 Replace entity extraction hint Input text in Corresponding content obtains the replaced entity extraction prompt and repeats steps S11 and S12.

[0039] The entity collection obtained above It is not only affected by the mask mechanism, but also relies on the reasonable construction of the prompt template. Since the subsequent prompt ideas of the present invention are mostly similar, they are only introduced in detail for the first time, and then incrementally introduced according to the differences in the prompt part. As shown in Table 1. The prompt mainly consists of six parts, namely "task description", "input text X", "category set C", "category explanation", "instructions" and "output". Among them, "task description" mainly introduces that the current prompt is divided into several modules, and introduces the specific meaning of each module, aiming to make the big model understand the specific utility of each part and the task requirements. "Input text X" is the text information of the current test data. "Category set C" and "category explanation" respectively contain the predefined category set and the text description of the corresponding interpretation of each category. It aims to make the big model fully understand the characteristics of each type and the differences between categories. Category set C is usually defined directly by the dataset. It comes from the entity categories manually delineated during the dataset annotation process, or it comes from the content extracted and annotated from the professional library in combination with data features and domain knowledge. "Instructions" vary according to different goals. For example, the core meaning of the current module is to extract as many entities as possible. "Output" requires the big model to output in a specified format to facilitate the extraction of the final result, such as dictionary, list and other formats.

[0040] Table 1 Hints for Entity Extraction

[0041]

[0042] Based on the above description, the technical advantages of this technical feature are as follows: the model's attention is limited and may miss some entities. The masking mechanism can reduce the occupation of the model's attention by the extracted entities, making it more inclined to mine the content that has not been recognized yet; the entity filtering strategy based on the text can effectively avoid errors caused by the model's free generation, and at the same time ensure that all output entities come from the input text, making the extraction result more reliable. In addition, it can also eliminate the situation where the model mistakenly regards the special annotation symbol "[MASK]" as an entity; and the above iterative method can make the large model focus on the text information in other positions, so as to mine more possible entities to improve the recall rate of the model. Moreover, this gradually convergent iterative method can not only effectively avoid omission, but also ensure that the extracted entity set can cover the potential key information in the text to the greatest extent. By guiding attention through masking, screening and avoiding model hallucinations, and automatically converging after iteration, this method improves the recall rate as much as possible while ensuring the reliability of the recognition result, providing an effective solution for entity extraction in complex text environments.

[0043] For example, for the input text "Typical generative model methods include Naive Bayes classifier, Gaussian mixture model, variational autoencoder, etc.", the preliminary named entity recognition result is {"algorithm": ["Naive Bayes classifier", "Gaussian mixture model"]}, then the input text for the next round will be "Typical generative model methods include [MASK], [MASK], variational autoencoder, etc.", which will make the model focus on the information in the remaining positions, such as "variational autoencoder".

[0044] S2: Adopt an entity filtering strategy based on a large model and use entity screening hints Filter the entity set obtained in step S1 to obtain the real entity set E that meets the definition of named entity recognition, where , (4) represents the th entity in the entity set , represents the size of the entity set, is the sequential index of the entities in the entity set E.

[0045] Although the above-mentioned entity extraction method based on the masking mechanism can effectively improve the recall rate of entity extraction by masking the extracted entity information, empirical analysis shows that there is a significant problem of accuracy decline. When the masked input text does not contain entities, the large model is prone to regarding some remaining common concepts or phrases as entities. For example, in the field of artificial intelligence, the intermediate result extracted by the model is as follows: "This task usually derives the [MASK] parameter of [MASK] given the output sequence". The model is prone to regarding the general concept of "task" as a named entity for extraction. In order to balance the trade-off between accuracy and recall and avoid the above phenomenon, the present invention also designs an entity filtering strategy based on the large model. This filtering strategy relies on entity screening hints to achieve, as shown in Table 2, The goal is to judge and filter out the wrong entities extracted in the previous step by introducing the normalized description and explanation of named entities.

[0046] Table 2 Entity Screening Hints

[0047]

[0048] Based on the above description, the technical advantages of this technical feature are as follows: by introducing the normalized description of named entities, it provides theoretical support for the model to judge whether the current extraction result is a named entity. Secondly, an explanation is added to the output format requirement, which not only enhances the interpretability of the model but also further improves the accuracy of the model judgment through the way of chain of thought.

[0049] For example, assume that the final extraction result of the entity extraction module based on coverage is: {"algorithm": ["Naive Bayes classifier", "Gaussian mixture model", "Variational autoencoder", "method"]}. Then the output result of "method" after going through this module may be {"explanation": "‘method’ does not conform to the definition of named entity recognition", "result": "no"}. In this case, we will remove it and leave the remaining results. {"algorithm": ["Naive Bayes classifier", "Gaussian mixture model", "Variational autoencoder"]}.

[0050] S3: Adopt an entity type matching strategy based on the large model, and use entity type matching hints to perform entity type matching on each named entity in the set E of real entities that conform to the definition of named entity recognition obtained in step S2 to form a structured classification output.

[0051] Table 3 Entity Type Matching Hints

[0052]

[0053] Step S3 mainly determines which category in the "category set" is closest to the entity extracted by the "category interpretation" judgment model, so as to verify the result of the entity extraction module. Among them, the entity type matching hint is shown in Table 3. Its model diagram is as Figure 1 shown in the lower part. The model will traverse the extraction results of the iterative entity extraction module based on coverage evaluation after filtering, and the goal is to match the entity with the type as accurately as possible by introducing category interpretation.

[0054] For each entity in the set of true entities E that conforms to the definition of named entity recognition obtained in step S2 is combined with the existing category set C to form an entity type matching hint which is provided as input to the large model, and then the matching category of each entity is obtained and denoted as . Then, it is judged whether the extraction result exists in the category set, as shown in formula (5).

[0055] (5) where, is the output result after category screening, represents the entity category corresponding to a certain entity, is the extraction result, and C is the existing category set.

[0056] If the output result is an empty set, it indicates that the entity does not correspond to any category in the category set, and it is discarded. Otherwise, the model output category is selected as the final category of the entity.

[0057] When the constraint condition for ending the traversal is not satisfied: (6), where, is the sequential index of the entity in the entity set E, is the number of entities in the entity set E; continue to loop through the entity set E, replace the entity , until the constraint condition is satisfied. After looping through all entities, each recognized entity will finally be assigned a clear category, forming a structured classification output.

[0058] The technical advantages of this technical feature are as follows: Due to the hallucination problem in large models, incorrect entities may be generated. By focusing on a single classification task, the present invention enables the large model to concentrate on entity type matching tasks, while ensuring that the classification results meet expectations, solving the problem of possible entity category classification errors in the output results of the prompt module, and improving classification accuracy and entity accuracy. The finally obtained entity category correspondence can be directly used in downstream tasks, making the entity extraction and classification process more stable and reliable.

[0059] Figure 2 It is the design flow chart of the zero-shot named entity recognition system based on large model feedback optimization described in the present invention. The input text obtains the final output result after going through the entity extraction module based on coverage, the entity filtering strategy based on the large model, and the entity type matching module based on accuracy. This output result reflects the structured named entity recognition information of the original output text. Among them, these three parts interact with the large model respectively to obtain the results of each module and use them as the input of the next module. In particular, there may be a phenomenon of multiple rounds of iteration in the entity extraction module based on coverage that is superior to the masking strategy.

[0060] Starting from the perspective of entity evaluation, the present invention constructs a zero-shot named entity recognition method based on large model feedback optimization: Due to the limited attention of the model, some entities are not recognized. Starting from the perspective of coverage, the present invention introduces a masking strategy based on to guide the model to focus on the remaining text fragments, enabling it to supplement the missing entities, improving the entity recall rate, and thus enhancing the coverage of entity recognition; Due to the hallucination problem of the large model itself, during the entity extraction process, entity screening based on text effectively avoids errors caused by the free generation of the model, while ensuring that all output entities are derived from the input text, thereby improving the accuracy of the extraction results; during the entity type matching process, by guiding the large model to focus on a single classification task, it effectively corrects possible classification errors in the extraction module and improves the precision of entity extraction.

[0061] Embodiment 2 As described above, a zero-shot named entity recognition method based on large model feedback optimization should evaluate its performance using relevant data. Therefore, the present invention uses the CrossNER dataset to evaluate this method, and the evaluation index uses Micro F1 as the evaluation index. Compared with Macro F1, it can more accurately reflect the overall performance of the model without being affected by class imbalance. Its specific calculation formula is as follows: (7) (8) (9) Among them represents the set of categories, represents the th category, represents the number of correctly predicted positive class samples in the th category, represents the number of samples mispredicted as positive in the th category, represents the number of positive class samples mispredicted as negative in the th category. represents the overall precision of the model across all categories, that is, the proportion of truly positive samples among all samples predicted as positive. Measures the overall recall ability of the model across all categories, that is, the proportion of truly positive samples correctly identified among all true positives. is the harmonic mean of precision and recall, comprehensively reflecting the global performance of the model. This evaluation is based on the Llama 3.1 - 8B large model. The experiment uses a single NVIDIA Tesla V100 for inference. Finally, on the CrossNER dataset in the fields of politics, natural science, music, literature, artificial intelligence, etc., the values are 60.2%, 58.3%, 64.9%, 50.8%, 52.9% respectively. To a certain extent, it proves the effectiveness of the method of the present invention.

[0062] The comparison of effects is shown in Table 4: Table 4 Comparison of the effects of each model

[0063] The PromptNER data is cited from "Ashok, Dhananjay, Zachary C. Lipton. Promptner:Prompting for named entity recognition[J]. arXiv preprint arXiv:2305.15444,2023." The USM data is cited from "Xiao Wang, Weikang Zhou, Can Zu, et al. Instructuie:Multi-task instruction tuning for unified information extraction[J]. arXivpreprint arXiv:2304.08085, 2023." The InstructUIE data is cited from "Jie Lou, Yaojie Lu, Dai Dai, et al. Universal information extraction as unified semantic matching[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2023, 37(11): 13318-13326." The Diluie data is cited from "Qian Guo, Yi Guo, Jin Zhao. Diluie: constructing diverse demonstrations of in-context learning with large language model for unified information extraction[J]. Neural Computing and Applications, 2024, 36(22): 13491-13512." Compared with other models, in the fields of politics, natural science, music, and artificial intelligence, the model of the method of the present invention has the highest value. Among the fields of politics, natural science, music, artificial intelligence, and literature, the average value is also the highest. In the field of named entity recognition, the higher the value, the better the global performance of the effect.

[0064] Example 3, This embodiment provides a device for a zero-shot named entity recognition method based on large model feedback optimization. The device includes: Prompt module: including entity extraction prompts and entity screening prompts and entity type matching prompts , which are respectively used for entity extraction, screening, and type matching; where the goal is to extract as many entities existing in the text as possible, the goal is to filter out the wrong entities extracted by the entity extraction module based on coverage, the goal is to match the entity and the type as accurately as possible; Entity extraction module based on coverage: built on the basis of large models and prompt learning, used to design entity extraction prompts according to expert knowledge Supplemented by a category set and category explanations, it guides the large model to extract as many words or phrases that may be named entities in the input text as possible; Entity type matching module based on accuracy: Combining the internal knowledge of the large model and the entity category explanations provided by experts, it conducts secondary verification on the extracted entities, performs type matching one by one, enabling the large model to select the most suitable type for the current entity from numerous category sets, thereby ensuring that the recognized entities conform to the expected types and can be flexibly adapted to named entity recognition tasks in different fields, thus enhancing the scalability of the method.

[0065] Starting from the perspective of entity evaluation, the present invention constructs a zero-shot named entity recognition device optimized based on the feedback of the large model: Entity extraction module based on coverage. Due to the limited attention scope of the model, some entities are not recognized. Starting from the perspective of coverage, the present invention introduces a masking strategy based on... Through the entity extraction prompt of the prompt module guides the model to focus on the remaining text fragments, enabling it to supplement the missing entities, thereby improving the coverage of entity recognition, increasing the recall rate of entity extraction, and adapting to more complex entity distribution patterns.

[0066] Due to the hallucination problem of the large model itself, in the entity extraction module, through text-based entity screening, it effectively avoids errors caused by the free generation of the model, while ensuring that all output entities are derived from the input text, thereby improving the accuracy of the extraction results.

[0067] In addition, since the masking mechanism extracts some noise, which to a certain extent leads to a decrease in accuracy, a special entity filtering strategy based on the large model is also designed. Through the entity screening prompt of the prompt module filters out the extraction results that do not conform to the definition of named entity recognition. This strategy also enables the large model to focus on judging whether the extracted content is an entity according to the definition of named entity through prompt design.

[0068] Entity type matching module. Due to the lack of examples in the zero-shot environment, the model may, due to its own hallucination problem, misidentify some non-entities or confuse entity types. Therefore, the present invention constructs an entity type matching module to perform a focused single classification task. Through the entity type matching prompt of the prompt module conducts secondary verification on the extracted entities, thereby ensuring that the recognized entities conform to the expected types and can be flexibly adapted to named entity recognition tasks in different fields, thus enhancing the scalability of the method and improving the precision of entity extraction.

[0069] Example 4 This embodiment also provides an electronic device, including: At least one processor; and, A memory that stores instructions which, when executed by the at least one processor, cause the at least one processor to execute the zero-shot named entity recognition method optimized based on large model feedback as described above.

[0070] In this embodiment, the electronic device may include, but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, and so on.

[0071] Embodiment 5 This embodiment also provides a machine-readable storage medium that stores executable instructions which, when executed, cause the machine to execute the zero-shot named entity recognition method optimized based on large model feedback as described above.

[0072] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system or device reads and executes the instructions stored in the readable storage medium.

[0073] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments, so the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.

[0074] Examples of the readable storage medium include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or a cloud via a communication network.

[0075] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A zero-shot named entity recognition method optimized based on large model feedback, characterized in that, The method includes: S1: Adopt an iterative entity extraction strategy based on large models and prompt learning, and utilize entity extraction prompts to extract entities from the input text X it contains, and obtain an entity set ; S2: Adopt an entity filtering strategy based on a large model and utilize entity screening prompts Filter the entity set obtained in step S1 to obtain a set of true entities that meet the definition of named entity recognition E , where , (1) Represents the $i$-th entity in the set of entities, where $|E|$ represents the size of the set of entities, and $i$ is the sequential index of the entities in the entity set $E$; S3: Adopt an entity type matching strategy based on a large model and utilize entity type matching prompts For each named entity in the set of true entities E that meets the definition of named entity recognition obtained in step S2 Perform entity type matching to form a structured classification output.

2. The zero-shot named entity recognition method based on large model feedback optimization according to claim 1, wherein Entity extraction prompt described in step S1 It also includes a category set C module and a category explanation module. The category set C contains a predefined category set, and the category explanation is used to explain and describe each category.

3. The zero-shot named entity recognition method based on large model feedback optimization according to claim 1, wherein, Specifically, step S1 is as follows: S11. Input the entity extraction prompt into the large model to obtain the extraction result , and perform text-based entity screening on the obtained extraction result to obtain the screened extraction result ; S12. Determine the filtered extraction results obtained in step S11 for new entities. If there are no new entities, that is , terminate the iteration and output the entity set , where is the extraction result after the -th round of entity filtering, and is the entity set composed of all entities in the previous i - 1 rounds. For the first round, is initialized as an empty set. If there are new entities, then execute the masking strategy based on to obtain the input text , execute step S13, and continue the iteration; S13. Replace the input text obtained in step S12 with the entity extraction prompt for the input text corresponding thereto, to obtain the replaced entity extraction prompt, and repeat steps S11 and S12 4. The zero-shot named entity recognition method based on large model feedback optimization according to claim 3, characterized in that The entity screening based on text for the obtained extraction result described in step S11 is specifically as follows: According to the extraction result Filter according to whether it exists in the input text X. The formula is as follows: (2) Among them, is the extraction result after the round of entity screening, represents an element in the set , represents the input text; the whole formula (2) means that for each element in , only if it exists in the input text will it be used as an element of the set.

5. The zero-shot named entity recognition method based on large model feedback optimization according to claim 3, wherein The output entity set described in step S12 , (3) Among them, is the entity set in the first i rounds, is the entity set in the first i - 1 rounds, is the extraction result after entity screening in the round; The execution described in step S12 is based on to obtain the input text according to the mask policy Specifically: In the -th iteration, according to the entity extraction results of the previous -th round, mask the input text : traverse the entities in and replace the entities in the input text X that correspond to those in with "[MASK]" to obtain the input text , as shown in Equation (4): ​ (4) Among them, is the input text obtained after masking, where e represents an element in the set and " " means replacement. The meaning of the entire formula is to traverse the elements in, and replace the entities existing in the original input text with the identifier "[MASK]".

6. The zero-shot named entity recognition method based on large model feedback optimization according to claim 1, characterized in that, Specifically, step S3 is as follows: Each entity in the set of true entities E that meets the definition of named entity recognition obtained in step S2 is combined with the existing category set C to form an entity type matching hint , and is provided as input to the large model to obtain the extraction result, i.e., the matching category, for each entity . It is judged whether the extraction result exists in the category set C, as shown in formula (5). (5) Among them, is the output result after category screening, represents the entity category corresponding to a certain entity, is the extraction result, and C is the existing category set; If the output result is an empty set, it indicates that the entity does not correspond to any category in the category set and is discarded; otherwise, the category output by the model is selected as the final category of the entity. When the constraint condition for traversal end is not satisfied: (6), Among them, is the sequential index of the entity in the entity set E, is the number of entities in the entity set E; Continue to loop through the entity set E and replace the entity , until the constraint condition is met. After looping through all entities, each identified entity will ultimately be assigned a clear category, forming a structured classification output.

7. An apparatus for implementing a zero-shot named entity recognition method based on large model feedback optimization, characterized in that, The device includes: Prompt module: including entity extraction prompts , entity screening prompts , entity type matching prompts , which are used for entity extraction, screening, and type matching respectively; among them aims to extract as many entities existing in the text as possible, aims to filter out the incorrect entities extracted by the entity extraction module based on coverage, aims to match entities and types as accurately as possible; Coverage-based entity extraction module: Built on large models and prompt learning, it is used to design entity extraction prompts based on expert knowledge Supplemented by a set of categories and category explanations, it guides the large model to extract as many words or phrases that may be named entities in the input text as possible; Accuracy-based entity type matching module: Through entity type matching prompts Perform secondary verification on the extracted entities, conduct type matching one by one, and let the large model select the most suitable type for the current entity from a large number of category sets, so as to ensure that the identified entities conform to the expected types and can be flexibly adapted to named entity recognition tasks in different fields, thereby enhancing the scalability of the method.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory that stores instructions, which when executed by the at least one processor, cause the at least one processor to execute the zero-shot named entity recognition method based on large model feedback optimization according to any one of claims 1 to 6.

9. A machine-readable storage medium, characterized in that, It stores executable instructions that, when executed, cause the machine to execute the zero-shot named entity recognition method based on large model feedback optimization according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Zero-sample named entity recognition method based on external knowledge enhancement

    CN116245104A

  • Medical named entity identification method and device based on large language model

    CN118114675A

  • Small sample named entity recognition method

    CN118133829A

  • Process industry text knowledge extraction data set automatic construction method and system

    CN119311874A

  • Large language model context learning method for few-sample named entity recognition and named entity recognition method

    CN119962534A