Zero-shot named entity recognition method and device based on large model feedback optimization

By using an iterative entity extraction and type matching strategy optimized by large model feedback, the problems of low recall and insufficient accuracy in named entity recognition tasks are solved, achieving efficient named entity recognition in zero-shot scenarios and improving the model's generalization ability and recognition performance.

CN120354855BActive Publication Date: 2025-10-21SHANDONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510855236.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-21
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Existing technologies suffer from low recall, insufficient accuracy, and inadequate model generalization ability in complex and specialized named entity recognition tasks, especially in zero-shot scenarios where accurate recognition and classification are difficult to achieve.

Method used

An iterative entity extraction strategy based on large model feedback optimization is adopted, combined with entity filtering and type matching strategies. A closed-loop recognition and correction mechanism is formed through iterative entity extraction and type matching. A masking strategy is used to complete missing entities, and an entity filtering strategy is used to remove erroneous entities. Accurate matching is achieved by combining expert knowledge and category interpretation.

Benefits of technology

It significantly improves the recall and accuracy of named entity recognition, enhances the model's adaptability to complex texts and different domains, and has good scalability and semantic consistency, making it suitable for open-domain named entity recognition tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354855B_ABST
    Figure CN120354855B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of electric digital data processing, and particularly relates to a zero-shot named entity recognition method and device based on large model feedback optimization. The method iterates an entity extraction strategy to force the large model to focus on the remaining part of the text to be extracted except the known entity, thereby improving the recall rate of entity recognition. In order to balance the recall rate and the precision rate, the method further proposes an entity filtering strategy based on the large model to realize the screening of candidate entities. After screening, the large model is guided to focus on the entity type classification task, effectively correcting the classification errors that may be generated in the extraction process. At the same time, a text-based entity filtering strategy is introduced in the entity extraction process to suppress the invalid or error entities caused by the large model illusion problem, thereby significantly improving the accuracy and stability of the overall recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electronic digital data processing, and in particular relates to a zero-sample named entity recognition method and device based on large model feedback optimization. Background Art

[0002] A named entity is one or more names that consistently represent a specific entity with a specific meaning. Named entity recognition (NER) originally focused on identifying names of people, places, and organizations in text. However, with the deep integration of technological research and industry needs, the scope of the definition of named entities has gradually expanded in both directions. There has been vertical differentiation at the basic conceptual level, such as from the general "personal name" to occupational attribute categories such as "researcher" and "biologist." Meanwhile, new entity types have been expanded horizontally, such as "compound" and "protein," which are field-specific. By accurately extracting entity information from text, NER provides structured data support for knowledge graph construction, question-answering, and information retrieval. It has significant practical value in vertical fields such as biomedical literature mining, financial risk prediction, and judicial document structuring. However, such datasets are often costly and difficult to label, and the labeling process is tedious, time-consuming, and dependent on domain knowledge. These issues have prompted researchers to investigate the problem of named entity recognition in zero-shot scenarios.

[0003] The task of zero-shot named entity recognition relies on the development of large language models (LLMs). Since the advent of ChatGPT in 2023, its rich knowledge base, strong generalization, and excellent contextual understanding capabilities have prompted scholars to study its learning and reasoning capabilities in zero-shot scenarios. Currently, there are two main approaches to solving zero-shot named entity recognition tasks based on large models. One is to mine the internal knowledge of the large model through prompts or system design, and the other is to focus on or supplement the internal knowledge of the large model by introducing external knowledge. The former can adapt to zero-shot named entity recognition tasks in different fields by reserving category options in the prompt template, but it relies on a large number of repeated prompt engineering attempts. The latter introduces external knowledge to compensate for the poor timeliness of the large model's knowledge and enhance its performance in specialized fields. However, its effectiveness is limited by the quality of external knowledge and the efficiency of the retrieval algorithm.

[0004] The main frontier work related to the problem of professional terminology extraction includes:

[0005] Chinese patent CN116245104A proposes a zero-shot named entity recognition method based on external knowledge enhancement. This method obtains sentences containing the target entity category name from an external knowledge base, uses a preprocessing model to extract the category semantic representation, and calculates semantic similarity with the entity to be identified to determine the entity category. Chinese patent CN118114675A discloses a medical named entity recognition method based on a large language model. This method identifies candidate entity categories under the guidance of multiple prompts, extracts arguments based on knowledge text, and evaluates the correctness of various opinions to determine the entity category.

[0006] While the prompt design and task flow described above are relatively simple, these methods face challenges when dealing with complex, ambiguous, or highly specialized text. Semantic ambiguity, polysemy, and domain-specific knowledge and expressions in complex text make it difficult to determine entity boundaries and categorize entities, making accurate recognition and classification difficult. This limits their application in specialized domains and complex scenarios. Summary of the Invention

[0007] In response to the shortcomings of the existing technology, the present invention discloses a zero-shot named entity recognition method based on large-model feedback optimization. The method introduces an iterative entity extraction strategy to force the large model to focus on the remaining parts of the text to be extracted except for known entities, thereby improving the recall rate of entity recognition; in order to take into account both recall rate and precision rate, the method further proposes an entity filtering strategy based on the large model to realize the screening of candidate entities; after screening, by guiding the large model to focus on the entity type classification task, it can effectively correct the classification errors that may occur during the extraction process.

[0008] The invention also discloses a device for realizing the method.

[0009] The invention also discloses an electronic device for implementing the method.

[0010] The present invention also discloses a machine-readable storage medium for implementing the above method.

[0011] In order to achieve the above object, the present invention adopts the following technical solutions:

[0012] A zero-shot named entity recognition method based on large model feedback optimization, the method comprising:

[0013] S1: Adopting an iterative entity extraction strategy based on large model and prompt learning, using entity extraction prompts Extract the entities contained in the input text X to obtain the entity set ;

[0014] S2: Adopt entity filtering strategy based on large model and use entity filtering prompts For the entity set obtained in step S1 Filter and obtain the real entity set E that meets the definition of named entity recognition, where

[0015] , (1)

[0016] Representing a collection of entities The entities, Indicates the size of the entity collection, is the sequential index of the entity in the entity set E;

[0017] S3: Adopting entity type matching strategy based on large model and using entity type matching hints For each named entity in the real entity set E that meets the definition of named entity recognition obtained in step S2 Perform entity type matching to form structured classification output.

[0018] Preferably, the entity extraction prompt in step S1 It also includes a category set C module and a category explanation module. The category set C contains a predefined category set, and the category explanation is used to explain and describe each category.

[0019] Further preferably, the entity extraction prompt It also includes a task description module, which is used to introduce the modules included in the entity extraction prompt and the meaning of each module, so that the big model can understand the specific utility and task requirements of each part; an explanation module, which is used to specifically describe the tasks that the big model should perform, so that the big model can extract as many entities as possible; and an output module, which requires the big model to output in a specified format.

[0020] Preferably, step S1 is specifically as follows:

[0021] S11. Extract entity prompts Input into the large model to obtain the extraction results , and extract the results Perform text-based entity screening to obtain the filtered extraction results ;

[0022] S12: Determine the extracted results after screening obtained in step S11 Is there a new entity? If there is no new entity, , terminate the iteration and output the entity set ,in, For the sutra The extraction results after round entity screening, is the entity set consisting of all entities in the first i-1 rounds. For the first round, Initialized to an empty set; if there are new entities, execute based on The mask strategy gets the input text , execute step S13 and continue iteration;

[0023] S13, the input text obtained in step S12 Replace entity extraction prompt Input text in The corresponding content obtains the replaced entity extraction prompt and repeats steps S11 and S12.

[0024] Further preferably, the extraction result obtained in step S11 is The text-based entity filtering is as follows:

[0025] Based on the extraction results Whether it exists in the input text X for filtering, the formula is as follows:

[0026] (2)

[0027] in, For the sutra The extraction results after round entity screening, Representing a collection The elements in Represents the input text; Formula (2) as a whole represents Each element in , only if it satisfies the input text Only in this way can we act An element of a collection.

[0028] Further preferably, the output entity set in step S12 is ,

[0029] (3)

[0030] in, is the entity collection of the first i rounds, It is the collection of entities in the first i-1 rounds. For the sutra The extraction results after round entity screening;

[0031] Step S12 is based on The mask strategy gets the input text Specifically:

[0032] In the In the round of iteration, according to the previous Entity extraction results of the wheel For input text Perform masking: traversal Entities in the input text X and The corresponding entity in is replaced with "[MASK]" to obtain the input text , as shown in formula (4):

[0033] (4)

[0034] in, is the input text after masking, e represents the set The elements in the " means replacement, the meaning of the whole formula is to traverse The element in will exist with the original input text Entities in are replaced with the identifier "[MASK]".

[0035] Preferably, the entity screening prompt in step S2 It includes a named entity definition module, which is used to standardize the description of named entities; an output format module, which contains an "explanation" entry, which is used to explain the reason for determining whether the extracted entity is a named entity.

[0036] Preferably, step S3 specifically comprises: each entity in the real entity set E that meets the definition of named entity recognition obtained in step S2 Combined with the existing category set C to form entity type matching hints ,Will Provided as input to the big model, the extraction results of each entity are obtained, that is, the matching category , to determine whether the extraction result exists in the category set C, as shown in formula (5),

[0037] (5)

[0038] in, is the output result after category filtering, Indicates the entity category corresponding to a certain entity, is the extraction result, and C is the existing category set.

[0039] If the output result If it is an empty set, it means that the entity does not correspond to any category in the category set and is discarded; otherwise, the model output category is selected as the final category of the entity;

[0040] When the constraints for traversal end are not met:

[0041] (6),

[0042] in, is the sequential index of the entity in the entity set E, is the number of entities in the entity set E;

[0043] Continue to loop through the entity set E and replace the entity , until the constraints are met, after looping through all entities, each identified entity will eventually be assigned a clear category to form a structured classification output.

[0044] Further preferably, entity type matching prompt It also includes a category interpretation module, which is used to determine which category in the "category set" the entity extracted by the model is closest to, thereby verifying the results of the entity extraction module.

[0045] In another aspect of the present invention, a device for a zero-shot named entity recognition method based on large model feedback optimization is provided, the device comprising:

[0046] Prompt module: contains entity extraction prompts , entity filtering prompts , entity type matching prompt , respectively used for entity extraction, screening and type matching; The goal is to extract as many entities as possible from the text. The goal is to filter out the wrong entities extracted by the coverage-based entity extraction module. The goal is to match entities to types as accurately as possible;

[0047] Coverage-based entity extraction module: built on the basis of large model and prompt learning, used to design entity extraction prompts based on expert knowledge This is supplemented by category sets and category explanations to guide the large model to extract as many words or phrases as possible that may be named entities in the input text;

[0048] Accuracy-based entity type matching module: hinted by entity type matching The extracted entities are checked twice and type matching is performed one by one, allowing the large model to select the most suitable type for the current entity from a large set of categories, thereby ensuring that the identified entities meet the expected type and can be flexibly adapted to named entity recognition tasks in different fields, thereby enhancing the scalability of the method.

[0049] In another aspect of the present invention, an electronic device is provided, comprising:

[0050] at least one processor; and,

[0051] A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to perform the zero-sample named entity recognition method based on large model feedback optimization as described above.

[0052] In another aspect of the present invention, a machine-readable storage medium is provided, which stores executable instructions, and when the instructions are executed, the machine performs the zero-sample named entity recognition method based on large model feedback optimization as described above.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] (1) From the perspective of entity evaluation, the present invention constructs a zero-shot named entity recognition method based on large-scale model feedback optimization. By linking the two modules of iterative entity extraction and type matching, a closed-loop entity recognition and correction mechanism is formed. Compared with the existing extraction process that lacks feedback channels and has a fixed structure, the present invention takes into account both semantic completion capabilities and output stability, and can effectively adapt to unlabeled resources and complex context environments. While improving recognition coverage and accuracy, this method has good scalability and model versatility, reflecting the generalization ability and application practicality for open-domain named entity recognition tasks.

[0055] (2) The present invention introduces a mask-based entity extraction strategy driven by large-scale model feedback. To address the problem of insufficient entity coverage in zero-shot named entity recognition, the present invention combines the large-scale model's ability to understand the semantic structure of language to construct a round-by-round feedback completion mechanism. Compared with existing methods that rely on static prompts or one-time extraction, the present invention uses entity omission prompts to generate a supplementary mask based on the first extraction and triggers the recognition process again, effectively expanding the entity boundary recognition ability and semantic extension ability. This technical solution can significantly improve the extraction coverage of complex sentences and implicit entities in zero-shot scenarios, which conforms to the regular characteristics of language redundant expression and implicit semantic transmission.

[0056] (3) Due to the hallucination phenomenon in the generation mechanism of large models, words or phrases outside the input text may be output, thus affecting the final recognition quality. Therefore, the present invention combines the characteristics of the named entity recognition task and designs a text-based entity filtering strategy to eliminate erroneous entities generated by the model. Specifically, the present invention establishes strict matching rules between the input text and the output entities to ensure that the entities finally retained are all real entities in the input text, significantly improving the accuracy and stability of the overall recognition.

[0057] (4) This paper proposes a large-scale model-based entity type matching technology. Compared with the traditional scheme based on single template comparison, this paper comprehensively considers the similarity judgment between the extraction results and the type semantics, constructs a dynamic adaptation relationship between entities and types, and further eliminates low-confidence interference items and type-inconsistent items. This module can achieve refined entity judgment under multi-source prompts, strengthen the certainty and semantic consistency of recognition results, and reveal the structural alignment rules and contextual semantic coordination mechanism in the named entity type matching process. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 This is an overall structural diagram of the zero-shot named entity recognition method and device based on large model feedback optimization according to the present invention;

[0059] Figure 2 This is a design flow chart of the zero-sample named entity recognition method and device based on large model feedback optimization described in the present invention. DETAILED DESCRIPTION

[0060] The present invention will be described in detail below with reference to the embodiments and the accompanying drawings, but is not limited thereto.

[0061] Technical term explanation:

[0062] 1. Named entity: In the present invention, it refers to one or more entity names that always represent a specific meaning.

[0063] 2. Input text: In the present invention, it refers to the text content to be processed and analyzed to identify named entities therein.

[0064] 3. Coverage: In the present invention, it corresponds to the recall rate in the named entity recognition evaluation index, which mainly measures the proportion of all real entities correctly recognized by the model.

[0065] 4. Accuracy: In the present invention, it corresponds to the precision rate in the named entity recognition evaluation index, which mainly measures the actual correct proportion of samples predicted to be entities.

[0066] 5. Large model: In this invention, it refers to a large language model, whose parameter quantity is usually measured in B, such as ChatGPT, Llama, Vicuna, etc.

[0067] 6. Prompted learning: In this invention, it refers to a learning method that guides the model to complete a specific task without parameter updates by inputting examples or instructions.

[0068] Example 1

[0069] The present invention provides a zero-sample named entity recognition method based on large model feedback optimization, such as Figure 1 As shown, the method includes:

[0070] S1: Adopting an iterative entity extraction strategy based on large model and prompt learning, using entity extraction prompts Extract the entities contained in the input text X to obtain the entity set ;

[0071] Step S1 of the present invention is based on a large model and prompt learning. Considering that directly using a large model for entity extraction will be affected by factors such as attention distribution and interference from prompt words, it is difficult to extract all entities at once. Therefore, this paper designs an iterative entity extraction strategy that allows the model to focus on the remaining text areas during the extraction process.

[0072] S11. Extract entity prompts Input into the large model to obtain the extraction results , the extraction results Perform text-based entity screening to obtain the filtered extraction results ;

[0073] Since the results of large model generation are difficult to strictly constrain, the output content may contain irrelevant information beyond the scope of the text, and may even fabricate entities and types out of thin air, which not only affects the final recognition effect, but also reduces the performance of downstream tasks. Therefore, it is necessary to introduce targeted strategies to avoid such interference. In this paper, this problem is solved by the entity filtering strategy of the text. Specifically, based on the extraction results Filter whether it exists in the input text X.

[0074] (1)

[0075] in, For the sutra The extraction results after round entity screening, Representing a collection The elements in Represents the input text; Formula (1) as a whole represents Each element in , only if it satisfies the input text Only in this way can we act An element of a collection.

[0076] S12: Determine the extracted results after screening obtained in step S11 Is there a new entity? If there is no new entity, , terminate the iteration and output the entity set ,in, For the sutra The extraction results after round entity screening, is the entity set consisting of all entities in the first i-1 rounds. For the first round, Initialized to an empty set; if there are new entities, execute based on The mask strategy gets the input text , execute step S13 and continue iteration;

[0077] Specifically, the output entity set Refers to the extraction results after the i-th round of entity screening The entity set obtained after accumulating the output results of the previous i-1 rounds is prepared for the next round of extraction, as shown in formula (2):

[0078] (2)

[0079] is the entity collection of the first i rounds, It is the collection of entities in the first i-1 rounds. is the extraction result after the i-th round of entity screening;

[0080] The execution is based on The mask strategy gets the input text Specifically:

[0081] Specifically, in In the round of iteration, according to the previous Entity extraction results of the wheel For input text Perform masking and traversal The corresponding entity in the input text X is replaced with "[MASK]" to obtain the input text , as shown in formula (3):

[0082] (3)

[0083] in, is the input text after masking, e represents the set The elements in the " means replacement, the meaning of the whole formula is to traverse The element in will exist with the original input text Entities in are replaced with the identifier "[MASK]".

[0084] S13, the input text obtained in step S12 Replace entity extraction prompt Input text in The corresponding content obtains the replaced entity extraction prompt and repeats steps S11 and S12.

[0085] The entity collection obtained above It is not only affected by the mask mechanism, but also relies on the reasonable construction of the prompt template. Since the subsequent design prompt ideas of this invention are mostly similar, they are only introduced in detail for the first time, and then incrementally introduced based on the differences in the prompt part. For the current module, the entity extraction prompt designed in this embodiment As shown in Table 1, the prompt consists of six main components: "Task Description," "Input Text X," "Category Set C," "Category Explanation," "Explanation," and "Output." The "Task Description" primarily introduces the modules of the prompt and explains the specific meaning of each module, helping the model understand the specific utility of each component and the task requirements. "Input Text X" is the textual information of the current test data. "Category Set C" and "Category Explanation" contain a predefined set of categories and a textual description of each category's corresponding meaning, respectively. This helps the model fully understand the characteristics of each category and the differences between them. Category Set C is typically defined directly by the dataset. It may be derived from manually defined entity categories during dataset annotation, or from content extracted and annotated from specialized libraries using a combination of data features and domain knowledge. "Explanation" varies depending on the goal. For example, the core purpose of the current module is to extract as many entities as possible. "Output" requires the model to output in a specified format, such as a dictionary or list, to facilitate extraction of the final result.

[0086] Table 1 Entity extraction tips

[0087]

[0088] Based on the above description, the technical advantages of this technical feature are: the model has limited attention and may miss some entities. The masking mechanism can reduce the model's attention occupied by extracted entities, making it more inclined to explore unrecognized content; the text-based entity filtering strategy can effectively avoid errors caused by the model's free generation, while ensuring that all output entities are derived from the input text, making the extraction results more reliable. In addition, it can also eliminate cases where the model mistakenly identifies the special marker "[MASK]" as an entity; and the above-mentioned iterative method allows the large model to focus on text information in other locations, thereby mining more possible entities and improving the model's recall rate. Moreover, this gradually convergent iterative method not only effectively avoids omissions, but also ensures that the extracted entity set maximizes coverage of potential key information in the text. By guiding attention through masks, filtering to avoid model hallucinations, and automatically converging after iterations, this method maximizes recall while ensuring the reliability of recognition results, providing an effective solution for entity extraction in complex text environments.

[0089] For example, for input text "Typical generative model methods include Naive Bayes classifier, Gaussian mixture model, variational autoencoder, etc." The result of the preliminary named entity recognition is {"algorithm": ["Naive Bayes classifier", "Gaussian mixture model"]}, then the input text of the next round It is "Typical generative model methods include [MASK], [MASK], variational autoencoders, etc." This will make the model pay attention to the information of the remaining positions, such as "variational autoencoders".

[0090] S2: Adopt entity filtering strategy based on large model and use entity filtering prompts For the entity set obtained in step S1 Filter and obtain the real entity set E that meets the definition of named entity recognition, where

[0091] , (4)

[0092] Representing a collection of entities The entities, Indicates the size of the entity collection, is the sequential index of the entity in the entity set E.

[0093] Although the above-mentioned entity extraction method based on the mask mechanism can effectively improve the recall rate of entity extraction by masking the extracted entity information, empirical analysis shows that it has a significant accuracy drop problem. When the masked input text does not contain entities, the large model tends to regard some remaining common concepts or phrases as entities. For example, in the field of artificial intelligence, the intermediate result of model extraction is as follows "The task is usually to derive the [MASK] parameter [MASK] given an output sequence." The model tends to regard the general concept of "task" as a named entity for extraction. In order to balance the trade-off between accuracy and recall and avoid the above phenomenon, the present invention also designs an entity filtering strategy based on a large model. This filtering strategy relies on entity screening prompts Implementation, as shown in Table 2, The goal is to filter out the erroneous entities extracted in the previous step by introducing the standardized description and interpretation of named entities.

[0094] Table 2 Entity screening tips

[0095]

[0096] Based on the above description, the technical advantages of this feature are: by introducing a standardized description of named entities, it provides theoretical support for the model to determine whether the current extraction result is a named entity. Secondly, it adds explanations to the output format requirements, which not only enhances the interpretability of the model but also further improves the accuracy of the model's judgment through a chain of thought process.

[0097] For example, suppose the final extraction result of the coverage-based entity extraction module is: {"algorithms":["Naive Bayes Classifier","Gaussian Mixture Model","Variational Autoencoder","Method"]}. Then, after the module passes "Method", the output result might be {"Explanation":"'Method' does not meet the definition of named entity recognition","Result":"No"}. In this case, we remove it, leaving the remaining result: {"algorithms":["Naive Bayes Classifier","Gaussian Mixture Model","Variational Autoencoder"]}.

[0098] S3: Adopting entity type matching strategy based on large model and using entity type matching hints For each named entity in the real entity set E that meets the definition of named entity recognition obtained in step S2 Perform entity type matching to form structured classification output.

[0099] Table 3 Entity type matching tips

[0100]

[0101] Step S3 mainly judges which category in the "category set" the entity extracted by the model is closest to based on the "category explanation", thereby verifying the results of the entity extraction module. As shown in Table 3. The model diagram is as follows Figure 1 As shown in the lower part, the model will traverse the extraction results of the filtered iterative entity extraction module based on coverage evaluation. The goal is to match entities to types as accurately as possible by introducing category interpretations.

[0102] Each entity in the real entity set E that meets the definition of named entity recognition obtained in step S2 is Combined with the existing category set C to form entity type matching hints Provided as input to the big model, the matching category of each entity is obtained and recorded as Then determine whether the extraction result exists in the category set, as shown in formula (5).

[0103] (5)

[0104] in, is the output result after category filtering, Indicates the entity category corresponding to a certain entity, is the extraction result, and C is the existing category set.

[0105] If the output result If it is an empty set, it means that the entity does not correspond to any category in the category set and is discarded. Otherwise, the model output category is selected as the final category of the entity.

[0106] When the constraints for traversal end are not met:

[0107] (6),

[0108] in, is the sequential index of the entity in the entity set E, is the number of entities in the entity set E;

[0109] Continue to loop through the entity set E and replace the entity , until the constraints are met, after looping through all entities, each identified entity will eventually be assigned a clear category to form a structured classification output.

[0110] The technical advantage of this feature is that, because large models can generate hallucinations and incorrect entities, this feature focuses on a single classification task, allowing the large model to focus on entity type matching. This ensures that the classification results meet expectations while addressing potential entity category misclassification issues in the prompt module output, improving classification accuracy and entity accuracy. The resulting entity category correspondence can be directly used in downstream tasks, making the entity extraction and classification process more stable and reliable.

[0111] Figure 2 This is a design flow chart for the zero-shot named entity recognition system based on large-model feedback optimization described in the present invention. The input text undergoes a coverage-based entity extraction module, a large-model-based entity filtering strategy, and an accuracy-based entity type matching module to obtain the final output result. This output reflects the structured named entity recognition information of the original output text. These three components interact with the large-model separately to obtain the results of each module, which serve as input for the next module. In particular, the coverage-based entity extraction module outperforms the masking strategy, which may involve multiple rounds of iteration.

[0112] From the perspective of entity evaluation, this paper constructs a zero-shot named entity recognition method based on large model feedback optimization:

[0113] Due to the limited attention of the model, some entities are not recognized. The masking strategy guides the model to focus on the remaining text fragments, allowing it to supplement the missing entities, improving the entity recall rate and thus improving the coverage of entity recognition;

[0114] Due to the hallucination problem of the large model itself, text-based entity screening is used during the entity extraction process to effectively avoid errors caused by the free generation of the model, while ensuring that all output entities are derived from the input text, thereby improving the accuracy of the extraction results; in the entity type matching process, by guiding the large model to focus on a single classification task, the classification errors that may occur in the extraction module are effectively corrected, thereby improving the accuracy of entity extraction.

[0115] Example 2

[0116] As described above, a zero-shot named entity recognition method based on large-scale model feedback optimization should use relevant data to evaluate its performance. Therefore, the present invention uses the CrossNER dataset to evaluate this method, using Micro F1 as the evaluation metric. Compared with Macro F1, it can more accurately reflect the overall performance of the model without being affected by class imbalance. Its specific calculation formula is as follows:

[0117] (7)

[0118] (8)

[0119] (9)

[0120] in Represents a collection of categories, Indicates the categories, Indicates the The number of correctly predicted positive samples in the category, Indicates the The number of samples in the category that are incorrectly predicted to be positive, Indicates the The number of positive samples in a category that are incorrectly predicted as negative. Represents the overall accuracy of the model on all categories, that is, the proportion of samples that are actually positive among all samples predicted to be positive. Measures the overall recall of the model across all categories, that is, the proportion of correctly identified examples among all true positive instances. It is the harmonic mean of precision and recall, which comprehensively reflects the overall performance of the model. This evaluation is based on the Llama 3.1-8B large model. The experiment uses a single NVIDIA Tesla V100 for inference. Finally, the CrossNER dataset is used in the fields of politics, natural sciences, music, literature, artificial intelligence, etc. The values ​​were 60.2%, 58.3%, 64.9%, 50.8% and 52.9% respectively, which proved the effectiveness of the method of the present invention to a certain extent.

[0121] The effect comparison is shown in Table 4:

[0122] Table 4 Comparison of the effects of various models

[0123]

[0124] PromptNER data is quoted from "Ashok, Dhananjay, Zachary C. Lipton. Promptner: Prompting for named entity recognition[J]. arXiv preprint arXiv:2305.15444,2023."

[0125] USM data is quoted from "Xiao Wang, Weikang Zhou, Can Zu, et al. Instructuie:Multi-task instruction tuning for unified information extraction[J]. arXivpreprint arXiv:2304.08085, 2023."

[0126] InstructUIE data is quoted from "Jie Lou, Yaojie Lu, Dai Dai, et al. Universal information extraction as unified semantic matching[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2023, 37(11): 13318-13326."

[0127] Diluie data is quoted from "Qian Guo, Yi Guo, Jin Zhao. Diluie: constructingdiverse demonstrations of in-context learning with large language model for unified information extraction[J]. Neural Computing and Applications, 2024,36(22): 13491-13512."

[0128] Compared with other models, the model of the method of the present invention is more effective in the fields of politics, natural science, music, and artificial intelligence. The highest value is in the fields of politics, natural sciences, music, artificial intelligence, and literature. The average value of is also the highest. In the field of named entity recognition, Higher values ​​result in better overall performance of the effect.

[0129] Example 3

[0130] This embodiment provides a device for a zero-shot named entity recognition method based on large-model feedback optimization, the device comprising:

[0131] Prompt module: contains entity extraction prompts , entity filtering prompts , entity type matching prompt , respectively used for entity extraction, screening and type matching; The goal is to extract as many entities as possible from the text. The goal is to filter out the wrong entities extracted by the coverage-based entity extraction module. The goal is to match entities to types as accurately as possible;

[0132] Coverage-based entity extraction module: built on the basis of large model and prompt learning, used to design entity extraction prompts based on expert knowledge This is supplemented by category sets and category explanations to guide the large model to extract as many words or phrases as possible that may be named entities in the input text;

[0133] Accuracy-based entity type matching module: Combining the internal knowledge of the large model and the entity category explanations provided by experts, the extracted entities are secondary verified and type matched one by one, allowing the large model to select the most suitable type for the current entity from a large set of categories, thereby ensuring that the identified entities meet the expected type and can be flexibly adapted to named entity recognition tasks in different fields, thereby enhancing the scalability of the method.

[0134] From the perspective of entity evaluation, the present invention constructs a zero-sample named entity recognition device based on large model feedback optimization:

[0135] Based on the coverage of the entity extraction module, due to the limited scope of the model, some entities are not recognized. The mask strategy is to extract the entity hint through the hint module Guiding the model to focus on the remaining text fragments allows it to supplement the missed entities, thereby improving the coverage of entity recognition, thereby improving the recall rate of entity extraction, and adapting to more complex entity distribution patterns.

[0136] Due to the hallucination problem of large models themselves, the entity extraction module effectively avoids errors caused by free model generation through text-based entity screening, while ensuring that all output entities are derived from the input text, thereby improving the accuracy of the extraction results.

[0137] In addition, since the masking mechanism will extract some noise, which will lead to a decrease in accuracy to a certain extent, a special entity filtering strategy based on the large model is designed. To filter out the extraction results that do not meet the definition of named entity recognition, this strategy also uses prompt design to enable the large model to focus on judging whether the extracted content is an entity based on the definition of named entities.

[0138] Entity type matching module, due to the lack of examples in the zero-sample environment, the model may mistakenly identify some non-entities or confuse entity types due to its own hallucination problem. Therefore, the present invention constructs an entity type matching module to perform a single classification task, and uses the entity type matching prompt of the prompt module to The extracted entities are verified twice to ensure that the identified entities meet the expected types and can be flexibly adapted to named entity recognition tasks in different fields, thereby enhancing the scalability of the method and improving the accuracy of entity extraction.

[0139] Example 4

[0140] This embodiment further provides an electronic device, including:

[0141] at least one processor; and,

[0142] A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to perform the zero-sample named entity recognition method based on large model feedback optimization as described above.

[0143] In this embodiment, electronic devices may include, but are not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, and the like.

[0144] Example 5

[0145] This embodiment also provides a machine-readable storage medium storing executable instructions, which, when executed, enable the machine to perform the zero-sample named entity recognition method based on large model feedback optimization as described above.

[0146] Specifically, a system or device equipped with a readable storage medium can be provided, on which software program codes that implement the functions of any of the above-mentioned embodiments are stored, and a computer or processor of the system or device can read and execute instructions stored in the readable storage medium.

[0147] In this case, the program code itself read from the machine-readable medium can implement the functions of any one of the above embodiments, and thus the machine-readable code and the machine-readable storage medium storing the machine-readable code constitute part of this specification.

[0148] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (e.g., CD-ROMs, CD-Rs, CD-RWs, DVD-ROMs, DVD-RAMs, DVD-RWs, DVD-RWs), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer or a cloud via a communication network.

[0149] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the technical solutions of the present invention, and are not intended to limit the specific implementation methods of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A zero-shot named entity recognition method based on large model feedback optimization, characterized by: The method comprises: S1: Adopting an iterative entity extraction strategy based on large model and prompt learning, using entity extraction prompts Extract the entities contained in the input text X to obtain the entity set ; S2: Adopt entity filtering strategy based on large model and use entity filtering prompts For the entity set obtained in step S1 Filter to obtain a set of real entities that meet the definition of named entity recognition E ,in, , (1) Representing a collection of entities The entities, Indicates the size of the entity collection, is the sequential index of the entity in the entity set E; S3: Adopting entity type matching strategy based on large model and using entity type matching hints For each named entity in the real entity set E that meets the definition of named entity recognition obtained in step S2 Perform entity type matching to form structured classification output; Step S1 is specifically as follows: S11. Extract entity prompts Input into the large model to obtain the extraction results , and extract the results Perform text-based entity screening to obtain the filtered extraction results ; S12: Determine the extracted results after screening obtained in step S11 Is there a new entity? If there is no new entity, , terminate the iteration and output the entity set ,in, For the sutra The extraction results after round entity screening, is the entity set consisting of all entities in the first i-1 rounds. For the first round, Initialized to an empty set; if there are new entities, execute based on The mask strategy gets the input text , execute step S13 and continue iteration; S13, the input text obtained in step S12 Replace entity extraction prompt Input text in The corresponding content obtains the replaced entity extraction prompt and repeats steps S11 and S12.

2. The zero-shot named entity recognition method based on large model feedback optimization according to claim 1 is characterized in that Entity extraction prompt in step S1 It also includes a category set C module and a category explanation module. The category set C contains a predefined category set, and the category explanation is used to explain and describe each category.

3. The zero-shot named entity recognition method based on large model feedback optimization according to claim 1 is characterized in that: Step S11 extracts the result The text-based entity filtering is as follows: Based on the extraction results Whether it exists in the input text X for filtering, the formula is as follows: (2) in, For the sutra The extraction results after round entity screening, Representing a collection The elements in Represents the input text; Formula (2) as a whole represents Each element in , only if it satisfies the input text Only in this way can we act An element of a collection.

4. The zero-shot named entity recognition method based on large model feedback optimization according to claim 1 is characterized in that: Output entity set in step S12 , (3) in, is the entity collection of the first i rounds, It is the collection of entities in the first i-1 rounds. For the sutra The extraction results after round entity screening; Step S12 is based on The mask strategy gets the input text Specifically: In the In the round of iteration, according to the previous Entity extraction results of the wheel For input text Perform masking: traversal Entities in the input text X and The corresponding entity in is replaced with "[MASK]" to obtain the input text , as shown in formula (4): (4) in, is the input text after masking, e represents the set Elements in " " means replacement, the meaning of the whole formula is to traverse The element in will exist with the original input text The entities in are replaced with the identifier "[MASK]".

5. The zero-shot named entity recognition method based on large model feedback optimization according to claim 1 is characterized in that: Step S3 is specifically as follows: Each entity in the real entity set E that meets the definition of named entity recognition obtained in step S2 is Combined with the existing category set C to form entity type matching hints ,Will Provided as input to the big model, the extraction results of each entity are obtained, that is, the matching category , to determine whether the extraction result exists in the category set C, as shown in formula (5), (5) in, is the output result after category filtering, Indicates the entity category corresponding to a certain entity, is the extraction result, C is the existing category set; If the output result If it is an empty set, it means that the entity does not correspond to any category in the category set and is discarded; otherwise, the model output category is selected as the final category of the entity; When the constraints for traversal end are not met: (6), in, is the sequential index of the entity in the entity set E, is the number of entities in the entity set E; Continue to loop through the entity set E and replace the entity , until the constraints are met, after looping through all entities, each identified entity will eventually be assigned a clear category to form a structured classification output.

6. A device for implementing the zero-shot named entity recognition method based on large model feedback optimization as described in any one of claims 1 to 5, characterized in that: The device comprises: Prompt module: contains entity extraction prompts , entity filtering prompts , entity type matching prompt , respectively used for entity extraction, screening and type matching; The goal is to extract as many entities as possible from the text. The goal is to filter out the wrong entities extracted by the coverage-based entity extraction module. The goal is to match entities to types as accurately as possible; Coverage-based entity extraction module: built on the basis of large model and prompt learning, used to design entity extraction prompts based on expert knowledge This is supplemented by category sets and category explanations to guide the large model to extract as many words or phrases as possible that may be named entities in the input text; Accuracy-based entity type matching module: hinted by entity type matching The extracted entities are checked twice and type matching is performed one by one, allowing the large model to select the most suitable type for the current entity from a large set of categories, thereby ensuring that the identified entities meet the expected type and can be flexibly adapted to named entity recognition tasks in different fields, thereby enhancing the scalability of the method.

7. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the zero-sample named entity recognition method based on large model feedback optimization as described in any one of claims 1 to 5.

8. A machine-readable storage medium, characterized in that It stores executable instructions, which, when executed, enable the machine to perform the zero-sample named entity recognition method based on large model feedback optimization as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Zero-sample named entity recognition method based on external knowledge enhancement

    CN116245104A

  • Medical named entity identification method and device based on large language model

    CN118114675A

  • Small sample named entity recognition method

    CN118133829A

  • Large model prompt project optimization system and method fusing domain knowledge graph

    CN120196734A