Evolved prompt-based low-resource named entity identification method and system
By using an evolutionary prompt-based method in low-resource named entity recognition, dynamically updating zero-sample instructions and pseudo-sample sets, the problem of insufficient optimization of initial prompt design and pseudo-sample sets in the prior art is solved, and the accuracy and recognition performance of entity recognition are improved.
Patent Information
- Application Number
- CN202510285767.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing low-resource named entity recognition method based on large language models has shortcomings in initial prompt design and pseudo-sample set optimization, which makes it difficult for the model to fully tap potential and adapt to data characteristics, affecting the recognition performance.
Using an evolutionary prompt method, the large language model is prompted by zero-sample instruction to generate initial entity prediction, and a sample library of definition of entity types and a sample library of differences in similar entity types are constructed. The pseudo-sample randomly selects pseudo-sample as context learning examples, dynamically updates the zero-sample instruction and pseudo-sample collection, and iteratively optimizes the model.
It improves the accuracy of entity recognition, makes full use of the potential of large language models, adapts to data characteristics and task needs, and continuously improves recognition performance.
Smart Images

Figure CN120163156A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a low-resource named entity recognition method and system based on evolutionary prompts, belonging to the technical field of prompt learning. Background Art
[0002] Named Entity Recognition (NER) is a basic and key task in the field of natural language processing (NLP). Its main goal is to identify entities with specific meanings from text, such as names of people, places, organizations, dates, etc. These entity information plays a vital role in many subsequent applications such as text understanding, information extraction, knowledge graph construction, etc. Traditional named entity recognition methods usually rely on a large amount of annotated data to learn the characteristics and patterns of entities by building specific models. However, in practical applications, obtaining large-scale, high-quality annotated data often faces many difficulties and challenges, such as high annotation costs, long annotation cycles, and scarce annotation resources. This seriously limits the application and development of named entity recognition technology in some low-resource scenarios, such as text processing in certain specific fields (medical, legal, etc.), or emerging languages or topic directions that do not have enough annotation accumulation.
[0003] To solve the above problems, researchers began to explore the use of large language models (LLMs) for named entity recognition in low-resource scenarios. Large language models, with their powerful language understanding and generation capabilities obtained through pre-training on massive texts, have shown potential application value in low-resource or even unsupervised situations. Through appropriate prompt design, large language models can be guided to make entity predictions for unlabeled data, thereby generating pseudo-labeled data, which can then be used for further training and optimization of the model.
[0004] However, existing low-resource named entity recognition methods based on large language models still have some shortcomings. On the one hand, the initial zero-shot instruction prompts are often relatively simple and fixed, making it difficult to fully tap the potential of the large language model, resulting in the limited quality of the generated initial pseudo-sample set, which in turn affects the training effect of subsequent models. On the other hand, in the process of using pseudo-sample sets for model training, there is a lack of effective mechanisms to dynamically update and optimize instruction prompts and the pseudo-sample set itself, making it difficult for the model to gradually adapt to data characteristics and task requirements, and difficult to continuously improve recognition performance.
[0005] The patent document with the patent number "CN107563444A" discloses a zero-shot image classification method and system. The problems existing in this method are as follows: It is necessary to pre-obtain semantic auxiliary information of class labels (such as artificially defined attributes, WordNet hierarchies, etc.), which may be difficult to achieve in low-resource scenarios or emerging fields (such as technical terms, minority languages), significantly increasing the application threshold. The optimization relies on fixed regularization terms (such as semantic consistency constraints), and the model update is only completed through convex optimization methods (such as gradient descent), lacking a dynamic adjustment mechanism. In the face of data distribution changes or noise, it is difficult to iteratively correct semantic embedding biases. Summary of the Invention
[0006] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes a low-resource named entity recognition method and system based on evolutionary prompts.
[0007] The technical solution of the present invention is as follows:
[0008] On the one hand, the present invention provides a low-resource named entity recognition method based on evolutionary prompts, including the following steps:
[0009] S1. Use zero-shot instruction prompts to the large language model LLMs to generate initial entity predictions for the unlabeled dataset, obtaining an initial pseudo-example set containing text and pseudo-labels;
[0010] S2. Based on the initial pseudo-example set, construct a definition example library and a difference example library of entity types, and generate a subset of definition examples and a subset of difference examples of entity types through a screening strategy of clustering or setting thresholds;
[0011] S3. Use the large language model LLMs to generate definitions and differences of entity types based on the subset of definition examples and the subset of difference examples;
[0012] S4. Randomly select 2 pseudo-examples from each subset of definition examples of entity types as context learning examples;
[0013] S5. Update the zero-shot instruction based on the definition, difference, and context learning examples, and use the updated zero-shot instruction to prompt the large language model LLMs for a new round of entity prediction;
[0014] S6. Update the initial pseudo-example set by integrating historical entity prediction results and a strategy of dynamically adjusting the weights of historical entity prediction results in the new round of entity prediction;
[0015] S7. Execute step S1 until a preset termination condition is met.
[0016] As a preferred implementation manner, the definition example library D l is expressed as:
[0017] D l = {s1, …, s i , …, s n};
[0018]
[0019] Among them, l represents the entity type, s i represents the i-th element of the defined example library D l , n represents the number of elements in the defined example library D l , x i represents the i-th text in the initial pseudo-example set, represents the i-th text x in the initial pseudo-example set i the j-th pseudo-label belonging to the entity type l, m represents the number of pseudo-labels.
[0020] As a preferred implementation, the similar entity type difference example library is expressed as:
[0021]
[0022]
[0023] Among them, l1l2 represents the similar entity types l1, l2, h i represents the i-th element of the similar entity type difference example library , n represents the number of elements in the similar entity type difference example library , r i represents the i-th pseudo-label in the initial pseudo-example set, represents the i-th pseudo-label r in the initial pseudo-example set i and the j-th text belonging to the entity type l1, represents the i-th pseudo-label r in the initial pseudo-example set i and the j-th text belonging to the entity type l2, m represents the number of texts belonging to the i-th pseudo-label r i and belonging to the entity type l1 in the initial pseudo-example set, b represents the number of texts belonging to the i-th pseudo-label r i and belonging to the entity type l2 in the initial pseudo-example set;
[0024] The method for judging similar entity types is:
[0025] The number of identical elements between entity type l1 and entity type l2 is greater than or equal to a preset similarity threshold.
[0026] As a preferred implementation, the method for obtaining the defined example subset and the difference example subset is:
[0027] For the defined example library D l , use the KMeans clustering algorithm to cluster the embedding representation of the i-th element s of the defined example library D l in the semantic space, and obtain the defined example subset of entity type l by demarcating the clustering center i
[0028] For the similar entity type difference example library Filter by setting an importance threshold:
[0029] If the i-th element h of the similar entity type difference example library belongs to the i-th pseudo-label r i in it and the number m of texts belonging to entity type l1 is greater than or equal to the importance threshold and belongs to the i-th pseudo-label r i and the number b of texts belonging to entity type l2 is also greater than or equal to the importance threshold, then retain h i , otherwise filter it to obtain the difference example subset i
[0030] As a preferred implementation, the method for obtaining the definition of the entity type is:
[0031] Based on the defined example subset instruct the large language model LLMs to generate the definition of the i-th entity type l i
[0032]
[0033] where P M represents the output probability distribution when the large language model performs inference, represents the task description, represents the definition of the entity type l in the previous iteration i represents the defined example subset of the entity type l i ;
[0034] The method for obtaining the difference of the entity type is:
[0035] Based on the difference example subset instruct the large language model LLMs to generate the difference between the i-th entity type l i and the j-th entity type l j
[0036]
[0037] where Represents a task description, Represents entity type l i With entity type l j The subset of difference example.
[0038] On the other hand, the present invention also provides a low-resource named entity recognition system based on evolutionary prompting, including:
[0039] Entity prediction module: Using zero-shot instruction prompting for the large language model LLMs to generate initial entity predictions for the unlabeled dataset, obtaining an initial pseudo-example set containing text and pseudo-labels;
[0040] Example library module: Based on the initial pseudo-example set, construct a definition example library and a difference example library for similar entity types, and generate a subset of definition examples and a subset of difference examples for entity types through a screening strategy of clustering or setting thresholds;
[0041] Definition and difference information module: Using the large language model LLMs to generate the definition and difference of entity types based on the subset of definition examples and the subset of difference examples;
[0042] Learning example module: Randomly select 2 pseudo-examples from the subset of definition examples of each entity type as context learning examples;
[0043] Iterative prediction module: Update the zero-shot instruction based on the definition, difference, and context learning examples, and use the updated zero-shot instruction to prompt the large language model LLMs for a new round of entity prediction; Update the initial pseudo-example set by integrating historical entity prediction results and a strategy of dynamically adjusting the weights of historical entity prediction results in the new round of entity prediction;
[0044] Iterative update module: Sequentially execute the entity prediction module, the example library module, the definition and difference information module, the learning example module, and the iterative prediction module and loop until a preset termination condition is met.
[0045] As a preferred embodiment, the definition example library D l Is represented as:
[0046] D l ={S1,…,S i ,…,s n};
[0047]
[0048] Wherein, l represents the entity type, s i Represents the i-th element of the definition example library D l Of, n represents the number of elements in the definition example library D l In, xi represents the i-th text in the initial pseudo-example set, represents the i-th text x in the initial pseudo-example set i the j-th pseudo-label belonging to entity type l in it, and m represents the number of pseudo-labels.
[0049] As a preferred embodiment, the similar entity type difference example library is represented as:
[0050]
[0051] where, l1l2 represents similar entity types l1, l2, h i represents the i-th element of the similar entity type difference example library and n represents the number of elements in the similar entity type difference example library in it, r i represents the i-th pseudo-label in the initial pseudo-example set, represents the i-th pseudo-label r in the initial pseudo-example set i and the j-th text belonging to entity type l1, represents the i-th pseudo-label r in the initial pseudo-example set i and the j-th text belonging to entity type l2, and m represents the number of texts belonging to the i-th pseudo-label r i in the initial pseudo-example set and belonging to entity type l1, b represents the number of texts belonging to the i-th pseudo-label r i in the initial pseudo-example set and belonging to entity type l2;
[0052] The method for judging similar entity types is:
[0053] The number of elements where entity type l1 and entity type l2 are the same is greater than or equal to a preset similarity threshold.
[0054] As a preferred embodiment, the method for obtaining the defined example subset and the difference example subset is:
[0055] For the defined example library D l , use the KMeans clustering algorithm to cluster the embedded representation of the i-th element s l of the defined example library D i in the semantic space, and obtain the defined example subset of entity type l by delimiting the cluster center
[0056] For the similar entity type difference example library perform screening by setting an importance threshold:
[0057] If the similar entity type difference example library the i-th element h i belonging to the i-th pseudo-label r i and the number of texts m belonging to entity type l1 is greater than or equal to the importance threshold and belonging to the i-th pseudo-label r i and the number of texts b belonging to entity type l2 is also greater than or equal to the importance threshold, then h is retained i , otherwise it is filtered to obtain a subset of different example instances
[0058] As a preferred implementation, the method for obtaining the definition of the entity type is as follows:
[0059] Based on the definition example subset instruct the large language model LLMs to generate the definition g of the i-th entity type l i : i1 :
[0060]
[0061] where P M represents the output probability distribution when the large language model performs inference, represents the task description, represents the definition of the entity type l in the previous iteration i ; represents the definition example subset of the entity type l i ;
[0062] The method for obtaining the difference of the entity type is as follows:
[0063] Based on the difference example subset instruct the large language model LLMs to generate the difference i between the i-th entity type l j and the j-th entity type l
[0064]
[0065] where represents the task description, represents the difference example subset between the entity type l i and the entity type l j ;
[0066] The present invention has the following beneficial effects:
[0067] The present invention generates initial entity predictions for an unlabeled dataset by zero-shot instruction prompting of a large language model, which can make full use of the knowledge learned by the large language model from a large amount of text to generate relatively accurate pseudo-labels, providing a basis for subsequent entity recognition. Constructing a definition example library of entity types and a difference example library of similar entity types based on the initial pseudo-example set can more comprehensively cover the characteristics of different entity types and the differences between similar entity types, providing data support for generating more accurate entity type definitions and differences. Using the large language model to generate entity type definitions and differences based on the definition example subset and the difference example subset can more accurately describe the characteristics of entity types and the differences between similar entity types, thereby improving the accuracy of entity recognition. Updating the zero-shot instruction based on the generated definitions, differences, and context learning examples can enable the large language model to more accurately identify entities in a new round of entity predictions, further improving the accuracy of entity recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 It is a flowchart of the method implementation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.
[0071] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0072] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0073] The term " / and / " refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0074] Embodiment 1:
[0075] SeeFigure 1 , the present invention provides a low-resource named entity recognition method based on evolutionary prompts, comprising the following steps:
[0076] S1. Use zero-shot instruction prompts to the large language model LLMs to generate initial entity predictions for the unlabeled dataset, obtaining an initial pseudo-example set containing text and pseudo-labels;
[0077] S2. Based on the initial pseudo-example set, construct a definition example library for entity types and a difference example library for similar entity types, and generate a subset of definition examples and a subset of difference examples for entity types through a screening strategy of clustering or setting thresholds;
[0078] The entity types include "PLOT", "GENRE", "AVERAGE" in the Movie dataset;
[0079] S3. Use the large language model LLMs to generate definitions and differences of entity types based on the subset of definition examples and the subset of difference examples;
[0080] S4. Randomly select 2 pseudo-examples from the subset of definition examples of each entity type as in-context learning examples;
[0081] S5. Update the zero-shot instruction based on the definition, difference, and in-context learning examples, and use the updated zero-shot instruction to prompt the large language model LLMs for a new round of entity prediction;
[0082] S6. Update the initial pseudo-example set by integrating historical entity prediction results and a strategy of dynamically adjusting the weights of historical entity prediction results in the new round of entity prediction;
[0083] S7. Execute step S1 until a preset termination condition is met, and the termination condition includes a preset number of iterations.
[0084] This embodiment uses four benchmark dialogue datasets for evaluation, namely CONLL03, ACE05, MIT-Movie dataset, and MIT-Restaurant dataset. The detailed information of the datasets is shown in Table 1.
[0085] Table 1 Dataset Information Table
[0086] Dataset Name CONLL03 ACE05 Movie Restaurant Number of Test Examples 3452 1060 2443 1521 Number of Entity Types 4 7 12 8
[0087] To prove the effectiveness of the present invention, the present invention compares it with five baseline frameworks:
[0088] 1. InstructUIE: Enhances the general ability of the model for information extraction tasks through instruction fine-tuning, achieving strong zero-shot generalization for unseen data.
[0089] 2. Gollie: Manual-designed entity type guidance is added to the prompt, and the model's ability to follow entity type guidance is enhanced through instruction fine-tuning.
[0090] 3. CodeIE: By using code-style prompts and large language models pre-trained on code, the model's ability to output structured content is improved, enabling it to more effectively complete information extraction tasks.
[0091] 4. Code4UIE: Selects context demonstrations based on sample similarity using annotated data.
[0092] 5. Self-Improving: Through the idea of ensemble learning, multiple inference paths are designed to stimulate the self-learning ability of large language models to prompt them to complete zero-shot named entity recognition tasks.
[0093] See Table 2 for the comparison results:
[0094] Table 2 Comparison Results Table
[0095]
[0096]
[0097] The #, *, and & in Table 2 represent that the large language model frameworks used by the model are gpt-3.5-turbo, Qwen2-72B-Chat-Int4, and Llama3.1-70B-Chat-Int4, respectively.
[0098] See Table 3 for the ablation results:
[0099] Table 3 Ablation Results Table
[0100]
[0101] As a preferred implementation, the defined example library D l is expressed as:
[0102] D l = {s1,…,s i ,…,s n};
[0103]
[0104] where l represents the entity type, s i represents the i-th element of the defined example library D l , n represents the number of elements in the defined example library D l , and x i represents the i-th text in the initial pseudo-example set. Denote the i-th text x in the initial pseudo-example set i The j-th pseudo-label belonging to entity type l in i , where m represents the number of pseudo-labels
[0105] As a preferred embodiment, the similar entity type difference example library Is denoted as
[0106]
[0107] Where l1l2 represent similar entity types l1, l2, and h i Denote the i-th element of the similar entity type difference example library Where n represents the number of elements in the similar entity type difference example library And r i Denote the i-th pseudo-label in the initial pseudo-example set Denote the r-th text in the initial pseudo-example set belonging to the i-th pseudo-label i And belonging to the j-th text of entity type l1 Denote the r-th text in the initial pseudo-example set belonging to the i-th pseudo-label i And belonging to the j-th text of entity type l2, where m represents the number of texts in the initial pseudo-example set belonging to the i-th pseudo-label r i And belonging to entity type l1, and b represents the number of texts in the initial pseudo-example set belonging to the i-th pseudo-label r i And belonging to entity type l2;
[0108] The method for judging similar entity types is as follows
[0109] The number of identical elements between entity type l1 and entity type l2 is greater than or equal to a preset similarity threshold
[0110] As a preferred embodiment, the method for obtaining the defined example subset and the difference example subset is as follows
[0111] For the defined example library D l , Use the KMeans clustering algorithm to cluster the embedding representation of the i-th element s of the defined example library D l In the semantic space, and obtain the defined example subset of entity type l by delimiting the cluster center i For the similar entity type difference example library
[0112] Filter by setting an importance threshold That is
[0113] If the i-th element h of the similar entity type difference example library Belonging to the i-th pseudo-label r in i In thei and the number m of texts belonging to entity type l1 is greater than or equal to the importance threshold and belongs to the i-th pseudo-label r i and the number b of texts belonging to entity type l2 is also greater than or equal to the importance threshold, then h is retained i , otherwise it is filtered to obtain a subset of different example instances
[0114] As a preferred implementation manner, the method for obtaining the definition of the entity type is as follows:
[0115] Based on the definition example subset instruct the large language model LLMs to generate the definition of the i-th entity type l i of
[0116]
[0117] where P M represents the output probability distribution when the large language model performs inference, represents the task description, represents the definition of the entity type l in the previous iteration i of represents the entity type l i of the definition example subset;
[0118] The method for obtaining the difference of the entity type is as follows:
[0119] Based on the difference example subset instruct the large language model LLMs to generate the difference between the i-th entity type l i and the j-th entity type l j of
[0120]
[0121] where represents the task description, represents the entity type l i and the entity type l j of the difference example subset.
[0122] Example Two:
[0123] The present invention also provides an evolving prompt-based low-resource named entity recognition system, including:
[0124] Entity prediction module: Use zero-shot instruction to prompt the large language model LLMs to generate initial entity predictions for the unlabeled data set, and obtain an initial pseudo-example set containing texts and pseudo-labels;
[0125] Example library module: Based on the initial pseudo-example set, construct a definition example library for entity types and a difference example library for similar entity types, and generate a subset of definition examples and a subset of difference examples for entity types through a screening strategy of clustering or setting thresholds;
[0126] Definition and difference information module: Based on the subset of definition examples and the subset of difference examples, use the large language model LLMs to generate the definitions and differences of entity types;
[0127] Learning example module: Randomly select 2 pseudo-examples from the subset of definition examples for each entity type as context learning examples;
[0128] Iterative prediction module: Update the zero-shot instruction based on the definition, difference, and context learning examples, and prompt the large language model LLMs to perform a new round of entity prediction through the updated zero-shot instruction; Update the initial pseudo-example set by integrating the historical entity prediction results and the strategy of dynamically adjusting the weights of the historical entity prediction results in the new round of entity prediction;
[0129] Iterative update module: Sequentially execute the entity prediction module, example library module, definition and difference information module, learning example module, and iterative prediction module and perform a loop until a preset termination condition is met.
[0130] As a preferred implementation, the definition example library D l is expressed as:
[0131] D l ={S1,…,S i ,…,s n};
[0132]
[0133] where l represents the entity type, s i represents the i-th element of the definition example library D l n represents the number of elements in the definition example library D l x i represents the i-th text in the initial pseudo-example set, represents the i-th text x i in the initial pseudo-example set, the j-th pseudo-label belonging to the entity type l, and m represents the number of pseudo-labels.
[0134] As a preferred implementation, the difference example library for similar entity types is expressed as:
[0135]
[0136] where l1l2 represents similar entity types l1, l2, hi Denote the sample library of differences in similar entity types The i-th element of, where n represents the sample library of differences in similar entity types The number of elements in, r i Denote the i-th pseudo-label in the initial pseudo-sample set, Denote the r that belongs to the i-th pseudo-label in the initial pseudo-sample set i And the j-th text that belongs to entity type l1, Denote the r that belongs to the i-th pseudo-label in the initial pseudo-sample set i And the j-th text that belongs to entity type l2, where m represents the r that belongs to the i-th pseudo-label in the initial pseudo-sample set i And the number of texts that belong to entity type l1, where b represents the r that belongs to the i-th pseudo-label in the initial pseudo-sample set i And the number of texts that belong to entity type l2;
[0137] The method for judging similar entity types is as follows:
[0138] The number of identical elements between entity type l1 and entity type l2 is greater than or equal to a preset similarity threshold.
[0139] As a preferred implementation, the method for obtaining the defined sample subset and the difference sample subset is as follows:
[0140] For the defined sample library D l , use the KMeans clustering algorithm to cluster the embedding representation of the i-th element s of the defined sample library D l in the semantic space, and obtain the defined sample subset of entity type l by demarcating the clustering center i
[0141] For the sample library of differences in similar entity types Filter by setting an importance threshold:
[0142] If the i-th element h of the sample library of differences in similar entity types where the number m of texts that belong to the i-th pseudo-label r i and belong to entity type l1 is greater than or equal to the importance threshold and the number b of texts that belong to the i-th pseudo-label r i and belong to entity type l2 is also greater than or equal to the importance threshold, then retain h i , otherwise filter it to obtain the difference sample subset i
[0143]
[0144] As a preferred implementation, the method for obtaining the definition of the entity type is as follows:
[0144] Based on the defined sample subset Instruct the large language model LLMs to generate the i-th entity type l i Definition of
[0145]
[0146] where P M represents the output probability distribution when the large language model performs inference, represents the task description, represents the definition of the entity type l in the previous iteration i Definition of represents the definition of the entity type l i The sample subset of the definition;
[0147] The method for obtaining the difference of entity types is as follows:
[0148] Based on the difference sample subset Instruct the large language model LLMs to generate the i-th entity type l i and the difference between the j-th entity type l j Difference of
[0149]
[0150] where represents the task description, represents the entity type l i and the entity type l j Difference sample subset.
[0151] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent the situation of A existing alone, A and B existing simultaneously, and B existing alone. Where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and rear associated objects. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.
[0152] Those of ordinary skill in the art will realize that the various units and algorithm steps described in the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0153] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0154] In several embodiments provided by this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (hereinafter referred to as ROMs), random access memories (hereinafter referred to as RAMs), magnetic disks, or optical discs that can store program codes.
[0155] The above are only the embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A low-resource named entity recognition method based on evolutionary hints, characterized in that: The following steps are involved: S1. Prompt the large language model LLMs to generate initial entity predictions for the unlabeled dataset through zero-shot instructions, and obtain an initial pseudo-sample set containing text and pseudo-labels; S2. Based on the initial pseudo sample set, construct a definition sample library of entity types and a difference sample library of similar entity types, and generate a subset of definition samples and a subset of difference samples of entity types through clustering or a screening strategy with a set threshold; S3, using large language models (LLMs) to generate definitions and differences of entity types based on a subset of definition samples and a subset of difference samples; S4, randomly select 2 pseudo samples from the definition sample set of each entity type as context learning samples; S5, updating the zero-shot instruction based on the definition, difference and context learning examples, and prompting the large language model LLMs to perform a new round of entity prediction through the updated zero-shot instruction; S6, updating the initial pseudo sample set by integrating historical entity prediction results in a new round of entity prediction and dynamically adjusting the weights of historical entity prediction results; S7. Execute step S1 until a preset termination condition is met.
2. The low-resource named entity recognition method based on evolutionary hints according to claim 1 is characterized in that: The definition sample library D l It is expressed as: D l ={s1,…,s i ,…,s n }; Among them, l represents the entity type, s i Indicates the definition sample library D l The i-th element of n represents the definition of the sample library D l The number of elements in x i represents the i-th text in the initial pseudo sample set, Represents the i-th text x in the initial pseudo sample set i is the jth pseudo-label of entity type l in , and m represents the number of pseudo-labels.
3. The low-resource named entity recognition method based on evolutionary hints according to claim 2 is characterized in that: The similar entity type difference sample library It is expressed as: Among them, l1l2 represents similar entity types l1, l2, h i A sample library representing differences in similar entity types The i-th element of n represents the sample library of similar entity type differences The number of elements in r i represents the i-th pseudo label in the initial pseudo sample set, Indicates the number of pseudo-labels r in the initial pseudo-sample set i and belongs to the jth text of entity type l1, Indicates the number of pseudo-labels r in the initial pseudo-sample set i And belongs to the jth text of entity type l2, m represents the i-th pseudo label r in the initial pseudo sample set i And the number of texts belonging to entity type l1, b represents the number of texts belonging to the i-th pseudo label r in the initial pseudo sample set i And the number of texts belonging to entity type l2; Similar entity type determination methods are: The number of elements that are the same between entity type l1 and entity type l2 is greater than or equal to a preset similarity threshold.
4. The low-resource named entity recognition method based on evolutionary hints according to claim 3 is characterized in that: The method for obtaining the subset of definition samples and the subset of difference samples is as follows: For defining sample library D l , use KMeans clustering algorithm to define sample library D l The i-th element s i Clustering is performed on the embedded representation in the semantic space, and a subset of definition examples of entity type l is obtained by defining the cluster center. Difference sample library for similar entity types Filter by setting a significance threshold: If similar entity types differ in sample libraries The i-th element h i The pseudo label r in i And the number of texts m belonging to entity type l1 is greater than or equal to the importance threshold and belongs to the i-th pseudo label r i And the number of texts b belonging to entity type l2 is also greater than or equal to the importance threshold, then h is retained i , otherwise filter and get a subset of difference samples 5. The low-resource named entity recognition method based on evolutionary hints according to claim 4 is characterized in that: The definition acquisition method of the entity type is: Based on the definition sample subset Instruct the large language model LLMs to generate the i-th entity type l i Definition Among them, P M represents the output probability distribution of the large language model when performing inference, Represents the task description, Indicates the entity type l of the last iteration i Definition, Represents entity type l i A collection of sample definitions of The difference obtaining method of entity type is: Based on the difference sample subset Instruct the large language model LLMs to generate the i-th entity type l i With the j-th entity type l j The Difference in, Represents the task description, Represents entity type l i With entity type l j A subset of difference samples.
6. A low-resource named entity recognition system based on evolutionary prompts, characterized in that: include: Entity prediction module: Prompts large language models (LLMs) to generate initial entity predictions for unlabeled datasets through zero-shot instructions, and obtains an initial set of pseudo-samples containing text and pseudo-labels; Sample library module: based on the initial pseudo sample set, constructs a definition sample library of entity types and a difference sample library of similar entity types, and generates a subset of definition samples and a subset of difference samples of entity types through clustering or a screening strategy with a set threshold; Definition and difference information module: Generates the definition and difference of entity types based on a subset of definition samples and a subset of difference samples using large language models (LLMs); Learning sample module: Randomly select 2 pseudo samples from the definition sample set of each entity type as context learning samples; Iterative prediction module: updates the zero-shot instructions based on the definitions, differences, and contextual learning examples, and prompts the large language models (LLMs) to perform a new round of entity prediction through the updated zero-shot instructions; Update the initial pseudo sample set by integrating historical entity prediction results in the new round of entity prediction and dynamically adjusting the weights of historical entity prediction results; Iterative update module: execute the entity prediction module, sample library module, definition and difference information module, learning sample module and iterative prediction module in sequence and loop until the preset termination condition is met.
7. The low-resource named entity recognition system based on evolutionary hints according to claim 6, characterized in that: The definition sample library D l It is expressed as: D l ={S1,…,S i ,…,s n }; Among them, l represents the entity type, s i Indicates the definition sample library D l The i-th element of n represents the definition of the sample library D l The number of elements in x i represents the i-th text in the initial pseudo sample set, Represents the i-th text x in the initial pseudo sample set i is the jth pseudo-label of entity type l in , and m represents the number of pseudo-labels.
8. The low-resource named entity recognition system based on evolutionary hints according to claim 7, characterized in that: The similar entity type difference sample library It is expressed as: D l1l2 ={h1,…,h i ,…,h n }; Among them, l1l2 represents similar entity types l1, l2, h i A sample library representing differences in similar entity types The i-th element of n represents the sample library of similar entity type differences The number of elements in r i represents the i-th pseudo label in the initial pseudo sample set, Indicates the number of pseudo-labels r in the initial pseudo-sample set i and belongs to the jth text of entity type l1, Indicates the number of pseudo-labels r in the initial pseudo-sample set i And belongs to the jth text of entity type l2, m represents the i-th pseudo label r in the initial pseudo sample set i And the number of texts belonging to entity type l1, b represents the number of texts belonging to the i-th pseudo label r in the initial pseudo sample set i And the number of texts belonging to entity type l2; The similar entity type determination method is: The number of elements that are the same between entity type l1 and entity type l2 is greater than or equal to a preset similarity threshold.
9. The low-resource named entity recognition system based on evolutionary hints according to claim 8, characterized in that: The method for obtaining the subset of definition samples and the subset of difference samples is as follows: For defining sample library D l , use KMeans clustering algorithm to define sample library D l The i-th element s i Clustering is performed on the embedded representation in the semantic space, and a subset of definition examples of entity type l is obtained by defining the cluster center. Difference sample library for similar entity types Filter by setting a significance threshold: If similar entity types differ in sample libraries The i-th element h i The pseudo label r in i And the number of texts m belonging to entity type l1 is greater than or equal to the importance threshold and belongs to the i-th pseudo label r i And the number of texts b belonging to entity type l2 is also greater than or equal to the importance threshold, then h is retained i , otherwise filter and get a subset of difference samples 10. The low-resource named entity recognition system based on evolutionary hints according to claim 9, characterized in that: The definition acquisition method of the entity type is: Based on the definition sample subset Instruct the large language model LLMs to generate the i-th entity type l i Definition Among them, P M represents the output probability distribution of the large language model when performing inference, Represents the task description, Indicates the entity type l of the last iteration i Definition, Represents entity type l i A collection of sample definitions of The difference obtaining method of entity type is: Based on the difference sample subset Instruct the large language model LLMs to generate the i-th entity type l i With the j-th entity type l j The Difference in, Represents the task description, Represents entity type l i With entity type l j A subset of difference samples.
Citation Information
Patent Citations
Zero sample image classification method and system
CN107563444A