Large language model context learning method for few-sample named entity recognition and named entity recognition method

Through the context learning method of large language model and the self-consistency verification mechanism, the problems of data scarcity and model generalization in the recognition of few-sample named entity are solved, and the accuracy and reliability of named entity recognition are significantly improved.

CN119962534APending Publication Date: 2025-05-09HENAN UNIVERSITY
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510040231.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing method of naming entity recognition of small samples has failed to fully stimulate the reasoning ability of large language models in the field of information extraction, and faces the challenges of data scarcity and model generalization.

Method used

The context learning method of large language model is adopted, and the naming entity recognition performance of the model is optimized by constructing a small sample support set and a test set, using the prompt template and thinking chain prompt for context learning, combined with the self-consistency verification mechanism.

Benefits of technology

It significantly enhances the model's naming entity understanding and recognition effect in resource scarce scenarios, improves the model's inference ability and recognition accuracy, and reduces the risk of wrong judgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962534A_ABST
    Figure CN119962534A_ABST
Patent Text Reader

Abstract

The invention provides a large language model context learning method for few-sample named entity recognition and a named entity recognition method. The context method comprises the following steps: constructing a few-sample support set and a test set for named entity recognition; loading a prompt template based on context learning, and inputting the few-sample support set into a preset large language model, so that the large language model generates a new text set according to the prompt template, and adding the new text set to at least the sample support set to form an expanded support set; and zero sample thinking chain prompt is introduced into the original prompt template, the test set and the extended support set are jointly input into a preset large language model for similarity comparison, a named entity recognition result about the test set is predicted, and a natural language text corresponding to the predicted correct named entity recognition result is fed back to at least the sample support set. According to the method, the problem of data scarcity of the named entity recognition task can be solved, and the generalization of the model on the named entity recognition task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a large language model context learning method and a named entity recognition method for small-sample named entity recognition. Background Art

[0002] Named entity recognition aims to identify entities with specific meanings from text, such as people, places, and units. As a basic and key language understanding task, named entity recognition is widely used in information extraction, knowledge graph construction, and other downstream applications of natural language understanding. The named entity recognition task requires not only identifying entity types, but also accurately locating entity boundaries, which requires the model to have a deep understanding of contextual semantics. Past research has mostly focused on named entity recognition with sufficient data samples, while ignoring the problem of limited number of samples in actual scenarios. In specific fields or with limited resources, it is difficult and expensive to obtain large-scale annotated data. Therefore, few-sample named entity recognition has become a challenging and practically important research problem.

[0003] There are three main research branches in the task of few-shot named entity recognition: prototype-based methods, knowledge transfer-based methods, and data enhancement-based methods. Prototype-based methods classify named entities based on the similarity between limited labeled data and the named entities, but they perform poorly when identifying unseen entities and lack generalization; knowledge transfer-based methods use high-resource external knowledge such as pre-trained language models and domain dictionaries to improve model performance, but due to the inconsistent distribution of training data and differences in label sets, noise may be introduced, leading to knowledge mismatch problems; common data enhancement-based methods increase the richness of the text while retaining the overall semantics of the sentence as much as possible, including enriching the context of each word and adding diverse grammar and vocabulary structures for similar semantic expressions. However, directly applying data enhancement methods in the absence of sufficient data can easily lead to overfitting risks.

[0004] Existing few-shot named entity recognition methods often fail to fully stimulate the reasoning ability of large language models (LLMs) in the field of information extraction. Although context-based learning methods have achieved excellent results in some complex reasoning tasks, there is still a gap in the field of information extraction. In general, few-shot named entity recognition mainly faces the following two challenges: data scarcity challenge and model generalization challenge.

[0005] The core of the data scarcity challenge lies in how to make full use of limited data. The feature limitation problem stems from the data scarcity challenge, which will limit the model's ability to understand the global feature distribution. Especially in natural language processing, words are closely connected in sentences to form a complex association network. Discrete feature representations are difficult to capture the continuous relationship between words. Although the Transformer model alleviates this problem to a certain extent through the attention mechanism, it is still challenging to obtain sufficient semantic and contextual information in the case of few samples.

[0006] The key to the model generalization challenge lies in how to solve the hallucination and imbalance problems inherent in large models. Insufficient generalization ability makes it difficult for the model to effectively use previously learned knowledge when facing new fields or different data distributions. For example, in different data sets, the same entity may be labeled as different types. This knowledge mismatch problem will cause a knowledge transfer disaster and seriously affect the reliability of few-sample named entity recognition tasks in practical applications. Summary of the invention

[0007] When faced with the challenges of data scarcity and model generalization, the present invention provides a large language model context learning method and a named entity recognition method for few-sample named entity recognition, which can effectively address the problems of entity type diversity and lack of high-quality labeled data, thereby significantly enhancing the model's understanding ability and recognition effect of named entities in resource-scarce scenarios.

[0008] In a first aspect, the present invention provides a large language model context learning method for few-sample named entity recognition, comprising:

[0009] Step 1: Construct a few-shot support set and test set for named entity recognition;

[0010] Step 2: Loading a prompt template based on context learning, and inputting the few-sample support set into a preset large language model, so that the large language model generates a new text set according to the prompt template, and adding the new text set to the few-sample support set to form an expanded support set;

[0011] Step 3: Introduce zero-sample thinking chain prompts into the original prompt template, input the test set and the expanded support set into the preset large language model for similarity comparison, predict the named entity recognition results about the test set, and feed back the natural language text corresponding to the predicted correct named entity recognition results to the few-sample support set.

[0012] Furthermore, in step 1, the process of constructing the few-sample support set specifically includes:

[0013] Collect texts, and perform word segmentation and entity annotation on each text; the entity annotation information includes: the type of each entity in the text, the index of the entity in the word segmentation sequence, and the natural language text corresponding to the entity.

[0014] Furthermore, the prompt template includes instructions, examples and text; wherein the instructions are used to indicate the entity type that needs to be identified; the examples are used to predefine the answer space mapping mechanism of the named entity recognition task, and the text refers to the target sentence that needs to identify the entity type.

[0015] Furthermore, after step 3, the method further includes:

[0016] Step 4: Adopting a self-consistency verification mechanism to iteratively optimize the large language model; the self-consistency verification mechanism specifically includes:

[0017] Step 4.1: Divide the incorrectly predicted named entity recognition results into coarse-grained and fine-grained ones, and annotate each incorrectly predicted named entity recognition result with a corresponding error type label;

[0018] Step 4.2: Feedback the predicted wrong named entity recognition results and the corresponding error type labels to the large language model for multiple rounds of reasoning, and obtain new named entity recognition results by adopting a voting strategy based on the multiple rounds of reasoning results;

[0019] Step 4.3: Verify the new named entity recognition result. If it is still wrong, repeat steps 4.1 to 4.2 until the large language model infers the correct named entity recognition result.

[0020] In a second aspect, the present invention provides a method for named entity recognition based on a large language model, comprising:

[0021] Step 1: Select an example that best matches the text to be recognized from a pre-built example template; the example is used for an answer space mapping mechanism for a predefined named entity recognition task;

[0022] Step 2: taking the entity type to be identified in the text to be identified as an instruction, and loading the instruction, the selected example and the text to be identified into a preset prompt template based on context learning;

[0023] Step 3: Input the loaded prompt template into a preset large language model to obtain a named entity recognition result of the text to be recognized; wherein the preset large language model is obtained by using the context learning method described in the first aspect.

[0024] Beneficial effects of the present invention:

[0025] (1) This paper redefines the named entity recognition task as a natural language generation task, making full use of the context learning idea. By introducing context information into the prompt template, the model can dynamically adjust its generation strategy for the named entity recognition task to better adapt to changes in the input text. This flexibility enables the model to quickly capture key information and make accurate judgments when faced with diverse texts.

[0026] (2) The present invention introduces thought chain prompts based on the generative model, which can greatly improve the model's reasoning ability in a simple and effective way. It can not only guide the model to perform step-by-step reasoning, but also maintain consistency during the reasoning process, thereby reducing the risk of misjudgment.

[0027] (3) The present invention combines contextual learning with the idea of ​​self-consistency verification, making full use of the advantages of generative pre-trained language models, which not only alleviates the problem of data scarcity, but also significantly enhances the model's ability to understand and express text semantics. In addition, combined with the self-consistency verification mechanism, the model can evaluate the rationality of its output in real time during the generation process, further improving the reliability of the results. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 An overall framework diagram of a large language model context learning method for few-sample named entity recognition provided by an embodiment of the present invention;

[0029] Figure 2 A schematic diagram of a support set provided by an embodiment of the present invention;

[0030] Figure 3 The inference phase and the evaluation phase of a large language model context learning method for few-sample named entity recognition provided by an embodiment of the present invention;

[0031] Figure 4 The present invention provides a flowchart of a method for named entity recognition based on a large language model. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] This paper aims to redefine the named entity recognition task from the traditional sequence labeling task to a natural language generation task. Figure 1 , Figure 2 and Figure 3 As shown, an embodiment of the present invention provides a large language model context learning method for few-sample named entity recognition, comprising the following steps:

[0034] S101: Constructing few-shot support and test sets for named entity recognition;

[0035] Specifically, this embodiment uses N-way K-shot to set the entire named entity recognition task. The support set contains N entity types, each entity type corresponds to K annotated texts, and a small amount of annotated texts are provided by the support set to improve the contextual learning ability of the large language model. The test set also contains N entity types, each entity type has Q texts, and the Q value is generally much larger than the K value. The N entity types in the support set and the test set are the same, but the text data is different.

[0036] In an exemplary embodiment, the process of constructing a few-sample support set includes: collecting texts, and performing word segmentation and entity labeling on each text; wherein the entity labeling information includes: the type of each entity in the text, the index of the entity in the word segmentation sequence, and the natural language text corresponding to the entity.

[0037] For example, given a text X, the text X is segmented to obtain a segmentation sequence, which is formally represented as X = {x1, x2…x n}, n represents the number of words; assuming that in the order of words, from the lth word to the rth word represents an entity, and the type of the entity is marked as t, then the annotation information of the entity can be formally expressed as: ["type":"t","offset":[l,l+1,l+2,…,r],"text":"x l …x r ”]; where 0≤l≤r≤n, type represents the entity type, offset represents the index, text represents the natural language text corresponding to the entity, and l and r represent the left and right boundary indexes of the entity respectively. It can be understood that a text sentence can include multiple entities of the same entity type or multiple entities of different entity types. Figure 2 shown.

[0038] S102: loading a prompt template based on context learning, and inputting the few-sample support set into a preset large language model, so that the large language model generates a new text set according to the prompt template, and adding the new text set to the few-sample support set to form an expanded support set;

[0039] Specifically, context-based learning of prompt templates is a technique for guiding large language models to generate specific outputs based on given contexts. Such prompt templates can help large language models better understand named entity recognition tasks and improve their performance on named entity recognition tasks.

[0040] In an exemplary embodiment, the context learning prompt template includes three parts: instruction (I), example (D) and text (S), which can be set in the following form:

[0041]

[0042] The instruction is used to indicate the entity type to be recognized. The example is used to predefine the answer space mapping mechanism of the named entity recognition task, so as to ensure the consistency between the generated named entity recognition results and the task requirements. The text refers to the target sentence for which the entity type needs to be recognized.

[0043] It should be noted that in the prompt template, the example can be one or more, and it should be ensured that one entity type in the instruction corresponds to at least one example. It can be understood that through the above prompt template, this embodiment can unify entity location and entity recognition into one round of prompt learning, effectively guide the large language model to learn and reason in a variety of contexts, enhance the ability of the large language model to "learn to learn", enable it to effectively imitate existing examples to solve new problems, and thus improve the processing ability of different entity types.

[0044] S103: Introduce a zero-sample thinking chain prompt into the original prompt template, input the test set and the expanded support set into the preset large language model for similarity comparison, predict the named entity recognition result of the test set, and feed back the natural language text corresponding to the predicted correct named entity recognition result to the few-sample support set.

[0045] Specifically, the zero-shot thinking chain prompt does not require manual writing of derivation examples, but only requires adding a sentence "Let's think step by step" to the end of the existing prompt word. With the help of zero-shot thinking prompts, the reasoning ability of large language models can be greatly improved, and the interpretability of large language models can also be improved.

[0046] The context learning method provided by the embodiment of the present invention makes full use of the context learning concept in the reasoning stage, introduces context information such as "instructions" and "examples" in the prompt template, so that the large language model can dynamically adjust its generation strategy to better adapt to the changes in the input text. This flexibility enables the model to quickly capture key information and make accurate judgments when facing diverse texts; in the testing stage, by further introducing thinking chain prompts in the prompt template, the model reasoning ability can be greatly improved in a simple and effective way, which can not only guide the model to perform step-by-step reasoning, but also maintain consistency during the reasoning process, reducing the risk of wrong judgment. The present invention solves the dilemma of few samples in the named entity recognition task from the perspective of model reasoning.

[0047] On the basis of the above embodiment, in order to solve the hallucination problem that may exist in the large language model, the embodiment of the present invention further adopts a self-consistency verification mechanism to iteratively optimize the reasoning of the large language model to ensure that the generated answer is more credible. The self-consistency verification mechanism specifically includes the following steps:

[0048] S201: dividing the named entity recognition results predicted incorrectly into coarse-grained and fine-grained categories, and marking a corresponding error type label for each named entity recognition result predicted incorrectly;

[0049] Specifically, in the process of named entity recognition, the prediction errors are divided into coarse-grained and fine-grained categories, aiming to guide the model to better understand and locate the errors that occur during the reasoning process, so as to optimize the subsequent processing flow. The coarse-grained division is mainly divided into two categories: unpredicted entities and prediction errors. Unpredicted entities refer to entities that the model fails to recognize, while prediction errors refer to entities that the model incorrectly recognizes or recognizes the wrong entity type. Further, in the category of prediction errors, the fine-grained division is: unpredicted (the model fails to predict any entity), wrong label (the entity label predicted by the model does not match the true label), wrong span (the left and right indexes of the entity recognized by the model do not match the left and right indexes of the true entity), wrong boundary (the model incorrectly recognizes the left or right boundary of the entity) and all errors (that is, the model has problems in multiple aspects). In practical applications, one-hot encoding can be used to encode the above six error types to generate error type labels.

[0050] S202: Feeding back the predicted wrong named entity recognition result and the corresponding error type label to the large language model for multiple rounds of reasoning, and obtaining a new named entity recognition result by adopting a voting strategy based on the multiple rounds of reasoning results;

[0051] Specifically, the large language model generates a prediction result in each round of reasoning. All prediction results are collected, the same prediction results are counted as one category, the number of prediction results in each category is determined, and the prediction results in the category with the largest number are taken as the final prediction results.

[0052] For example, "Zhengzhou is person" is a named entity recognition result with an incorrect prediction, and the error type is "wrong label". A token sequence of "Zhengzhou is person" is generated, and the token sequence and "wrong label" information are input into the large language model. The large language model is set to perform three rounds of reasoning. Assume that in the first round of reasoning, the prediction result of the large language model is: "Zhengzhou is organization"; in the second round of reasoning, the prediction result of the large language model is "Zhengzhou is location"; in the third round of reasoning, the prediction result of the large language model is "Zhengzhou is location". According to the voting strategy, the number of votes for "Zhengzhou is location" is 2, and the number of votes for "Zhengzhou is organization" is 1, so the new named entity recognition result is determined to be: "Zhengzhou is location".

[0053] In this embodiment, multiple rounds of reasoning are combined with a voting strategy to determine the final prediction result, which can reduce the impact of accidental errors or noise in a single reasoning process.

[0054] S203: Verify the new named entity recognition result. If it is still wrong, repeat steps S201 to S202 until the large language model infers a correct named entity recognition result.

[0055] Specifically, in order to verify the rationality of the prediction results, it can be verified in combination with existing template information. During the voting process, the template information includes context information and known entity distribution rules, etc. This information will be used to check and verify the prediction results to ensure that the selected entities are consistent with the context and effectively avoid irrelevant entities from being misidentified.

[0056] Still taking the above named entity recognition result "Zhengzhou is location" as an example, by comparing it with the real named entity information, it can be seen that the recognition result is correct. At this time, the above iterative process can be ended to obtain the optimized large language model.

[0057] Furthermore, the verified prediction results can be further fed back to the model and enter the loop optimization process. Specifically, the verified prediction results are input into the model as new input information to help the model adjust its reasoning strategy and improve its ability to identify similar errors. This feedback mechanism is an iterative process, and the model is continuously optimized in each cycle, thereby improving its consistency and recognition accuracy in practical applications.

[0058] like Figure 4 As shown, the embodiment of the present invention also provides a method for named entity recognition based on a large language model, comprising the following steps:

[0059] S301: Selecting an example that best matches the text to be recognized from a pre-built example template; the example is used to predefine an answer space mapping mechanism for a named entity recognition task;

[0060] Specifically, the similarity between the text to be recognized and the example template is compared, and the example template that best matches the current text to be recognized is adaptively selected to enhance the learning effect of the model.

[0061] The CoNLL dataset is used as an example to illustrate the specific process of similarity comparison. From the exemplary context learning prompt template given in step S102 above, it can be seen that the context learning prompt template includes three parts: instruction (I), example (D) and text (S), where the instruction is an entity type and can be formalized as: instruction + = 'Entity types:"location","person","organization","miscellaneous"\n\n'; In the constructed support set, in addition to the formalization of entity annotations, an entity type mapping set entity.schema is provided, which can be formalized as ["location","person","organization","miscellaneous"]. The model automatically loads the entity type mapping set and parses the entity type in the instruction, compares the entity type mapping set in the text to be recognized with the example template, and uses the set intersection method to calculate the similarity between the instruction and the entity type mapping set. If the intersection returns intersection = {'location','person','organization','miscellaneous'}, a similarity score similarity_score = 1.0 with a value of 1 will be obtained. When the similarity reaches the specified threshold of 1.0, the model believes that the template matches the current task, and thus uses the template for subsequent processes such as entity recognition. It can be understood that the matching threshold of similarity can be adjusted according to needs.

[0062] S302: taking the entity type to be recognized in the text to be recognized as an instruction, and loading the instruction, the selected example and the text to be recognized into a preset prompt template based on context learning;

[0063] S303: Input the loaded prompt template into a preset large language model to obtain a named entity recognition result of the text to be recognized; wherein the preset large language model is obtained by using the context learning method of the above embodiment.

[0064] The method for named entity recognition based on a large language model provided by the present invention aims to effectively deal with the low-resource dilemma of named entity recognition from the perspective of model reasoning. This method can significantly improve the entity recognition performance in a few-sample scenario through context learning without updating the model parameters, thereby effectively alleviating the challenge of data scarcity; at the same time, combined with the zero-sample thinking chain and self-consistency verification mechanism, it not only enhances the interpretability of the large model, but also effectively alleviates the problems of generative hallucinations and sample imbalance. In general, the method of the present invention significantly improves the accuracy of the few-sample named entity recognition task, which is significantly better than the traditional "pre-training + fine-tuning" paradigm.

[0065] In order to verify the effectiveness of the context learning method and the named entity recognition method provided by the present invention, an experimental simulation is performed using the T5 model as an example, and compared with a traditional named entity recognition scheme.

[0066] This experiment uses the pre-trained model T5 to encode the input text. Different from the traditional sequence labeling task, there is no need for labeling and complex feature engineering. The pre-trained model T5 can process the input text in an end-to-end manner and convert the input information into context-related encoding representations, thereby effectively capturing contextual information and improving the accuracy and robustness of named entity recognition.

[0067] The T5 model obtains the entity category and the position index on both sides of the entity according to the reasoning strategy, decodes the entity in the sentence, and finally converts an entity triple E i = [l, t, r] is organized into natural language output, such as <X l,r >isa / an <t>entity.

[0068] The present invention uses the pre-trained model T5-large (parameter size is about 770M) as the backbone network, and also uses it for the encoding and decoding process of the text, and uses a 409080G GPU for large-scale pre-training on Wikipedia data. The pre-training other hyperparameter settings are shown in Table 1.

[0069] Table 1 shows the hyperparameter settings.

[0070]

[0071]

[0072] Evaluation criteria: Precision P, recall R, and F1 score are selected as the evaluation criteria for the experiment. For the named entity recognition task, it is stipulated that the predicted entity is considered correct only when the type and position offset of the predicted entity completely match the true result. Based on the true label of the text data and the predicted label generated by the method of the present invention, the prediction results can be divided into four categories: true positive TP, true negative TN, false positive FP, and false negative TN. The calculation formulas of each evaluation index are as follows:

[0073] 1) Accuracy:

[0074] 2) Recall rate:

[0075] 3) F1 score:

[0076] Table 2 lists the detailed information of the six public datasets involved in this experiment.

[0077] Table 2 shows the statistical information of the data set.

[0078]

[0079] In the first group of experiments, the model proposed in the present invention was compared on 6 data sets, and the few-sample condition was set to the extreme 1-shot. The comparison models selected by the present invention include the text-to-text generation model T5-base and its extended version T5-large, the generative pre-trained Transformer model GPT-2, the generation model OPT based on GPT-3, the meta-training named entity recognition method MetaNER-base based on contextual learning and its extended model MetaNER.

[0080] Table 3 shows the 1-shot experimental results of the proposed method on 6 public datasets.

[0081]

[0082] The first set of experiments aims to verify the superiority of the contextual learning method and named entity recognition method proposed in this paper, especially the performance under extreme sample conditions. The experimental results show that the proposed model COT-SCNER significantly outperforms other models on all data sets, and the F1 score is improved by an average of 19.78% compared with the OPT model. In addition, compared with the previous strongest MetaNER method, the F1 score of CoT-SCNER is improved by an average of 6.23%. This result fully demonstrates that CoT-SCNER performs better in knowledge transfer, demonstrating its ability to effectively capture and utilize contextual information in a few-sample learning environment, and further proves the potential and value of this method in practical applications.

[0083] Table 4 shows the 5-shot experimental results of the proposed method on 6 public datasets.

[0084]

[0085] From the perspective of reasoning generalization, under the extreme 1-shot and 5-shot settings, CoT-SCNER achieved the best performance on all datasets, with F1 scores increased by 19.78% and 11.10% respectively compared to OPT; and compared to MetaNER, F1 scores increased by 6.23% and 2.62% respectively. These results fully demonstrate that CoT-SCNER has stronger generalization ability when facing few-shot learning, and can more effectively capture the key contextual information between examples, thereby improving the performance of few-shot named entity recognition.

[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.< / t>

Claims

1. A large language model context learning method for few-shot named entity recognition, characterized in that: include: Step 1: Construct a few-shot support set and test set for named entity recognition; Step 2: Loading a prompt template based on context learning, and inputting the few-sample support set into a preset large language model, so that the large language model generates a new text set according to the prompt template, and adding the new text set to the few-sample support set to form an expanded support set; Step 3: Introduce zero-sample thinking chain prompts into the original prompt template, input the test set and the expanded support set into the preset large language model for similarity comparison, predict the named entity recognition results about the test set, and feed back the natural language text corresponding to the predicted correct named entity recognition results to the few-sample support set.

2. The large language model context learning method for few-sample named entity recognition according to claim 1, characterized in that: In step 1, the process of constructing the few-sample support set specifically includes: Collect texts, and perform word segmentation and entity annotation on each text; the entity annotation information includes: the type of each entity in the text, the index of the entity in the word segmentation sequence, and the natural language text corresponding to the entity.

3. The self-consistency verification thought chain context learning method for few-sample named entity recognition according to claim 1 is characterized in that: The prompt template includes instructions, examples and text; wherein the instructions are used to indicate the entity type that needs to be identified; the examples are used to predefine the answer space mapping mechanism of the named entity recognition task, and the text refers to the target sentence that needs to identify the entity type.

4. A self-consistency verification thought chain context learning method for few-sample named entity recognition according to any one of claims 1 to 3, characterized in that: After step 3, also include: Step 4: Adopting a self-consistency verification mechanism to iteratively optimize the large language model; the self-consistency verification mechanism specifically includes: Step 4.1: Divide the incorrectly predicted named entity recognition results into coarse-grained and fine-grained ones, and annotate each incorrectly predicted named entity recognition result with a corresponding error type label; Step 4.2: Feedback the predicted wrong named entity recognition results and the corresponding error type labels to the large language model for multiple rounds of reasoning, and obtain new named entity recognition results by adopting a voting strategy based on the multiple rounds of reasoning results; Step 4.3: Verify the new named entity recognition result. If it is still wrong, repeat steps 4.1 to 4.2 until the large language model infers the correct named entity recognition result.

5. A method for named entity recognition based on a large language model, characterized in that: include: Step 1: Select an example that best matches the text to be recognized from a pre-built example template; the example is used for an answer space mapping mechanism for a predefined named entity recognition task; Step 2: taking the entity type to be identified in the text to be identified as an instruction, and loading the instruction, the selected example and the text to be identified into a preset prompt template based on context learning; Step 3: Input the loaded prompt template into a preset large language model to obtain a named entity recognition result of the text to be recognized; wherein the preset large language model is obtained by using the context learning method according to any one of claims 1 to 4.

Citation Information

Cited By

  • Zero-sample named entity recognition method and device based on large model feedback optimization

    CN120354855A

  • Emergency distortion information identification method and system

    CN120744342A

  • Small-sample contrast enhancement fine tuning method and system based on large language model

    CN121052322A

  • A few-shot contrastive augmentation fine-tuning method and system based on a large language model

    CN121052322B

  • Method for carrying out noise label learning from code annotation data

    CN121117590A