Social media-oriented open type named entity identification method and platform

Through the combination of Monte Carlo tree search algorithm and large language model, the prompts and reconstructing social media texts are optimized, which solves the difficulty of naming entity recognition in noise and new word processing in the prior art, and achieves a highly accurate naming entity recognition effect.

CN120124630APending Publication Date: 2025-06-10SHANDONG WOMENS UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510186709.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing large language models do not perform well in open named entity recognition tasks in social media texts, especially when dealing with noise and new words, making it difficult to achieve accurate named entity recognition.

Method used

Using a combination of Monte Carlo tree search algorithm and large language model, prompts are automatically optimized to reconstruct social media short texts, converting them into low-noise formal texts with standard syntactic structures, thereby realizing named entity recognition.

Benefits of technology

Through the combination of text reconstruction and concept correction thinking chain templates, the accuracy and efficiency of naming entity recognition are significantly improved, and the naming entities in social media text can be effectively identified under zero sample settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124630A_ABST
    Figure CN120124630A_ABST
Patent Text Reader

Abstract

The invention discloses an open type named entity recognition method and platform for social media, relates to the technical field of natural language processing, and aims to automatically optimize prompts for a large model based on a Monte Carlo tree search algorithm and reconstruct texts according to the prompts. A social media short text is converted into a low-noise formal text with a standard syntactic structure, and then a concept correction thinking chain template is utilized to extract named entities from a reconstructed text and classify the named entities. The method is used for carrying out open naming body recognition on the disordered social media short text under the condition that fine adjustment is not carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and more specifically, to an open named entity recognition method and platform for social media. Background Art

[0002] The task of named entity recognition (NER) is to identify named entities from a large amount of text and classify them into predefined types. Traditional NER methods mainly focus on entity boundary recognition and entity classification tasks, usually for grammatically and structurally complete sentences. With the popularity and widespread use of social platforms such as Twitter and Facebook, how to accurately and efficiently extract key information from a large amount of unstructured or semi-structured data poses a severe challenge to traditional NER methods. This challenge is twofold: on the one hand, new words constantly appear in social media texts, which requires the model to have strong generalization ability to handle new entities and new entity types beyond the predefined ones, which is in line with the characteristics of open NER; on the other hand, the characteristics of social media texts, such as noise and abbreviations, make methods relying on grammar and structure extremely ineffective and inefficient, which requires the model to have strong understanding ability for denoising and text reconstruction. And large language models (LLMs) have brought impetus to natural language processing (NLP) tasks in various fields due to their strong understanding and reasoning abilities. The emergence of large language models has demonstrated excellent performance in entity recognition, with strong understanding and generalization abilities.

[0003] However, recent research has shown that existing LLMs still face difficulties in information extraction tasks. For example, the F1 score of GPT-3.5-turbo on the Ontonotes dataset is not ideal and cannot meet the requirements of zero-shot named entity recognition (NER) tasks in the real world. Instruction fine-tuning is an effective method to improve the performance of large language models in specific tasks or domains, but it requires a large amount of time and computing resources. Generally speaking, large language models (LLMs) that can complete NER tasks can be divided into two categories: 1) NER-oriented LLMs, which are specifically fine-tuned for NER tasks, such as TEMPGEN, EnTDA, GPT-NER, LLMaAA, PromptNER, Cp-NER, UniNER; 2) General information extraction LLMs, which are unified generation models for multiple tasks including NER, relation extraction (RE), event extraction (EE), etc., such as UIE, InstructUIE, ChatIE, GIELLM, Set, CODEIE, CodeKGC, GoLLIE, Code4UIE. However, these two types of models mainly face the following difficulties: 1) Designing few-shot tasks is both time-consuming and laborious, and it is difficult to accurately reflect the sample distribution of the dataset. 2) The vast majority of models need to be fine-tuned, but it is still difficult to avoid catastrophic forgetting, resulting in insufficient generalization ability and zero-shot ability. Due to the lack of ability to handle few-shot and zero-shot scenarios, most paradigms applicable to NER tasks face difficulties in dealing with open NER tasks.

[0004] Therefore, how to improve the performance of open named entity recognition in social media and enhance the recognition accuracy is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides an open named entity recognition method and platform for social media, which can convert social media short texts into formal texts with low noise and standard syntactic structures to achieve accurate named entity recognition.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] An open named entity recognition method for social media, comprising the following steps:

[0008] Step 1: Given an initial prompt as the root node;

[0009] Step 2: Construct a Monte Carlo tree according to the root node using Monte Carlo tree search, and screen out the current optimal node from the Monte Carlo tree according to a preset screening rule to obtain the prompt corresponding to the optimal node;

[0010] Step 3: The large language model combines the selected prompts to reconstruct the recognition training text to obtain the reconstructed training text;

[0011] Step 4: Based on the concept correction thought chain template, use the large language model to extract predicted named entities from the reconstructed training text and classify the predicted named entities;

[0012] Step 5: Input the extracted predicted entities and the true entities of the recognition training text into the large language model for comparison to generate the reasons for errors and optimization prompts;

[0013] Step 6: Update the Monte Carlo tree with the optimization prompts as new nodes, and return to Step 2. Terminate the growth of the Monte Carlo tree according to the preset termination conditions, and use the current optimal node as the optimal solution to obtain the corresponding optimal prompt;

[0014] Step 7: The large language model combines the optimal prompt to reconstruct the text to be recognized to obtain the reconstructed recognition text, extracts predicted named entities from the reconstructed recognition text based on the concept correction thought chain template, and classifies the predicted named entities.

[0015] Preferably, when constructing the Monte Carlo tree according to the root node using Monte Carlo tree search in Step 2, the UCB1-Improve formula is used to select the node with the most opportunity to expand from all current leaf nodes. Each search process starts from the root node s 0 and calculates the reward value Q of each node in each layer; The UCB1-Improve formula is expressed as:

[0016]

[0017] where A(s t ) represents the action set of node s t ; represents the current optimal action selected from the action set; represents the number of visits to node s t ; ch(s t , a′ t ) represents the child node generated by node s t after applying the action a′ t in the action set; represents the total number of growth expansions; c represents the exploration weight, which tends to preferentially visit less visited nodes in the initial stage, and then gradually preferentially utilize existing knowledge to select nodes with greater value; Q(s t , a′ t ) represents the reward value after node s t executes the action a′ t ; represents the exploration degree of the node.

[0018] Preferably, the preset screening rule includes selecting the node with the highest reward value or the node with the most access times.

[0019] Preferably, after calculating all the child nodes that are most likely to be expanded from the selection step, denote it as node s t and action a' t , add a new child node s t to this node. The expansion process is to calculate all the child nodes that are most likely to be selected for expansion, and then achieve state transition by applying the meta - hint multiple times, finally generating an optimized hint. The meta - hint is the error reason and the optimized hint generated in step 5. First, error feedback is performed by comparing the downstream entity extraction results to identify deficiencies in text reconstruction; subsequently, based on this error feedback, the large - language model generates an optimized hint according to the error feedback; finally, the large - language model makes adjustments according to the optimized hint to correct the error task. Multiple training batches can be sampled to obtain different error feedbacks.

[0020] Preferably, when the large - language model conducts the comparison process in step 5, it also generates an accuracy rate; the preset termination conditions include an accuracy rate fluctuation threshold or a tree size; calculate the accuracy rate fluctuation based on the accuracy rate and determine whether it is within the accuracy rate fluctuation threshold; the tree size includes the Monte Carlo tree width and depth.

[0021] Preferably, take the optimized hint as a new node. To reduce the computational cost of simulation and simplify the process, directly use the optimized hint of the new node to reconstruct a batch of data, and then perform entity extraction, and calculate the F1 metric as the reward value of the new node.

[0022] Preferably, after the Monte Carlo tree stops growing, back - propagate the reward value from the root node to the path of the new node, and update the reward values of all nodes on this path, expressed as:

[0023]

[0024] where, Q * (s t , a t ) represents the updated reward value of node s t , each node s t includes a corresponding action a t ; M represents the number of paths starting from node s t ; represents the sequence of nodes at the j - th layer starting from node s t ; represents the sequence of actions at the j - th layer starting from the action a t corresponding to node s t , that is, the sequence of actions at the j - th layer starting from node s tAll action sequences corresponding to the j-th layer child nodes at the start; s′ is the node s t One of the j-th layer child nodes at the start, a′ represents the node s t One of all action sequences of the j-th layer child nodes at the start; r represents the reward function.

[0025] Preferably, the concept correction thought chain template is used in the execution process of entity classification, including:

[0026] Step 11: Set few-shot prompts;

[0027] Step 12: Use the large language model to extract all entities from the reconstructed training text or the reconstructed recognition text;

[0028] Step 13: According to the few-shot prompts, use the large language model to divide all entities into named entities and non-named entities;

[0029] Step 14: The large language model classifies the named entities.

[0030] Preferably, the few-shot prompts include that the named entities must be proper nouns based on Wikipedia, follow the thinking process outlined above, and format the output results according to the embedding template.

[0031] An open named entity recognition platform for social media, based on the large language model, including: a prompt optimization module, a text reconstruction module, a thought chain extraction module, and a contrast optimization module;

[0032] The prompt optimization module, given the initial prompt as the root node, constructs a Monte Carlo tree using Monte Carlo tree search according to the root node, filters out the current optimal node from the Monte Carlo tree according to the preset screening rules, obtains the prompt corresponding to the optimal node, transmits the prompt to the text reconstruction module, and receives the named entities of the thought chain extraction module for comparative evaluation, generating an optimized prompt to update the Monte Carlo tree;

[0033] The text reconstruction module, where the large language model reconstructs the text according to the prompt;

[0034] The thought chain extraction module extracts and classifies the named entities from the reconstructed text.

[0035] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses an open named entity recognition method and platform for social media, designs a framework applicable to noisy texts in the real world, can be applied to open NER, performs open named entity recognition on messy short social media texts without fine-tuning, automatically optimizes the prompts for large language models based on the Monte Carlo tree search algorithm, thereby converting short social media texts into formal texts with low noise and standard syntactic structures, realizing self-optimized text reconstruction, and then optimizing the performance of large language models based on the concept correction chain of thought template, accurately identifying named entities from the reconstructed formal texts and classifying them.

[0036] The large language model adopted in the present invention contains rich prior knowledge and can be used as an expert for text reconstruction to extract or analyze the required elements from unstructured texts to generate low-noise texts, helping most downstream NLP tasks to be applied to noisy texts, and using the combination of Monte Carlo tree search (MCTS) and prompt optimization to explore expert-level prompts for text reconstruction; the zero-shot chain of thought (CoT) can improve the model performance and guide the large language model to complete the thinking process step by step.

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.

[0038] Figure 1 Schematic diagram of the self-optimized text reconstruction process for open named entity recognition for social media provided by the present invention;

[0039] Figure 2 Schematic diagram of the text reconstruction process provided by the present invention;

[0040] Figure 3 Schematic diagram of entity classification based on text reconstruction and chain of thought provided by the present invention;

[0041] Figure 4 Schematic diagram of retrieval through entity recognition provided by the present invention. Detailed implementation manners

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0043] An embodiment of the present invention discloses an open named entity recognition method for social media. For short social media texts containing noise, a large language model is guided by natural language prompts to perform an open named entity recognition (NER) task. First, an automatic optimization prompt method is adopted to enable the large language model to generate optimal prompts and effectively cope with the noise space. Second, a zero-shot chain of thought (Zero-shot CoT) is introduced to construct a concept correction chain of thought template to guide the large language model to perform the named entity recognition task according to standard logical thinking. As Figure 1 shown, the specific steps are as follows:

[0044] S1: Given an initial prompt as the root node;

[0045] S2: According to the root node, a Monte Carlo tree is constructed using Monte Carlo tree search. The current optimal node is selected from the Monte Carlo tree according to a preset screening rule, and the prompt corresponding to the optimal node is obtained; the preset screening rule includes selecting the node with the highest reward or the node with the most visits;

[0046] S3: The large language model reconstructs the recognition training text in combination with the selected prompt to obtain a reconstructed training text; as Figure 2 shown is the text reconstruction process;

[0047] S4: Based on the concept correction chain of thought template, the large language model extracts predicted named entities from the reconstructed training text and classifies the predicted named entities;

[0048] S5: The extracted predicted entities and the true entities of the recognition training text are input into the large language model for comparison to generate the reason for the error and an optimized prompt;

[0049] S6: The optimized prompt is used as a new node to update the Monte Carlo tree, and S2 is returned. The growth of the Monte Carlo tree is terminated according to a preset termination condition, and the current optimal node is used as the optimal solution to obtain the corresponding optimal prompt; after calculating all the child nodes that are most likely to be selected for expansion, state transition is achieved by applying the meta-prompt multiple times to generate the optimal prompt;

[0050] S7: The large language model reconstructs the text to be recognized in combination with the optimal prompt to obtain a reconstructed recognition text, extracts predicted named entities from the reconstructed recognition text based on the concept correction chain of thought template, and classifies the predicted named entities.

[0051] Further, when constructing the Monte Carlo tree using Monte Carlo tree search according to the root node in S2, the UCB1-Improve formula is used to select the node with the most opportunity to expand from all current leaf nodes. Each search process starts from the root node s 0Start by calculating the reward value Q for each node in each layer; The UCB1-Improve formula is expressed as:

[0052]

[0053] where A(s t ) represents the set of actions of node s t ; represents the current optimal action selected from the set of actions; represents the number of visits to node s t ; ch(s t , a′ t ) represents the child node generated by node s t after applying the action a′ t in the set of actions; represents the total number of growth expansions; c represents the exploration weight. In the initial stage, it tends to preferentially visit less visited nodes, and then gradually gives priority to using existing knowledge and selects nodes with greater value; W(s t , a′ t ) represents the reward value after node s t executes the action a′ t ; represents the exploration degree of the node.

[0054] Furthermore, after calculating all the child nodes that are most likely to be expanded from the selection S, denoted as node s t and action a′ t , a new child node s′ t is added to this node. The expansion process is to calculate all the child nodes that are most likely to be selected for expansion, and then achieve state transition by applying the meta-hints multiple times, and finally generate optimized hints. The meta-hints are the error reasons and optimized hints generated in step 5. First, error feedback is performed by comparing the downstream entity extraction results to identify deficiencies in text reconstruction; subsequently, based on this error feedback, the large language model generates optimized hints according to the error feedback; finally, the large language model makes adjustments according to the optimized hints to correct the error task. Multiple training batches can be sampled to obtain different error feedbacks.

[0055] Furthermore, during the comparison process in S5, the large language model also generates an accuracy rate; The preset termination conditions include the accuracy rate fluctuation threshold or the tree size; Calculate the accuracy rate fluctuation based on the accuracy rate and determine whether it is within the accuracy rate fluctuation threshold; The tree size includes the Monte Carlo tree width and depth.

[0056] Furthermore, the optimization hint is used as a new node. In order to reduce the computational cost of the simulation and simplify the process, a batch of data is directly reconstructed using the optimization hint of the new node, and then entity extraction is performed to calculate the F1 indicator as the reward value of the new node.

[0057] Furthermore, after the Monte Carlo tree stops growing, the reward value is back-propagated from the root node to the new node, and the reward values ​​of all nodes on the path are updated, which is expressed as:

[0058]

[0059] Among them, Q * (s t , a t ) represents node s t The updated reward value of each node s t includes a corresponding action a t ; M represents the slave node s t The number of paths to start with; Represents slave node s t The starting j-th layer node sequence; Represents slave node s t The corresponding action a t The j-th layer action sequence starts at node s t All action sequences corresponding to the j-th layer child nodes at the beginning; s′ is the node s t Start with one of the j-th layer child nodes, a′ represents node s t Start one of all action sequences of the j-th layer child node; r represents the reward function.

[0060] Furthermore, the concept correction thinking chain template is used in the execution process of entity classification, such as Figure 3 As shown, Figure 4 The original text (Origin Text) is reconstructed after text reconstruction (Text Reconstruction), the entity (Vanilla Prompt) is extracted, and the entity result (Zero-Shot CoT with ConceptCorrection) is obtained after correction; including:

[0061] S11: Set a few sample prompt;

[0062] S12: Extract all entities from the reconstructed training text or the reconstructed recognition text using a large language model;

[0063] S13: Based on the few-sample prompts, use the large language model to classify all entities into named entities and unnamed entities;

[0064] S14: The large language model classifies named entities;

[0065] Based on the named entities extracted in S13, use the text retrieval tool contriever-msmarco model to retrieve the knowledge base (Wikipedia's explanations of all nouns), obtain each named entity and its best-matched definition, and use this as context information to assist the large language model in classifying named entities. As Figure 4 shown in the information retrieval or recommendation schematic diagram.

[0066] Furthermore, few-shot prompts include that the named entity must be a proper noun based on Wikipedia, follow the thinking process outlined above, and format the output result according to the embedding template.

[0067] On the other hand, in a specific embodiment, an open named entity recognition platform for social media is based on a large language model and includes: a prompt optimization module, a text reconstruction module, a thought chain extraction module, and a contrast optimization module;

[0068] The prompt optimization module takes the initial prompt as the root node, constructs a Monte Carlo tree using Monte Carlo tree search based on the root node, screens out the current optimal node from the Monte Carlo tree according to the preset screening rules, obtains the prompt corresponding to the optimal node, transmits the prompt to the text reconstruction module, and receives the named entities from the thought chain extraction module for contrast evaluation to generate an optimized prompt to update the Monte Carlo tree;

[0069] The text reconstruction module reconstructs the text according to the prompt by the large language model;

[0070] The thought chain extraction module extracts and classifies named entities from the reconstructed text.

[0071] In a specific embodiment, ChatGPT is selected as the large language model of the platform, and extensive experiments are carried out under the zero-shot setting to verify the effectiveness of the method of the present invention.

[0072] I. Datasets:

[0073] In order to verify the effectiveness of the self-optimizing text reconstruction and concept correction thought chain template of the present invention on different basic large language models, the following two datasets are mainly used for comparative and ablation experiments:

[0074] (1) Tweebank-NER-v1.0: Tweebank-NER V1.0 is an annotated named entity recognition (NER) dataset based on Tweebank V2. Tweebank V2 is the main UD treebank for natural language processing (NLP) tasks for English tweets. This dataset contains four types of entities: person names, organizations, place names, and other categories.

[0075] (2) WNUT17: This benchmark dataset consists of 1,008 development documents and 1,287 test documents, containing nearly 2,000 entity mentions. The selection criteria for these documents are that they contain a large number of rare and novel entities. The entity types include: personal names, place names, companies, products, creative works, and groups.

[0076] Exclude the data labeled as "None" from the above two datasets, and only retain the samples with valid labels to reduce the impact of invalid labels on the experiment and ensure the accuracy and reliability of the evaluation results.

[0077] II. Evaluation Metrics:

[0078] Use the span-based offset Micro-F1 score as the main metric for evaluating the model. In the Named Entity Recognition (NER) task, follow the span-level evaluation settings. Since a general-domain large language model is used, there may be a certain degree of instability. Therefore, the strict requirements for entity boundaries are relaxed. If the predicted entity type is correct and the predicted entity contains the true entity, it is regarded as a correct recognition and counted as a True Positive.

[0079] For example, if the true label is 'WHO:ORG' and the prediction result is 'the WHO:ORG', it is still considered a correct named entity recognition.

[0080] III. Comparison Models:

[0081] Compare the open named entity recognition platform that combines the ChatGPT model with the concept-corrected chain of thought template with the following existing large language models.

[0082] (1) ChatIE: Constructed a two-stage multi-round Q&A framework and was extensively evaluated on three Information Extraction (IE) tasks, namely overall relation triple extraction, Named Entity Recognition (NER), and event extraction.

[0083] (2) UniNER: Utilized ChatGPT to generate NER instruction tuning data from a wide range of unlabeled web text and performed instruction fine-tuning optimization on the LLaMA model.

[0084] (3) InstructUIE: This is a multi-task instruction tuning system that shows significant advantages in multiple tasks.

[0085] (4) GoLLIE: In the zero-shot information extraction task, this method outperformed previous methods and supported users to perform reasoning according to real-time defined annotation patterns.

[0086] (5) BERT-Large-NER: A fine-tuned BERT model that can be directly used for NER tasks and has achieved advanced performance on this task.

[0087] In this embodiment, GPT-3.5 is specifically selected as the default base large language model (LLM) for optimization to reconstruct the text, and GPT-4 is selected as the default optimizer LLM to generate error feedback and optimization prompts.

[0088] In the inference stage, the temperature of the base LLM is set to 0.0 to ensure the certainty of the prediction results; while in other stages, the temperature is set to 1.0 to increase the generation diversity.

[0089] When implementing open named entity recognition using the platform of the present invention, the number of iterations of Monte Carlo tree search (MCTS) is set to 12, and the exploration weight c is set to 2.5. In the growth and expansion step of the Monte Carlo tree, by extracting a small batch of data from the training samples (extracting one data from the recognition training text), corresponding operations are generated based on the model feedback errors to obtain optimization prompts. The maximum depth of each path is limited by the depth limit, and the width and depth of the Monte Carlo tree are set to control the scale and complexity of the search space.

[0090] IV. Comparison Results

[0091] Table 1 Comparison of Model Experimental Results

[0092]

[0093] As shown in Table 1, the present invention has achieved the expected performance on the Tweebank-NER-v1.0 dataset, outperforming the mainstream large NER models by 5%-25% on the Tweebank-NER-v1.0 dataset and obtaining competitive results on the WNUT17 dataset. Specifically, there are significant performance differences between fine-tuning methods such as BERT-Large-NER, UniNER, InstructUIE, and GoLLIE and ChatIE based on multi-turn question and answer (Q&A). ChatIE decomposes the NER task into multi-turn question and answer in a zero-shot scenario, but does not substantially change ChatGPT's understanding of named entity concepts or knowledge gaps, and its performance is similar to that of the vanilla prompt; GoLLIE highly depends on the user's explanations of different entity types, and the simpler and clearer the explanations are, the better the performance of the model; InstructUIE improves the model performance through multi-task fine-tuning; UniNER further optimizes the model performance through dialogue-style instruction fine-tuning and negative sampling. The processes of the above models for named entity recognition do not perform text cleaning in text preprocessing, while the present invention effectively reduces the impact of bad texts on the performance of open NER through text reconstruction, thus achieving good performance.

[0094] V. Evaluation Experiment of Text Reconstruction

[0095] In the supervised learning setting, test set samples usually appear in the training set, which often leads to high performance metrics. However, in the zero-shot setting, it is required that the test set does not contain or contains as few samples from the training set as possible, which poses a challenge to the model. In this embodiment, on the Tweebank-NER-v1.0 and WNUT17 datasets, multiple basic large language models are used for named entity recognition to explore the performance differences of basic large language models in zero-shot tasks and evaluate the role of text reconstruction in zero-shot named entity recognition (NER). The evaluation results are shown in Table 2 below.

[0096] Table 2 Evaluation and Comparison of NER Performance of Baseline Large Language Models

[0097]

[0098] The method of the present invention was evaluated according to the parameter settings designed in the third part. Table 2 shows the experimental results of all zero-shot methods. The BERT-Large-NER model was fine-tuned on the CoNLL2003 dataset and can only recognize four types of entities: person, location, organization, and miscellaneous. Therefore, it is not suitable for the WNUT17 dataset. The Micro-F1 scores of most basic large language models on the two datasets are between 45 and 65, showing a significant performance decline compared to the performance under the supervised learning setting, indicating that there is still much room for improvement in the generalization ability of the models. After text reconstruction, the F1 scores of most basic large language models increased by 1-5 percentage points, which shows that text reconstruction helps to clearly represent the components of named entities, especially free entities, making it easier for large language models to recognize them, thus improving the precision. ChatIE, which uses ChatGPT as the backbone model, performs significantly lower than other basic large language models on the two public datasets. ChatIE performs poorly under zero-shot and few-shot conditions and is difficult to accurately identify entity types in the first stage. Therefore, it is not suitable for the entity extraction task of zero-shot social media short texts.

[0099] VI. Generalization Experiment and Ablation Experiment

[0100] To verify the generalization ability of the platform framework of the present invention, the backbone models of downstream tasks (NER) were replaced with different large language models such as LLaMA-2, Yi, and ChatGPT. At the same time, to clearly verify that the above performance improvement comes from each module, ablation experiments were conducted on the two datasets. Specifically, the effects of the prompt optimization module, text reconstruction module, and thought chain extraction module based on the concept-corrected thought chain template in the present invention were experimentally verified based on different large language models.

[0101] Baseline, different large language models only use the vanilla prompt to implement NER.

[0102] Text Reconstruction, different large language models only perform text reconstruction.

[0103] CoT, different large language models only use the concept-corrected thought chain template to implement NER.

[0104] By comparing the model performance under different configurations, the contribution of each improved module to the overall performance is revealed, and the applicability and effectiveness in different downstream tasks and dataset scenarios are further explored. The ablation results are shown in Table 3, and the values in the table represent the F1 score.

[0105] Table 3 Model ablation and generalization results for Tweebank-NER-v1.0 and WNUT17 datasets

[0106]

[0107] Generalizability analysis: The method of the present invention can be applied to any general large language model, such as ChatGPT, Llama, Yi, etc., and can also be embedded in pre-trained language models fine-tuned by instructions, such as InstructUIE, BERT-Large-NER, etc.

[0108] Effectiveness analysis:

[0109] 1. With the support of text reconstruction and the concept correction chain of thought template, the model achieves the best performance. This shows a significant synergistic gain effect in the open-domain named entity recognition (NER) task.

[0110] 2. The reconstructed text optimized by prompts can standardize non-standard entities. This standardization significantly affects the named entity recognition performance. The optimized prompts (expert-level prompts) can effectively eliminate social elements (such as USER and URL), contribute to text denoising and reconstruction, and rephrase it into a format suitable for the zero-shot named entity recognition task.

[0111] 3. The introduction of the concept correction chain of thought template can effectively improve the model performance. Adopting the zero-shot chain of thought (Zero-shot CoT) can correct the inference path of the model in the named entity recognition process, enabling it to reason and recognize according to the expected steps, improving the named entity recognition performance, and minimizing the interference of non-proper nouns (i.e., common nouns) on the recognition results to the greatest extent.

[0112] The experimental results show that the method of the present invention achieves state-of-the-art performance in the zero-shot setting and performs excellently in the NER task of processing noisy social media short texts.

[0113] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0114] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An open named entity recognition method for social media, characterized in that: The following steps are involved: Step 1: Given the initial prompt as the root node; Step 2: Construct a Monte Carlo tree using Monte Carlo tree search based on the root node, filter out the current optimal node from the Monte Carlo tree according to the preset filtering rules, and obtain the prompt corresponding to the optimal node; Step 3: The large language model reconstructs the recognition training text in combination with the screened prompts to obtain the reconstructed training text; Step 4: Based on the concept correction thinking chain template, the large language model is used to extract predicted named entities from the reconstructed training text, and the predicted named entities are classified; Step 5: Input the extracted predicted entities and the real entities of the recognized training text into the large language model for comparison, and generate error reasons and optimization tips; Step 6: Update the Monte Carlo tree with the optimization hint as a new node, and return to step 2 to terminate the Monte Carlo tree growth according to the preset termination condition, and take the current optimal node as the optimal solution to obtain the corresponding optimal hint; Step 7: The large language model reconstructs the text to be recognized in combination with the optimal prompt to obtain the reconstructed recognized text, extracts predicted named entities from the reconstructed recognized text based on the concept correction thinking chain template, and classifies the predicted named entities.

2. The open named entity recognition method for social media according to claim 1, characterized in that: When the Monte Carlo tree is constructed based on the root node using Monte Carlo tree search in step 2, the UCB1-Improve formula is used to select the node with the best chance of expansion from all current leaf nodes. Each search process starts from the root node s0 and calculates the reward value Q of each node in each layer. The UCB1-Improve formula is expressed as: Among them, A(s t ) represents node s t A collection of actions; Represents the current optimal action selected from the action set; Represents node s t The number of visits; ch(s t , a′ t ) represents the action a′ in the set of applied actions t Post nodes t The generated child nodes; represents the total number of growth and expansion; c represents the exploration weight, which tends to give priority to visiting less visited nodes in the initial stage, and then gradually gives priority to using existing knowledge and selecting nodes with greater value; Q(s t , a′ t ) represents node s t Execute action a′ t The reward value after Indicates the exploration degree of the node.

3. The open named entity recognition method for social media according to claim 2, characterized in that: The preset filtering rules include selecting the node with the highest reward value or the most visits in the optimal path.

4. The open named entity recognition method for social media according to claim 1, characterized in that: In step 5, the comparison process of the large language model also generates an accuracy rate; the preset termination conditions include an accuracy rate fluctuation threshold or a tree size; the accuracy rate fluctuation is calculated based on the accuracy rate to determine whether it is within the accuracy rate fluctuation threshold; The tree size includes the Monte Carlo tree width and depth.

5. The open named entity recognition method for social media according to claim 2, characterized in that: After the Monte Carlo tree stops growing, the reward value is back-propagated from the root node to the new node, and the reward values ​​of all nodes on the path are updated, which is expressed as: Among them, Q * (s t , a t ) represents node s t The updated reward value of each node s t includes a corresponding action a t ; M represents the slave node s t The number of paths to start with; Represents slave node s t The starting j-th layer node sequence; Represents slave node s t The corresponding action a t The j-th layer action sequence starts; s′ is the node s t Start with one of the j-th layer child nodes, a′ represents node s t Start one of all action sequences of the j-th layer child node; r represents the reward function.

6. The open named entity recognition method for social media according to claim 1, characterized in that: The concept correction thinking chain template is used in the implementation process of entity classification, including: Step 11: Set the few sample prompt; Step 12: Extract all entities from the reconstructed training text or the reconstructed recognition text using the large language model; Step 13: Based on the few-sample prompts, use the large language model to divide all entities into named entities and unnamed entities; Step 14: The large language model classifies the named entities.

7. The open named entity recognition method for social media according to claim 1, characterized in that: The few-shot prompts include that the named entities must be proper nouns based on Wikipedia, follow the thought process outlined above, and format the output results according to the embedded template.

8. An open named entity recognition platform for social media, characterized in that: The open named entity recognition method for social media as described in any one of claims 1 to 7 is applied, based on a large language model, and includes: a prompt optimization module, a text reconstruction module, a thought chain extraction module, and a comparison optimization module; The prompt optimization module is given an initial prompt as the root node, and a Monte Carlo tree is constructed based on the root node using a Monte Carlo tree search. The current optimal node is selected from the Monte Carlo tree according to the preset screening rules, and the prompt corresponding to the optimal node is obtained. The prompt is transmitted to the text reconstruction module, and the named entity of the thought chain extraction module is received for comparative evaluation, and the optimized prompt is generated to update the Monte Carlo tree. Text reconstruction module: the large language model reconstructs the text according to the prompts; The thought chain extraction module extracts named entities from the reconstructed text and classifies them.

Citation Information

Cited By

  • Text classification method and device, equipment, medium and product

    CN121144520A