Entity extraction agent construction method and system based on large language model

By building a label vector library and knowledge base in specific fields, selecting adaptive large language models, and developing verification tools, the problem of poor generalization and poor performance in downstream applications of large language models is solved, and the rapid construction of agents is achieved, and the accuracy and efficiency of entity extraction is improved.

CN120197686AActive Publication Date: 2025-06-24CENT SOUTH UNIV

Patent Information

Application Number
CN202510614742.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-24
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

When used in downstream specific domain tasks, existing large language models have problems such as poor general effectiveness, cumbersome use, and long-term fine-tuning.

Method used

By obtaining entity label files in specific fields, classifying and slicing processing, building a label vector library and knowledge base; writing preset paradigms prompt words; selecting an adapted large language model and evaluating accuracy and recall; developing tools and verifying; inputting knowledge base and prompt words into the large language model for entity extraction, and calling the tool to process the results to verify the correctness of the output, and finally building an agent.

Benefits of technology

It realizes the rapid construction of agents in different fields, reduces the time of model fine-tuning, improves the accuracy and efficiency of entity extraction, and simplifies user operation processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197686A_ABST
    Figure CN120197686A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of entity extraction, and discloses an entity extraction agent construction method based on a large language model, which comprises the following steps: classifying and slicing entity tag files, converting tag names and corresponding entity words into vectors to be queried, organizing into a tag vector library, and finally merging into a knowledge base; writing prompt words according to a preset normal form; selecting an adaptive large language model from a model library, performing accuracy and recall rate evaluation on the candidate model through a test sample set, and selecting an optimal model based on an evaluation result; a tool is developed according to downstream task requirements, and service URL testing, tool availability testing and automatic calling process verification are completed; inputting the knowledge base and the cue words into a large language model to perform entity extraction, calling a tool to process an extraction result, verifying output correctness, and iteratively optimizing configuration of each stage according to a debugging result; according to the method, the problems of poor model general effect and relatively tedious use in an existing agent construction method are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of entity extraction, and particularly to a method and system for constructing an entity extraction intelligent agent based on a large language model. Background Art

[0002] The development process of entity extraction (Named Entity Recognition, NER) technology can be divided into several main stages.

[0003] (1) Early stage: Rule- and dictionary-based methods Rule-based methods: In the early stage of natural language processing, entity extraction mainly relied on manually written rules. These rules were usually based on the syntactic structure and lexical features of the language, and a series of patterns were defined to identify specific entity types. For example, a rule was used to identify a phrase starting with "Mr." as a person's name.

[0004] Dictionary matching: Another common method is to use a dictionary for matching. The dictionary contains known entity names, and entity extraction is achieved by searching for these known entities in the text. The advantage of this method is simplicity and directness, but the disadvantage is that the coverage of the dictionary is limited, and it is difficult to handle newly emerging entity names. (2) Statistical learning method stage Feature engineering: With the development of machine learning, statistical-based methods began to appear. These methods learned the features of entities from labeled training data. Common features included word context information, affixes, part of speech, context features, etc. Feature engineering was a key step in this stage, and researchers needed to carefully design and select features according to the task requirements and data characteristics.

[0005] Model application: Commonly used statistical models included Hidden Markov Model (HMM), Maximum Entropy (MaxEnt), Support Vector Machine (SVM), and Conditional Random Field (CRF), etc. These models could effectively capture the context information and sequence dependencies of entities, improving the accuracy of entity extraction.

[0006] Sequence labeling: The entity extraction task was usually transformed into a sequence labeling problem, and annotation methods such as IO, BIO, and BIOES were used to annotate entities. For each word or character in the text, there were several candidate labels, and the model needed to predict the label of each word or character to identify the boundaries and types of entities.

[0007] (3) Deep learning method stage Neural network models: In recent years, the rapid development of deep learning technology has brought new breakthroughs to entity extraction. Deep neural networks such as Convolutional Neural Networks (CNN), Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), and Transformer have been widely applied to entity extraction tasks. These models can automatically learn the deep semantic features and complex context information in the text, reducing the dependence on manual feature engineering.

[0008] Pre-trained large language models: The emergence of pre-trained language models (such as BERT and GPT) has further promoted the development of entity extraction technology. These models learn rich language features and knowledge through pre-training on large-scale text corpora, and then adapt to downstream NER tasks through fine-tuning. Pre-trained large language models have achieved good accuracy in entity extraction tasks.

[0009] However, there are some problems when applying large language models to downstream specific domain tasks: (1) Based on prior knowledge, there is currently no general large model, that is, a model that can achieve the best results in all fields, which actually brings great limitations to downstream task processing. The corpora to be processed in each specific domain are different, and the entity words to be extracted and the given labels also vary. Therefore, when applying large language models to entity extraction tasks in different domains, the final extraction accuracy is not high, and it also takes a long time to fine-tune the model. (2) In some fields with entity extraction requirements, practitioners may not be familiar with large language models, and they just want an intelligent method to quickly use large language models to process downstream tasks. (3) After using large language models to complete entity extraction tasks, it is generally necessary to further process the extracted entity words manually and then process downstream tasks. Summary of the Invention

[0010] The present invention provides a method and system for constructing an entity extraction agent based on a large language model to solve the problems of poor general model effect and cumbersome use in existing agent construction methods.

[0011] To achieve the above object, the present invention is implemented through the following technical solutions: In a first aspect, the present invention provides a method for constructing an entity extraction agent based on a large language model, including: Obtain the entity label file for a specific domain, classify and slice the entity label file, convert the label names and the entity words corresponding to the label names in the entity label file into query vectors to be queried, organize the query vectors to form a label vector library, and finally merge the label vector library into a knowledge base; Write prompt words according to a preset paradigm, where the prompt words include: defined role, task content, Chinese and English labels, extraction instructions, and standard output format; Select a large language model adapted to the specific domain from the model library, generate a test case set based on the specific domain, evaluate the accuracy and recall rate of the selected large language model through the test case set, and select the optimal model based on the evaluation results; Develop tools according to the downstream task requirements of the specific domain, and conduct service URL tests, usability tests, and automated call process verification on the developed tools. When all tests and verifications pass, regard the developed tools as the optimal tools; Input the knowledge base and prompt words into the optimal model for entity extraction, call the optimal tool to process the results of entity extraction, verify the output correctness, and when the correctness meets the requirements, construct the optimal model and the optimal tool into an intelligent agent.

[0012] Optionally, the entity label file includes: text file or table file.

[0013] Optionally, the preset paradigm includes: The preset paradigm of the prompt word for defining the role is: the model role starting with a hash character plus the role name; The preset paradigm of the prompt word for task content is: the entity extraction task starting with a hash character plus the task name; The preset paradigm of the prompt word for Chinese and English labels is: the example label mapping relationship starting with a hash character plus the Chinese label and its English name; The preset paradigm of the prompt word for extraction instructions is: the extraction rule starting with a hash character plus the instructions for entity extraction; The preset paradigm of the prompt word for the standard output format is: the output format starting with a hash character plus the standard output.

[0014] Optionally, generating a test case set based on the specific domain includes: Write a test case library containing the domain problems and standard answers in the specific domain; Randomly sample the test case library according to a preset ratio to generate a test case set; Evaluating the accuracy and recall rate of the selected large language model through the test case set and selecting the optimal model based on the evaluation results includes: Input each domain problem in the test sample set into the selected large language model one by one, and record the running results of the large language model; Compare the running results of the large language model with the standard answers in the test sample set, calculate the accuracy and recall rate of the running results and the standard answers, and use the large language model with the highest accuracy and recall rate as the optimal model.

[0015] Optionally, the automated call process verification of the tool includes: Ensure that the service URL response status code is 200 and the returned data format meets the requirements; Verify that the tool executes its functions without anomalies and the results meet expectations; Integrate the tool into the intelligent agent and verify the end-to-end call process.

[0016] Optionally, the process of verifying the output correctness includes: verifying the output result, verifying the output format, and verifying the tool processing result; When inputting a problem in the specific domain, if the output result is the correct answer, then verify that the output result is correct. If the correct answer output conforms to the format, then verify that the output format is correct. If the correct answer output is not the entity extraction result, then verify that the tool processing result is correct; When verifying that the output result is correct, the output format is correct, and the tool processing result is correct, the correctness meets the requirements; When there are incorrectness in verifying the output result, the output format, and the tool processing result, if the output result is incorrect, then reconfigure the knowledge base. If the output format does not conform, then adjust the model or prompt. If the output result is the entity extraction result, then re-develop the tool.

[0017] In a second aspect, an embodiment of the present application provides a system for constructing an entity extraction intelligent agent based on a large language model, and the system includes: A knowledge base module for storing a label vector library and a knowledge base in a specific domain; A prompt module for generating entity extraction instructions according to a preset paradigm; A model adaptation module for selecting an adapted large language model from a model library; A tool integration module for accessing downstream task tools and verifying the tools; An execution engine module for calling the large language model to perform entity extraction based on the label vector library and the entity extraction instructions, and inputting the entity extraction results into the tool to complete downstream tasks; A debugging and optimization module for iteratively optimizing the knowledge base, prompts, large language model, and tool configuration according to the output results.

[0018] Optionally, the tool integration module includes a tool development sub-module and a tool verification sub-module; The tool development sub-module is used to develop tools according to the requirements of downstream tasks in the specific field; The tool verification sub-module is used to perform service URL testing, usability testing, and automated call process verification on the developed tools. When all tests and verifications pass, the developed tools are used as tools for accessing downstream tasks.

[0019] Optionally, the execution engine module includes a model processing sub-module and a result output sub-module; The model processing sub-module is used to input the knowledge base and prompt words into the large language model for entity extraction; The result output sub-module is used to call the tools to process the results of entity extraction and output the results as output results.

[0020] Optionally, the debugging and optimization module includes a verification sub-module and an iterative debugging sub-module; The verification sub-module is used to verify the correctness of the output results in the execution engine module, and when the correctness meets the requirements, the large language model and tools are constructed into an intelligent agent; The iterative debugging sub-module is used to send iterative instructions to the knowledge base module when the correctness does not meet the requirements. The knowledge base module reconfigures the label vector library and knowledge base based on the iterative instructions, and / or sends iterative instructions to the prompt word module. The prompt word module regenerates entity extraction instructions based on the iterative instructions, and / or sends iterative instructions to the model adaptation module. The model adaptation module reselects the adapted large language model based on the iterative instructions, and / or sends iterative instructions to the tool integration module. The tool integration module redevelops tools based on the iterative instructions.

[0021] Optionally, the system is compatible with different large language model APIs and extends the model library and tool library through a plug-in mechanism.

[0022] Advantageous effects: For the method for constructing an entity extraction intelligent agent based on a large language model provided by the present invention, there is no restriction on the file type of the file used in the knowledge base. The intelligent agent can process different types of data, and has strong practicability; for writing prompt words, we have given the specific form of the paradigm, and only need to rewrite according to the paradigm for different fields; the models used can be flexibly selected and do not require long-term fine-tuning, saving a lot of time; the constructed intelligent agent is convenient for users to use. After the user inputs the content to be extracted in the dialog box, the intelligent agent will directly return the results after calling the large language model and tools, eliminating the intermediate manual operation process and realizing the intelligence of downstream task processing; this intelligent agent can be embedded in the original systems of enterprises in different fields, saving development time. Description of the Drawings

[0023] Figure 1 One of the flowcharts for constructing an entity extraction agent based on a large language model according to a preferred embodiment of the present invention; Figure 2 Two of the flowcharts for constructing an entity extraction agent based on a large language model according to a preferred embodiment of the present invention; Figure 3 The flowchart for using the agent according to a preferred embodiment of the present invention. Detailed implementation manners

[0024] The technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0025] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms such as "a" or "one" do not denote a quantity limitation, but mean that there is at least one. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship also changes accordingly.

[0026] Embodiment 1 Please refer to Figure 1 , an embodiment of the present application provides a method for constructing an entity extraction agent based on a large language model, including: Obtain an entity label file in a specific domain, classify and slice the entity label file, convert the label name and the corresponding entity word into a vector to be queried, and organize them into a label vector library, and finally merge them into a knowledge base; Write a prompt word according to a preset paradigm, where the prompt word includes: defining a role, task content, Chinese and English labels, extraction instructions, and a standard output format; Select a suitable large language model from the model library, evaluate the accuracy rate and recall rate of the candidate model through a test sample set, and select the optimal model based on the evaluation results; Develop tools according to the downstream task requirements, and complete the service URL reachability test, tool usability test, and automated call process verification; Input the knowledge base and prompt words into the large language model for entity extraction, call tools to process the extraction results, verify the correctness of the output, and iteratively optimize the configurations of each stage according to the debugging results.

[0027] As Figure 2 shown, in this embodiment, the construction method may specifically include the following steps: (1) Knowledge base acquisition stage: First, it is necessary to obtain files related to entity extraction. These files should include both Chinese and English label information and give all entity words corresponding to each label. The file types of these files are not restricted and can be text files, table files, etc. After obtaining these files, first classify the files belonging to different labels, slice the classified files, and then convert the label names and corresponding entity words into query vectors to be queried , organize the query vectors to be queried for each label together to form a label vector library. After all the label vector libraries are merged, they jointly form a knowledge base. The knowledge base is the knowledge reserve for the large language model to process tasks in related fields, which helps the large language model learn the features of each label. After the large language model obtains and learns the knowledge base, it can perform entity extraction on the questions of users in a specific field.

[0028] (2)Prompt Writing Stage: The prompt is a small set of text instructions for the large language model to perform entity extraction, which can be adjusted according to the requirements of downstream task processing. The writing of the prompt is not arbitrary and needs to follow a certain paradigm. First, the role that the large language model plays needs to be given, such as "#Role: You are a rigorous expert in the field of entity extraction". Then, the content of the entity extraction task needs to be defined. The task content is adjusted according to the actual usage requirements, starting with "#Task:", and the specific extraction task is written after the colon. It should be noted that the task content should be summary and not be described in detail. Next, the Chinese labels (and English labels if any) to be extracted need to be given. This step is very crucial and is related to the success rate of the large language model in performing entity extraction tasks. The paradigm content for this part starts with "#Chinese Labels and Their English Names:", and then the actual label information is added after the colon. Different labels are separated by commas, for example, "Call time corresponds to current_time, personnel code corresponds to user_id". After all the above information is given, the specific extraction instructions should be given directly next. The content of the instructions can be flexibly changed according to the requirements. The format requirement is to start with "#Instructions for Entity Extraction:", and the actual instructions are added after the colon. A specific example is "1. Work strictly according to the instructions and output in strict accordance with the standard format. 2. If no entity words are extracted, just output an empty string. 3. After learning the content of the knowledge base, perform entity extraction". Finally, it should be declared what the standard output format is. This is related to the user experience of the system users, so it is best to design the final output format in advance. The prompt for this part must start with "#Standard Output:", and then the content is written according to the actual required format.

[0029] (3) Model Selection Phase: The models used in this system can be flexibly selected according to the actual application field, and there are no major restrictions on the selection and use of models. The design architecture of this system is compatible with different large language models. That is to say, people in different fields can use different models to build agents. The models need to be selected from the model library, which is a collection of APIs that can be called by the builders. In order to find the most suitable model for the entity extraction task in this field, it is recommended to test all the models in the model library. The testing steps are as follows: First, the actual users of the system write a test case library. The writing rule of the test cases is that it must be ensured that there are correct (not necessarily unique) answers to the questions with the test cases as the input in the application field. After the test case library is written, random sampling can be carried out in the test case library according to a certain proportion, and the proportion can also be determined according to actual needs. After the selection phase is completed, it is automatically merged into the test case set. Then, the questions in the test case set are used as the input of the large language model one by one, and the running results of different models on the same test case set are recorded, that is, the final entity extraction results. After comparing and calculating with the correct answers, important indicators such as accuracy and recall rate are obtained respectively, and finally the model with the best effect is selected for use.

[0030] (4) Tool Development Phase: The downstream tasks that need to be completed by personnel in each field are different, which also brings certain challenges to the construction of agents. The differences in tasks directly determine the need to use different tools. Therefore, the agent must be compatible with different forms of tools. The tools are changed according to the downstream task requirements of a specific field. Based on this characteristic, the agent created by this invention has relatively weak restrictions on the use of tools. The staff in a specific field first independently develop the tools according to the needs of their own field, and then can first conduct a service URL reachability test, that is, test the reachability of the service URL specified in the configuration file. The passing standard for the test is that the URL can successfully respond, and the response status code is 200, and the data format returned meets the predetermined requirements. Then conduct a tool usability test, that is, the tool can execute the specified function as expected, and no exceptions are thrown during the execution process, and the function execution result meets the expected effect. After passing the two tests, it is also necessary to check its automated call process, which must be carried out after completing all the above steps.

[0031] (5)Agent Debugging Phase: The agent first reads the provided knowledge base as the knowledge reserve for processing domain problems, uses the written prompts as the prompts for the large language model, and then uses the selected large language model to perform entity extraction tasks. The results of entity extraction are obtained and stored in the agent's workflow. However, the results of entity extraction are not output on the front-end interface. Instead, they are directly used as the input for downstream tools to directly call the tools and obtain the final results after the problems are solved. That is to say, the results output on the final front-end interface are the answers to the user's questions. If the entire process runs without errors and the agent can return the correct answers to the corresponding questions, then the debugging phase ends. If the returned results are incorrect, return to step (1) to rebuild. If the returned results do not conform to the format, return to step (3) to rebuild. If the returned results are the results of entity extraction, return to step (4) to rebuild.

[0032] Example 2 If using an agent on the Dify platform, first create a blank application - Agent on the Dify platform. Business personnel also need to slice 10 text files such as operation_type.txt and system_platform.txt to form query vectors to be searched, {<operation type, initiate rollback>}, {<system platform, cloud video platform>}. Then organize the query vectors of all label categories together {<operation type, initiate rollback>, <system platform, cloud video platform>} to form a label vector library. Next, proceed to the second step - write prompts according to the paradigm. According to the actual requirements of the telecommunications business, the prompt written according to the paradigm is "#Role: You are a meticulous assistant in the field of entity extraction, but you are only allowed to extract the content within []. The other content is just instructions and does not need to be extracted. Do not be interfered by irrelevant information. Strictly obey all instructions. An entity word belongs to only one label and duplicate extraction is not allowed. #Task: Extract the entity words within [] that belong to specific labels. #Chinese labels and their English names: call start time corresponding to api_inst_create_date_start, call end time corresponding to api_inst_create_date_end, operation type corresponding to oper_type, order status corresponding to order_state, personnel code corresponding to user_id, system platform corresponding to error_link_type_value, large order number corresponding to cust_order_code, small order number corresponding to service_order_code #Instructions for entity extraction: 1. Work strictly according to the instructions and output in strict standard format. 2. If no entity words are extracted, output an empty string. 3. Business code entity words must contain uppercase English letters. 4. If there is information about a time range and it is not specific to a certain day, it needs to be extracted into two parameters: call time and callback time. 5. If there is information about a specific day, it is extracted as 'order creation time'. Each label can have at most one entity word. If the customer does not mention the year or month, reasonably complete it using the system's current time. #Standard output: Only when an entity word is extracted for a certain label, add the English label and its entity word in the output. Labels without entity words are directly ignored in the output and are not allowed to appear in the output". After writing the prompt, proceed to the third step - select a model. First, query which selectable models are available in the telecommunications model library. After querying, it is known that the available models are "Qwen, Qwen-2b, Chatglm4". Business personnel first write a test case library, such as "I need to query all business orders in March; I want to know all the business situations of employee No. 4031; directly give me all the orders on the cloud video platform".Next, the three available models can be tested using the test case library to obtain important metrics such as the accuracy and recall rate of each model on this extraction task. Through comparison, the Qwen large model is finally selected to handle this task. Since the downstream task that the telecommunications party wants to perform is report query, the intelligent agent needs to be equipped with a report query tool. This tool takes the entity words extracted by the large language model as parameter input, calls the report query ability interface, and then returns the queried URL information. Clicking on the URL link to download the report allows one to obtain the desired data. At this time, the prompt words can be adjusted according to the response speed of the intelligent agent. It is okay to delete or add prompt words, and everything depends on the final effect. After determining the final version of the prompt words, the construction process of the intelligent agent is completely finished. Figure 3 This is the usage process of the intelligent agent of the present invention. The user only needs to input the sentence to be extracted in the dialog box and click send, and the intelligent agent will return the result processed by the large language model and the tool.

[0033] Embodiment 3 If one wants to use the intelligent agent on the telecommunications' own platform, first, the business personnel need to slice 5 files such as order_status.xlsx and personnel_code.xlsx to form query vectors to be queried {<order status, CancelN>}, {<personnel code, 4031>}. Then, organize all the query vectors of the label categories together {<order status, CancelN>, <personnel code, 4031>} to form a label vector library. Next, perform the second step - write the prompt words according to the paradigm. According to the actual requirements of the telecommunications business, the prompt words written according to the paradigm are: "Robot: Hello, this is the Hunan Big Data Analysis Center. We are analyzing a piece of data information generated in the Hunan system and need your help. #Role: You are a rigorous little assistant in the field of entity extraction, but you are only allowed to extract the specified content. Other content is just instructions and does not need to be extracted. Do not be interfered by irrelevant information. Strictly obey all instructions. An entity word belongs to only one label and repeated extraction is not allowed. #Task: Extract the entity words belonging to specific labels from the specified content. #Chinese labels and their English names: Call start time corresponding to api_inst_create_date_start, call end time corresponding to api_inst_create_date_end, operation type corresponding to oper_type, order status corresponding to order_state, personnel code corresponding to user_id, system platform corresponding to error_link_type_value, large order number corresponding to cust_order_code, small order number corresponding to service_order_code # Instructions for entity extraction: 1. Divergent thinking is strictly prohibited. Work strictly according to the instructions and output in strict accordance with the standard format. Do not organize the language by yourself.

[0034] 2. If no entity words are extracted, just output an empty string. Don't make them up.

[0035] 3. The large order number corresponds to the cust_order_code, which is a pure number starting with "8".

[0036] 4. The small order number corresponds to the service_order_code, which is a pure number starting with "7". 5. If information about the time range appears and it is not specific to a certain day, two parameters, namely the call time and the callback time, need to be extracted. For example: if the word "May" appears but not specific to a certain day, the "call time" is extracted as "2024-05-01" and the "callback time" is extracted as "2024-05-31". 6. If information about a specific day appears, it is extracted as the "order creation time". There can be at most one entity word in one tag. If the customer does not mention the year or month, the current system time is used to reasonably complete it. For example: if the date "November 30th" is mentioned in the conversation, the order creation time is extracted as the current system year "2024-11-30".

[0037] 7. The entity words of the operation type tag only need to perfectly match the two strings "initiate rollback" and "force receipt". Perfect match means that every character is exactly the same. Finally, only the corresponding codes "10R" or "1FF" are retained.

[0038] 8. The entity word of the personnel code is the name of the person.

[0039] 9. The entity word of the order status has no Chinese characters. "Force receipt" and "initiate rollback" have nothing to do with the order status. Do not extract wrongly.

[0040] It only needs to perfectly match with "10C", "10CE", "10F", "10N", "CancelN", "CancelF", "FAILURE", "Rollback", "RollbackN", "AwaitOrder". If the specified content is "query all business orders in May", the output sample is: "api_inst_create_date_start=2024-05-01|api_inst_create_date_end=2024-05-31" After writing the prompt, the third step is to select a model. First, it is necessary to query which selectable models are available in the telecommunications model library. After querying, it is known that the available models are "Qwen, Qwen-2b, Chatglm4". The business personnel first compile a test case library, such as "I need to query all business orders in December last year; I want to know all the business situations of employee No. 1000; Give me the work order records of the cloud video platform". Next, the three available models can be tested with the test case library to obtain important indicators such as the accuracy and recall rate of each model in this extraction task. Through comparison, the Qwen large model is finally selected to handle this task. The telecommunications party wants to use an agent to summarize the work situation, so an intelligent summary tool needs to be equipped. This tool takes the entity words extracted by the large language model as parameter input, calls the intelligent summary ability interface, and then returns the queried URL information. Clicking on the URL link to download the report can obtain the desired data, and then the large language model uses the obtained data as input to summarize the work situation. At this time, the prompt can be adjusted according to the response speed of the agent until a satisfactory effect is achieved. After determining the final version of the prompt, all capabilities are packaged and uploaded to the telecommunications platform, and this ability is successfully called on the front-end page. At this time, the agent is created successfully.

[0041] The embodiment of the present application also provides a system for constructing an entity extraction agent based on a large language model, and the system includes: A knowledge base module for storing a label vector library in a specific domain; A prompt module for generating entity extraction instructions according to a preset paradigm; A model adaptation module for selecting an adapted large language model from a model library; A tool integration module for accessing downstream task processing tools and verifying their availability; An execution engine module for calling a large language model to perform entity extraction and inputting the result into a tool to complete downstream tasks; A debugging and optimization module for iteratively optimizing the knowledge base, prompt, large language model, and tool configuration according to the output result.

[0042] Optionally, the tool integration module supports dynamic access to multi-domain tools and specifies the service URL and data format through a configuration file.

[0043] Optionally, the output result of the execution engine module is the final processing result of the downstream task, and the front-end interface directly displays the answer to the user's question, hiding the intermediate entity extraction process.

[0044] Optionally, the system is compatible with different large language model APIs and expands the model library and tool library through a plug-in mechanism.

[0045] The entity extraction agent construction system based on a large language model provided in this embodiment can implement each embodiment of the above-mentioned entity extraction agent construction method based on a large language model, and can achieve the same beneficial effects, which will not be elaborated here.

[0046] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A method for constructing an entity extraction agent based on a large language model, characterized in that: include: Obtain entity label files in a specific field, classify and slice the entity label files, convert label names and entity words corresponding to the label names in the entity label files into vectors to be queried, organize the vectors to be queried into a label vector library, and finally merge the label vector library into a knowledge base; Write prompt words according to the preset paradigm, where prompt words include: define role, task content, Chinese and English labels, extract instructions and standard output format; Selecting a large language model adapted to the specific domain from a model library, generating a test sample set based on the specific domain, evaluating the accuracy and recall of the selected large language model through the test sample set, and selecting an optimal model based on the evaluation result; Develop tools according to the downstream task requirements in the specific field, and conduct service URL testing, usability testing, and automated call process verification on the developed tools. When all tests and verifications pass, the developed tools will be considered the optimal tools. The knowledge base and prompt words are input into the optimal model to extract entities, the optimal tool is called to process the results of entity extraction, and the correctness of the output is verified. When the correctness meets the requirements, the optimal model and the optimal tool are constructed into an intelligent entity.

2. The entity extraction agent construction method based on a large language model according to claim 1 is characterized in that: The entity label file includes: a text file or a table file.

3. The entity extraction agent construction method based on a large language model according to claim 1 is characterized in that: The preset paradigms include: The default pattern for defining a role is: a model role that starts with the pound sign plus the role name; The default pattern of the prompt words for the task content is: an entity extraction task that starts with the pound sign plus the task name; The default pattern of prompt words for Chinese and English labels is: the example label mapping relationship starting with the pound character plus the Chinese label and its English name; The default pattern of the prompt word for the extraction instruction is: the extraction rule starts with the pound character plus the instruction for entity extraction; The default format for the prompt word of the standard output format is: an output format that begins with the pound sign and standard output.

4. The entity extraction agent construction method based on a large language model according to claim 1 is characterized in that: Generate a test sample set based on the specific field, including: Compile a test sample library containing domain questions and standard answers in the specific domain; Randomly sampling the test sample library according to a preset ratio to generate a test sample set; The accuracy and recall of the selected large language model are evaluated through the test sample set, and the optimal model is selected based on the evaluation results, including: Input the domain questions in the test sample set into the selected large language model one by one to record the operation results of the large language model; The running results of the large language model are compared with the standard answers in the test sample set, the accuracy and recall of the running results and the standard answers are calculated, and the large language model with the highest accuracy and recall is taken as the optimal model.

5. The entity extraction agent construction method based on a large language model according to claim 1 is characterized in that: The automated call flow verification of the tool includes: Ensure that the service URL response status code is 200 and the returned data format meets the requirements; Verify that the tool performs its functions without exception and the results are as expected; Integrate tools into agents and validate end-to-end call flows.

6. The entity extraction agent construction method based on a large language model according to claim 1 is characterized in that: The process of verifying the correctness of the output includes: verifying the output result, verifying the output format and verifying the tool processing result; When a question in the specific field is input, if the output result is a correct answer, then the output result is verified to be correct; if the output correct answer conforms to the format, then the output format is verified to be correct; if the output correct answer is not an entity extraction result, then the processing result of the verification tool is correct; Correctness meets the requirements when the verification output is correct, the verification output format is correct, and the verification tool processes the result correctly; When there are inaccuracies in the verification output results, verification output format, and verification tool processing results, if the output result is wrong, reconfigure the knowledge base; if the output format does not match, adjust the model or prompt words; if the output result is an entity extraction result, redevelop the tool.

7. A system for constructing an entity extraction agent based on a large language model, characterized in that: The system comprises: The knowledge base module is used to store the label vector library and knowledge base of a specific field; The prompt word module is used to generate entity extraction instructions according to the preset paradigm; Model adaptation module, used to select an adapted large language model from the model library; Tool integration module, used to access downstream task tools and verify the tools; The execution engine module is used to call the large language model to extract entities based on the label vector library and entity extraction instructions, and input the entity extraction results into the tool to complete downstream tasks; The debugging and optimization module is used to iteratively optimize the knowledge base, prompt words, large language model and tool configuration according to the output results.

8. The entity extraction agent construction system based on a large language model according to claim 7 is characterized in that: The tool integration module includes a tool development submodule and a tool verification submodule; The tool development submodule is used to develop tools according to the downstream task requirements of the specific field; The tool verification submodule is used to perform service URL testing, usability testing, and automated call process verification on the developed tools. When all tests and verifications pass, the developed tools are used as tools for downstream tasks.

9. The entity extraction agent construction system based on a large language model according to claim 7 is characterized in that: The execution engine module includes a model processing submodule and a result output submodule; The model processing submodule is used to input the knowledge base and prompt words into the large language model to extract entities; The result output submodule is used to call the tool to process the result of entity extraction and output the result as an output result.

10. The entity extraction agent construction system based on a large language model according to claim 7, characterized in that: The debugging and optimization module includes a verification submodule and an iterative debugging submodule; The verification submodule is used to verify the correctness of the output results in the execution engine module, and to construct the large language model and the tool into an intelligent agent when the correctness meets the requirements; The iterative debugging submodule is used to send iterative instructions to the knowledge base module when the correctness does not meet the requirements, and the knowledge base module reconfigures the label vector library and the knowledge base based on the iterative instructions, and / or sends iterative instructions to the prompt word module, and the prompt word module regenerates the entity extraction instructions based on the iterative instructions, and / or sends iterative instructions to the model adaptation module, and the model adaptation module reselects the adapted large language model based on the iterative instructions, and / or sends iterative instructions to the tool integration module, and the tool integration module redevelops the tool based on the iterative instructions.

Citation Information

Patent Citations

  • Knowledge graph question and answer method based on deep learning and similarity matching

    CN112765310A

  • Method for providing high-quality data for multi-mode large model system

    CN117743315A

  • Named entity recognition method combining large model and dynamic sample

    CN119378552A

  • Information extraction method and device based on large language model, equipment and storage medium

    CN119415669A

  • Data query statement generation method and device, storage medium and electronic equipment

    CN119513326A

Cited By

  • Test verification evaluation method based on AI intelligent agent and related device

    CN120873522A

  • Adaptation verification method and system for large language model and medium

    CN121390082A