Information extraction method based on artificial intelligence, model training method and related device

By combining the information extraction model and the big model, evaluating and optimizing the information extraction results, the problem of poor generalization ability of information extraction in the existing technology is solved, and higher accuracy and efficiency are achieved.

CN119988977APending Publication Date: 2025-05-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122864.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology has poor generalization ability in information extraction, and the accuracy of information extraction needs to be improved.

Method used

Using an artificial intelligence-based method, combining information extraction model and large model, the output of the information extraction model is evaluated and optimized to improve the accuracy and generalization ability of information extraction.

Benefits of technology

It significantly improves the accuracy and efficiency of information extraction results, and improves the generalization ability of information extraction methods, which can better adapt to diverse dialogue scenarios and professional terms in different fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988977A_ABST
    Figure CN119988977A_ABST
Patent Text Reader

Abstract

The invention provides an artificial intelligence-based information extraction method, a model training method and a related device, relates to the technical field of data processing, in particular to the technical fields of artificial intelligence, large models, deep learning and the like, and can be applied to application scenes of intelligent e-commerce, intelligent assistants, virtual assistants, intelligent search and the like. According to the specific implementation scheme, an information extraction task is executed on a to-be-processed text based on an information extraction model, so that initial information is extracted from the to-be-processed text; and evaluating an execution result of the information extraction task of the information extraction model for the to-be-processed text based on the large model to optimize the initial information, and obtaining an information extraction result of the to-be-processed text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to technical fields such as artificial intelligence, large models, and deep learning, and can be used in application scenarios such as smart e-commerce, smart assistants, virtual assistants, and smart searches. Background Art

[0002] Information Extraction (IE) is a technology that automatically extracts structured information from text data. Information extraction supports the identification of valuable and specific information from unstructured or semi-structured text for subsequent analysis, storage and application. Summary of the invention

[0003] The present invention provides an information extraction method, a model training method and related devices based on artificial intelligence.

[0004] According to one aspect of the present disclosure, there is provided an information extraction method based on artificial intelligence, comprising:

[0005] Performing an information extraction task on the text to be processed based on the information extraction model to extract initial information from the text to be processed;

[0006] Based on the large model, the execution result of the information extraction model for the information extraction task of the text to be processed is evaluated to optimize the initial information and obtain the information extraction result of the text to be processed.

[0007] According to another aspect of the present disclosure, there is provided an information extraction device based on artificial intelligence, comprising:

[0008] An extraction module, used for performing an information extraction task on the text to be processed based on the information extraction model, so as to extract initial information from the text to be processed;

[0009] The first optimization module is used to evaluate the execution result of the information extraction task of the information extraction model for the text to be processed based on the large model, so as to optimize the initial information and obtain the information extraction result of the text to be processed.

[0010] According to one aspect of the present disclosure, a method for training an information extraction model is provided, comprising:

[0011] For information extraction tasks, a large model is used to annotate the sample set to be annotated to obtain the initial sample set;

[0012] Optimizing the initial sample set to obtain a training sample set;

[0013] The information extraction model is trained based on the training sample set so that the information extraction model supports the information extraction task.

[0014] According to another aspect of the present disclosure, there is provided a training device for an information extraction model, comprising:

[0015] The annotation module is used to annotate the sample set to be annotated using a large model for information extraction tasks to obtain an initial sample set;

[0016] A fourth optimization module, used for optimizing the initial sample set to obtain a training sample set;

[0017] A training module is used to train the information extraction model based on the training sample set so that the information extraction model supports the information extraction task.

[0018] According to another aspect of the present disclosure, there is provided an artificial intelligence-based information extraction system, comprising:

[0019] An intelligent agent, wherein the intelligent agent integrates an information extraction model and a large model, wherein:

[0020] The information extraction model performs an information extraction task on the text to be processed to extract initial information from the text to be processed;

[0021] The large model is used to evaluate the execution result of the information extraction model on the information extraction task of the text to be processed, so as to optimize the initial information and obtain the information extraction result of the text to be processed.

[0022] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0023] at least one processor; and

[0024] a memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any method in the embodiments of the present disclosure.

[0026] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0027] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements any method according to the embodiments of the present disclosure when executed by a processor.

[0028] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0030] Figure 1 is a flowchart of an information extraction method based on artificial intelligence provided according to an embodiment of the present disclosure;

[0031] Figure 2 is another flow chart of an information extraction method based on artificial intelligence provided according to an embodiment of the present disclosure;

[0032] Figure 3 is a schematic diagram of an interface of an information extraction method based on artificial intelligence provided according to an embodiment of the present disclosure;

[0033] Figure 4 is a flow chart of an optimized information extraction model provided according to an embodiment of the present disclosure;

[0034] Figure 5 is a flowchart of a training method for an information extraction model provided according to an embodiment of the present disclosure;

[0035] Figure 6 is another flowchart of a method for training an information extraction model provided according to an embodiment of the present disclosure;

[0036] Figure 7 is another flowchart of a method for training an information extraction model provided according to an embodiment of the present disclosure;

[0037] Figure 8 is a flow chart of an artificial intelligence-based information extraction system provided according to an embodiment of the present disclosure;

[0038] Fig. 9 is a schematic diagram of the structure of an information extraction device based on artificial intelligence provided according to an embodiment of the present disclosure;

[0039] Fig.10 is a structural schematic diagram of a training device for an information extraction model provided according to an embodiment of the present disclosure;

[0040] Fig.11 It is a block diagram of an electronic device used to implement the artificial intelligence-based information extraction method and / or the information extraction model training method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0041] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0042] Traditional information extraction methods use methods such as NLP (Natural Language Processing) and machine learning. This method has poor generalization ability and the accuracy of information extraction needs to be improved.

[0043] In view of this, the present disclosure provides an information extraction method based on artificial intelligence, such as Figure 1 As shown, it is a flow chart of the method, which includes the following contents:

[0044] S101, performing an information extraction task on the text to be processed based on the information extraction model to extract initial information from the text to be processed.

[0045] S102, evaluating the execution result of the information extraction model for the information extraction task of the text to be processed based on the large model to optimize the initial information and obtain the information extraction result of the text to be processed.

[0046] In the disclosed embodiment, the information extraction model is a model that focuses on extracting structured information, and the scale of the model is much smaller than that of the large model. When implemented, the UIE (Universal Information Extraction) model can be selected as the information extraction model. The number of parameters of the UIE model is about 1b.

[0047] The large model in the disclosed embodiment may be a large language model, which refers to a specific type of large model specifically used to process text data. This type of model is a natural language processing model based on a neural network, which can be used to generate, understand, and process text data. Like a large language model, it can have tens of billions of parameters, can generate high-quality text, and can be used for various natural language processing tasks, such as question answering, text generation, dialogue systems, etc.

[0048] In the disclosed embodiment, the information extraction model focuses on extracting structured information, so its information extraction accuracy is guaranteed and higher than that of the large model. In order to further ensure the accuracy and generalization ability of information extraction, in the disclosed embodiment, the large model is further used to evaluate the execution results of the information extraction task of the information extraction model to optimize the initial information, thereby improving the quality of information extraction, and at the same time, the generalization ability of the large model can be used to adapt the information that the information extraction model is not good at extracting. In addition, the number of parameters of the information extraction model is much smaller than that of the large model, and it can be compared with the simple large model for information extraction. The combination of the two can reduce the model processing time through the information extraction model and improve the efficiency of information extraction. In summary, compared with the use of only a single model in the related art, the semantic understanding ability of the large model in the disclosed embodiment is combined with the structured extraction advantages of the information extraction model, which significantly improves the accuracy and efficiency of the information extraction results, and has the corresponding generalization ability to cope with diverse dialogue scenarios and professional terms in different fields.

[0049] In the embodiments of the present disclosure, in order to further improve the efficiency of information extraction, during implementation, multiple information extraction models can be deployed on different devices according to the performance of the devices, and multiple information extraction models can be used to simultaneously execute information extraction tasks. This can cope with large-scale business scenarios and improve the efficiency of information extraction.

[0050] In the embodiment of the present disclosure, in order to further improve the accuracy of the information extraction results, the execution results of the information extraction task of the information extraction model for the to-be-processed text are evaluated based on the large model to optimize the initial information and obtain the information extraction results of the to-be-processed text, which can be implemented as follows: Figure 2 As shown:

[0051] S201, for the information extraction task, construct a large model prompt word based on the text to be processed and the initial information; the large model prompt word indicates the evaluation task that the large model needs to perform; the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, and / or correcting the initial information.

[0052] Among them, constructing the big model prompt words can enable the big model to understand and evaluate the input (text to be processed) and output (initial information) of the information extraction model. Through the evaluation task, the big model can accurately understand its role requirements.

[0053] S202, inputting the large model prompt words into the large model so that the large model performs the evaluation task and obtains the information extraction result of the text to be processed.

[0054] In the disclosed embodiment, by constructing a large model prompt word, the large model can accurately perform the evaluation task according to the large model prompt word, so as to facilitate targeted evaluation of the execution results of the information extraction task of the information extraction model for the text to be processed, thereby providing an accurate basis for the optimization of the initial information and improving the accuracy of the information extraction results.

[0055] In the disclosed embodiment, the big model prompt includes the task requirements of the evaluation task. The task requirements define the execution specifications of the big model to perform the evaluation task. Among them, since the big model has a strong semantic understanding ability, the semantic understanding ability of the big model is used to fully understand the task requirements of the evaluation task, and complete the evaluation of the accuracy of the execution results of the information extraction model for the information extraction task of the text to be processed to determine whether its execution results meet the corresponding specification requirements; and / or, the semantic understanding ability of the big model can also be used to correct the initial information extracted by the information extraction model based on the task requirements.

[0056] With the semantic understanding ability of the big model, the execution specifications of the task requirements of the evaluation task can be broadly expressed without having to enumerate all situations in a rule-based manner. Therefore, with the task requirements and semantic understanding ability of the big model, the task requirements can be accurately understood, and the execution results of the information extraction task of the information extraction model can be reasonably evaluated. This can be used as a basis to optimize the initial information, provide a high-quality optimization benchmark, and improve the accuracy of the information extraction results.

[0057] In an embodiment of the present disclosure, when the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, the task requirements include at least one of the following:

[0058] a1) The information to be extracted must meet basic requirements, which are determined based on the information extraction task.

[0059] Taking the extraction of product terms as an example, the basic requirement is that the extracted product terms should include specific product names. In other words, product terms cannot be general categories, but specific product names, so that subsequent business based on product names can be carried out, such as intelligent customer service dialogue, such as searching based on product terms.

[0060] It is understandable that the specific basic requirements are determined based on the specific information extraction task, and the embodiments of the present disclosure do not limit this.

[0061] a2) In certain circumstances, the information items included in the information to be extracted.

[0062] Taking the extraction of product words as an example, it can be required that if possible, the extracted product words should also contain key information such as brand and model. This can accurately locate specific products, so as to provide accurate information for downstream businesses.

[0063] a3) Evaluate the consistency between the information to be extracted and the text to be processed.

[0064] In the disclosed embodiment, information consistency means that the information extraction result is consistent with the corresponding text segment in the text to be processed. For example, if the text to be processed is: "Diamond heat sink AA manufacturer", the extracted product word is "diamond heat sink", not "diamond heat dissipation".

[0065] In summary, in the disclosed embodiments, the task requirements in the large model prompt words can be expressed in a limited and broad manner, and the semantic understanding ability of the large model can be used to accurately understand the execution specifications, thereby effectively evaluating the accuracy of the execution results of the information extraction task of the information extraction model. In addition, compared to a large number of rules and traditional machine learning methods, the broad expression of task requirements can make the way of combining the information extraction model with the large model more flexible, easier to deal with new scenarios and needs, and improve the generalization ability of information extraction.

[0066] In some embodiments, when the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, the task requirements may also include:

[0067] a4) If the initial information is incorrect, provide the reasons and / or explanation for the correction of the initial information.

[0068] The reason refers to the direct factors or conditions that lead to inaccurate initial information extracted by the information extraction model. It is a factual statement, usually concise and specific.

[0069] An explanation is a detailed explanation of the inaccuracy of the initial information extracted by the information extraction model, usually including the reasons, but also involving more background information, logical reasoning or inference. Explanations focus more on "why" and "how" the inaccuracy occurs.

[0070] In the embodiments of the present disclosure, providing inaccurate reasons and / or explanations can, on the one hand, provide necessary information for optimizing the initial information to improve the accuracy of the information extraction results, and on the other hand, can also collect sample instances that the information extraction model cannot handle well, so as to iteratively optimize the information extraction model. In addition, through the reasons and / or explanations, the information extraction method in the embodiments of the present disclosure is explainable, so as to improve the user experience and optimize the information extraction method.

[0071] In the disclosed embodiment, no matter which of the above specifications a1)-a4) is included in the execution specification for evaluating accuracy, when it is determined that the initial information is inaccurate, a rule-based method can be further used to optimize the initial information. For example, the initial information may include multiple sub-information, such as product words and services. If the product words are inaccurate, they can be re-extracted from the text to be processed in a regularized manner based on the execution rules for extracting product words. This improves the accuracy of the information extraction results.

[0072] Of course, the semantic understanding ability of the large model itself can also be used to correct inaccurate sub-information in the initial information. Based on the evaluation results of the accuracy of the information extraction task of the information extraction model, the large model can extract accurate information from the text to be processed to optimize the initial information.

[0073] In the disclosed embodiment, correcting the initial information may include correcting the originally inaccurate information and / or supplementing the information that is not extracted by the information extraction model. Therefore, when the evaluation task includes correcting the initial information, the task requirements include at least one of the following:

[0074] b1) Supplement the missing information in the initial information.

[0075] The information extraction model is usually trained based on training samples. It can accurately and quickly extract structured information for the fields it has learned well. However, it may be difficult to accurately extract information for fields it has not learned. Moreover, the information extraction model is good at extracting structured information. For some special cases, such as general descriptions and long texts, it is difficult to give a general summary. Therefore, in cases where the information extraction model is not good at it, the semantic understanding ability of the large model (such as the ability to summarize) can be used to extract additional information to supplement the initial information. For example, in the scenario of intelligent customer service and user dialogue, the user describes a lot of things, including the cause and process, but the purpose is not clearly expressed. Using the summarization ability of the large model, it can be extracted that it expects to provide legal consulting services. In this way, the text to be processed provided by the user does not have an explicit expression such as "legal consultation", and the information extraction model cannot extract "legal consultation", but the semantic understanding ability of the large model can be used to additionally extract the keyword to improve the generalization ability and accuracy of information extraction.

[0076] b2) Correct inaccurate information in the initial information based on the text to be processed.

[0077] Among them, if there is inaccurate information in the initial information extracted by the information extraction model, it may not meet the execution specifications of the information extraction. Therefore, the large model can re-extract information from the text to be processed based on the execution specifications to correct the inaccurate information and improve the accuracy of the information extraction results.

[0078] It can be seen that in the embodiments of the present disclosure, the semantic understanding ability of the large model and the task requirements of the evaluation task can be used to accurately correct errors in the initial information and provide additional supplementary information to ultimately improve the accuracy of the information extraction results.

[0079] Taking the extraction of product words as an example, the large model prompt words may include the following:

[0080] You are a professional data annotator. You need to read the input and output of the following product word extraction model (i.e., information extraction model) to determine whether the extracted product word results are accurate. If not, please correct them to get the correct results.

[0081] Task requirements:

[0082] ******

[0083] ******

[0084] At least one example

[0085] Output format requirements

[0086] Input content

[0087] Assume that Figure 3 As shown, the information extraction model is the UIE model, and the input content in the large model extraction word is "UIE model input: ** PP toughened plastic filler granulator plastic PE general talcum powder granulator

[0088] UIE model output: [“granulator”, “talcum powder granulator”]”.

[0089] Under the guidance of the large model prompt words, after the large model performed the evaluation task and optimized it, the information extraction results were corrected to "PP toughened plastic filler granulator" and "Plastic PE general talcum powder granulator".

[0090] In addition, the big model gave an explanation for the correction, namely, "the model output omitted the specific description information of the product, such as 'PP toughened plastic filler' and 'Plastic PE general talc', which is necessary to fully describe the product, so they need to be included in the extraction results."

[0091] In addition, if the initial information of the information extraction model is accurate, the large model will also give an evaluation result. As shown in the following example:

[0092] Information extraction model input: Diamond heat sink Shenzhen manufacturer

[0093] Information extraction model output: ["diamond heat sink"]

[0094] The output after large model evaluation optimization is:

[0095] {

[0096] "correctedInfo":["Diamond heat sink"],

[0097] "explaination":"The extraction is correct and does not need correction"

[0098] }

[0099] In the disclosed embodiment, in addition to improving the accuracy of the information extraction results by combining the information extraction model with a large model while taking into account the generalization ability, in order to further improve the accuracy of the information extraction results, the disclosed embodiment can also adopt a self-consistency strategy based on the large model to optimize the information extraction results.

[0100] In large models, self-consistency strategies are used to improve the model's reasoning ability. By sampling multiple reasoning results and selecting the most consistent answer, the model can solve complex problems more accurately, thereby improving the accuracy of information extraction results.

[0101] In some possible implementations, the large model may repeatedly execute S102 to obtain multiple information extraction results, and select the information extraction result with the highest frequency or probability as the optimized information extraction result. Thus, the large model can adopt a self-consistency strategy by executing S102 multiple times to improve the accuracy of information extraction.

[0102] In some other possible implementations, the large model may extract results multiple times for the sub-information that needs to be corrected based on the evaluation results, and select the sub-information with the highest frequency of occurrence or probability to update the information extraction results obtained in S102, so as to optimize the information extraction results.

[0103] For example, for the content that needs to be supplemented, multiple inferences are performed to obtain multiple supplementary contents, and then the supplementary contents with the highest frequency of occurrence or probability are selected and updated to the information extraction results.

[0104] For the information that needs to be corrected, multiple inferences are performed to obtain multiple correction results, and then the correction result with the highest frequency or probability is selected and updated to the information extraction result.

[0105] In addition, in order to further improve the accuracy of the information extraction results, the embodiments of the present disclosure may also perform post-processing optimization on the information extraction results based on the extraction rules corresponding to the information extraction task.

[0106] During implementation, for specific application fields, the extracted information is structured, and corresponding rules can be formulated based on the parts that may still be insufficient in the information extraction results after the information extraction model and the large model are combined to further optimize the information extraction results. For example, for an ordered string of numbers, the length of the characters contained in the string is limited, and the information extraction results can be checked based on the rules. If some numbers are missing, they can be re-extracted from the text to be processed based on the rules to further improve the accuracy of the information extraction results of this part of the information.

[0107] In the disclosed embodiment, after the large model uses a self-consistency strategy to optimize the information extraction results, the optimization results can be post-processed and optimized based on the extraction rules, thereby systematically solving the problem that some information extraction may be incorrect and improving the accuracy of the information extraction results.

[0108] In summary, in the embodiments of the present disclosure, the accuracy of the information extraction results and the generalization ability of the information extraction method are improved by combining the information extraction model and the big model. On this basis, the local information is further optimized and verified through the extraction rules, so as to further improve the accuracy of the information extraction results.

[0109] In addition, in the disclosed embodiment, the information extraction model is obtained through sample training, in order to be able to continuously adapt to new business needs or changes in language expression. In the disclosed embodiment, the information extraction model can also be monitored and continuously optimized, such as Figure 4 As shown, including the following:

[0110] S401, perform quality monitoring on the information extraction result to obtain a quality monitoring result.

[0111] During implementation, the required number of original texts and corresponding information extraction results for the current batch can be collected, and they can be scored using a scoring model and / or spot-checked by manual sampling to obtain quality monitoring results.

[0112] Among them, a score value higher than the preset threshold indicates that the quality monitoring result given by the scoring model is qualified. If the quality monitoring result given by the scoring model is unqualified, the quality monitoring result of the scoring model can be further reviewed by manual sampling inspection.

[0113] When implemented, the scoring model can evaluate and score on multiple evaluation dimensions, including at least one of the following evaluation dimensions: information accuracy, information completeness, and format compliance.

[0114] The scoring model can be a large language model. The scoring model scores according to the evaluation dimensions, which include information accuracy, information completeness, and format compliance for some tasks. The output of the scoring model includes scores and scoring explanations. For ease of understanding, the definitions and examples of each evaluation dimension are as follows:

[0115] C1), Evaluation dimension name: Information accuracy

[0116] Definition: "Evaluate whether the [initial information] (such as key information such as product name) extracted from the [text to be processed] is accurate. The system needs to ensure that the extracted information is completely consistent with the description in the text to be processed. Pay attention to check whether there are errors in the extracted information. If the extracted initial information such as product name is inconsistent with the text to be processed, it is considered inaccurate.

[0117] "Evaluation result value":{

[0118] "0":"The information extracted from [initial information] is completely inconsistent with the original text in [to be processed] or there are differences in format or expression. The output conclusion is \"0\". Example: When the product name in [dialogue history] is \"Phone iPro Max\", the [initial information] is extracted as \"phone I promax\", there is a difference in uppercase and lowercase letters, and the [scoring result] is \"0 points\".,

[0119] "1":"The information extracted from [initial information] is exactly the same as [text to be processed], and the output conclusion is \"1\". Example: When the product name in [text to be processed] is \"Phone i Pro Max\", the [initial information] extracted is \"Phonei Pro Max\", and the [scoring result] is \"1 point\".

[0120] C2), Evaluation dimension name: Information integrity

[0121] Definition: "Evaluate whether the [initial information] extracted from the [text to be processed] completely contains all the required key information. The system needs to extract all target information types (such as all product names involved, etc.) that appear in the text to be processed. If some key information that appears in the text to be processed is omitted, or only part of the information is extracted, the information is considered incomplete.

[0122] "Evaluation result value":{

[0123] "0":"[Initial information] omits key information in [Text to be processed], and the output conclusion is \"0\". Example: When three product names are mentioned in [Text to be processed], but [Initial information] only extracts one, the [Score Result] is \"0 points\".,

[0124] "1":"[Initial information] completely extracts key information, and the output conclusion is \"1\". Example: When [Text to be processed] mentions \"I want to buy mobile phone A and a mobile phone case\", [Initial information] is \"mobile phone A, mobile phone case\", [Score result] is \"1 point\".

[0125] C3), Evaluation dimension name: Format compliance

[0126] Definition: "Evaluate whether [initial information] simultaneously meets the following requirements: 1) strictly follows the expected output format (such as JSON format); 2) contains all required key fields. For example, the initial information extracted by the information extraction model includes two fields, action and slots, in the format of (action: string, slots: array). Note that the check requires not only the correct format, but also the inclusion of all specified key fields. Missing any requirement will be considered unqualified.

[0127] "Evaluation result value":{

[0128] "0":"[Initial information] If any of the following problems exist, the output conclusion is \"0\". Example: \n1. Format error: using single quotes such as {'action':'Query product information'}; \n2. Missing required fields: such as {\"slots\":[\"Mobile phone case\"]} (missing the required action field); \n3. Field name error: such as {\"title\":\"Mobile phone case\",\"category\":\"Accessories\"} (using non-required field names).",

[0129] "1":"[Initial information] satisfies both correct format and complete fields, and the output conclusion is \"1\". Example: When the JSON format is met and the action and slots fields are included, [Initial information] is {\"action\":\"Query product information\",\"slots\":[\"Accessories\"]}, and [Score result] is \"1 point\".

[0130] S402, when the quality monitoring result indicates unqualified, collect cases and construct an iterative sample set.

[0131] During implementation, the collected cases may include cases with qualified quality and cases with unqualified quality, so that the information extraction model can fully learn the knowledge in the new field.

[0132] S403, optimizing the information extraction model based on the iterative sample set.

[0133] That is, the information extraction model is supervisedly trained based on the iterative sample set to iteratively optimize the information extraction model.

[0134] S404, automatically evaluating the iterative optimization results of the information extraction model based on the scoring model.

[0135] In the disclosed embodiment, the information extraction model can be continuously iteratively optimized through an automatic monitoring mechanism, so that the information extraction model can continuously adapt to new scenarios and business needs, and improve the generalization ability of the information extraction method provided by the disclosed embodiment. In addition, the optimization results of the information extraction model can be automatically evaluated through the scoring model, which improves the automation level of model optimization and can reduce the evaluation cost and efficiency of model optimization.

[0136] The scoring model is consistent with the previous description. Based on the scoring model, the iterative optimization results of the information extraction model can be scored in at least one of the following evaluation dimensions to obtain the model optimization results: information accuracy, information completeness, and format compliance.

[0137] The scoring model can automatically evaluate the iterative optimization results of the information extraction model, improve the degree of automation of the information extraction method optimization, and thus improve the efficiency and accuracy of the information extraction results.

[0138] In summary, the disclosed embodiments provide a hybrid information extraction method, which organically combines the information extraction model and the big model to better leverage the advantages of both to improve the efficiency and accuracy of information extraction. The accuracy of information extraction can be further improved through self-consistency strategies and rule-based verification and supplementation methods.

[0139] Based on the same technical concept, the embodiment of the present disclosure also provides a method for training an information extraction model, which is applicable to any need to train and optimize an information extraction model. Figure 5 As shown, including the following:

[0140] S501, for the information extraction task, a large model is used to annotate the sample set to be annotated to obtain an initial sample set.

[0141] During implementation, the semantic understanding ability of the big model is relied upon to annotate each sample in the annotated sample set based on the information extraction task. For example, if the information extraction task is to extract product words, the big model needs to annotate the product words in the sample; if the information extraction task is to extract service words, the big model needs to annotate the service words in the sample. Similarly, according to the actual needs of the information extraction task, the big model completes the corresponding annotation.

[0142] S502, optimizing the initial sample set to obtain a training sample set.

[0143] That is, the labeling result of the large model depends entirely on the ability of the large model itself. In order to obtain an efficient and accurate information extraction model, the initial sample set obtained by labeling the large model is optimized in the embodiment of the present disclosure to obtain a relatively high-quality training sample set. That is, the quality of the training sample set is higher than that of the initial sample set.

[0144] S503: Train the information extraction model based on the training sample set, so that the information extraction model supports the information extraction task.

[0145] In the disclosed embodiment, an initial sample set is obtained by initial annotation based on a large model, and then a training sample set is obtained by optimizing and improving the quality of the sample set. In the whole process, the cost of manual annotation can be reduced, and even manual annotation is not required, and the information extraction model can be trained using a high-quality training sample set to improve the quality of information extracted by the information extraction model.

[0146] In some embodiments, the initial sample set may be optimized by manual review and correction to obtain a training sample set.

[0147] In order to further reduce labor costs and improve the efficiency of model optimization, in the embodiment of the present disclosure, the initial sample set is optimized to obtain a training sample set, which can be implemented as follows: Figure 6 As shown:

[0148] S601, based on the information processing rules corresponding to the information extraction task, the initial sample set is cleaned and corrected to obtain a sample set to be verified.

[0149] The specific information processing rules can be determined according to the actual information extraction task. For example, if the extracted target information has a fixed format, a regular matching method can be used to verify whether the extracted target information is correct and complete.

[0150] In addition, information processing rules can be constructed based on the characteristics of the annotation results output by the large model to complete the data cleaning task. During implementation, the experience summarized a posteriori from the output results of the large model can be used to generate corresponding information processing rules. For example, through information processing rules, annotation results whose output formats do not meet the requirements can be screened out, thereby cleaning such data. For another example, in the scenario where the large model fails to extract, it will output keywords such as "not mentioned" and "cannot be analyzed". By analyzing specific keywords of this type, data cleaning can be performed.

[0151] In addition, for specific application scenarios, some input sources have certain patterns, so regular matching can be used to verify and correct such information. For example, the content format submitted through a form is fixed as "Hope to obtain xxxx", and the labeling results of the large model can be verified through regular matching. If there is a problem with the labeling of the large model, the correct labeling results can be obtained from the original sentence of the sample through regular matching to correct the labeling results of the large model.

[0152] S602: Output the sample set to be verified to the target terminal, so as to correct the initial sample set through human-computer interaction to obtain a training sample set.

[0153] In the disclosed embodiment, information extraction rules are first used to automatically clean and optimize the annotation results of the large model, and then corrections are made using human-computer interaction to further improve the quality of training samples, thereby reducing manual annotation costs and improving model training efficiency.

[0154] In some embodiments, in S602, correcting the initial sample set by human-computer interaction may be implemented as follows:

[0155] S6021, perform preliminary correction on the initial sample set through human-computer interaction to obtain an intermediate sample set.

[0156] That is, the corresponding initial sample set can be pushed to the target terminal of the relevant labeler so that the labeler can perform spot checks and corrections on it, thereby reducing the cost of manual labeling and obtaining an intermediate sample set of manual labeling quality.

[0157] S6022, evaluate the intermediate sample set and obtain an evaluation result.

[0158] During implementation, the evaluation can be based on the scoring model described above, and can be implemented as follows: based on the scoring model, the annotation results of the intermediate sample set are scored in at least one of the following evaluation dimensions to obtain an evaluation result: information accuracy, information completeness, and format compliance.

[0159] The scoring model can automatically evaluate the sample quality, thereby improving the training efficiency of the information extraction model.

[0160] S6023: When the evaluation result meets the preset requirements, the intermediate sample set is determined as the training sample set.

[0161] In the disclosed embodiment, the sample quality is improved through human-computer interaction, and whether the sample quality meets the requirements can be determined through automatic evaluation, thereby facilitating the acquisition of a high-quality training sample set, thereby improving the training efficiency of the information extraction model.

[0162] In some embodiments, Figure 6 As shown, it also includes:

[0163] S603, when the evaluation result does not meet the preset requirements, filter out the sample set with inaccurate annotations from the intermediate sample set to construct a sample set to be corrected.

[0164] S604: construct a labeling prompt word based on the reason why the labeling result of each sample in the to-be-corrected sample set contained in the evaluation result is inaccurate.

[0165] For example, the scoring model not only evaluates whether the sample is qualified, but also gives the reason for failure if it is not qualified.

[0166] S605, using the large model under the guidance of the annotation prompt words, return to execute S501 to use the large model to annotate the sample set to be annotated, and obtain the step of obtaining the initial sample set, until a high-quality training sample set is obtained.

[0167] In the disclosed embodiment, in the process of automatic labeling, if the sample quality does not meet the preset requirements, the reasons for the inaccurate labeling results can be analyzed and the semantic understanding ability of the large model can be used to re-label in order to obtain high-quality training samples. By automatically improving the sample quality, the cost of manual labeling can be reduced and the information extraction ability of the information extraction model can be improved.

[0168] In some embodiments, in order to improve the training efficiency of the information extraction model, the information extraction model is trained based on the training sample set, which can be implemented as follows: Figure 7 As shown:

[0169] S701, for any target training sample in the training sample set, convert the target training sample into a target format, wherein the target format includes a sample sentence and an extracted label in the target training sample, wherein the extracted label includes the start and end positions of the annotation result of the target training sample in the sample sentence; the extracted label is used to calculate the training loss.

[0170] For example, the large model annotation results are output according to the output format of the large model, but they cannot be directly used to train the information extraction model. In the embodiment of the present disclosure, a regular matching method can be used to determine the start and end positions of the annotation results of the large model in the sample sentence of the target training sample. Then, the extraction label of the information extraction model is constructed.

[0171] S702, constructing training prompt words based on the target format.

[0172] In a possible example, the prompt words constructed may include sample sentences of the target training sample, so that the information extraction model can perform the information extraction task on it. It can also include the aforementioned target format to facilitate the determination of training loss. It can also include a brief task prompt to enable the information extraction model to perform the task category through the brief task prompt, for example, to let the information extraction model know whether to extract product words or service words. Of course, it is determined according to the business needs of the information extraction task what kind of content to extract, so as to facilitate the construction of training prompt words.

[0173] S703: training an information extraction model based on the training prompt words.

[0174] Among them, the information extraction model is supervisedly trained by training prompt words, and the training loss includes the predicted start and end positions of the extraction results predicted by the information extraction model, and the start and end positions in the extracted labels are compared and learned to determine the training loss. Among them, the starting position predicted by the information extraction model and the starting position of the extracted label determine the cross entropy loss, and the end position predicted by the information extraction model and the end position of the extracted label determine the cross entropy loss. These two cross entropy losses can be further fused, for example, by summing them up to obtain the predicted position loss. Furthermore, the model parameters of the information extraction model can be optimized based on the predicted position loss.

[0175] In summary, in the embodiments of the present disclosure, the samples annotated by the large model are converted into corresponding formats to construct training prompt words, so as to optimize the training information extraction model and improve the training efficiency of the information extraction model.

[0176] Based on the same technical concept, such as Figure 8 As shown, the embodiment of the present disclosure also provides an artificial intelligence-based information extraction system 800, comprising:

[0177] Agent 801, in which an information extraction model and a large model are integrated, wherein:

[0178] The information extraction model performs information extraction tasks on the text to be processed to extract initial information from the text to be processed;

[0179] The large model is used to evaluate the execution results of the information extraction model on the information extraction task of the text to be processed, so as to optimize the initial information and obtain the information extraction results of the text to be processed.

[0180] In some embodiments, the large model is also used to optimize the information extraction results using a self-consistency strategy.

[0181] In some embodiments, it also includes:

[0182] The post-processing system 802 is used to perform post-processing optimization on the information extraction results based on the extraction rules corresponding to the information extraction task.

[0183] In some embodiments, an optimization system 803 is further included for:

[0184] Perform quality monitoring on the information extraction results to obtain quality monitoring results;

[0185] Based on the quality monitoring results, collect cases with unqualified quality and build an iterative sample set;

[0186] Optimize the information extraction model based on iterative sample sets;

[0187] The iterative optimization results of the information extraction model are automatically evaluated based on the scoring model.

[0188] It should be noted that the operations performed by the aforementioned information extraction model and the large model refer to the aforementioned method embodiments, and the embodiments of the present disclosure will not be repeated here.

[0189] Based on the same technical concept, the present disclosure also proposes an information extraction device 900 based on artificial intelligence, such as Fig. 9 As shown, including:

[0190] An extraction module 901 is used to perform an information extraction task on the text to be processed based on the information extraction model to extract initial information from the text to be processed;

[0191] The first optimization module 902 is used to evaluate the execution result of the information extraction task of the information extraction model for the text to be processed based on the large model, so as to optimize the initial information and obtain the information extraction result of the text to be processed.

[0192] In some embodiments, the first optimization module includes:

[0193] A construction unit is used to construct a large model prompt word based on the to-be-processed text and the initial information for the information extraction task; the large model prompt word indicates the evaluation task to be performed by the large model; the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, and / or correcting the initial information;

[0194] The optimization unit is used to input the large model prompt words into the large model so that the large model performs the evaluation task and obtains the information extraction result of the text to be processed.

[0195] In some embodiments, the large model prompt includes task requirements of the assessment task.

[0196] In some embodiments, where the evaluation task includes verifying the accuracy of initial information extracted by the information extraction model, the task requirements include at least one of the following:

[0197] The information to be extracted must meet basic requirements, which are determined based on the information extraction task;

[0198] In certain cases, the information items to be extracted include:

[0199] Evaluate the consistency between the information to be extracted and the text to be processed.

[0200] In some embodiments, where the evaluation task includes verifying the accuracy of initial information extracted by the information extraction model, the task requirements include:

[0201] Where the initial information is incorrect, provide reasons and / or explanation why the initial information needs to be corrected.

[0202] In some embodiments, where the assessment task includes correcting initial information, the task requirements include at least one of the following:

[0203] Supplement the missing information in the initial information;

[0204] Corrections to inaccurate information in the initial message will be made based on the text to be processed.

[0205] In some embodiments, a second optimization module is also included for optimizing the information extraction results by adopting a self-consistency strategy based on the large model.

[0206] In some embodiments, a third optimization module is also included, which is used to perform post-processing optimization on the information extraction results based on the extraction rules corresponding to the information extraction task.

[0207] In some embodiments, an iterative optimization module is further included, which is used to:

[0208] Perform quality monitoring on the information extraction results to obtain quality monitoring results;

[0209] When the quality monitoring results indicate failure, collect cases and construct an iterative sample set;

[0210] Optimize the information extraction model based on iterative sample sets;

[0211] The iterative optimization results of the information extraction model are automatically evaluated based on the scoring model.

[0212] In some embodiments, the iterative optimization module is specifically used to score the iterative optimization result of the information extraction model in at least one of the following evaluation dimensions based on the scoring model to obtain the model optimization result;

[0213] Information accuracy, information completeness, and format compliance.

[0214] Based on the same technical concept, the present disclosure also proposes a training device 1000 for an information extraction model. Fig.10 As shown, including:

[0215] The labeling module 1001 is used to label the sample set to be labeled using a large model for information extraction tasks to obtain an initial sample set;

[0216] The fourth optimization module 1002 is used to optimize the initial sample set to obtain a training sample set;

[0217] The training module 1003 is used to train the information extraction model based on the training sample set so that the information extraction model supports the information extraction task.

[0218] In some embodiments, the fourth optimization module includes:

[0219] A processing unit, used to clean and correct the initial sample set based on the information processing rules corresponding to the information extraction task to obtain a sample set to be verified;

[0220] The correction unit is used to output the sample set to be verified to the target terminal, so as to correct the initial sample set through human-computer interaction to obtain the training sample set.

[0221] In some embodiments, the correction unit is specifically configured to:

[0222] The initial sample set is preliminarily corrected by human-computer interaction to obtain an intermediate sample set;

[0223] Evaluate the intermediate sample set and obtain the evaluation result;

[0224] When the evaluation results meet the preset requirements, the intermediate sample set is determined as the training sample set.

[0225] In some embodiments, the correction unit is specifically configured to:

[0226] Based on the scoring model, the annotation results of the intermediate sample set are scored in at least one of the following evaluation dimensions to obtain an evaluation result;

[0227] Information accuracy, information completeness, and format compliance.

[0228] In some embodiments, a fifth optimization module is further included, configured to:

[0229] When the evaluation results do not meet the preset requirements, the sample set with inaccurate labels is selected from the intermediate sample set to construct the sample set to be corrected;

[0230] Constructing annotation prompt words based on the reasons why the annotation results of each sample in the to-be-corrected sample set included in the evaluation results are inaccurate;

[0231] Using the large model under the guidance of the labeling prompt word, return to execute the step of labeling the sample set to be labeled using the large model to obtain the initial sample set.

[0232] In some embodiments, the training module includes:

[0233] A conversion unit, for converting any target training sample in the training sample set into a target format, wherein the target format includes a sample sentence and an extracted label in the target training sample, and the extracted label includes the start and end positions of the labeled result of the target training sample in the sample sentence; the extracted label is used to calculate the training loss;

[0234] A construction unit, used to construct training prompt words based on the target format;

[0235] The training unit is used to train the information extraction model based on the training prompt words.

[0236] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0237] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0238] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0239] Fig.11 A schematic block diagram of an example electronic device 1100 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0240] like Fig.11As shown, the device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the device 1100 can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0241] A number of components in the device 1100 are connected to the I / O interface 1105, including: an input unit 1106, such as a keyboard, a mouse, etc.; an output unit 1107, such as various types of displays, speakers, etc.; a storage unit 1108, such as a disk, an optical disk, etc.; and a communication unit 1109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1109 allows the device 1100 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0242] The computing unit 1101 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1101 performs the various methods and processes described above, such as information extraction methods based on artificial intelligence and / or training methods for information extraction models. For example, in some embodiments, information extraction methods based on artificial intelligence and / or training methods for information extraction models may be implemented as computer software programs, which are tangibly included in machine-readable media, such as storage units 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by computing unit 1101, one or more steps of the information extraction method based on artificial intelligence and / or the training method of the information extraction model described above may be performed. Alternatively, in other embodiments, computing unit 1101 may be configured to perform the information extraction method based on artificial intelligence and / or the training method of the information extraction model by any other appropriate means (e.g., by means of firmware).

[0243] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0244] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0245] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0246] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0247] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0248] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0249] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0250] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. An information extraction method based on artificial intelligence, comprising: Performing an information extraction task on the text to be processed based on the information extraction model to extract initial information from the text to be processed; Based on the large model, the execution result of the information extraction model for the information extraction task of the text to be processed is evaluated to optimize the initial information and obtain the information extraction result of the text to be processed.

2. The method according to claim 1, wherein: The step of evaluating the execution result of the information extraction task of the information extraction model for the text to be processed based on the large model to optimize the initial information and obtain the information extraction result of the text to be processed includes: For the information extraction task, a large model prompt word is constructed based on the text to be processed and the initial information; the large model prompt word indicates an evaluation task that the large model needs to perform; the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, and / or correcting the initial information; The large model prompt words are input into the large model so that the large model performs the evaluation task and obtains the information extraction result of the text to be processed.

3. The method according to claim 2, wherein: The large model prompt words include the task requirements of the evaluation task.

4. The method according to claim 3, wherein: In the case where the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, the task requirements include at least one of the following: The information to be extracted must meet basic requirements, where the basic requirements are determined based on the information extraction task; In certain cases, the information items to be extracted include: Evaluate the consistency between the information to be extracted and the information of the text to be processed.

5. The method according to claim 3, wherein: In the case where the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, the task requirements include: In the event that the initial information is incorrect, provide the reason and / or explanation why the initial information needs to be corrected.

6. The method according to claim 3, wherein: In the case where the assessment task includes correcting the initial information, the task requirements include at least one of the following: Supplement the missing information in the initial information; The inaccurate information in the initial information is corrected based on the text to be processed.

7. The method according to any one of claims 1 to 6, further comprising: The information extraction result is optimized by adopting a self-consistency strategy based on the large model.

8. The method according to any one of claims 1 to 7, further comprising: Based on the extraction rules corresponding to the information extraction task, the information extraction result is post-processed and optimized.

9. The method according to any one of claims 1 to 8, further comprising: Performing quality monitoring on the information extraction result to obtain a quality monitoring result; When the quality monitoring result indicates failure, collecting cases and constructing an iterative sample set; Optimizing the information extraction model based on the iterative sample set; The iterative optimization results of the information extraction model are automatically evaluated based on the scoring model.

10. The method according to claim 9, wherein: Automatically evaluating the iterative optimization results of the information extraction model based on the scoring model includes: Based on the scoring model, scoring the iterative optimization result of the information extraction model in at least one of the following evaluation dimensions to obtain a model optimization result; Information accuracy, information completeness, and format compliance.

11. A method for training an information extraction model, comprising: For information extraction tasks, a large model is used to annotate the sample set to be annotated to obtain the initial sample set; Optimizing the initial sample set to obtain a training sample set; The information extraction model is trained based on the training sample set so that the information extraction model supports the information extraction task.

12. The method according to claim 11, wherein: The step of optimizing the initial sample set to obtain a training sample set includes: Based on the information processing rules corresponding to the information extraction task, the initial sample set is cleaned and corrected to obtain a sample set to be verified; The sample set to be verified is output to a target terminal, so as to correct the initial sample set through human-computer interaction to obtain the training sample set.

13. The method according to claim 12, wherein: The correcting the initial sample set by human-computer interaction to obtain the training sample set includes: Performing preliminary correction on the initial sample set by means of the human-computer interaction to obtain an intermediate sample set; Evaluating the intermediate sample set to obtain an evaluation result; When the evaluation result meets the preset requirement, the intermediate sample set is determined as the training sample set.

14. The method according to claim 13, wherein: The step of evaluating the intermediate sample set to obtain an evaluation result includes: Based on the scoring model, scoring the annotation results of the intermediate sample set in at least one of the following evaluation dimensions to obtain the evaluation result; Information accuracy, information completeness, and format compliance.

15. The method according to claim 13, further comprising: When the evaluation result does not meet the preset requirement, filter out the sample set with inaccurate annotations from the intermediate sample set to construct a sample set to be corrected; Constructing a labeling prompt word based on the reason why the labeling result of each sample in the to-be-corrected sample set contained in the evaluation result is inaccurate; Using the large model under the guidance of the labeling prompt word, return to the step of labeling the sample set to be labeled using the large model to obtain an initial sample set.

16. The method according to any one of claims 11 to 15, wherein: The step of training the information extraction model based on the training sample set includes: For any target training sample in the training sample set, convert the target training sample into a target format, wherein the target format includes a sample sentence and an extraction label in the target training sample, and the extraction label includes the start and end positions of the labeling result of the target training sample in the sample sentence; the extraction label is used to calculate the training loss; Based on the target format, construct training prompt words; Based on the training prompt words, the information extraction model is trained.

17. An information extraction system based on artificial intelligence, comprising: An intelligent agent, wherein the intelligent agent integrates an information extraction model and a large model, wherein: The information extraction model performs an information extraction task on the text to be processed to extract initial information from the text to be processed; The large model is used to evaluate the execution result of the information extraction model on the information extraction task of the text to be processed, so as to optimize the initial information and obtain the information extraction result of the text to be processed.

18. The system of claim 17, wherein: The large model is also used to optimize the information extraction results by adopting a self-consistency strategy.

19. The system according to claim 17 or 18, further comprising: A post-processing system is used to perform post-processing optimization on the information extraction result based on the extraction rules corresponding to the information extraction task.

20. The system according to any one of claims 17 to 19, further comprising an optimization system for: Performing quality monitoring on the information extraction result to obtain a quality monitoring result; Based on the quality monitoring results, cases with unqualified quality are collected to construct an iterative sample set; Optimizing the information extraction model based on the iterative sample set; The iterative optimization results of the information extraction model are automatically evaluated based on the scoring model.

21. An information extraction device based on artificial intelligence, comprising: An extraction module, used for performing an information extraction task on the text to be processed based on the information extraction model, so as to extract initial information from the text to be processed; The first optimization module is used to evaluate the execution result of the information extraction task of the information extraction model for the text to be processed based on the large model, so as to optimize the initial information and obtain the information extraction result of the text to be processed.

22. The device according to claim 21, wherein The first optimization module comprises: A construction unit is used to construct a large model prompt word based on the text to be processed and the initial information for the information extraction task; the large model prompt word indicates an evaluation task that the large model needs to perform; the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, and / or correcting the initial information; The optimization unit is used to input the large model prompt words into the large model so that the large model performs the evaluation task and obtains the information extraction result of the text to be processed.

23. The device according to claim 22, wherein: The large model prompt words include the task requirements of the evaluation task.

24. The device according to claim 23, wherein: In the case where the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, the task requirements include at least one of the following: The information to be extracted must meet basic requirements, where the basic requirements are determined based on the information extraction task; In certain cases, the information items to be extracted include: Evaluate the consistency between the information to be extracted and the information of the text to be processed.

25. The device according to claim 23, wherein: In the case where the evaluation task includes verifying the accuracy of the initial information extracted by the information extraction model, the task requirements include: In the event that the initial information is incorrect, provide the reason and / or explanation why the initial information needs to be corrected.

26. The device according to claim 23, wherein In the case where the assessment task includes correcting the initial information, the task requirements include at least one of the following: Supplement the missing information in the initial information; The inaccurate information in the initial information is corrected based on the text to be processed.

27. The device according to any one of claims 21-26 further includes a second optimization module for optimizing the information extraction result by adopting a self-consistency strategy based on the large model.

28. The device according to any one of claims 21-27 further includes a third optimization module, which is used to perform post-processing optimization on the information extraction result based on the extraction rules corresponding to the information extraction task.

29. The apparatus according to any one of claims 21 to 28, further comprising an iterative optimization module, configured to: Performing quality monitoring on the information extraction result to obtain a quality monitoring result; When the quality monitoring result indicates failure, collecting cases and constructing an iterative sample set; Optimizing the information extraction model based on the iterative sample set; The iterative optimization results of the information extraction model are automatically evaluated based on the scoring model.

30. The device according to claim 29, wherein: The iterative optimization module is specifically used to score the iterative optimization result of the information extraction model in at least one of the following evaluation dimensions based on the scoring model to obtain a model optimization result; Information accuracy, information completeness, and format compliance.

31. A training device for an information extraction model, comprising: The annotation module is used to annotate the sample set to be annotated using a large model for information extraction tasks to obtain an initial sample set; A fourth optimization module, used for optimizing the initial sample set to obtain a training sample set; A training module is used to train the information extraction model based on the training sample set so that the information extraction model supports the information extraction task.

32. The device according to claim 31, wherein The fourth optimization module comprises: A processing unit, configured to clean and correct the initial sample set based on the information processing rule corresponding to the information extraction task to obtain a sample set to be verified; The correction unit is used to output the sample set to be verified to the target terminal, so as to correct the initial sample set in a human-computer interaction manner to obtain the training sample set.

33. The device according to claim 32, wherein: The correction unit is specifically used for: Performing preliminary correction on the initial sample set by means of the human-computer interaction to obtain an intermediate sample set; Evaluating the intermediate sample set to obtain an evaluation result; When the evaluation result meets the preset requirement, the intermediate sample set is determined as the training sample set.

34. The device according to claim 33, wherein: The correction unit is specifically used for: Based on the scoring model, scoring the annotation results of the intermediate sample set in at least one of the following evaluation dimensions to obtain the evaluation result; Information accuracy, information completeness, and format compliance.

35. The apparatus according to claim 33, further comprising a fifth optimization module, configured to: When the evaluation result does not meet the preset requirement, filter out the sample set with inaccurate annotations from the intermediate sample set to construct a sample set to be corrected; Constructing a labeling prompt word based on the reason why the labeling result of each sample in the to-be-corrected sample set contained in the evaluation result is inaccurate; Using the large model under the guidance of the labeling prompt word, return to the step of labeling the sample set to be labeled using the large model to obtain an initial sample set.

36. The device according to any one of claims 31 to 35, wherein: The training module comprises: a conversion unit, configured to convert any target training sample in the training sample set into a target format, wherein the target format includes a sample sentence and an extraction label in the target training sample, and the extraction label includes the start and end positions of the labeling result of the target training sample in the sample sentence; the extraction label is used to calculate the training loss; A construction unit, used for constructing a training prompt word based on the target format; A training unit is used to train the information extraction model based on the training prompt words.

37. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 16.

38. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-16.

39. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 16.