Task processing method, information extraction model training method, and classification task processing method
By constructing target prompt information and inputting information extraction models, the resource consumption and model limitations of complex and diverse text processing tasks in the prior art are solved, and unified and efficient processing of multi-type text processing is achieved.
Patent Information
- Application Number
- PCT/CN2024/124563
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-10-12
- Publication Date
- 2025-06-05
AI Technical Summary
When handling complex and diverse text processing tasks, the prior art needs to utilize different machine learning models, resulting in large consumption of model training resources and high model limitations, and lack of general task processing solutions.
A task processing method is proposed. By obtaining the pending text and task description information, building the target prompt information, and extracting the model with the text input information, and generating an information extraction matrix to determine the task processing result. This method is suitable for multi-type text processing tasks, improving the scalability and versatility of the model.
It realizes unified processing of multi-type text processing tasks, improves the scalability of the information extraction model and the universality of task processing, and ensures the accuracy and stability of task processing results.
Smart Images

Figure CN2024124563_05062025_PF_FP_ABST
Abstract
Description
Task processing, information extraction model training, and classification task processing methods
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311608227X and application name “Task processing, information extraction model training and classification task processing method”, the entire contents of which are incorporated by reference into this disclosure. Technical Field
[0002] The embodiments of this specification relate to the field of computer technology, and in particular to task processing, information extraction model training, and classification task processing methods. Background Art
[0003] With the development of computer technology, text processing has become increasingly dependent on the Internet. Text processing is the process of analyzing, understanding, and extracting text, and has been widely used in various fields of people's daily lives.
[0004] Currently, different processing models are typically used for different text processing tasks, such as using information extraction models for information extraction and text classification models for text classification. However, complex and diverse text processing tasks require the use of different machine learning models, which results in high resource consumption during model training and high limitations in the trained models. Therefore, a universal task processing solution is urgently needed.
[0005] Summary of the Invention
[0006] In view of this, embodiments of this specification provide a task processing method. One or more embodiments of this specification also relate to an information extraction model training method, a classification task processing method, a task processing device, an information extraction model training device, a classification task processing device, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.
[0007] According to a first aspect of an embodiment of this specification, a task processing method is provided, including:
[0008] Get the pending text and task description information of the target task;
[0009] According to the task type of the target task, extract the category information to be extracted from the task description information, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type;
[0010] Concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the text to be processed and the target prompt information;
[0011] According to the information extraction matrix, the task processing result of the target task is determined.
[0012] According to a second aspect of an embodiment of this specification, a method for training an information extraction model is provided, which is applied to a cloud-side device and includes:
[0013] Obtaining a sample set, wherein the sample set includes sample texts and sample task description information of multiple sample tasks, the sample texts carry sample matrix labels, and the sample matrix labels are obtained based on sample processing results of the sample texts;
[0014] For any sample task, according to the sample task type of the sample task, the sample extraction category information is extracted from the sample task description information, and the sample prompt information is constructed according to the preset prompt format corresponding to the sample extraction category information and the sample task type;
[0015] Splicing the sample prompt information and the sample text, and inputting the spliced sample information into the information extraction model to obtain a prediction information extraction matrix, wherein the prediction information extraction matrix represents the correspondence between the sample text and the sample prompt information;
[0016] According to the sample matrix labels and the predicted information extraction matrix, the model parameters of the information extraction model are adjusted to obtain a trained information extraction model.
[0017] According to a third aspect of the embodiments of this specification, a classification task processing method is provided, including:
[0018] Obtain the pending text and task description information of the target classification task;
[0019] According to the task type of the target classification task, the category information to be extracted is extracted from the task description information, and the target prompt information is constructed according to the preset prompt format corresponding to the category information to be extracted and the task type;
[0020] Concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the text to be processed and the target prompt information;
[0021] According to the information extraction matrix, the classification result of the target classification task is determined.
[0022] According to a fourth aspect of the embodiments of this specification, there is provided a task processing device, including:
[0023] A first acquisition module is configured to acquire the to-be-processed text and task description information of the target task;
[0024] The first construction module is configured to extract the category information to be extracted from the task description information according to the task type of the target task, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type;
[0025] A first input module is configured to concatenate target prompt information and the text to be processed, and input the concatenated information into an information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the text to be processed and the target prompt information;
[0026] The first generation module is configured to determine a task processing result of the target task according to the information extraction matrix.
[0027] According to a fifth aspect of the embodiments of this specification, there is provided an information extraction model training device, which is applied to a cloud-side device, including:
[0028] A second acquisition module is configured to acquire a sample set, wherein the sample set includes sample texts and sample task description information of multiple sample tasks, the sample texts carry sample matrix labels, and the sample matrix labels are obtained based on sample processing results of the sample texts;
[0029] The second construction module is configured to extract sample extraction category information from the sample task description information according to the sample task type of the sample task for any sample task, and construct sample prompt information according to the preset prompt format corresponding to the sample extraction category information and the sample task type;
[0030] A second input module is configured to concatenate the sample prompt information and the sample text, and input the concatenated sample information into the information extraction model to obtain a prediction information extraction matrix, wherein the prediction information extraction matrix represents the correspondence between the sample text and the sample prompt information;
[0031] The adjustment module is configured to adjust the model parameters of the information extraction model according to the sample matrix label and the prediction information extraction matrix to obtain a trained information extraction model.
[0032] According to a sixth aspect of the embodiments of this specification, there is provided a classification task processing apparatus, comprising:
[0033] A third acquisition module is configured to acquire the to-be-processed text and task description information of the target classification task;
[0034] The third construction module is configured to extract the category information to be extracted from the task description information according to the task type of the target classification task, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type;
[0035] a third input module configured to concatenate the target prompt information and the text to be processed, and input the concatenated information into an information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the text to be processed and the target prompt information;
[0036] The second generation module is configured to determine the classification result of the target classification task according to the information extraction matrix.
[0037] According to a seventh aspect of the embodiments of this specification, a computing device is provided, including:
[0038] memory and processor;
[0039] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.
[0040] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.
[0041] According to a ninth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in the first aspect, the second aspect, or the third aspect above.
[0042] The task processing method provided by one embodiment of the present specification obtains the to-be-processed text and task description information of the target task; extracts the category information to be extracted from the task description information according to the task type of the target task, and constructs the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type; splices the target prompt information and the to-be-processed text, and inputs the spliced information into the information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the to-be-processed text and the target prompt information; and determines the task processing result of the target task according to the information extraction matrix. By constructing the target prompt information based on the task type, the target prompt information can be used to prompt the information extraction model with the text processing task type, so that the information extraction model can process multiple types of text processing tasks, thereby improving the scalability of the information extraction model and the versatility of task processing. Moreover, since the task processing result is derived from the task description information and the to-be-processed text, the accuracy and stability of the task processing result are guaranteed. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] FIG1 is an architecture diagram of a task processing system provided by one embodiment of this specification;
[0044] FIG2 is an architecture diagram of another task processing system provided by one embodiment of this specification;
[0045] FIG3 is a flowchart of a task processing method provided by one embodiment of this specification;
[0046] FIG4 is a framework diagram of an information extraction model in a task processing method provided in one embodiment of this specification;
[0047] FIG5 is a flow chart of an information extraction model training method provided by one embodiment of this specification;
[0048] FIG6 is a flowchart of a classification task processing method provided by one embodiment of this specification;
[0049] FIG7 a is a flowchart of a processing process of a first task processing method provided by an embodiment of this specification;
[0050] FIG7 b is a flowchart of a second task processing method provided in one embodiment of this specification;
[0051] FIG7 c is a flowchart of a third task processing method provided in an embodiment of this specification;
[0052] FIG8 is a flowchart of a fourth task processing method according to an embodiment of the present disclosure;
[0053] FIG9 is a flowchart of a fifth task processing method according to an embodiment of the present disclosure;
[0054] FIG10 is a schematic diagram of the structure of a task processing device provided by one embodiment of this specification;
[0055] FIG11 is a schematic diagram of the structure of an information extraction model training device provided by one embodiment of this specification;
[0056] FIG12 is a schematic diagram of the structure of a classification task processing device provided by one embodiment of this specification;
[0057] FIG13 is a structural block diagram of a computing device provided in one embodiment of this specification. DETAILED DESCRIPTION
[0058] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0059] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0060] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0061] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0062] First, the terms involved in one or more embodiments of this specification are explained.
[0063] Linking mechanism: The Linking mechanism is a decoding mechanism in the information extraction task, which extracts entities by connecting the entity label with the entity fragment content in the original text in some way.
[0064] Schema: Schema is a structured template expression of the expected output content, usually in JSON format, that is, it is constructed as a set of key-value pairs.
[0065] Prompt: Prompt refers to the use of a template, which is usually a prompt message composed of natural language. By inserting it before the input text, it will form training data together with the original input. The same template is used during inference to unify the training and inference forms, thereby fully realizing the potential of the pre-trained model.
[0066] Natural Language Understanding: Natural Language Understanding (NLP) refers to non-generative natural language processing (NLP) tasks. It typically takes a piece of text as input and outputs externally given text labels or parts of the text. Natural language understanding includes named entity recognition, relationship extraction, event extraction, attribute sentiment extraction, sentiment classification, text classification, text matching, reading comprehension, and natural language reasoning.
[0067] Zero-shot learning: Zero-shot learning involves making predictions on data distributions that the model has not learned before, without any further training. Zero-shot learning is often used in cold-start scenarios.
[0068] Few-shot learning: The model is fine-tuned on only a very small number of examples (e.g., 1, 5, or 10), and then predictions are made on that data. Few-shot learning is often used in cold-start or low-resource scenarios.
[0069] Classification tasks: Classification tasks involve determining which label type an input text belongs to, given an external label set and input text. In addition to typical classification tasks, text matching, coreference resolution, natural language inference, and some reading comprehension tasks can also be converted into classification tasks.
[0070] Extraction tasks: Given an external label set and input text, an extraction task involves extracting the portion of the input text that satisfies the label conditions. This includes tasks such as entity recognition, relationship extraction, event extraction, and attribute sentiment extraction.
[0071] Generation task: A generation task is a task in which the model autonomously generates the expected answer given a context.
[0072] Named Entity Recognition: Named entity recognition refers to the task of extracting entity fragments from input text.
[0073] Relation extraction: Relation extraction refers to the task of extracting subject and object entity fragments with a specific relationship type in the input text.
[0074] Event extraction: Event extraction refers to the task of extracting interconnected event argument fragments from the input text.
[0075] Attribute sentiment extraction: Attribute sentiment extraction refers to the task of extracting relevant sentiment word fragments of specific objects in the input text.
[0076] Co-reference resolution: Co-reference resolution is the task of determining whether a given pronoun refers to an object in the input text.
[0077] Sentiment classification: Sentiment classification refers to the task of returning a label that can appropriately represent the emotional tendency of the input text given an input text and a candidate sentiment label.
[0078] Text classification: Text classification refers to the task of returning a label that can appropriately represent the category to which the input text belongs, given an input text and a candidate category label.
[0079] Text matching: Text matching refers to the task of returning a label that can appropriately represent the similarity between two input text segments given two input text segments and corresponding similarity labels. Among them, text similarity labels are usually "similar" and "dissimilar".
[0080] Natural language reasoning: Natural language reasoning refers to the task of returning a label that can appropriately represent the relationship between two input texts given two input texts and corresponding text relationship labels. Among them, text relationship labels are usually "implied", "contradictory" and "neutral".
[0081] Reading comprehension: Reading comprehension refers to the task of inputting questions and reference text and returning answers based on the reference text. Generally, reading comprehension can be divided into selective reading comprehension and extraction reading comprehension.
[0082] Natural language understanding involves various types of tasks in different forms, and different processing models can usually be used to handle different text processing tasks. However, for complex and diverse text processing tasks, different machine learning models need to be used, resulting in high resource consumption in the model training process and high limitations of the trained models. Previous work has spliced classification labels to the front of the original text, and then adopted an extraction framework for task processing in order to achieve the unification of classification and extraction models. However, due to the method of splicing labels to the front of the original text, the original text and labels will share the model's maximum input text length of 512. Therefore, when the number of labels is very large, the original text will be truncated, which will affect the classification effect of the model; at the same time, the splicing method means that complex hierarchical structures cannot be represented, so the above solution cannot handle hierarchical classification tasks.
[0083] To address these issues, the embodiments of this specification propose a universal natural language understanding solution for zero-shot and few-shot learning based on the extraction paradigm. Leveraging the Linking mechanism used for the extraction model, this solution constructs target prompt information by presetting prompt formats and corresponding data. This unlocks the information extraction model's capabilities for multi-category, multi-label, and hierarchical classification, and provides a unified solution for all natural language understanding tasks except generation tasks, truly achieving a "one model for all natural language understanding tasks" approach. Furthermore, the information extraction model is deeply and extensively trained on high-quality annotated data, endowing it with powerful zero-shot and few-shot learning capabilities.
[0084] Specifically, the method obtains the target task's to-be-processed text and task description information; extracts the to-be-extracted category information from the task description information based on the target task's task type, and constructs target prompt information based on a preset prompt format corresponding to the to-be-extracted category information and the task type; concatenates the target prompt information and the to-be-processed text, and inputs the concatenated information into an information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the to-be-processed text and the target prompt information; and determines the task processing result of the target task based on the information extraction matrix. By constructing target prompt information based on the task type, the target prompt information can be used to prompt the information extraction model with the text processing task type, enabling the information extraction model to handle multiple types of text processing tasks, improving the scalability of the information extraction model and the versatility of task processing. Furthermore, since the task processing results are derived from the task description information and the to-be-processed text, the accuracy and stability of the task processing results are guaranteed.
[0085] In this specification, a task processing method is provided. This specification also involves an information extraction model training method, a classification task processing method, a task processing device, an information extraction model training device, a classification task processing device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.
[0086] Referring to FIG1 , FIG1 shows an architecture diagram of a task processing system provided by an embodiment of this specification. The task processing system may include a client 100 and a server 200;
[0087] The client 100 is used to send the target task's pending text and task description information to the server 200;
[0088] The server 200 is configured to extract the category information to be extracted from the task description information based on the task type of the target task, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type; concatenate the target prompt information with the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the text to be processed and the target prompt information; determine the task processing result of the target task based on the information extraction matrix; and send the task processing result to the client 100;
[0089] The client 100 is also used to receive the task processing result sent by the server 200.
[0090] By applying the solution of the embodiments of this specification, by constructing target prompt information based on the task type, the target prompt information can be used to prompt the information extraction model of the text processing task type, so that the information extraction model can handle multiple types of text processing tasks, thereby improving the scalability of the information extraction model and the versatility of task processing. Moreover, since the task processing results are derived from the task description information and the text to be processed, the accuracy and stability of the task processing results are guaranteed.
[0091] Referring to FIG. 2 , FIG. 2 shows an architecture diagram of another task processing system provided in accordance with one embodiment of this specification. The task processing system may include multiple clients 100 and a server 200. The clients 100 may include end-side devices, and the server 200 may include cloud-side devices. Multiple clients 100 may establish communication connections via the server 200. In a task processing scenario, the server 200 is used to provide task processing services between the multiple clients 100. The multiple clients 100 may act as either senders or receivers, communicating via the server 200.
[0092] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100. In the task processing scenario, users can publish data streams to the server 200 through the client 100. The server 200 generates task processing results based on the data stream and pushes the task processing results to other clients with which communication has been established.
[0093] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.
[0094] The client 100 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The client 100 can be based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device to run or certain APPs in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0095] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that provide background training to support models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server that is integrated with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0096] It is worth noting that the task processing methods provided in the embodiments of this specification are generally executed by the server. However, in other embodiments of this specification, the client may also have similar functions to the server and thus execute the task processing methods provided in the embodiments of this specification. In other embodiments, the task processing methods provided in the embodiments of this specification may also be executed jointly by the client and the server.
[0097] Referring to FIG3 , FIG3 shows a flowchart of a task processing method provided by an embodiment of this specification, which specifically includes the following steps:
[0098] Step 302: Obtain the to-be-processed text and task description information of the target task.
[0099] In one or more embodiments of this specification, during task processing, the to-be-processed text and task description information of the target task may be obtained, so as to process the to-be-processed text based on the task description information and obtain a task processing result.
[0100] Specifically, the target task can be a task in different scenarios, such as a document processing task in a document self-learning scenario, a text dialogue task in a dialogue scenario, and so on. The language type and length of the text to be processed are set according to the actual situation. For example, the text to be processed can be long or short, and can be in English or Chinese. The task description information is used to tell the information extraction model the category information to be extracted and the structure of the category information. The task description information includes but is not limited to the task type and the category information to be extracted, where the category information to be extracted can be understood as the category label to be extracted. The task description information can be in text format or in JSON format. For example, if the target task is a text classification task, the task description information may include the category information to be extracted {"culture": None, "entertainment": None, "sports": None, "finance": None}.
[0101] In actual applications, there are many ways to obtain the to-be-processed text and task description information of the target task, which can be selected based on actual conditions. The embodiments of this specification do not limit this task. In one possible implementation of this specification, the to-be-processed text and task description information of the target task sent by the user through the client can be received. In another possible implementation of this specification, the to-be-processed text and task description information of the target task can be read from other data acquisition devices or databases. The text to be processed can be obtained by performing optical character recognition (OCR) on an image, or by performing audio-to-text conversion on audio data.
[0102] Step 304: extracting the category information to be extracted from the task description information according to the task type of the target task, and constructing target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type.
[0103] In one or more embodiments of the present specification, after obtaining the to-be-processed text and task description information of the target task, further, the category information to be extracted can be extracted from the task description information according to the task type of the target task, and the target prompt information can be constructed according to the preset prompt format corresponding to the category information to be extracted and the task type.
[0104] Specifically, the task types of the target task include but are not limited to extraction tasks, classification tasks, non-hierarchical tasks, and multi-level tasks. Extraction tasks include but are not limited to relationship extraction tasks, event extraction tasks, and attribute sentiment extraction tasks. Classification tasks include but are not limited to sentiment classification tasks and industry classification tasks. Non-hierarchical tasks are ordinary classification tasks, such as sentiment classification tasks. Multi-level classification tasks refer to tasks with multiple label levels and multiple category labels in each label level, such as hierarchical classification tasks. The category information to be extracted refers to the category label corresponding to the target task. There can be multiple category information to be extracted. For example, if the target task is an attribute sentiment extraction task, the corresponding two category information to be extracted are sentiment classification labels {"positive": None, "negative": None}. Target prompt information (ESI, Explicit Schema Instructor) is a prompt used to prompt the information processing model of the current target task.
[0105] The preset prompt formats include but are not limited to the category prompt [TYPE], the result prompt [PREFIX], and the classification identifier used only in the classification task, where the classification identifier includes the single-label classification identifier [CLASSIFY] and the multi-label classification identifier [MULTICLASSIFY].
[0106] [PREFIX] is used to identify the task processing results of the previous level in a multi-level task. For example, in a hierarchical classification task, the current level is the third level, that is, the current round is the third round. The target prompt information for the current level is "[PREFIX] Level 1 result - Level 2 result [TYPE] Level 3 category information to be extracted 1 [TYPE] Level 3 category information to be extracted 2..."
[0107] [TYPE] is used to identify each category of information to be extracted. When the information extraction model is processed, it is also decoded through the value corresponding to each [TYPE] and the text to be processed.
[0108] [CLASSIFY] and [MULTICLASSIFY] are used to identify the target task as a classification task. This not only tells the information extraction model that the task is classification, not extraction, but also establishes a connection between the classification identifier and the category information to be extracted. While the overall framework of the information extraction model is extraction-based, unlike extraction tasks, which have corresponding segments to be extracted in the text to be processed, classification tasks require an additional extractable field in the text to be connected to the preceding [TYPE] and decoded. Specifically, in extraction tasks, [TYPE] is used to obtain the category information to be extracted, such as names of people and places, and then connected to the corresponding segments in the text to be processed, such as "Xiao Wang" and "Hangzhou." In classification tasks, [TYPE] is also used to obtain the category information to be extracted, such as the category "positive sentiment" in sentiment classification tasks, and is then connected to [CLASSIFY] or [MULTICLASSIFY] to convert the classification task into an extraction task, thus enabling the use of the information extraction model to implement both classification and extraction tasks. Among them, [CLASSIFY] refers to the single-label classification task, that is, the information extraction model will only output the category information with the highest probability; [MULTICLASSIFY] refers to the multi-label classification task, and the information extraction model will output the category information whose probability exceeds the preset threshold.
[0109] It should be noted that, based on the task type of the target task, the task type of the target task can be determined before extracting the category information to be extracted from the task description information. There are many ways to determine the task type of the target task, and the specific selection is made according to the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, the task description information includes the task type, and the task type of the target task is obtained by parsing the task description information. In another possible implementation of this specification, if the task description information does not include the task type, the task description information can be input into a type recognition model to obtain the task type of the target task. For example, if the task description information is "Please extract the character information below", the task description information is input into the information recognition model to determine that the task type is an extraction task.
[0110] In actual applications, when extracting the category information to be extracted from the task description information according to the task type of the target task, it is possible to determine whether the target task is a multi-level task based on the task type. If the target task is a non-hierarchical task, the category information to be extracted can be directly extracted from the task description information; if the target task is a multi-level task, the category information to be extracted can be extracted from the task description information based on the task processing results of the previous level of the current level.
[0111] In an optional embodiment of the present specification, the target task includes multi-level tasks; and the above-mentioned step of extracting the category information to be extracted from the task description information according to the task type of the target task may include the following steps:
[0112] According to the task type of the multi-level task, extract the category information to be extracted at the current level from the task description information. If the current level is not the first level, the category information to be extracted at the current level is obtained based on the task processing results of the previous level of the current level.
[0113] Construct target prompt information based on the preset prompt format corresponding to the category information to be extracted and the task type, including:
[0114] Target prompt information is constructed based on the category information to be extracted, the preset prompt format corresponding to the task type, and the task processing results of the previous level of the current level.
[0115] It should be noted that, assuming that the task description information is {"Result":{"Winning Bid Announcement":None, "Transaction Announcement":None}, "Notice":{"Prequalification":None, "Opinion Collection":None}, it can be seen that the target task is a multi-level task, and its labels are "Result-Winning Bid Announcement", "Result-Transaction Announcement", "Notice-Prequalification", and "Notice-Opinion Collection". In the embodiment of this specification, in order to improve the processing efficiency and prediction accuracy of the information extraction model, it is hoped that the information extraction model will first predict the labels of the first label level, and then continue to predict the labels of the second label level after obtaining the prediction results of the first label level. Specifically, when the current level is the first level, the labels "result" and "forecast" of the first label level are extracted from the task description information as the category information to be extracted at the first level; assuming that the prediction result of the information extraction model for the first label level is "result", when the current level is the second level, prediction can only be made among the labels "winning bid announcement" and "transaction announcement" of the second label level corresponding to "result". Therefore, the category information to be extracted at the second level is "winning bid announcement" and "transaction announcement".
[0116] Furthermore, after determining that the first-level category information to be extracted is "result" and "preview," the preset prompt format corresponding to the task type can be obtained as "[PREFIX] previous-level task processing result [TYPE] category information to be extracted [TYPE] category information to be extracted ... [TYPE] category information to be extracted." Since the first level does not have the previous-level task processing result, the target prompt information corresponding to the first level is "[PREFIX] [TYPE] result [TYPE] preview." Based on the first-level target prompt information and the to-be-processed text, the first-level task processing result is determined to be "result," and after determining that the second-level category information to be extracted is "bid winning announcement" and "transaction announcement," the second-level target prompt information can be constructed as "[PREFIX] result [TYPE] bid winning announcement [TYPE] transaction announcement" based on the second-level category information to be extracted, the preset prompt format corresponding to the task type, and the first-level task processing result.
[0117] Applying the solution of the embodiment of this specification, the category information to be extracted at the current level is extracted from the task description information according to the task type of the multi-level task, wherein, when the current level is not the first level, the category information to be extracted at the current level is obtained based on the task processing results of the previous level of the current level; and target prompt information is constructed according to the category information to be extracted, the preset prompt format corresponding to the task type, and the task processing results of the previous level of the current level. When constructing the target prompt information of the second level, the prior information obtained at the first level is fully utilized, thereby improving the prediction accuracy of the information extraction model, and by limiting the candidate category information, the information extraction model is prevented from predicting non-existent labels, such as "result-prequalification", thereby improving the processing efficiency of the information extraction model.
[0118] In an optional embodiment of the present specification, the target task includes an extraction task; and constructing the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type may include the following steps:
[0119] Obtaining a preset prompt format corresponding to the extraction task, wherein the preset prompt format includes a category prompt;
[0120] The spelling receptionist extracts category information and category prompts to obtain target prompt information.
[0121] It should be noted that, since the target task is an extraction task, the preset prompt format corresponding to the extraction task includes a category prompt [TYPE]. Therefore, the target prompt information can be obtained by concatenating the category information to be extracted and the category prompt.
[0122] For example, assuming the categories to be extracted are "person," "place," and "organization," the target prompt is "[PREFIX][TYPE]person[TYPE]place[TYPE]organization" after the extracted categories and prompts. Since this task is a non-hierarchical classification task, the value after [PREFIX] is empty, and [PREFIX] does not affect the output of the information extraction model.
[0123] Furthermore, since the extraction task may be a multi-level extraction task, such as relationship extraction, which includes relationships such as "person-age", "person-position", and "person-nationality", the first level can construct the "[PREFIX][TYPE]person" prompt information to extract the "person" entity in the text to be processed. After the first-level extraction is completed and the "person" entity (for example, "Xiao Ming") is obtained, the second-level prompt information "[PREFIX]person: Xiao Ming [TYPE]age [TYPE]position [TYPE]nationality" can be constructed to further extract the text to be processed.
[0124] Applying the solution of the embodiments of this specification, a preset prompt format corresponding to an extraction task is obtained, wherein the preset prompt format includes a category prompt. The target category information and the category prompt are combined to obtain the target prompt information. The category prompt provides candidate labels to the information extraction model for prediction, thereby ensuring the stability of the information extraction model.
[0125] In another optional embodiment of the present specification, the target task includes a classification task; and constructing the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type may include the following steps:
[0126] Obtaining a preset prompt format corresponding to the classification task, wherein the preset prompt format includes a category prompt and a category identifier;
[0127] The spelling reception device extracts category information, category prompts and category identifiers to obtain target prompt information.
[0128] It should be noted that since the target task is a classification task, the preset prompt format corresponding to the classification task includes the category prompt [TYPE] and the classification identifier [CLASSIFY] or [MULTICLASSIFY]. Therefore, the target prompt information can be obtained by splicing the category information to be extracted, the category prompt and the classification identifier.
[0129] For example, assuming the classification information to be extracted is "positive sentiment" and "negative sentiment," and the classification task is a single-label classification task, then the category information to be extracted, the category prompt, and the category identifier are obtained, and the target prompt information is "[PREFIX][TYPE]Positive Sentiment[TYPE]Negative Sentiment[CLASSIFY]". Since this task is a non-hierarchical classification task, the value after [PREFIX] is empty, and [PREFIX] does not affect the output of the information extraction model.
[0130] Furthermore, since the classification task may be a multi-level classification task, for example, the hierarchical classification labels are "Result - Winning Bid Announcement", "Result - Transaction Announcement", "Preview - Prequalification", and "Preview - Opinion Collection". The first level constructs the prompt information "[PREFIX][TYPE]Result[TYPE]Preview" to first classify whether the processed text belongs to "Result" or "Preview". After obtaining the first-level classification result (for example, "Result"), the second-level prompt information "[PREFIX]Result[TYPE]Winning Bid Announcement[TYPE]Transaction Announcement" can be constructed to further extract the processed text.
[0131] By applying the solution of the embodiments of this specification, a preset prompt format corresponding to a classification task is obtained, wherein the preset prompt format includes a category prompt and a category identifier; the category information, category prompt, and category identifier are then extracted to obtain the target prompt information. The category prompt provides candidate labels to the information extraction model, ensuring the stability of the information extraction model. Furthermore, by converting the classification task into an extraction task through [CLASSIFY] or [MULTICLASSIFY], a single information extraction model is used to implement both the classification and extraction tasks, improving the scalability of the information extraction model and the versatility of task processing.
[0132] Step 306: Concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the text to be processed and the target prompt information.
[0133] In one or more embodiments of the present specification, the to-be-processed text and task description information of the target task are obtained; based on the task type of the target task, the category information to be extracted is extracted from the task description information, and after the target prompt information is constructed according to the preset prompt format corresponding to the category information to be extracted and the task type, the target prompt information and the to-be-processed text can be further spliced, and the spliced information can be input into the information extraction model to obtain an information extraction matrix.
[0134] Specifically, the information extraction model is trained based on sample texts from multiple sample tasks, sample task descriptions, and sample matrix labels carried by the sample texts. The sample matrix labels are derived based on the sample processing results of the sample texts. The information extraction matrix is an N*N score matrix, where N is the length of the input concatenated information.
[0135] It should be noted that when splicing the target prompt information and the text to be processed, the target prompt information and the text to be processed can be spliced together in order to obtain the spliced information. Among them, the order of the target prompt information and the text to be processed is selected according to the actual situation, and the embodiments of this specification do not impose any restrictions on this. Furthermore, in order to distinguish the target prompt information from the text to be processed, when splicing the target prompt information and the text to be processed, the information distinguisher [Text] can be added, and the spliced information is "target prompt information + [Text] + text to be processed".
[0136] In an optional embodiment of the present specification, the information extraction model includes an encoding unit, a hint separation unit, and a feature processing unit; the above-mentioned inputting the spliced information into the information extraction model to obtain the information extraction matrix may include the following steps:
[0137] The spliced information is input into the information extraction model, and the target prompt information is separated by position through the prompt separation unit to obtain the prompt separation matrix;
[0138] The hint separation matrix and the concatenated information are encoded by the encoding unit to obtain an encoding feature matrix;
[0139] The feature processing unit fuses the prompt separation matrix and the encoding feature matrix to obtain the information extraction matrix.
[0140] It should be noted that the prompt separation unit is used to generate a prompt separation matrix, and the prompt separation matrix is used to mark the information that does not need to be extracted by the information extraction model as zero. In the embodiment of this specification, since the information extraction model does not extract from the category identifier and the result identifier in the target prompt information when extracting the task processing result, the spliced information can be input into the information extraction model, and the target prompt information is positionally separated by the prompt separation unit to obtain the prompt separation matrix, and the prompt separation matrix is input into the encoding unit so that the encoding unit can quickly encode the spliced information to obtain the encoding feature matrix. Furthermore, after obtaining the encoding feature matrix, the prompt separation matrix and the encoding feature matrix can be multiplied by the feature processing unit to obtain the information extraction matrix. At this time, the positions corresponding to the category identifier and the result identifier in the information extraction matrix are all zero.
[0141] Referring to FIG4 , FIG4 shows a framework diagram of an information extraction model in a task processing method provided by an embodiment of the present specification. As shown in FIG4 , during the i-th round of information extraction, the start symbol [CLS], the target prompt information ESIi, the information distinguisher [Text], and the text to be processed are spliced together, and the spliced information Qi is input into the information extraction model. In the information extraction model, the target prompt information is positionally separated by the prompt separation unit to obtain a prompt separation matrix; the prompt separation matrix and the spliced information are encoded by the encoding unit to obtain an encoding feature matrix; the encoding feature matrix is processed by the feedforward layer in the feature processing unit, and the prompt separation matrix and the processed encoding feature matrix are multiplied by the fusion layer in the feature processing unit to obtain an information extraction matrix. The information extraction matrix is decoded based on the encoding mechanism including the Linking mechanism and the "handshake mechanism" to generate the task processing result Yi of the i-th round.
[0142] After the i-th round of information extraction is complete, the start symbol [CLS], the target prompt information ESIi+1, the information distinguisher [Text], and the text to be processed are concatenated in the same manner as in the i-th round. The concatenated information Qi+1 is then fed into the information extraction model for the i+1-th round of information extraction. Finally, based on the task processing results ...Yi-1, Yi, Yi+1..., the target task processing result is output.
[0143] It is worth noting that a triple loop structure can be adopted in the process of constructing target prompt information:
[0144] The first recursive loop: loops through the schema hierarchy to determine the category information to be extracted. For example, hierarchical classification, relationship extraction, and event extraction require multiple loops at this layer.
[0145] Second recursive loop: ensure that the length of the target prompt information is less than or equal to the preset prompt length threshold. If the length of the target prompt information is greater than the preset prompt length threshold, split the target prompt information to obtain multiple target sub-prompt information;
[0146] The third recursive loop: The second recursive loop ensures the length of the text to be processed, preventing it from being too short and affecting the information extraction effect. However, if the sum of the length of the target prompt information and the length of the text to be processed exceeds 512, the text to be processed is split into multiple sub-texts based on the category information to be extracted.
[0147] The solutions implemented in the embodiments of this specification utilize a loop structure, enabling the information extraction model to effectively handle large-scale and hierarchical classification tasks. The design of the linking mechanism and the "handshake mechanism" also make the information extraction model highly scalable, enabling it to handle not only classification tasks but also extraction tasks. Furthermore, because the model utilizes an extraction framework, task processing results are derived from both the task description information and the text to be processed, ensuring the accuracy and stability of the results.
[0148] Step 308: Determine the task processing result of the target task according to the information extraction matrix.
[0149] In one or more embodiments of the present specification, the to-be-processed text and task description information of the target task are obtained; based on the task type of the target task, the category information to be extracted is extracted from the task description information, and the target prompt information is constructed according to the preset prompt format corresponding to the category information to be extracted and the task type; the target prompt information and the to-be-processed text are spliced together, and the spliced information is input into the information extraction model to obtain the information extraction matrix. Further, based on the information extraction matrix, the task processing result of the target task can be determined.
[0150] Specifically, if the target task is an extraction task, the task processing result is the extracted field name and the corresponding field fragment; if the target task is a classification task, the task processing result is the classification label.
[0151] By applying the solution of the embodiments of this specification, by constructing target prompt information based on the task type, the target prompt information can be used to prompt the information extraction model of the text processing task type, so that the information extraction model can handle multiple types of text processing tasks, thereby improving the scalability of the information extraction model and the versatility of task processing. Moreover, since the task processing results are derived from the task description information and the text to be processed, the accuracy and stability of the task processing results are guaranteed.
[0152] In an optional embodiment of the present specification, the above-mentioned determination of the task processing result of the target task based on the information extraction matrix may include the following steps:
[0153] Transforming the information extraction matrix to obtain a target information extraction matrix, wherein the target information extraction matrix represents the information extraction probability of the text to be processed;
[0154] According to the target information extraction matrix, the task processing result of the target task is determined.
[0155] In practical applications, there are many ways to transform the information extraction matrix to obtain the target information extraction matrix, and the specific method is selected according to the actual situation. The embodiments of this specification do not impose any limitation on this.
[0156] In one possible implementation of this specification, a matrix element threshold can be obtained, and the information extraction matrix can be transformed based on the matrix element threshold to generate a 0-1 information extraction matrix. Specifically, the matrix element threshold can be used to determine the size relationship of the matrix element values at each position in the information extraction matrix. If the matrix element value is greater than the matrix element threshold, the corresponding position is determined to be 1; if the matrix element value is less than or equal to the matrix element threshold, the corresponding position is determined to be 0, thereby obtaining a 0-1 information extraction matrix. After obtaining the 0-1 information extraction matrix, the task processing result of the target task can be determined based on the 0-1 information extraction matrix based on the Linking mechanism.
[0157] In another possible implementation of this specification, since all positions may be 0 or all positions may be 1 in the 0-1 information extraction matrix, the 0-1 information extraction matrix may be invalid and cannot construct a corresponding relationship between the text to be processed and the target prompt information, resulting in the information extraction model not returning a classification label and being unable to determine the task processing result of the target task. Therefore, the information extraction matrix can be input into the activation layer (sigmod layer) for conversion, and the probability information extraction matrix is obtained after processing by the activation layer, wherein the matrix elements in the probability information extraction matrix represent a probability value. After obtaining the probability information extraction matrix, the task processing result of the target task can be determined based on the handshake mechanism and the probability information extraction matrix.
[0158] It should be noted that, based on the handshake mechanism, when determining the task processing result of the target task according to the probability information extraction matrix, since the classification identifiers [CLASSIFY] and [MULTICLASSIFY] appear fixedly before the text to be processed, and the position of the category identifier [TYPE] is known information, it is easy to determine the probability value of the category identifier [TYPE] of each category information to be extracted mapped to the classification identifier, that is, the probability value of the "[CLASSIFY]-[TYPE] pair" or the "[MULTICLASSIFY]-[TYPE] pair". For single-label classification tasks, the category information to be extracted corresponding to the "[CLASSIFY]-[TYPE] pair" with the largest average probability value in the probability information extraction matrix can be used as the classification result; for multi-label classification tasks, the category information to be extracted corresponding to the "[MULTICLASSIFY]-[TYPE] pair" with the largest average probability value in the probability information extraction matrix, and the category information to be extracted corresponding to the "[MULTICLASSIFY]-[TYPE] pair" with an average probability value above the extraction probability threshold (such as 0.9) can be used as the classification result.
[0159] By applying the solution of the embodiment of this specification, the information extraction matrix is transformed to obtain a 0-1 information extraction matrix or a probability information extraction matrix; the task processing result of the target task is further determined based on the Linking mechanism or the handshake mechanism, thereby improving the task processing efficiency and accuracy.
[0160] In an optional embodiment of the present specification, a text matrix can be constructed based on the text to be processed and the target prompt information, so that the target extraction result can be extracted from the text matrix based on the valid values in the information extraction matrix. That is, the above-mentioned determination of the task processing result of the target task based on the information extraction matrix can include the following steps:
[0161] Construct a text matrix based on the text to be processed and the target prompt information;
[0162] According to the information extraction matrix, the task processing results of the target task are extracted from the text matrix.
[0163] It should be noted that, when constructing a text matrix based on the text to be processed and the target prompt information, the text to be processed and the target prompt information can be spliced together, and the spliced text information is used as the rows and columns of the text matrix respectively to obtain the text matrix.
[0164] For example, assume the text to be extracted is the English text "Today is a sunny day, I like it" and the target prompt information is "[PREFIX][TYPE]good[TYPE]bad[CLASSIFY]". Concatenate the two to obtain "[PREFIX][TYPE]good[TYPE]bad[Text][CLASSIFY]Today is a sunny day, I like it". Further, use "[PREFIX][TYPE]good[TYPE]bad[Text][CLASSIFY]Today is a sunny day, I like it" as the rows and columns of the text matrix to obtain a 16*16 text matrix. Based on the information extraction matrix, the task processing result of the target task is extracted from the text matrix.
[0165] In an optional embodiment of the present specification, before extracting the task processing result of the target task from the text matrix according to the information extraction matrix, the text matrix and the information extraction matrix can be aligned, so as to accurately extract the task processing result of the target task from the text matrix according to the information extraction matrix.
[0166] By applying the solution of the embodiment of this specification, a text matrix is constructed based on the text to be processed and the target prompt information; based on the information extraction matrix, the task processing results of the target task are extracted from the text matrix. Since the task processing results are derived from the task description information and the text to be processed, the accuracy and stability of the task processing results are guaranteed.
[0167] In practical applications, since the category information to be extracted may include a large number of labels, if all the category information to be extracted is spliced together to obtain the target prompt information, the length of the target prompt information may be too long. The maximum input length of the information extraction model is generally only 512, and the target prompt information and the text to be extracted will share the maximum input length of the model. When the target prompt information is too long, the available length of the text to be extracted will be compressed to a very short length or even to zero, which will greatly affect the model effect of the information extraction model. Therefore, a preset prompt length threshold can be introduced to limit the length of the target prompt information to ensure that the text to be processed has a reasonable length to provide sufficient extraction basis for the information extraction model. That is, after constructing the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type, the following steps can also be included:
[0168] When the length of the target prompt information is greater than a preset prompt length threshold, splitting the target prompt information to obtain multiple target sub-prompt information, wherein the lengths of the multiple target sub-prompt information are less than or equal to the preset prompt length threshold;
[0169] The target prompt information and the text to be processed are concatenated, and the concatenated information is input into the information extraction model to obtain the information extraction matrix. The following steps may be included:
[0170] Extracting first target sub-prompt information from the plurality of target sub-prompt information, wherein the first target sub-prompt information is the target sub-prompt information that is not spliced with the to-be-processed text among the plurality of target sub-prompt information;
[0171] The first target sub-prompt information and the text to be processed are spliced together, and the spliced information is input into the information extraction model until the multiple target sub-prompt information do not include target sub-prompt information that is not spliced with the text to be processed, thereby obtaining an information extraction matrix.
[0172] Specifically, the preset prompt length threshold is set according to actual conditions, such as 256. The embodiment of this specification does not impose any limitation on the preset prompt length threshold.
[0173] It should be noted that after splicing the first target sub-prompt information and the text to be processed and inputting the spliced information into the information extraction model, the first information extraction matrix can be obtained. At this time, the step of extracting the first target sub-prompt information from multiple target sub-prompt information can be returned to obtain the first information extraction matrix until the multiple target sub-prompt information does not include the target sub-prompt information that is not spliced with the text to be processed. According to the multiple first information extraction matrices, the information extraction matrix is generated.
[0174] In an optional embodiment of the present specification, if the length of the target prompt information is less than or equal to a preset prompt length threshold, such as if the length of the target prompt information is 30, the length of the text to be processed is 482. In this case, the target prompt information and the text to be processed can be directly spliced together. If the length of the target prompt information is greater than the preset prompt length threshold, such as if the length of the target prompt information is 300, the target prompt information can be split to obtain target sub-prompt information 1 with a length of 250 and target sub-prompt information 2 with a length of 50. Furthermore, the target sub-prompt information 1 and the text to be processed can be spliced together to obtain spliced information 1, and the target sub-prompt information 2 and the text to be processed can be spliced together to obtain spliced information 2. Then, the spliced information 1 is processed using the information extraction model to obtain information extraction matrix 1, and the spliced information 2 is processed to obtain information extraction matrix 2. Finally, the information extraction matrix 1 and the information extraction matrix 2 are summarized to obtain the information extraction matrix.
[0175] By applying the solution of the embodiment of this specification, when the length of the target prompt information is greater than the preset prompt length threshold, the target prompt information is split to obtain multiple target sub-prompt information, thereby ensuring that the length of the prompt information spliced with the text to be processed is less than or equal to the preset prompt length threshold, leaving sufficient input space for the text to be processed, and ensuring the extraction effect of the information extraction model.
[0176] In actual applications, there are many ways to split the target prompt information to obtain multiple target sub-prompt information. The specific method is selected according to the actual situation. The embodiments of this specification do not impose any restrictions on this. In one possible implementation of this specification, when the length of the target prompt information is greater than a preset prompt length threshold, the target prompt information can be directly split at the position of the preset prompt length threshold to obtain multiple target sub-prompt information.
[0177] In another possible implementation of the present specification, since directly truncating the target prompt information at the position of the preset prompt length threshold may result in incomplete category information to be extracted, causing errors in the information extraction model, the above-mentioned splitting of the target prompt information to obtain multiple target sub-prompt information when the length of the target prompt information is greater than the preset prompt length threshold may include the following steps:
[0178] When the length of the target prompt information is greater than a preset prompt length threshold, the target prompt information is split based on the category information to be extracted to obtain a plurality of target sub-prompt information.
[0179] By applying the solution of the embodiment of this specification, the target prompt information is split into units of category information to be extracted to obtain multiple target sub-prompt information, ensuring that each category information to be extracted in each target sub-prompt information is complete, thereby ensuring the accuracy and stability of the information extraction model.
[0180] In another optional embodiment of the present specification, similar to the case where the length of the target prompt information may be greater than the preset prompt length threshold, the length of the text to be processed may also be very long, exceeding the maximum input sequence length of the information extraction model. Therefore, a preset text length threshold may be introduced to limit the length of the text to be processed. That is, before the above-mentioned concatenation of the target prompt information and the text to be processed and inputting the concatenated information into the information extraction model to obtain the information extraction matrix, the following steps may be further included:
[0181] When the length of the text to be processed is greater than a preset text length threshold, the text to be processed is split to obtain multiple sub-texts to be processed, wherein the lengths of the multiple sub-texts to be processed are less than or equal to the preset text length threshold;
[0182] The target prompt information and the text to be processed are concatenated, and the concatenated information is input into the information extraction model to obtain the information extraction matrix. The following steps may be included:
[0183] Extracting a first subtext to be processed from the plurality of subtexts to be processed, wherein the first subtext to be processed is a subtext to be processed that is not spliced with the target prompt information among the plurality of subtexts to be processed;
[0184] The first subtext to be processed and the target prompt information are spliced together, and the spliced information is input into the information extraction model until the plurality of subtexts to be processed do not include the subtext to be processed that is not spliced with the target prompt information, thereby obtaining an information extraction matrix.
[0185] Specifically, the preset text length threshold is set according to actual conditions, and the embodiments of this specification do not impose any limitation on the preset prompt length threshold.
[0186] It should be noted that after splicing the first sub-text to be processed and the target prompt information and inputting the spliced information into the information extraction model, a second information extraction matrix can be obtained. At this time, the step of extracting the first sub-text to be processed from multiple sub-texts to be processed can be returned to execute until the multiple sub-texts to be processed do not include the sub-text to be processed that is not spliced with the target prompt information, and an information extraction matrix is generated based on multiple second information extraction matrices.
[0187] By applying the solution of the embodiment of this specification, when the length of the text to be processed is greater than the preset text length threshold, the text to be processed is split to obtain multiple sub-texts to be processed, thereby ensuring that the length of the sub-text to be processed spliced with the target prompt information is less than or equal to the preset text length threshold, thereby ensuring the extraction effect of the information extraction model.
[0188] It is worth noting that the target prompt information and the text to be processed can be split at the same time. Referring to the above example of splitting the target prompt information, after the target prompt information is divided into target sub-prompt information 1 and target sub-prompt information 2, the length of the text to be processed is guaranteed to be between 256 and 511, avoiding the text to be processed being too short in the spliced information, which affects the information extraction effect. In actual applications, the length of the target sub-prompt information and the text to be processed after splicing may still be greater than 512. Therefore, the text to be processed can also be split into sub-text to be processed 1 and sub-text to be processed 2, so that the information extraction model processes "target sub-prompt information 1 + sub-text to be processed 1, target sub-prompt information 1 + sub-text to be processed 2, target sub-prompt information 2 + sub-text to be processed 1, target sub-prompt information 2 + sub-text to be processed 2" in turn, and finally the processing results corresponding to the four spliced information are combined to obtain the final task processing result.
[0189] In an optional embodiment of the present specification, before the above-mentioned splicing target prompt information and the text to be processed and inputting the spliced information into the information extraction model to obtain the information extraction matrix, the following steps may also be included:
[0190] Obtaining a sample set, wherein the sample set includes sample texts and sample task description information of multiple sample tasks, the sample texts carry sample matrix labels, and the sample matrix labels are obtained based on sample processing results of the sample texts;
[0191] For any sample task, according to the sample task type of the sample task, the sample extraction category information is extracted from the sample task description information, and the sample prompt information is constructed according to the preset prompt format corresponding to the sample extraction category information and the sample task type;
[0192] Splicing the sample prompt information and the sample text, and inputting the spliced sample information into the information extraction model to obtain a prediction information extraction matrix, wherein the prediction information extraction matrix represents the correspondence between the sample text and the sample prompt information;
[0193] According to the sample matrix labels and the predicted information extraction matrix, the model parameters of the information extraction model are adjusted to obtain a trained information extraction model.
[0194] Specifically, the training method of the information extraction model is supervised training based on prompt learning, that is, each sample text in the sample set carries a real sample matrix label, and the sample matrix label is the extraction target of the information extraction model, which is used to guide the training process of the information extraction model. The method of obtaining the sample set can be to read a large number of sample texts carrying sample matrix labels and sample task description information from other data acquisition devices or databases to form a sample set. It can also be to receive a large number of sample texts carrying sample matrix labels and sample task description information input by the user to form a sample set. The method of obtaining the sample set is selected according to the actual situation, and the embodiments of this specification do not impose any restrictions on this.
[0195] It should be noted that the sample processing results are the actual processing results of the sample text. When generating the sample matrix label based on the sample processing results, the sample extraction category information can be extracted from the sample task description information according to the sample type of the sample task, and the sample prompt information can be constructed according to the preset prompt format corresponding to the sample extraction category information and the sample type. The sample matrix label is constructed based on the sample prompt information and the sample text. For example, an initial sample matrix containing all zeros is constructed based on the sample prompt information and the sample text, and the position corresponding to the sample processing result in the initial sample matrix is modified to one to obtain the sample matrix label.
[0196] It is worth noting that the implementation method of "extracting sample extraction category information from the sample task description information according to the sample task type of the sample task, and constructing sample prompt information according to the preset prompt format corresponding to the sample extraction category information and the sample task type; splicing the sample prompt information and the sample text, and inputting the spliced sample information into the information extraction model to obtain a prediction information extraction matrix" is the same as the implementation method of the above-mentioned "extracting category information to be extracted from the task description information according to the task type of the target task, and constructing target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type; splicing the target prompt information and the text to be processed, and inputting the spliced information into the information extraction model to obtain an information extraction matrix", and the embodiments of this specification will not be repeated. During the prediction process, the prediction information extraction matrix can be transformed to obtain a prediction target information extraction matrix. The implementation method of "transforming the prediction information extraction matrix to obtain a prediction target information extraction matrix" is the same as the implementation method of "transforming the information extraction matrix to obtain a target information extraction matrix" mentioned above, and the embodiments of this specification will not be repeated.
[0197] In practical applications, when adjusting the model parameters of the information extraction model according to the sample matrix labels and the predicted information extraction matrix, the loss value can be calculated according to the sample matrix labels and the predicted information extraction matrix, and the model parameters of the information extraction model can be adjusted based on the loss value until the preset stopping condition is reached to obtain a trained information extraction model, wherein the preset stopping condition includes but is not limited to the loss value being less than or equal to the preset threshold and the number of iterations reaching the preset number of iterations. Specifically, the loss value can be calculated using the following formula (1):
[0198] Among them, L represents the loss value, i represents the i-th sample, and j represents the matrix zg i and z i The jth position in the matrix zg is represented by k. i and z i In the kth position, e represents the exponential function, zg i represents the sample matrix label, z i Represents the prediction information extraction matrix.
[0199] Using the solutions of the embodiments of this specification, a loss value is calculated based on the sample matrix labels and the predicted information extraction matrix. This loss value is then compared with a preset stopping condition. If the preset stopping condition is not met, the information extraction model is trained continuously until the preset stopping condition is met, completing the training and obtaining the information extraction model. By continuously adjusting the model parameters of the information extraction model, the resulting information extraction model can be made more accurate.
[0200] Referring to FIG5 , FIG5 shows a flowchart of an information extraction model training method provided by one embodiment of this specification. The information extraction model training method is applied to a cloud-side device and specifically includes the following steps:
[0201] Step 502: Obtain a sample set, wherein the sample set includes sample texts and sample task description information of multiple sample tasks, and the sample texts carry sample matrix labels, which are obtained based on sample processing results of the sample texts.
[0202] Step 504: For any sample task, according to the sample task type of the sample task, sample extraction category information is extracted from the sample task description information, and sample prompt information is constructed according to the preset prompt format corresponding to the sample extraction category information and the sample task type.
[0203] Step 506: Concatenate the sample prompt information and the sample text, and input the concatenated sample information into the information extraction model to obtain a prediction information extraction matrix, wherein the prediction information extraction matrix represents the correspondence between the sample text and the sample prompt information.
[0204] Step 508: According to the sample matrix labels and the predicted information extraction matrix, the model parameters of the information extraction model are adjusted to obtain a trained information extraction model.
[0205] It should be noted that the implementation of steps 502 to 508 is detailed in the training method of the information extraction model in the above-mentioned task processing method, and the embodiments of this specification do not impose any limitation on this.
[0206] In actual applications, after obtaining the trained information extraction model, the model parameters of the trained information extraction model can be sent to the end-side device, so that the user can build the information extraction model locally based on the model parameters to complete the information extraction task.
[0207] By applying the solution of the embodiment of this specification, the model parameters of the information extraction model are adjusted according to the sample matrix labels and the predicted information extraction matrix to obtain a trained information extraction model. By continuously adjusting the model parameters of the information extraction model, the final information extraction model can be made more accurate.
[0208] The following further illustrates the task processing method provided in this specification using the application of the task processing method in a text classification scenario as an example, in conjunction with FIG6 . FIG6 shows a flowchart of a classification task processing method provided in one embodiment of this specification, specifically comprising the following steps:
[0209] Step 602: Obtain the to-be-processed text and task description information of the target classification task.
[0210] Step 604: extracting the category information to be extracted from the task description information according to the task type of the target classification task, and constructing target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type.
[0211] Step 606: Concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the text to be processed and the target prompt information.
[0212] Step 608: Determine the classification result of the target classification task according to the information extraction matrix.
[0213] It should be noted that the implementation of steps 602 to 608 is the same as the implementation of steps 302 to 308 described above, and the embodiments of this specification do not impose any limitation on this.
[0214] By applying the solution of the embodiments of this specification, by constructing target prompt information based on the task type, the target prompt information can be used to prompt the information extraction model that the task is a classification task, so that the information extraction model can handle the classification task, thereby improving the scalability of the information extraction model and the versatility of task processing. Moreover, since the classification results are derived from the task description information and the text to be processed, the accuracy and stability of the classification results are guaranteed.
[0215] Refer to Figure 7a, which shows a processing flow chart of the first task processing method provided by an embodiment of the present specification. Taking the single-label classification task as an example, assuming that the text to be extracted is "Today is a sunny day, I like it", and the target prompt information is "[PREFIX][TYPE]good[TYPE]bad[CLASSIFY]", the target prompt information and the text to be processed are spliced together, and the spliced information "[PREFIX][TYPE]good[TYPE]bad[Text][CLASSIFY]Today is a sunny day, I like it" is input into the information extraction model to obtain a 0-1 information extraction matrix as shown in Figure 7a, where [P] is [PREFIX], [T] is [TYPE], and [CLST] is [CLASSIFY].
[0216] In Figure 7a, region 1 indicates that [TYPE] is connected to the head of the corresponding segment in the text to be processed, i.e., the type-token-head type; region 2 indicates that [TYPE] is connected to the tail of the corresponding segment in the text to be processed, i.e., the type-token-tail type; and region 3 indicates that it is connected to both the head and tail of the corresponding segment in the text to be processed, i.e., the token-head-tail type. Since a single token is used here to mark the category, both the head and tail point to the same location, i.e., [CLST]. As shown in Figure 7a, since [CLST] points to the identifier [T] belonging to the good label (the portion with a value of 1 in Figure 7a), the linking mechanism can decode the sentiment label corresponding to "Today is a sunny day, I like it" as "good". Furthermore, due to the high scalability of the linking mechanism, this framework can be used for extraction tasks without the [CLST] that marks the category.
[0217] Refer to Figure 7b, which shows a processing flow chart of the second task processing method provided by an embodiment of the present specification. Taking the multi-label classification task as an example, assuming that the text to be extracted is "Today is a sunny day, I like it", and the target prompt information is "[PREFIX][TYPE]weather[TYPE]mood[MULTICLASSIFY]", the target prompt information and the text to be processed are spliced to obtain "[PREFIX][TYPE]weather[TYPE]mood[Text][MULTICLASSIFY]Today is a sunny day, I like it", and the spliced information is input into the information extraction model to obtain a 0-1 information extraction matrix as shown in Figure 7b, where [P] is [PREFIX], [T] is [TYPE], and [CLST] is [MULTICLASSIFY]. Further, according to the 0-1 information extraction matrix, the Linking mechanism can be used to decode "Today is a sunny day, I like it". The topics corresponding to "it" are "weather" and "mood". For the description of area 1, area 2 and area 3 in Figure 7b, please refer to Figure 7a, and will not be repeated in this embodiment of the specification.
[0218] Refer to Figure 7c, which shows a processing flow chart of the third task processing method provided by an embodiment of the present specification. Referring to the example of Figure 7b, assuming that the 0-1 information extraction matrix is invalid and the model cannot return the classification label, at this time, the handshake mechanism shown in Figure 7c can be triggered, and the information extraction matrix is passed through a sigmod layer to obtain a probability information extraction matrix. According to the probability value of each "[MULTICLASSIFY]-[TYPE] pair" in the probability information extraction matrix, the topic corresponding to "Today is a sunny day, I like it" is determined to be "weather" and "mood", where each "[MULTICLASSIFY]-[TYPE] pair" is shown as the shaded squares and double arrows in Figure 7c. For the description of area 1, area 2 and area 3 in Figure 7c, please refer to Figure 7a, and the embodiment of this specification will not be repeated.
[0219] Refer to Figure 8, which shows a processing flow chart of the fourth task processing method provided by an embodiment of this specification. Taking the information extraction task as an example, assuming that the text to be extracted is "In 1997, AB returned to M as CEO", the target prompt information is "[PREFIX][TYPE]person[TYPE]organization", the target prompt information and the text to be processed are spliced together, and the spliced information "[PREFIX][TYPE]person[TYPE]organization[Text]In 1997, AB returned to M as CEO" is input into the information extraction model to obtain the information extraction matrix shown in Figure 8, where [P] is [PREFIX], [T] is [TYPE], person represents a person, and organization represents an organization.
[0220] As shown in Figure 8, since the first 1 in area 1 points to the identifier [T] to which the A and person tags belong, and the second 1 points to the identifier [T] to which the M and organization tags belong, the first 1 in area 2 points to the identifier [T] to which the B and person tags belong, and the second 1 points to the identifier [T] to which the M and organization tags belong, and the first 1 in area 3 points to A and B, and the second 1 points to M, it can be decoded through the Linking mechanism that person corresponds to AB and organization corresponds to M.
[0221] Referring to FIG9 , FIG9 shows a flowchart of the processing process of the fifth task processing method provided by one embodiment of this specification. Referring to the example of FIG8 , FIG9 illustrates how, in a relationship extraction task, after the extracted person is AB, the affiliation relationship between AB and the work organization M is extracted. Assuming that the text to be extracted is "In 1997, AB returned to M as CEO" and the target prompt information is "[PREFIX]person:AB[TYPE]work for(organization)", the target prompt information and the text to be processed are concatenated, and the concatenated information "[PREFIX]person:AB[TYPE]work for (organization)[Text]In 1997, AB returned to M as CEO" is input into the information extraction model. The information extraction matrix obtained is shown in FIG9 , where [P] is [PREFIX], [T] is [TYPE], person represents a person, and organization represents an organization.
[0222] As shown in Figure 9, since the 1 in area 1 points to the identifier [T] to which M and the work label belong, the 1 in area 2 points to the identifier [T] to which M and the work label belong, and the 1 in area 3 points to M, the Linking mechanism can be used to decode the affiliation between AB and the work organization M as "work".
[0223] By applying the solution of the embodiments of this specification, by constructing target prompt information based on the task type, the model can process multiple labels at one time in the extraction task, which improves the efficiency by about 30%. The target prompt information can also be used to prompt the information extraction model of the text processing task type, so that the information extraction model can process multiple types of text processing tasks, such as large-scale classification tasks, multi-label classification tasks and hierarchical classification tasks, thereby improving the scalability of the information extraction model and the versatility of task processing. Moreover, since the task processing results are derived from the task description information and the text to be processed, the accuracy and stability of the task processing results are guaranteed.
[0224] In practical applications, a 100 million-parameter base model and a 300 million-parameter large model were trained on high-quality supervised data covering a variety of task types, excluding generation tasks. This gave the information extraction model powerful zero-shot and few-shot learning capabilities. Furthermore, the information extraction model's F1 (F-measure) performance was significantly improved on 17 test datasets encompassing both extraction and classification tasks. Furthermore, both the base and large models demonstrated strong performance in reading comprehension and natural language inference tasks.
[0225] Corresponding to the above-mentioned task processing method embodiment, this specification also provides a task processing device embodiment. FIG10 shows a schematic diagram of the structure of a task processing device provided in one embodiment of this specification. As shown in FIG10 , the device includes:
[0226] The first acquisition module 1002 is configured to acquire the to-be-processed text and task description information of the target task;
[0227] The first constructing module 1004 is configured to extract the category information to be extracted from the task description information according to the task type of the target task, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type;
[0228] The first input module 1006 is configured to concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the text to be processed and the target prompt information;
[0229] The first generating module 1008 is configured to determine a task processing result of the target task according to the information extraction matrix.
[0230] Optionally, the target task includes multi-level tasks; the first construction module 1004 is further configured to extract the category information to be extracted at the current level from the task description information according to the task type of the multi-level task, wherein, when the current level is not the first level, the category information to be extracted at the current level is obtained based on the task processing result of the previous level of the current level; the target prompt information is constructed according to the category information to be extracted, the preset prompt format corresponding to the task type and the task processing result of the previous level of the current level.
[0231] Optionally, the device also includes: a first splitting module, configured to split the target prompt information when the length of the target prompt information is greater than a preset prompt length threshold, to obtain multiple target sub-prompt information, wherein the length of the multiple target sub-prompt information is less than or equal to the preset prompt length threshold; a first input module 1006, further configured to extract the first target sub-prompt information from the multiple target sub-prompt information, wherein the first target sub-prompt information is the target sub-prompt information in the multiple target sub-prompt information that is not spliced with the text to be processed; splicing the first target sub-prompt information and the text to be processed, and inputting the spliced information into the information extraction model until the multiple target sub-prompt information does not include the target sub-prompt information that is not spliced with the text to be processed, to obtain an information extraction matrix.
[0232] Optionally, the device also includes: a second splitting module, configured to split the text to be processed when the length of the text to be processed is greater than a preset text length threshold, to obtain multiple sub-texts to be processed, wherein the lengths of the multiple sub-texts to be processed are less than or equal to the preset text length threshold; a first input module 1006, further configured to extract a first sub-text to be processed from the multiple sub-texts to be processed, wherein the first sub-text to be processed is the sub-text to be processed that is not spliced with the target prompt information in the multiple sub-texts to be processed; splicing the first sub-text to be processed and the target prompt information, and inputting the spliced information into the information extraction model, until the multiple sub-texts to be processed do not include the sub-text to be processed that is not spliced with the target prompt information, to obtain an information extraction matrix.
[0233] Optionally, the target task includes an extraction task; the first construction module 1004 is further configured to obtain a preset prompt format corresponding to the extraction task, wherein the preset prompt format includes a category prompt; and assemble the extracted category information and the category prompt to obtain the target prompt information.
[0234] Optionally, the target task includes a classification task; the first construction module 1004 is further configured to obtain a preset prompt format corresponding to the classification task, wherein the preset prompt format includes a category prompt and a category identifier; the splicing reception device extracts the category information, the category prompt and the category identifier to obtain the target prompt information.
[0235] Optionally, the information extraction model includes an encoding unit, a prompt separation unit and a feature processing unit; the first input module 1006 is further configured to input the spliced information into the information extraction model, and through the prompt separation unit, perform position separation on the target prompt information to obtain a prompt separation matrix; through the encoding unit, encode the prompt separation matrix and the spliced information to obtain a coding feature matrix; through the feature processing unit, fuse the prompt separation matrix and the coding feature matrix to obtain an information extraction matrix.
[0236] Optionally, the device also includes: a model training module, configured to obtain a sample set, wherein the sample set includes sample texts and sample task description information of multiple sample tasks, the sample text carries a sample matrix label, and the sample matrix label is obtained based on the sample processing result of the sample text; for any sample task, according to the sample task type of the sample task, the sample extraction category information is extracted from the sample task description information, and the sample prompt information is constructed according to the preset prompt format corresponding to the sample extraction category information and the sample task type; the sample prompt information and the sample text are spliced, and the spliced sample information is input into the information extraction model to obtain a predicted information extraction matrix, wherein the predicted information extraction matrix represents the correspondence between the sample text and the sample prompt information; according to the sample matrix label and the predicted information extraction matrix, the model parameters of the information extraction model are adjusted to obtain a trained information extraction model.
[0237] Optionally, the first generating module 1008 is further configured to construct a text matrix according to the text to be processed and the target prompt information; and extract the task processing result of the target task from the text matrix according to the information extraction matrix.
[0238] Optionally, the first generation module 1008 is further configured to transform the information extraction matrix to obtain a target information extraction matrix, wherein the target information extraction matrix represents the information extraction probability of the text to be processed; and determine the task processing result of the target task based on the target information extraction matrix.
[0239] By applying the solution of the embodiments of this specification, by constructing target prompt information based on the task type, the target prompt information can be used to prompt the information extraction model of the text processing task type, so that the information extraction model can handle multiple types of text processing tasks, thereby improving the scalability of the information extraction model and the versatility of task processing. Moreover, since the task processing results are derived from the task description information and the text to be processed, the accuracy and stability of the task processing results are guaranteed.
[0240] The above is a schematic scheme of a task processing device of this embodiment. It should be noted that the technical scheme of the task processing device and the technical scheme of the task processing method described above are of the same concept. For details not described in detail in the technical scheme of the task processing device, please refer to the description of the technical scheme of the task processing method described above.
[0241] Corresponding to the above-mentioned information extraction model training method embodiment, this specification also provides an information extraction model training device embodiment. Figure 11 shows a schematic diagram of the structure of an information extraction model training device provided in one embodiment of this specification. As shown in Figure 11, the device is applied to a cloud-side device and includes:
[0242] The second acquisition module 1102 is configured to acquire a sample set, wherein the sample set includes sample texts and sample task description information of multiple sample tasks, the sample texts carry sample matrix labels, and the sample matrix labels are obtained based on sample processing results of the sample texts;
[0243] The second constructing module 1104 is configured to extract sample extraction category information from the sample task description information according to the sample task type of any sample task, and construct sample prompt information according to the preset prompt format corresponding to the sample extraction category information and the sample task type;
[0244] The second input module 1106 is configured to concatenate the sample prompt information and the sample text, and input the concatenated sample information into the information extraction model to obtain a prediction information extraction matrix, wherein the prediction information extraction matrix represents the correspondence between the sample text and the sample prompt information;
[0245] The adjustment module 1108 is configured to adjust the model parameters of the information extraction model according to the sample matrix labels and the predicted information extraction matrix to obtain a trained information extraction model.
[0246] By applying the solution of the embodiment of this specification, the model parameters of the information extraction model are adjusted according to the sample matrix labels and the predicted information extraction matrix to obtain a trained information extraction model. By continuously adjusting the model parameters of the information extraction model, the final information extraction model can be made more accurate.
[0247] The above is a schematic diagram of an information extraction model training device according to this embodiment. It should be noted that the technical solution of this information extraction model training device and the technical solution of the aforementioned information extraction model training method are based on the same concept. For details not described in detail in the technical solution of the information extraction model training device, please refer to the description of the technical solution of the aforementioned information extraction model training method.
[0248] Corresponding to the above-mentioned classification task processing method embodiment, this specification also provides a classification task processing device embodiment. FIG12 shows a schematic structural diagram of a classification task processing device provided by one embodiment of this specification. As shown in FIG12, the device includes:
[0249] The third acquisition module 1202 is configured to acquire the to-be-processed text and task description information of the target classification task;
[0250] The third constructing module 1204 is configured to extract the category information to be extracted from the task description information according to the task type of the target classification task, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type;
[0251] The third input module 1206 is configured to concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the correspondence between the text to be processed and the target prompt information;
[0252] The second generating module 1208 is configured to determine the classification result of the target classification task according to the information extraction matrix.
[0253] By applying the solution of the embodiments of this specification, by constructing target prompt information based on the task type, the target prompt information can be used to prompt the information extraction model that the task is a classification task, so that the information extraction model can handle the classification task, thereby improving the scalability of the information extraction model and the versatility of task processing. Moreover, since the classification results are derived from the task description information and the text to be processed, the accuracy and stability of the classification results are guaranteed.
[0254] The above is a schematic diagram of a classification task processing device according to this embodiment. It should be noted that the technical solution of the classification task processing device and the technical solution of the classification task processing method described above are based on the same concept. For details not described in detail in the technical solution of the classification task processing device, please refer to the description of the technical solution of the classification task processing method described above.
[0255] Figure 13 shows a block diagram of a computing device according to one embodiment of the present disclosure. Components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 via a bus 1330, and a database 1350 is used to store data.
[0256] The computing device 1300 also includes an access device 1340 that enables the computing device 1300 to communicate via one or more networks 1360. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a World Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0257] In one embodiment of the present specification, the aforementioned components of computing device 1300 and other components not shown in FIG13 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG13 is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0258] Computing device 1300 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1300 may also be a mobile or stationary server.
[0259] Among them, the processor 1320 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned task processing method or information extraction model training method or classification task processing method.
[0260] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the aforementioned task processing method, information extraction model training method, and classification task processing method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned task processing method, information extraction model training method, or classification task processing method.
[0261] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned task processing method or information extraction model training method or classification task processing method.
[0262] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the aforementioned task processing method, information extraction model training method, and classification task processing method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the aforementioned task processing method, information extraction model training method, or classification task processing method.
[0263] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned task processing method or information extraction model training method or classification task processing method.
[0264] The above is a schematic diagram of a computer program according to this embodiment. It should be noted that the technical solution of this computer program is based on the same concept as the technical solutions of the aforementioned task processing method, information extraction model training method, and classification task processing method. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the aforementioned task processing method, information extraction model training method, or classification task processing method.
[0265] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0266] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0267] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0268] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0269] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A task processing method, comprising: Obtain the pending text and task description information of the target task; According to the task type of the target task, extract the category information to be extracted from the task description information, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type; splicing the target prompt information and the text to be processed, and inputting the spliced information into an information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the corresponding relationship between the text to be processed and the target prompt information; A task processing result of the target task is determined according to the information extraction matrix.
2. According to the method of claim 1, the target task includes multi-level tasks; The step of extracting the category information to be extracted from the task description information according to the task type of the target task includes: According to the task type of the multi-level task, extracting the category information to be extracted of the current level from the task description information, wherein, when the current level is not the first level, the category information to be extracted of the current level is obtained based on the task processing result of the previous level of the current level; The step of constructing target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type includes: Target prompt information is constructed according to the category information to be extracted, the preset prompt format corresponding to the task type, and the task processing result of the previous level of the current level.
3. The method according to claim 1, after constructing the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type, further comprising: In the case where the length of the target prompt information is greater than a preset prompt length threshold, splitting the target prompt information to obtain a plurality of target sub-prompt information, wherein the lengths of the plurality of target sub-prompt information are less than or equal to the preset prompt length threshold; The step of splicing the target prompt information and the to-be-processed text, and inputting the spliced information into an information extraction model to obtain an information extraction matrix includes: Extracting first target sub-prompt information from the multiple target sub-prompt information, wherein the first target sub-prompt information is the target sub-prompt information among the multiple target sub-prompt information that is not spliced with the to-be-processed text; The first target sub-prompt information and the text to be processed are spliced together, and the spliced information is input into an information extraction model until the multiple target sub-prompt information do not include target sub-prompt information that is not spliced with the text to be processed, so as to obtain an information extraction matrix.
4. The method according to claim 1, before splicing the target prompt information and the to-be-processed text and inputting the spliced information into the information extraction model to obtain the information extraction matrix, further comprises: If the length of the text to be processed is greater than the preset text length threshold, the text to be processed is processed. Splitting the lines to obtain a plurality of sub-texts to be processed, wherein the lengths of the plurality of sub-texts to be processed are less than or equal to the preset text length threshold; The step of splicing the target prompt information and the to-be-processed text, and inputting the spliced information into an information extraction model to obtain an information extraction matrix includes: Extracting a first subtext to be processed from the plurality of subtexts to be processed, wherein the first subtext to be processed is a subtext to be processed that is not spliced with the target prompt information among the plurality of subtexts to be processed; The first sub-text to be processed and the target prompt information are concatenated, and the concatenated information is input into an information extraction model until the plurality of sub-texts to be processed do not include sub-texts to be processed that are not concatenated with the target prompt information, thereby obtaining an information extraction matrix.
5. The method according to claim 1, wherein the target task comprises an extraction task; The step of constructing target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type includes: Obtaining a preset prompt format corresponding to the extraction task, wherein the preset prompt format includes a category prompt; The category information to be extracted and the category prompt are concatenated to obtain target prompt information.
6. The method according to claim 1, wherein the target task comprises a classification task; The step of constructing target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type includes: Acquire a preset prompt format corresponding to the classification task, wherein the preset prompt format includes a category prompt and a category identifier; The category information to be extracted, the category prompt and the classification identifier are concatenated to obtain target prompt information.
7. The method according to claim 1, wherein the information extraction model comprises an encoding unit, a hint separation unit and a feature processing unit; The step of inputting the spliced information into the information extraction model to obtain the information extraction matrix includes: The spliced information is input into the information extraction model, and the target prompt information is positionally separated by the prompt separation unit to obtain a prompt separation matrix; The coding unit encodes the hint separation matrix and the concatenated information to obtain a coding feature matrix; The feature processing unit performs fusion processing on the prompt separation matrix and the encoding feature matrix to obtain an information extraction matrix.
8. The method according to claim 1, before the step of concatenating the target prompt information and the to-be-processed text and inputting the concatenated information into an information extraction model to obtain an information extraction matrix, further comprises: Obtain a sample set, wherein the sample set includes sample texts of multiple sample tasks and sample task description information, the sample texts carry sample matrix labels, and the sample matrix labels are based on sample processing of the sample texts. The result is obtained; For any sample task, according to the sample task type of the sample task, sample extraction category information is extracted from the sample task description information, and sample prompt information is constructed according to the preset prompt format corresponding to the sample extraction category information and the sample task type; splicing the sample prompt information and the sample text, and inputting the spliced sample information into an information extraction model to obtain a prediction information extraction matrix, wherein the prediction information extraction matrix represents the corresponding relationship between the sample text and the sample prompt information; According to the sample matrix labels and the prediction information extraction matrix, the model parameters of the information extraction model are adjusted to obtain a trained information extraction model.
9. The method according to claim 1, wherein determining the task processing result of the target task according to the information extraction matrix comprises: Constructing a text matrix according to the text to be processed and the target prompt information; According to the information extraction matrix, the task processing result of the target task is extracted from the text matrix.
10. The method according to claim 1, wherein determining the task processing result of the target task according to the information extraction matrix comprises: Transforming the information extraction matrix to obtain a target information extraction matrix, wherein the target information extraction matrix represents the information extraction probability of the text to be processed; A task processing result of the target task is determined according to the target information extraction matrix.
11. A method for training an information extraction model, applied to a cloud-side device, comprising: Acquire a sample set, wherein the sample set includes sample texts and sample task description information of a plurality of sample tasks, the sample texts carry sample matrix labels, and the sample matrix labels are obtained based on sample processing results of the sample texts; For any sample task, according to the sample task type of the sample task, sample extraction category information is extracted from the sample task description information, and sample prompt information is constructed according to the preset prompt format corresponding to the sample extraction category information and the sample task type; splicing the sample prompt information and the sample text, and inputting the spliced sample information into an information extraction model to obtain a prediction information extraction matrix, wherein the prediction information extraction matrix represents the corresponding relationship between the sample text and the sample prompt information; According to the sample matrix labels and the prediction information extraction matrix, the model parameters of the information extraction model are adjusted to obtain a trained information extraction model.
12. A classification task processing method, comprising: Obtain the text to be processed and task description information of the target classification task; According to the task type of the target classification task, extract the category information to be extracted from the task description information, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type; splicing the target prompt information and the text to be processed, and inputting the spliced information into an information extraction model to obtain an information extraction matrix, wherein the information extraction matrix represents the corresponding relationship between the text to be processed and the target prompt information; According to the information extraction matrix, a classification result of the target classification task is determined.
13. A computing device comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1 to 10 or claim 11 or claim 12 are implemented.
14. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method described in any one of claims 1 to 10 or claim 11 or claim 12.
15. A computer program, when executed in a computer, causes the computer to execute the steps of the method according to any one of claims 1 to 10 or claim 11 or claim 12.
Citation Information
Patent Citations
Language processing method and device based on automatic prompt recommendation and terminal
CN114238629A
Information extraction method, article identification method and information extraction model training method
CN116644743A
Multi-task training method and device, storage medium and electronic equipment
CN116957070A
Unified information extraction method and device
CN117113998A
Multi-task learning framework for multi-context machine learning
US20210390390A1
Cited By
Model task classification method and device, model selection method and electronic equipment
CN121093185A