Task processing method, information extraction model training method and classification task processing method

By constructing target prompt information and inputting information extraction models, the problems of large resource consumption and model limitations in the existing technology are solved, and efficient and general text processing task processing is achieved.

CN120067315APending Publication Date: 2025-05-30HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311608227.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, different machine learning models are needed to use for complex and diverse text processing tasks, resulting in large resource consumption during model training and high limitations of the model obtained by training.

Method used

A task processing method is provided, by obtaining the pending text and task description information, extracting the category information to be extracted, and constructing target prompt information according to the task type, inputting the information extraction model to obtain the information extraction matrix, and determining the task processing result.

Benefits of technology

It improves the scalability of the information extraction model and the universality of task processing, and ensures the accuracy and stability of task processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067315A_ABST
    Figure CN120067315A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a task processing method, an information extraction model training method and a classification task processing method. The task processing method comprises the steps of obtaining a to-be-processed text and task description information of a target task; according to the task type of the target task, to-be-extracted category information is extracted from the task description information, and target prompt information is constructed according to the to-be-extracted category information and a preset prompt format corresponding to the task type; the target prompt information and the to-be-processed text are spliced, the spliced information is input into an information extraction model, an information extraction matrix is obtained, and the information extraction matrix represents the corresponding relation between the to-be-processed text and the target prompt information; and determining a task processing result of the target task according to the information extraction matrix. According to the method, the information extraction model is prompted with the text processing task type by using the target prompt information, so that the information extraction model can process multiple types of text processing tasks, and the expansibility of the information extraction model and the universality of task processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and particularly to task processing, information extraction model training, and classification task processing methods. Background Art

[0002] With the development of computer technology, text processing increasingly relies on the Internet. Text processing is a process of analyzing, understanding, extracting, etc. text, and has been widely applied to various fields of people's daily lives.

[0003] Currently, different processing models can generally be used for different text processing tasks. For example, an information extraction model is used for information extraction tasks, and a text classification model is used for text classification tasks. However, for complex and diverse text processing tasks, different machine learning models need to be used, resulting in a large consumption of resources in the model training process and a high limitation of the trained model. Therefore, a general task processing solution is urgently needed. Summary of the Invention

[0004] In view of this, the embodiments of this specification provide a task processing method. One or more embodiments of this specification also relate to an information extraction model training method, a classification task processing method, a task processing device, an information extraction model training device, a classification task processing device, a computing device, a computer-readable storage medium, and a computer program, so as to solve the technical defects existing in the prior art.

[0005] According to the first aspect of the embodiments of this specification, a task processing method is provided, including:

[0006] Obtain the text to be processed and task description information of the target task;

[0007] Extract the category information to be extracted from the task description information according to the task type of the target task, and construct target prompt information according to the category information to be extracted and the preset prompt format corresponding to the task type;

[0008] Concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the corresponding relationship between the text to be processed and the target prompt information;

[0009] Determine the task processing result of the target task according to the information extraction matrix.

[0010] According to the second aspect of the embodiments of this specification, an information extraction model training method is provided, which is applied to a cloud-side device and includes:

[0011] Obtain a sample set, where the sample set includes sample texts of multiple sample tasks and sample task description information, the sample texts carry sample matrix tags, and the sample matrix tags are obtained based on the sample processing results of the sample texts;

[0012] For any sample task, according to the sample task type of the sample task, extract sample extraction category information from the sample task description information, and construct sample prompt information according to the sample extraction category information and the preset prompt format corresponding to the sample task type;

[0013] Concatenate the sample prompt information and the sample text, and input the concatenated sample information into the information extraction model to obtain a predicted information extraction matrix, where the predicted information extraction matrix represents the corresponding relationship between the sample text and the sample prompt information;

[0014] According to the sample matrix tag and the predicted information extraction matrix, adjust the model parameters of the information extraction model to obtain a trained information extraction model.

[0015] According to the third aspect of the embodiments of the present specification, a classification task processing method is provided, including:

[0016] Obtain the text to be processed and task description information of the target classification task;

[0017] According to the task type of the target classification task, extract the category information to be extracted from the task description information, and construct target prompt information according to the category information to be extracted and the preset prompt format corresponding to the task type;

[0018] Concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the corresponding relationship between the text to be processed and the target prompt information;

[0019] Determine the classification result of the target classification task according to the information extraction matrix.

[0020] According to the fourth aspect of the embodiments of the present specification, a task processing device is provided, including:

[0021] A first acquisition module configured to acquire the text to be processed and task description information of the target task;

[0022] A first construction module configured to extract the category information to be extracted from the task description information according to the task type of the target task, and construct target prompt information according to the category information to be extracted and the preset prompt format corresponding to the task type;

[0023] A first input module, configured to splice a target prompt message and a text to be processed, and input the spliced message into an information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the correspondence between the text to be processed and the target prompt message;

[0024] A first generation module, configured to determine a task processing result of a target task according to the information extraction matrix.

[0025] According to the fifth aspect of the embodiments of this specification, an information extraction model training device is provided, which is applied to a cloud-side device and includes:

[0026] A second acquisition module, configured to acquire a sample set, where the sample set includes sample texts and sample task description information of multiple sample tasks, the sample texts carry sample matrix tags, and the sample matrix tags are obtained based on the sample processing results of the sample texts;

[0027] A second construction module, configured to, for any sample task, extract sample extraction category information from the sample task description information according to the sample task type of the sample task, and construct a sample prompt message according to the sample extraction category information and a preset prompt format corresponding to the sample task type;

[0028] A second input module, configured to splice the sample prompt message and the sample text, and input the spliced sample information into the information extraction model to obtain a predicted information extraction matrix, where the predicted information extraction matrix represents the correspondence between the sample text and the sample prompt message;

[0029] An adjustment module, configured to adjust the model parameters of the information extraction model according to the sample matrix tag and the predicted information extraction matrix to obtain a trained information extraction model.

[0030] According to the sixth aspect of the embodiments of this specification, a classification task processing device is provided, including:

[0031] A third acquisition module, configured to acquire the text to be processed and task description information of a target classification task;

[0032] A third construction module, configured to extract to-be-extracted category information from the task description information according to the task type of the target classification task, and construct a target prompt message according to the to-be-extracted category information and a preset prompt format corresponding to the task type;

[0033] A third input module, configured to splice the target prompt message and the text to be processed, and input the spliced information into the information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the correspondence between the text to be processed and the target prompt message;

[0034] A second generation module, configured to determine a classification result of a target classification task according to an information extraction matrix.

[0035] According to a seventh aspect of the embodiments of the present specification, there is provided a computing device, including:

[0036] A memory and a processor;

[0037] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in the above first aspect or second aspect or third aspect are implemented.

[0038] According to an eighth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium storing computer-executable instructions, and when the instructions are executed by a processor, the steps of the method provided in the above first aspect or second aspect or third aspect are implemented.

[0039] According to a ninth aspect of the embodiments of the present specification, there is provided a computer program, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the method provided in the above first aspect or second aspect or third aspect.

[0040] The task processing method provided in an embodiment of the present specification obtains a text to be processed and task description information of a target task; extracts category information to be extracted from the task description information according to the task type of the target task, and constructs target prompt information according to the category information to be extracted and a preset prompt format corresponding to the task type; splices the target prompt information and the text to be processed, and inputs the spliced information into an information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the correspondence between the text to be processed and the target prompt information; determines a task processing result of the target task according to the information extraction matrix. By constructing target prompt information based on the task type, the text processing task type can be prompted to the information extraction model by using the target prompt information, so that the information extraction model can process multi-type text processing tasks, improving the expandability of the information extraction model and the generality of task processing. Moreover, since the task processing result is derived from the task description information and the text to be processed, the accuracy and stability of the task processing result are ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is an architecture diagram of a task processing system provided in an embodiment of the present specification;

[0042] Figure 2 is an architecture diagram of another task processing system provided in an embodiment of the present specification;

[0043] Figure 3It is a flowchart of a task processing method provided by an embodiment of this specification;

[0044] Figure 4 It is a framework diagram of an information extraction model in a task processing method provided by an embodiment of this specification;

[0045] Figure 5 It is a flowchart of a method for training an information extraction model provided by an embodiment of this specification;

[0046] Figure 6 It is a flowchart of a classification task processing method provided by an embodiment of this specification;

[0047] Figure 7a It is a flowchart of the processing process of the first task processing method provided by an embodiment of this specification;

[0048] Figure 7b It is a flowchart of the processing process of the second task processing method provided by an embodiment of this specification;

[0049] Figure 7c It is a flowchart of the processing process of the third task processing method provided by an embodiment of this specification;

[0050] Figure 8 It is a flowchart of the processing process of the fourth task processing method provided by an embodiment of this specification;

[0051] Figure 9 It is a flowchart of the processing process of the fifth task processing method provided by an embodiment of this specification;

[0052] Figure 10 It is a schematic structural diagram of a task processing device provided by an embodiment of this specification;

[0053] Figure 11 It is a schematic structural diagram of an information extraction model training device provided by an embodiment of this specification;

[0054] Figure 12 It is a schematic structural diagram of a classification task processing device provided by an embodiment of this specification;

[0055] Figure 13 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners

[0056] Numerous specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.

[0057] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0058] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0059] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to select to authorize or refuse.

[0060] First, the noun terms involved in one or more embodiments of this specification are explained.

[0061] Linking mechanism: The Linking mechanism is a decoding mechanism in the information extraction task, which realizes the extraction of entities by connecting entity tags with the entity fragment content in the original text in a certain way.

[0062] Schema: Schema is a structured template expression of the expected output content, usually in json format, that is, constructed as a collection of key-value pairs.

[0063] Prompt: A prompt is a template, usually composed of natural language, which is inserted before the input text to form training data together with the original input. The same template is used during inference to unify the training and inference formats, thereby fully leveraging the potential of the pre-trained model.

[0064] Natural Language Understanding: Natural Language Understanding refers to non-generative natural language processing (NLP) tasks. Usually, the input is a piece of text, and the output is an externally given text label or a part of the content in the text. Natural Language Understanding includes named entity recognition, relation extraction, event extraction, attribute sentiment extraction, sentiment classification, text classification, text matching, reading comprehension, natural language inference, etc.

[0065] Zero-shot Learning: Zero-shot learning means that the model directly makes predictions on data distributions that it has not learned without any further training. Zero-shot learning is often used in cold-start scenarios.

[0066] Few-shot Learning: The model is fine-tuned on only a very small number of example data (such as 1, 5, 10), and then makes predictions on this type of data. Few-shot learning is often used in cold-start or low-resource scenarios.

[0067] Classification Task: A classification task refers to the task of judging which label type the input text belongs to given an external label set and the input text. In addition to typical classification tasks, text matching, coreference resolution, natural language inference, and some reading comprehension tasks can also be transformed into classification tasks.

[0068] Extraction Task: An extraction task refers to the task of extracting part of the input text content that meets the label conditions from the input text given an external label set. It includes entity recognition, relation extraction, event extraction, attribute sentiment extraction tasks, etc.

[0069] Generation Task: A generation task refers to the task of having the model autonomously generate the expected answer given the context.

[0070] Named Entity Recognition: Named entity recognition refers to the task of extracting entity fragments from the input text.

[0071] Relation Extraction: Relation extraction refers to the task of extracting the subject and object entity fragments with specific relation types from the input text.

[0072] Event Extraction: Event extraction refers to the task of extracting the interrelated event argument fragments from the input text.

[0073] Attribute sentiment extraction: Attribute sentiment extraction refers to the task of extracting relevant sentiment word fragments of specific objects in the input text.

[0074] Pronoun resolution: Pronoun resolution refers to the task of determining whether a given pronoun refers to an object in the input text.

[0075] Sentiment classification: Sentiment classification refers to the task of returning a label that can appropriately represent the sentiment tendency of the input text given the input text and candidate sentiment labels.

[0076] Text classification: Text classification refers to the task of returning a label that can appropriately represent the category to which the input text belongs given the input text and candidate category labels.

[0077] Text matching: Text matching refers to the task of returning a label that can appropriately represent the similarity of two input texts given two input texts and corresponding similarity labels. Among them, the text similarity labels are usually "similar" and "dissimilar".

[0078] Natural language inference: Natural language inference refers to the task of returning a label that can appropriately represent the relationship between two input texts given two input texts and corresponding text relationship labels. Among them, the text relationship labels are usually "entailment", "contradiction" and "neutral".

[0079] Reading comprehension: Reading comprehension refers to the task of inputting a question and reference text and returning the answer to the question based on the reference text. Usually, reading comprehension can be divided into multiple-choice reading comprehension and extraction reading comprehension.

[0080] Natural language understanding includes various types of tasks in different forms. Usually, different processing models can be used to handle different text processing tasks. However, for complex and diverse text processing tasks, different machine learning models are needed, resulting in a large consumption of resources in the model training process and relatively high limitations of the trained models. In previous work, classification labels were spliced in front of the original text, and then an extraction framework was adopted for task processing in order to unify the classification and extraction models. However, due to the method of splicing the labels in front of the original text, the original text and the labels will share the maximum input text length of 512 of the model. Therefore, when the number of labels is very large, the original text will be truncated, thus affecting the classification effect of the model; at the same time, the splicing method means that complex hierarchical structures cannot be represented. Therefore, the above scheme cannot handle hierarchical classification tasks.

[0081] To solve the above problems, an embodiment of this specification proposes a general natural language understanding solution for zero-shot learning and few-shot learning based on the extraction paradigm. Relying on the Linking mechanism for the extraction model, target prompt information is constructed by presetting the prompt format and corresponding data, unlocking the capabilities of the information extraction model for multi-class classification, multi-label classification, and hierarchical classification, and providing a unified solution for all natural language understanding tasks except for generation tasks, achieving the true "one model for all natural language understanding tasks". At the same time, the information extraction model is deeply and extensively trained on high-quality labeled data, endowing the information extraction model with powerful zero-shot and few-shot learning capabilities.

[0082] Specifically, obtain the text to be processed and task description information of the target task; according to the task type of the target task, extract the category information to be extracted from the task description information, and construct target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type; splice the target prompt information and the text to be processed, and input the spliced information into the information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the correspondence between the text to be processed and the target prompt information; determine the task processing result of the target task according to the information extraction matrix. By constructing target prompt information based on the task type, the text processing task type can be prompted to the information extraction model using the target prompt information, enabling the information extraction model to process multiple types of text processing tasks, improving the extensibility of the information extraction model and the generality of task processing. Moreover, since the task processing result is derived from the task description information and the text to be processed, the accuracy and stability of the task processing result are ensured.

[0083] In this specification, a task processing method is provided. This specification also relates to an information extraction model training method, a classification task processing method, a task processing device, an information extraction model training device, a classification task processing device, a computing device, and a computer-readable storage medium, which will be described in detail one by one in the following embodiments.

[0084] See Figure 1 , Figure 1 shows an architecture diagram of a task processing system provided by an embodiment of this specification. The task processing system may include a client 100 and a server 200;

[0085] The client 100 is used to send the text to be processed and task description information of the target task to the server 200;

[0086] The server 200 is configured to extract the category information to be extracted from the task description information according to the task type of the target task, and construct the target prompt information according to the category information to be extracted and the preset prompt format corresponding to the task type; splice the target prompt information and the text to be processed, and input the spliced information into the information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the corresponding relationship between the text to be processed and the target prompt information; determine the task processing result of the target task according to the information extraction matrix; send the task processing result to the client 100.

[0087] The client 100 is further configured to receive the task processing result sent by the server 200.

[0088] By applying the solution of the embodiments of this specification, the target prompt information is constructed based on the task type, so that the text processing task type can be prompted to the information extraction model by using the target prompt information, enabling the information extraction model to process multi-type text processing tasks, improving the scalability of the information extraction model and the generality of task processing. Moreover, since the task processing result is derived from the task description information and the text to be processed, the accuracy and stability of the task processing result are ensured.

[0089] See Figure 2 , Figure 2 shows an architecture diagram of another task processing system provided by an embodiment of this specification. The task processing system may include multiple clients 100 and a server 200. Among them, the client 100 may include a terminal device, and the server 200 may include a cloud device. Communication connections can be established between multiple clients 100 through the server 200. In a task processing scenario, the server 200 is used to provide task processing services between multiple clients 100. Multiple clients 100 can be used as senders or receivers respectively to achieve communication through the server 200.

[0090] The user can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100, etc. In a task processing scenario, it may be that the user publishes a data stream to the server 200 through the client 100, and the server 200 generates a task processing result according to the data stream and pushes the task processing result to other clients that have established communication.

[0091] Among them, a connection is established between the client 100 and the server 200 through a network. The network provides a medium for the communication link between the client 100 and the server 200. The network may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The data transmitted by the client 100 may need to be processed such as encoded, transcoded, compressed, etc. before being published to the server 200.

[0092] The client 100 can be a browser, an APP (Application), a web application such as an H5 (HyperText Markup Language 5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application, etc. The client 100 can be developed based on the software development kit (SDK) of the corresponding service provided by the server 200, such as developed based on the real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0093] The server 200 can include servers that provide various services. For example, a server that provides communication services for multiple clients, or a server for background training that provides support for the models used on the client, or a server that processes the data sent by the client, etc. It should be noted that the server 200 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.

[0094] It is worth noting that the task processing method provided in the embodiments of this specification is generally executed by the server. However, in other embodiments of this specification, the client can also have similar functions as the server, so as to execute the task processing method provided in the embodiments of this specification. In other embodiments, the task processing method provided in the embodiments of this specification can also be jointly executed by the client and the server.

[0095] See Figure 3 , Figure 3 which shows a flowchart of a task processing method provided by an embodiment of this specification, specifically including the following steps:

[0096] Step 302: Obtain the text to be processed and the task description information of the target task.

[0097] In one or more embodiments of this specification, when processing a task, the text to be processed and the task description information of the target task can be obtained to process the text to be processed based on the task description information and obtain a task processing result.

[0098] Specifically, the target task can be a task in different scenarios, such as a document processing task in a document self-learning scenario, a text dialogue task in a dialogue scenario, and so on. The text length of the language type of the text to be processed is specifically set according to the actual situation. For example, the text to be processed can be a long text or a short text, and the text to be processed can also be an English text or a Chinese text. The task description information is used to tell the information extraction model the category information to be extracted and the structure of the category information. The task description information includes but is not limited to the task type and the category information to be extracted. Among them, the category information to be extracted can be understood as the category label to be extracted. The task description information can be in text format or json format Schema information. For example, when the target task is a text classification task, the task description information can include the category information to be extracted {"Culture": None, "Entertainment": None, "Sports": None, "Finance and Economics": None}.

[0099] In practical applications, there are various ways to obtain the text to be processed and the task description information of the target task, which are specifically selected according to the actual situation, and this specification embodiment does not make any restrictions on this. In one possible implementation manner of this specification, the text to be processed and the task description information of the target task sent by the user through the client can be received. In another possible implementation manner of this specification, the text to be processed and the task description information of the target task can be read from other data acquisition devices or databases. Among them, the text to be processed can be obtained by performing optical character recognition (OCR, Optical Character Recognition) on a picture, or can be obtained by performing audio text conversion on audio data.

[0100] Step 304: Extract the category information to be extracted from the task description information according to the task type of the target task, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type.

[0101] In one or more embodiments of this specification, after obtaining the text to be processed and the task description information of the target task, further, the category information to be extracted can be extracted from the task description information according to the task type of the target task, and the target prompt information can be constructed according to the preset prompt format corresponding to the category information to be extracted and the task type.

[0102] Specifically, the task types of the target tasks include, but are not limited to, extraction tasks, classification tasks, non-hierarchical tasks, and multi-level tasks. Extraction tasks include, but are not limited to, relation extraction tasks, event extraction tasks, and attribute sentiment extraction tasks. Classification tasks include, but are not limited to, sentiment classification tasks and industry classification tasks. Non-hierarchical tasks are ordinary classification tasks, such as sentiment classification tasks. Multi-level classification tasks refer to tasks with multiple label levels and multiple category labels in each label level, such as hierarchical classification tasks. The category information to be extracted refers to the category labels corresponding to the target tasks. There can be multiple pieces of category information to be extracted. For example, if the target task is an attribute sentiment extraction task, the two pieces of category information to be extracted are the sentiment classification labels {"positive": None, "negative": None}. The target prompt information (ESI, Explicit Schema Instructor) is a type of Prompt used to prompt the information processing model about the current target task.

[0103] The preset prompt formats include, but are not limited to, the category prompt symbol [TYPE], the result prompt symbol [PREFIX], and the classification identifiers used only in classification tasks. Among them, the classification identifiers include the single-label classification identifier [CLASSIFY] and the multi-label classification identifier [MULTICLASSIFY].

[0104] [PREFIX] is used to identify the task processing results of the previous-level tasks in multi-level tasks. Suppose in a hierarchical classification task, the current level is the third level, that is, the current round is the third round, and the target prompt information at the current level is "[PREFIX] Level 1 result - Level 2 result [TYPE] Level 3 category information to be extracted 1 [TYPE] Level 3 category information to be extracted 2…".

[0105] [TYPE] is used to identify each piece of category information to be extracted. When the information extraction model processes, it also decodes through the values corresponding to each [TYPE] and the text to be processed.

[0106] [CLASSIFY] and [MULTICLASSIFY] are used to identify that the target task is a classification task. On the one hand, it tells the information extraction model that the task is a classification task rather than an extraction task. On the other hand, it establishes the connection between the classification identifier and the category information to be extracted. Since the overall framework of the information extraction model is an extraction-based framework, but different from the extraction task where there are corresponding fragments to be extracted in the text to be processed, the classification task requires an additional extractable field in the text body to be connected with the previous [TYPE] and achieve decoding. That is, in the extraction task, the category information to be extracted, such as person name, place name, etc., is obtained through [TYPE], and then connected with the corresponding fragments in the text to be processed, such as Xiaowang, Hangzhou, etc.; while in the classification task, the category information to be extracted, such as the category "positive sentiment" in the sentiment classification task, is also obtained through [TYPE], and connected with [CLASSIFY] or [MULTICLASSIFY] to convert the classification task into an extraction task, so as to realize the classification task and the extraction task using the information extraction model. Among them, [CLASSIFY] refers to a single-label classification task, that is, the information extraction model will only output the category information with the highest probability; [MULTICLASSIFY] refers to a multi-label classification task, and the information extraction model will output the category information with a probability exceeding the preset threshold.

[0107] It should be noted that before extracting the category information to be extracted from the task description information according to the task type of the target task, the task type of the target task can be determined. There are various ways to determine the task type of the target task, which are specifically selected according to the actual situation, and this specification embodiment does not make any limitation on this. In one possible implementation of this specification, the task description information includes the task type, and by parsing the task description information, the task type of the target task is obtained. In another possible implementation of this specification, if the task type is not included in the task description information, the task description information can be input into the type recognition model to obtain the task type of the target task. For example, if the task description information is "Please extract the person information in the following text", inputting the task description information into the information recognition model can determine that the task type is an extraction task.

[0108] In practical applications, when extracting the category information to be extracted from the task description information according to the task type of the target task, it can be determined whether the target task is a multi-level task based on the task type. In the case where the target task is a non-level task, the category information to be extracted is directly extracted from the task description information; in the case where the target task is a multi-level task, the category information to be extracted can be extracted from the task description information based on the task processing result of the previous level of the current level.

[0109] In an alternative embodiment of this specification, the target task includes multi-level tasks; the step of extracting the category information to be extracted from the task description information according to the task type of the target task may include the following steps:

[0110] Extract the category information to be extracted at the current level from the task description information according to the task type of the multi-level task, where, when the current level is not the first level, the category information to be extracted at the current level is obtained based on the task processing result of the previous level of the current level;

[0111] Construct the target prompt information according to the category information to be extracted and the preset prompt format corresponding to the task type, including:

[0112] Construct the target prompt information according to the category information to be extracted, the preset prompt format corresponding to the task type, and the task processing result of the previous level of the current level.

[0113] It should be noted that assuming the task description information is {"result": {"winning bid announcement": None, "transaction announcement": None}, "notice": {"prequalification": None, "opinion solicitation": None}}, it can be seen that the target task is a multi-level task, and its labels are "result-winning bid announcement", "result-transaction announcement", "notice-prequalification", "notice-opinion solicitation". In the embodiment of this specification, in order to improve the processing efficiency and prediction accuracy of the information extraction model, it is desired that the information extraction model first predicts the labels of the first label level, and then continues to predict the labels of the second label level after obtaining the prediction result of the first label level. Specifically, when the current level is the first level, extract the labels "result" and "notice" of the first label level from the task description information as the category information to be extracted at the first level; assuming that the prediction result of the information extraction model for the first label level is "result", when the current level is the second level, it can only predict among the labels "winning bid announcement" and "transaction announcement" corresponding to the second label level of "result", so the category information to be extracted at the second level is "winning bid announcement" and "transaction announcement".

[0114] Further, after determining that the category information to be extracted at the first level is "result" and "announcement", the preset prompt format corresponding to the task type can be obtained as "[PREFIX]The task processing result of the previous level[TYPE]The category information to be extracted[TYPE]The category information to be extracted...[TYPE]The category information to be extracted". Since there is no task processing result of the previous level at the first level, the target prompt information corresponding to the first level is "[PREFIX][TYPE]Result[TYPE]Announcement". After determining that the task processing result at the first level is "result" based on the target prompt information at the first level and the text to be processed, and determining that the category information to be extracted at the second level is "winning bid announcement" and "transaction announcement", the target prompt information at the second level can be constructed as "[PREFIX]Result[TYPE]Winning bid announcement[TYPE]Transaction announcement" according to the category information to be extracted at the second level, the preset prompt format corresponding to the task type, and the task processing result at the first level.

[0115] Applying the solution of the embodiment of the present specification, according to the task type of the multi-level task, the category information to be extracted at the current level is extracted from the task description information. Among them, when the current level is not the first level, the category information to be extracted at the current level is obtained based on the task processing result of the previous level of the current level; according to the category information to be extracted, the preset prompt format corresponding to the task type, and the task processing result of the previous level of the current level, the target prompt information is constructed. When constructing the target prompt information at the second level, the prior information obtained at the first level is fully utilized, which improves the prediction accuracy of the information extraction model. Moreover, by restricting the candidate category information, the information extraction model is prevented from predicting non-existent labels, such as "result - prequalification", which improves the processing efficiency of the information extraction model.

[0116] In an optional embodiment of the present specification, the target task includes an extraction task; the above-mentioned construction of the target prompt information according to the category information to be extracted and the preset prompt format corresponding to the task type may include the following steps:

[0117] Obtain the preset prompt format corresponding to the extraction task, where the preset prompt format includes a category prompt symbol;

[0118] Concatenate the category information to be extracted and the category prompt symbol to obtain the target prompt information.

[0119] It should be noted that since the target task is an extraction task and the preset prompt format corresponding to the extraction task includes the category prompt symbol [TYPE], the target prompt information can be obtained by concatenating the category information to be extracted and the category prompt symbol.

[0120] Exemplarily, assume that the category information to be extracted is "person", "location", and "organization". Concatenate the category information to be extracted and the category prompt symbol to obtain the target prompt information as "[PREFIX][TYPE]person[TYPE]location[TYPE]organization". Since this task is a non-hierarchical classification task, [PREFIX] is empty, and [PREFIX] will not affect the result output by the information extraction model.

[0121] Furthermore, since the extraction task may be a multi-level extraction task, such as relationship extraction, the relationships it contains are "person - age", "person - position", "person - nationality". At the first level, the prompt information "[PREFIX][TYPE]person" can be constructed to extract the "person" entity in the text to be processed. After obtaining the "person" entity (such as "Xiaoming") at the end of the first-level extraction, the second-level prompt information "[PREFIX]person: Xiaoming[TYPE]age[TYPE]position[TYPE]nationality" can be constructed to further extract the text to be processed.

[0122] Applying the solution of the embodiment of this specification, obtain the preset prompt format corresponding to the extraction task, where the preset prompt format includes the category prompt symbol; concatenate the category information to be extracted and the category prompt symbol to obtain the target prompt information. Provide candidate labels to the information extraction model through the category prompt symbol for the information extraction model to make predictions, ensuring the stability of the information extraction model.

[0123] In another optional embodiment of this specification, the target task includes a classification task; the above-mentioned construction of the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type may include the following steps:

[0124] Obtain the preset prompt format corresponding to the classification task, where the preset prompt format includes the category prompt symbol and the classification identifier;

[0125] Concatenate the category information to be extracted, the category prompt symbol, and the classification identifier to obtain the target prompt information.

[0126] It should be noted that since the target task is a classification task, the preset prompt format corresponding to the classification task includes the category prompt symbol [TYPE] and the classification identifier [CLASSIFY] or [MULTICLASSIFY], so concatenating the category information to be extracted, the category prompt symbol, and the classification identifier can obtain the target prompt information.

[0127] Exemplarily, assume that the classification category information to be extracted is "positive sentiment" and "negative sentiment", and the classification task is a single-label classification task. Then, by extracting the category information, category delimiters, and classification identifiers, the target prompt information obtained is "[PREFIX][TYPE]Positive sentiment[TYPE]Negative sentiment[CLASSIFY]". Since this task is a non-hierarchical classification task, [PREFIX] is empty, and [PREFIX] will not affect the result output by the information extraction model.

[0128] Furthermore, since the classification task may be a multi-level classification task. For example, the hierarchical classification labels are "Result - Winning Bid Announcement", "Result - Transaction Announcement", "Notice - Prequalification", "Notice - Solicitation of Opinions". The first level constructs the prompt information "[PREFIX][TYPE]Result[TYPE]Notice" to first classify whether the text to be processed belongs to "Result" or "Notice". After obtaining the first-level classification result (such as "Result"), the second-level prompt information "[PREFIX]Result[TYPE]Winning Bid Announcement[TYPE]Transaction Announcement" can be constructed to further extract the text to be processed.

[0129] Applying the solution of the embodiments of this specification, obtain the preset prompt format corresponding to the classification task, where the preset prompt format includes category delimiters and classification identifiers; concatenate the category information to be extracted, category delimiters, and classification identifiers to obtain the target prompt information. By providing candidate labels to the information extraction model through the category delimiters, the stability of the information extraction model is ensured. Moreover, by using [CLASSIFY] or [MULTICLASSIFY] to convert the classification task into an extraction task, it is realized to use one information extraction model to implement the classification task and the extraction task, improving the scalability of the information extraction model and the generality of task processing.

[0130] Step 306: Concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the corresponding relationship between the text to be processed and the target prompt information.

[0131] In one or more embodiments of this specification, obtain the text to be processed and task description information of the target task; according to the task type of the target task, extract the category information to be extracted from the task description information, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type. Further, the target prompt information and the text to be processed can be concatenated, and the concatenated information can be input into the information extraction model to obtain an information extraction matrix.

[0132] Specifically, the information extraction model is trained based on the sample texts of multiple sample tasks, the sample task description information, and the sample matrix labels carried by the sample texts. The sample matrix labels are obtained based on the sample processing results of the sample texts. The information extraction matrix is an N*N score matrix, where N is the length of the concatenated information input.

[0133] It should be noted that when concatenating the target prompt information and the text to be processed, the target prompt information and the text to be processed can be concatenated in sequence before and after to obtain the concatenated information. Among them, the front-back order of the target prompt information and the text to be processed is specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard. Further, in order to distinguish the target prompt information and the text to be processed, when concatenating the target prompt information and the text to be processed, an information separator [Text] can also be added, and the concatenated information is "target prompt information + [Text] + text to be processed".

[0134] In an optional embodiment of this specification, the information extraction model includes an encoding unit, a prompt separation unit, and a feature processing unit; the above-mentioned step of inputting the concatenated information into the information extraction model to obtain the information extraction matrix may include the following steps:

[0135] Input the concatenated information into the information extraction model. Through the prompt separation unit, the position of the target prompt information is separated to obtain a prompt separation matrix;

[0136] Through the encoding unit, the prompt separation matrix and the concatenated information are encoded to obtain an encoded feature matrix;

[0137] Through the feature processing unit, the prompt separation matrix and the encoded feature matrix are fused to obtain the information extraction matrix.

[0138] It should be noted that the prompt separation unit is used to generate a prompt separation matrix, and the prompt separation matrix is used to mark the information that does not need to be extracted by the information extraction model as zero. In the embodiments of this specification, since the information extraction model does not extract from the category identifier and the result identifier in the target prompt information when extracting the task processing result, therefore, the concatenated information can be input into the information extraction model. Through the prompt separation unit, the position of the target prompt information is separated to obtain a prompt separation matrix, and the prompt separation matrix is input into the encoding unit to enable the encoding unit to quickly encode the concatenated information to obtain an encoded feature matrix. Further, after obtaining the encoded feature matrix, the prompt separation matrix and the encoded feature matrix can be multiplied through the feature processing unit to obtain the information extraction matrix. At this time, the positions corresponding to the category identifier and the result identifier in the information extraction matrix are both zero.

[0139] See Figure 4 , Figure 4The following shows a framework diagram of an information extraction model in a task processing method provided by an embodiment of this specification. As Figure 4 shown, during the i-th round of information extraction, the start symbol [CLS], the target prompt information ESI i , the information discriminator [Text], and the text to be processed are concatenated, and the concatenated information Q i is input into the information extraction model. In the information extraction model, through the prompt separation unit, the position of the target prompt information is separated to obtain a prompt separation matrix; through the encoding unit, the prompt separation matrix and the concatenated information are encoded to obtain an encoded feature matrix; through the feed-forward layer in the feature processing unit, the encoded feature matrix is processed, and through the fusion layer in the feature processing unit, the prompt separation matrix and the processed encoded feature matrix are multiplied to obtain an information extraction matrix, and the information extraction matrix is decoded based on an encoding mechanism including a Linking mechanism and a "handshake mechanism" to generate the task processing result Y i of the i-th round.

[0140] Further, after the i-th round of information extraction is completed, in the same way as the i-th round of information extraction, the start symbol [CLS], the target prompt information ESI i+1 , the information discriminator [Text], and the text to be processed are concatenated, and the concatenated information Q i+1 is input into the information extraction model for the (i + 1)-th round of information extraction. Finally, according to the task processing results...Y i-1 、Y i 、Y i+1 ... the target task processing result is output.

[0141] It should be noted that a triple-loop structure can be adopted in the construction process of the target prompt information:

[0142] The first recursive loop: loop according to the schema hierarchical structure to determine the category information to be extracted. For example, hierarchical classification, relation extraction, and event extraction need to go through multiple loops at this level;

[0143] The second recursive loop: ensure that the length of the target prompt information is less than or equal to the preset prompt length threshold. If the length of the target prompt information is greater than the preset prompt length threshold, the target prompt information is split to obtain multiple target sub-prompt information;

[0144] The third recursive loop: The second recursive loop ensures the length of the text to be processed, thus avoiding the information extraction effect being affected by the text to be processed being too short. However, if the sum of the length of the target prompt information and the length of the text to be processed still exceeds 512, the text to be processed is split into multiple sub-texts to be extracted based on the category information to be extracted.

[0145] Applying the solution of the embodiments of this specification, through a loop structure, the information extraction model can well handle large-scale classification tasks and hierarchical classification tasks. The design of the Linking mechanism and the "handshake mechanism" also makes the information extraction model highly extensible. In addition to being able to handle classification tasks, it can also handle extraction tasks. Moreover, since the model adopts an extraction framework, the task processing result is derived from the task description information and the text to be processed, ensuring the accuracy and stability of the task processing result.

[0146] Step 308: Determine the task processing result of the target task according to the information extraction matrix.

[0147] In one or more embodiments of this specification, obtain the text to be processed and the task description information of the target task; extract the category information to be extracted from the task description information according to the task type of the target task, and construct the target prompt information according to the preset prompt format corresponding to the category information to be extracted and the task type; splice the target prompt information and the text to be processed, and input the spliced information into the information extraction model. After obtaining the information extraction matrix, further, the task processing result of the target task can be determined according to the information extraction matrix.

[0148] Specifically, if the target task is an extraction task, the task processing result is the extraction field name and the corresponding field segment; if the target task is a classification task, the task processing result is the classification label.

[0149] Applying the solution of the embodiments of this specification, by constructing the target prompt information based on the task type, the text processing task type can be prompted to the information extraction model by using the target prompt information, so that the information extraction model can process multi-type text processing tasks, improving the extensibility of the information extraction model and the generality of task processing. Moreover, since the task processing result is derived from the task description information and the text to be processed, the accuracy and stability of the task processing result are ensured.

[0150] In an optional embodiment of this specification, the above determining the task processing result of the target task according to the information extraction matrix may include the following steps:

[0151] Convert the information extraction matrix to obtain a target information extraction matrix, where the target information extraction matrix represents the information extraction probability of the text to be processed;

[0152] Determine the task processing result of the target task according to the target information extraction matrix.

[0153] In practical applications, there are various ways to convert the information extraction matrix to obtain the target information extraction matrix, which is specifically selected according to the actual situation. The embodiments of this specification do not make any limitations on this.

[0154] In a possible implementation of this specification, a matrix element threshold can be obtained, and the information extraction matrix can be transformed according to the matrix element threshold to generate a 0-1 information extraction matrix. Specifically, the size relationship of the matrix element values at each position in the information extraction matrix can be judged by using the matrix element threshold. If the matrix element value is greater than the matrix element threshold, the corresponding position is determined to be 1. If the matrix element value is less than or equal to the matrix element threshold, the corresponding position is determined to be 0, and thus the 0-1 information extraction matrix can be obtained. After obtaining the 0-1 information extraction matrix, based on the Linking mechanism, the task processing result of the target task can be determined according to the 0-1 information extraction matrix.

[0155] In another possible implementation of this specification, since all positions in the 0-1 information extraction matrix may be 0 or all positions may be 1, the 0-1 information extraction matrix may be invalid and cannot establish the corresponding relationship between the text to be processed and the target prompt information, resulting in that the information extraction model does not return a classification label and the task processing result of the target task cannot be determined. Therefore, the information extraction matrix can be input into the activation layer (sigmod layer) for transformation, and after being processed by the activation layer, a probability information extraction matrix is obtained, where the matrix elements in the probability information extraction matrix represent a probability value. After obtaining the probability information extraction matrix, based on the handshake mechanism, the task processing result of the target task can be determined according to the probability information extraction matrix.

[0156] It should be noted that when determining the task processing result of the target task based on the handshake mechanism according to the probability information extraction matrix, since the classification identifiers [CLASSIFY] and [MULTICLASSIFY] always appear before the text to be processed, and the position of the category identifier [TYPE] is known information, thus, it is easy to determine the probability value of the mapping of the category identifier [TYPE] of each category information to be extracted to the classification identifier, that is, the probability value of the "[CLASSIFY]-[TYPE] pair" or "[MULTICLASSIFY]-[TYPE] pair". For a single-label classification task, the category information to be extracted corresponding to the "[CLASSIFY]-[TYPE] pair" with the largest average probability value in the probability information extraction matrix can be used as the classification result; for a multi-label classification task, the category information to be extracted corresponding to the "[MULTICLASSIFY]-[TYPE] pair" with the largest average probability value in the probability information extraction matrix, and the category information to be extracted corresponding to the "[MULTICLASSIFY]-[TYPE] pair" with an average probability value above the extraction probability threshold (such as 0.9) can be used as the classification result together.

[0157] By applying the solution of the embodiment of this specification, the information extraction matrix is ​​transformed to obtain a 0-1 information extraction matrix or a probability information extraction matrix; the task processing result of the target task is further determined according to the Linking mechanism or the handshake mechanism, thereby improving the task processing efficiency and accuracy.

[0158] In an optional embodiment of the present specification, a text matrix can be constructed according to the text to be processed and the target prompt information, so as to extract the target extraction result from the text matrix according to the valid value in the information extraction matrix. That is, the above-mentioned determination of the task processing result of the target task according to the information extraction matrix can include the following steps:

[0159] Construct a text matrix based on the text to be processed and the target prompt information;

[0160] According to the information extraction matrix, the task processing results of the target task are extracted from the text matrix.

[0161] It should be noted that, when constructing a text matrix based on the text to be processed and the target prompt information, the text to be processed and the target prompt information can be spliced, and the spliced ​​text information is used as the rows and columns of the text matrix respectively to obtain the text matrix.

[0162] For example, assuming that the text to be extracted is the English text "Today is a sunny day, I like it", and the target prompt information is "[PREFIX][TYPE]good[TYPE]bad[CLASSIFY]", the two are concatenated to obtain "[PREFIX][TYPE]good[TYPE]bad[Text][CLASSIFY]Today is a sunny day, I like it", and further use "[PREFIX][TYPE]good[TYPE]bad[Text][CLASSIFY]Today is a sunny day, I like it" as the rows and columns of the text matrix to obtain a 16*16 text matrix. According to the information extraction matrix, the task processing result of the target task is extracted from the text matrix.

[0163] In an optional embodiment of the present specification, before extracting the task processing result of the target task from the text matrix according to the information extraction matrix, the text matrix and the information extraction matrix can be aligned, so as to accurately extract the task processing result of the target task from the text matrix according to the information extraction matrix.

[0164] Applying the solution of the embodiments of this specification, a text matrix is constructed according to the text to be processed and the target prompt information; according to the information extraction matrix, the task processing result of the target task is extracted from the text matrix. Since the task processing result is derived from the task description information and the text to be processed, the accuracy and stability of the task processing result are ensured.

[0165] In practical applications, since the information to be extracted may include a large number of tags, if all the information to be extracted is concatenated to obtain the target prompt information, the length of the target prompt information may be too long. Generally, the maximum input length of the information extraction model is only 512, and the target prompt information and the text to be extracted will share the maximum input length of the model. When the length of the target prompt information is too long, the available length of the text to be extracted will be compressed very short or even non-existent, which will greatly affect the model effect of the information extraction model. Therefore, a preset prompt length threshold can be introduced to limit the length of the target prompt information, ensuring that the text to be processed has a reasonable length to provide sufficient extraction basis for the information extraction model. That is, after constructing the target prompt information according to the preset prompt format corresponding to the information to be extracted and the task type, the following steps may also be included:

[0166] When the length of the target prompt information is greater than the preset prompt length threshold, the target prompt information is split to obtain multiple target sub-prompt information, where the lengths of the multiple target sub-prompt information are less than or equal to the preset prompt length threshold;

[0167] Concatenating the target prompt information and the text to be processed, and inputting the concatenated information into the information extraction model to obtain the information extraction matrix may include the following steps:

[0168] Extract the first target sub-prompt information from the multiple target sub-prompt information, where the first target sub-prompt information is the target sub-prompt information that has not been concatenated with the text to be processed among the multiple target sub-prompt information;

[0169] Concatenate the first target sub-prompt information and the text to be processed, and input the concatenated information into the information extraction model until there is no target sub-prompt information that has not been concatenated with the text to be processed among the multiple target sub-prompt information to obtain the information extraction matrix.

[0170] Specifically, the preset prompt length threshold is specifically set according to the actual situation, such as 256, and the embodiments of this specification do not make any limitations on the preset prompt length threshold.

[0171] It should be noted that after splicing the first target sub - prompt information and the text to be processed and inputting the spliced information into the information extraction model, a first information extraction matrix can be obtained. At this time, the step of extracting the first target sub - prompt information from multiple target sub - prompt information can be returned to obtain the first information extraction matrix until there is no target sub - prompt information in the multiple target sub - prompt information that has not been spliced with the text to be processed. According to the multiple first information extraction matrices, an information extraction matrix is generated and obtained.

[0172] In an optional embodiment of this specification, if the length of the target prompt information is less than or equal to the preset prompt length threshold, for example, the length of the target prompt information is 30, then the length left for the text to be processed is 482. At this time, the target prompt information and the text to be processed can be directly spliced. If the length of the target prompt information is greater than the preset prompt length threshold, for example, the length of the target prompt information is 300, then the target prompt information can be split to obtain a target sub - prompt information 1 with a length of 250 and a target sub - prompt information 2 with a length of 50. Further, the target sub - prompt information 1 and the text to be processed can be spliced to obtain the spliced information 1, and the target sub - prompt information 2 and the text to be processed can be spliced to obtain the spliced information 2. Then, the information extraction model is used to process the spliced information 1 respectively to obtain the information extraction matrix 1, and the spliced information 2 is processed to obtain the information extraction matrix 2. Finally, the information extraction matrix 1 and the information extraction matrix 2 are summarized to obtain the information extraction matrix.

[0173] Applying the solution of the embodiment of this specification, when the length of the target prompt information is greater than the preset prompt length threshold, the target prompt information is split to obtain multiple target sub - prompt information, so as to ensure that the length of the prompt information spliced with the text to be processed is less than or equal to the preset prompt length threshold, leaving sufficient input positions for the text to be processed and ensuring the extraction effect of the information extraction model.

[0174] In practical applications, there are various ways to split the target prompt information to obtain multiple target sub - prompt information, which is specifically selected according to the actual situation, and this specification does not make any limitations in this regard. In a possible implementation manner of this specification, when the length of the target prompt information is greater than the preset prompt length threshold, the target prompt information can be directly split at the position of the preset prompt length threshold to obtain multiple target sub - prompt information.

[0175] In another possible implementation of this specification, directly truncating the target prompt information from the position of the preset prompt length threshold may result in incomplete category information to be extracted, leading to errors in the information extraction model. Therefore, when the length of the target prompt information is greater than the preset prompt length threshold, splitting the target prompt information to obtain multiple target sub-prompt information may include the following steps:

[0176] When the length of the target prompt information is greater than the preset prompt length threshold, split the target prompt information in units of the category information to be extracted to obtain multiple target sub-prompt information.

[0177] Applying the solution of the embodiment of this specification, splitting the target prompt information in units of the category information to be extracted to obtain multiple target sub-prompt information ensures that each category information to be extracted in each target sub-prompt information is complete, thus ensuring the accuracy and stability of the information extraction model.

[0178] In another optional embodiment of this specification, similar to the length of the target prompt information may be greater than the preset prompt length threshold, the length of the text to be processed may also be very long, exceeding the maximum input sequence length of the information extraction model. Therefore, a preset text length threshold can be introduced to limit the length of the text to be processed. That is, before splicing the target prompt information and the text to be processed and inputting the spliced information into the information extraction model to obtain the information extraction matrix, the following steps may also be included:

[0179] When the length of the text to be processed is greater than the preset text length threshold, split the text to be processed to obtain multiple sub-texts to be processed, where the lengths of the multiple sub-texts to be processed are less than or equal to the preset text length threshold;

[0180] Splicing the target prompt information and the text to be processed and inputting the spliced information into the information extraction model to obtain the information extraction matrix may include the following steps:

[0181] Extract the first sub-text to be processed from the multiple sub-texts to be processed, where the first sub-text to be processed is the sub-text to be processed that has not been spliced with the target prompt information among the multiple sub-texts to be processed;

[0182] Splice the first sub-text to be processed and the target prompt information, and input the spliced information into the information extraction model until there is no sub-text to be processed that has not been spliced with the target prompt information among the multiple sub-texts to be processed, to obtain the information extraction matrix.

[0183] Specifically, the preset text length threshold is specifically set according to the actual situation, and this specification embodiment does not make any limitation on the preset prompt length threshold.

[0184] It should be noted that after splicing the first sub-text to be processed and the target prompt information and inputting the spliced information into the information extraction model, a second information extraction matrix can be obtained. At this time, the step of extracting the first sub-text to be processed from multiple sub-texts to be processed can be returned until there is no sub-text to be processed that has not been spliced with the target prompt information among the multiple sub-texts to be processed, and an information extraction matrix is generated based on the multiple second information extraction matrices.

[0185] Applying the solution of the embodiment of this specification, when the length of the text to be processed is greater than the preset text length threshold, the text to be processed is split to obtain multiple sub-texts to be processed, so as to ensure that the length of each sub-text to be processed spliced with the target prompt information is less than or equal to the preset text length threshold, ensuring the extraction effect of the information extraction model.

[0186] It is worth noting that the target prompt information and the text to be processed can be split simultaneously. Referring to the above example of splitting the target prompt information, after dividing the target prompt information into target sub-prompt information 1 and target sub-prompt information 2, it is ensured that the length of the text to be processed is between 256 and 511, avoiding the situation that the text to be processed in the spliced information is too short and affecting the information extraction effect. In practical applications, the length of the spliced information of the target sub-prompt information and the text to be processed may still be greater than 512. Therefore, the text to be processed can also be split into sub-text to be processed 1 and sub-text to be processed 2, so that the information extraction model processes "target sub-prompt information 1 + sub-text to be processed 1, target sub-prompt information 1 + sub-text to be processed 2, target sub-prompt information 2 + sub-text to be processed 1, target sub-prompt information 2 + sub-text to be processed 2" in sequence. Finally, by comprehensively combining the processing results corresponding to the four spliced information, the final task processing result can be obtained.

[0187] In an optional embodiment of this specification, before splicing the target prompt information and the text to be processed and inputting the spliced information into the information extraction model to obtain the information extraction matrix, the following steps may further be included:

[0188] Obtain a sample set, where the sample set includes the sample texts and sample task description information of multiple sample tasks. The sample texts carry sample matrix labels, and the sample matrix labels are obtained based on the sample processing results of the sample texts;

[0189] For any sample task, according to the sample task type of the sample task, extract the sample extraction category information from the sample task description information, and construct the sample prompt information according to the sample extraction category information and the preset prompt format corresponding to the sample task type;

[0190] Concatenate the sample prompt information and the sample text, and input the concatenated sample information into the information extraction model to obtain a predicted information extraction matrix, where the predicted information extraction matrix represents the correspondence between the sample text and the sample prompt information;

[0191] According to the sample matrix label and the predicted information extraction matrix, adjust the model parameters of the information extraction model to obtain a trained information extraction model.

[0192] Specifically, the training method of the information extraction model is supervised training based on prompt learning, that is, each sample text in the sample set carries a true sample matrix label, and the sample matrix label is the extraction target of the information extraction model, which is used to guide the training process of the information extraction model. The way to obtain the sample set can be to read a large number of sample texts carrying sample matrix labels and sample task description information from other data acquisition devices or databases to form a sample set. It can also be to receive a large number of sample texts carrying sample matrix labels and sample task description information input by the user to form a sample set. The way to obtain the sample set is specifically selected according to the actual situation, and the embodiments of this specification do not make any limitations in this regard.

[0193] It should be noted that the sample processing result is the true processing result of the sample text. When generating the sample matrix label based on the sample processing result, the sample extraction category information can be extracted from the sample task description information according to the sample type of the sample task, and the sample prompt information can be constructed according to the sample extraction category information and the preset prompt format corresponding to the sample type. Construct the sample matrix label based on the sample prompt information and the sample text. For example, construct an initial sample matrix full of zeros based on the sample prompt information and the sample text, and modify the position corresponding to the sample processing result in the initial sample matrix to one to obtain the sample matrix label.

[0194] It should be noted that the implementation method of "extracting sample extraction category information from the sample task description information according to the sample task type of the sample task, and constructing sample prompt information according to the sample extraction category information and the preset prompt format corresponding to the sample task type; splicing the sample prompt information and the sample text, and inputting the spliced sample information into the information extraction model to obtain a predicted information extraction matrix" is the same as the above "extracting the category information to be extracted from the task description information according to the task type of the target task, and constructing target prompt information according to the category information to be extracted and the preset prompt format corresponding to the task type; splicing the target prompt information and the text to be processed, and inputting the spliced information into the information extraction model to obtain an information extraction matrix", and this will not be elaborated in the embodiments of this specification. During the prediction process, the predicted information extraction matrix can be transformed to obtain a predicted target information extraction matrix. The implementation method of "transforming the predicted information extraction matrix to obtain a predicted target information extraction matrix" is the same as the above "transforming the information extraction matrix to obtain a target information extraction matrix", and this will not be elaborated in this specification.

[0195] In practical applications, when adjusting the model parameters of the information extraction model according to the sample matrix label and the predicted information extraction matrix, the loss value can be calculated based on the sample matrix label and the predicted information extraction matrix, and the model parameters of the information extraction model can be adjusted based on the loss value until the preset stop condition is reached to obtain a trained information extraction model. The preset stop condition includes but is not limited to the loss value being less than or equal to a preset threshold, and the number of iterations reaching a preset number of iterations. Specifically, the loss value can be calculated by the following formula (1):

[0196]

[0197] where L represents the loss value, i represents the i-th sample, j represents the j-th position in the matrices zg i and z i and k represents the k-th position in the matrices zg i and z i e represents the exponential function, zg i represents the sample matrix label, and z i represents the predicted information extraction matrix.

[0198] Applying the solution of the embodiments of this specification, the loss value is calculated according to the sample matrix label and the predicted information extraction matrix, the loss value is compared with the preset stop condition, and the information extraction model is continuously trained when the preset stop condition is not met until the preset stop condition is reached, and the trained information extraction model is obtained. By continuously adjusting the model parameters of the information extraction model, the finally obtained information extraction model can be made more accurate.

[0199] SeeFigure 5 , Figure 5 shows a flowchart of a method for training an information extraction model provided by an embodiment of this specification. The information extraction model training method is applied to a cloud-side device and specifically includes the following steps:

[0200] Step 502: Obtain a sample set, where the sample set includes sample texts and sample task description information of multiple sample tasks. The sample texts carry sample matrix labels, and the sample matrix labels are obtained based on the sample processing results of the sample texts.

[0201] Step 504: For any sample task, extract sample extraction category information from the sample task description information according to the sample task type of the sample task, and construct sample prompt information according to the sample extraction category information and the preset prompt format corresponding to the sample task type.

[0202] Step 506: Concatenate the sample prompt information and the sample text, and input the concatenated sample information into the information extraction model to obtain a predicted information extraction matrix, where the predicted information extraction matrix represents the correspondence between the sample text and the sample prompt information.

[0203] Step 508: Adjust the model parameters of the information extraction model according to the sample matrix label and the predicted information extraction matrix to obtain a trained information extraction model.

[0204] It should be noted that the implementation manners of steps 502 to 508 are detailed in the training manner of the information extraction model in the above task processing method, and this embodiment of this specification does not make any limitation thereto.

[0205] In practical applications, after obtaining the trained information extraction model, the model parameters of the trained information extraction model can be sent to the end-side device so that the user can construct an information extraction model locally based on the model parameters to complete the information extraction task.

[0206] Applying the solution of this embodiment of this specification, adjusting the model parameters of the information extraction model according to the sample matrix label and the predicted information extraction matrix to obtain a trained information extraction model. By continuously adjusting the model parameters of the information extraction model, the finally obtained information extraction model can be made more accurate.

[0207] The following combines the attached Figure 6 , taking the application of the task processing method provided in this specification in the text classification scenario as an example, to further illustrate the task processing method. Among them, Figure 6 shows a flowchart of a classification task processing method provided by an embodiment of this specification, which specifically includes the following steps:

[0208] Step 602: Obtain the text to be processed and task description information of the target classification task.

[0209] Step 604: Extract the category information to be extracted from the task description information according to the task type of the target classification task, and construct the target prompt information according to the to-be-extracted category information and the preset prompt format corresponding to the task type.

[0210] Step 606: Concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the corresponding relationship between the text to be processed and the target prompt information.

[0211] Step 608: Determine the classification result of the target classification task according to the information extraction matrix.

[0212] It should be noted that the implementation manners of steps 602 to 608 are the same as those of steps 302 to 308 above, and the embodiments of this specification do not make any limitations on this.

[0213] Applying the solution of the embodiments of this specification, by constructing the target prompt information based on the task type, the task can be prompted to the information extraction model as a classification task by using the target prompt information, so that the information extraction model can process the classification task, improving the extensibility of the information extraction model and the generality of task processing. Moreover, since the classification result is derived from the task description information and the text to be processed, the accuracy and stability of the classification result are ensured.

[0214] See Figure 7a , Figure 7a shows the process flow chart of the first task processing method provided by an embodiment of this specification. Taking the single-label classification task as an example, assuming that the text to be extracted is "Today is a sunny day, I like it" and the target prompt information is "[PREFIX][TYPE]good[TYPE]bad[CLASSIFY]", concatenate the target prompt information and the text to be processed, and input the concatenated information "[PREFIX][TYPE]good[TYPE]bad[Text][CLASSIFY]Today is a sunny day, I like it" into the information extraction model to obtain a 0-1 information extraction matrix as Figure 7a shown, where [P] is [PREFIX], [T] is [TYPE], and [CLST] is [CLASSIFY].

[0215] In Figure 7aIn it, Region 1 indicates that [TYPE] is connected to the head of the corresponding segment in the text to be processed, that is, the type-token-head type; Region 2 indicates that [TYPE] is connected to the tail of the corresponding segment in the text to be processed, that is, the type-token-tail type; Region 3 indicates being connected to the head and tail of the corresponding segment in the text to be processed, that is, the token-head-tail type. Since a single token is used here to mark the classification, both the head and the tail point to the same position, that is, [CLST]. As Figure 7a shown, since [CLST] points to the identifier [T] to which the good label belongs ( Figure 7a the part with a value of 1 in it), therefore, through the Linking mechanism, the emotion label corresponding to "Today is a sunny day, I like it" can be decoded as "good". And, due to the high scalability of the Linking mechanism, this framework can be used for extraction tasks by removing the [CLST] for marking classification.

[0216] See Figure 7b , Figure 7b shows the processing procedure flowchart of the second task processing method provided by an embodiment of this specification. Taking the multi-label classification task as an example, assume the text to be extracted is "Today is a sunny day, I like it", and the target prompt information is "[PREFIX][TYPE]weather[TYPE]mood[MULTICLASSIFY]". Concatenate the target prompt information and the text to be processed to obtain "[PREFIX][TYPE]weather[TYPE]mood[Text][MULTICLASSIFY]Today is a sunny day, I like it". Input the concatenated information into the information extraction model to obtain a 0-1 information extraction matrix as Figure 7b shown, where [P] is [PREFIX], [T] is [TYPE], [CLST] is [MULTICLASSIFY]. Further, according to the 0-1 information extraction matrix, through the Linking mechanism, the topics corresponding to "Today is a sunny day, I like it" can be decoded as "weather" and "mood". For the Figure 7b description of Region 1, Region 2, and Region 3 in it, see Figure 7a , and this will not be elaborated in the embodiments of this specification.

[0217] See Figure 7c , Figure 7c shows the processing procedure flowchart of the third task processing method provided by an embodiment of this specification. Refer toFigure 7b As an example, assume that the 0-1 information extraction matrix is invalid and the model cannot return classification labels. At this time, a handshaking mechanism as shown in Figure 7c can be triggered. The information extraction matrix is passed through a sigmoid layer to obtain a probability information extraction matrix. According to the probability values of each "[MULTICLASSIFY]-[TYPE]" pair in the probability information extraction matrix, it is determined that the topics corresponding to "Today is a sunny day, I like it" are "weather" and "mood". Among them, each "[MULTICLASSIFY]-[TYPE]" pair is as shown by the shaded squares and double arrows in Figure 7c . For the descriptions of regions 1, 2, and 3 in Figure 7c , reference can be made to Figure 7a , and the embodiments of this specification will not be elaborated here.

[0218] Refer to Figure 8 . Figure 8 FIG. shows a flowchart of the processing procedure of the fourth task processing method provided by an embodiment of this specification. Taking the information extraction task as an example, assume that the text to be extracted is "In 1997, A B returned to M as CEO", and the target prompt information is "[PREFIX][TYPE]person[TYPE]organization". The target prompt information and the text to be processed are concatenated, and the concatenated information "[PREFIX][TYPE]person[TYPE]organization[Text]In 1997, AB returned to M as CEO" is input into the information extraction model to obtain an information extraction matrix as shown in Figure 8 . Among them, [P] is [PREFIX], [T] is [TYPE], person represents a person, and organization represents an organization.

[0219] As shown in Figure 8 , since the first 1 in region 1 points to the identifier [T] to which A and the person label belong, the second 1 points to the identifier [T] to which M and the organization label belong, the first 1 in region 2 points to the identifier [T] to which B and the person label belong, the second 1 points to the identifier [T] to which M and the organization label belong, and the first 1 in region 3 points to A and B, and the second 1 points to M, therefore, through the Linking mechanism, it can be decoded that person corresponds to A B and organization corresponds to M.

[0220] Refer to Figure 9 . Figure 9The flowchart of the processing procedure of the fifth task processing method provided by an embodiment of this specification is shown. Referring to the example of Figure 8 shows Figure 9 how to extract the subordination relationship between AB and the work organization M after person A B has been extracted in the relation extraction task. Suppose the text to be extracted is "In 1997, A B returned to M as CEO", and the target prompt information is "[PREFIX]person:A B[TYPE]work for(organization)". Concatenate the target prompt information and the text to be processed, and input the concatenated information "[PREFIX]person:A B[TYPE]work for(organization)[Text]In 1997, AB returned to M as CEO" into the information extraction model to obtain the information extraction matrix as Figure 9 shown, where [P] is [PREFIX], [T] is [TYPE], person represents a person, and organization represents an organization.

[0221] As Figure 9 shown, since 1 in region 1 points to the identifier [T] to which M and the work label belong, 1 in region 2 points to the identifier [T] to which M and the work label belong, and 1 in region 3 points to M, the subordination relationship between AB and the work organization M can be decoded as "work" through the Linking mechanism.

[0222] Applying the solution of the embodiment of this specification, by constructing the target prompt information based on the task type, the model can process multiple labels at one time in the extraction task, and the efficiency is improved by about 30%. Moreover, the target prompt information can be used to prompt the text processing task type to the information extraction model, enabling the information extraction model to process multi-type text processing tasks, such as large-scale classification tasks, multi-label classification tasks, and hierarchical classification tasks, improving the expandability of the information extraction model and the generality of task processing. Also, since the task processing result is derived from the task description information and the text to be processed, the accuracy and stability of the task processing result are ensured.

[0223] In practical applications, a base model with 100 million parameters and a large model with 300 million parameters are trained on high-quality supervised data covering various task types except for generation tasks, endowing the information extraction model with powerful zero-shot learning and few-shot learning capabilities. Moreover, on 17 test data sets including extraction tasks and classification tasks, the F1 (F-measure) index of the information extraction model has been greatly improved. In addition, the base model and the large model are extremely capable in reading comprehension and natural language inference tasks.

[0224] Corresponding to the above embodiments of the task processing method, this specification also provides an embodiment of a task processing device, Figure 10 showing a schematic structural diagram of a task processing device provided by an embodiment of this specification. As Figure 10 shown, the device includes:

[0225] A first acquisition module 1002, configured to acquire the text to be processed and task description information of a target task;

[0226] A first construction module 1004, configured to extract the category information to be extracted from the task description information according to the task type of the target task, and construct target prompt information according to the category information to be extracted and the preset prompt format corresponding to the task type;

[0227] A first input module 1006, configured to splice the target prompt information and the text to be processed, and input the spliced information into an information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the corresponding relationship between the text to be processed and the target prompt information;

[0228] A first generation module 1008, configured to determine the task processing result of the target task according to the information extraction matrix.

[0229] Optionally, the target task includes a multi-level task; the first construction module 1004 is further configured to extract the category information to be extracted at the current level from the task description information according to the task type of the multi-level task, where, when the current level is not the first level, the category information to be extracted at the current level is obtained based on the task processing result of the previous level of the current level; construct target prompt information according to the category information to be extracted, the preset prompt format corresponding to the task type, and the task processing result of the previous level of the current level.

[0230] Optionally, the device further includes: a first splitting module, configured to split the target prompt information to obtain a plurality of target sub-prompt information when the length of the target prompt information is greater than a preset prompt length threshold, where the lengths of the plurality of target sub-prompt information are less than or equal to the preset prompt length threshold; the first input module 1006 is further configured to extract a first target sub-prompt information from the plurality of target sub-prompt information, where the first target sub-prompt information is the target sub-prompt information that has not been spliced with the text to be processed among the plurality of target sub-prompt information; splice the first target sub-prompt information and the text to be processed, and input the spliced information into the information extraction model until there is no target sub-prompt information that has not been spliced with the text to be processed among the plurality of target sub-prompt information, to obtain an information extraction matrix.

[0231] Optionally, the device further includes: a second splitting module configured to split the text to be processed to obtain multiple sub-texts to be processed when the length of the text to be processed is greater than a preset text length threshold, where the lengths of the multiple sub-texts to be processed are less than or equal to the preset text length threshold; a first input module 1006 further configured to extract a first sub-text to be processed from the multiple sub-texts to be processed, where the first sub-text to be processed is a sub-text to be processed that has not been concatenated with the target prompt information among the multiple sub-texts to be processed; concatenate the first sub-text to be processed and the target prompt information, and input the concatenated information into the information extraction model until there is no sub-text to be processed that has not been concatenated with the target prompt information among the multiple sub-texts to be processed, to obtain an information extraction matrix.

[0232] Optionally, the target task includes an extraction task; a first construction module 1004 further configured to obtain a preset prompt format corresponding to the extraction task, where the preset prompt format includes a category prompt symbol; concatenate the category information to be extracted and the category prompt symbol to obtain the target prompt information.

[0233] Optionally, the target task includes a classification task; a first construction module 1004 further configured to obtain a preset prompt format corresponding to the classification task, where the preset prompt format includes a category prompt symbol and a classification identifier; concatenate the category information to be extracted, the category prompt symbol, and the classification identifier to obtain the target prompt information.

[0234] Optionally, the information extraction model includes an encoding unit, a prompt separation unit, and a feature processing unit; a first input module 1006 further configured to input the concatenated information into the information extraction model, and through the prompt separation unit, perform position separation on the target prompt information to obtain a prompt separation matrix; through the encoding unit, encode the prompt separation matrix and the concatenated information to obtain an encoded feature matrix; through the feature processing unit, perform fusion processing on the prompt separation matrix and the encoded feature matrix to obtain an information extraction matrix.

[0235] Optionally, the device further includes: a model training module configured to obtain a sample set, where the sample set includes sample texts and sample task description information of multiple sample tasks, the sample texts carry sample matrix tags, and the sample matrix tags are obtained based on the sample processing results of the sample texts; for any sample task, extract sample extraction category information from the sample task description information according to the sample task type of the sample task, and construct sample prompt information according to the sample extraction category information and a preset prompt format corresponding to the sample task type; splice the sample prompt information and the sample text, and input the spliced sample information into an information extraction model to obtain a predicted information extraction matrix, where the predicted information extraction matrix represents the correspondence between the sample text and the sample prompt information; adjust the model parameters of the information extraction model according to the sample matrix tag and the predicted information extraction matrix to obtain a trained information extraction model.

[0236] Optionally, the first generation module 1008 is further configured to construct a text matrix according to the text to be processed and the target prompt information; extract the task processing result of the target task from the text matrix according to the information extraction matrix.

[0237] Optionally, the first generation module 1008 is further configured to convert the information extraction matrix to obtain a target information extraction matrix, where the target information extraction matrix represents the information extraction probability of the text to be processed; determine the task processing result of the target task according to the target information extraction matrix.

[0238] By applying the solution of the embodiment of this specification, by constructing the target prompt information based on the task type, the text processing task type can be prompted to the information extraction model by using the target prompt information, so that the information extraction model can process multiple types of text processing tasks, improving the expandability of the information extraction model and the generality of task processing. Moreover, since the task processing result is derived from the task description information and the text to be processed, the accuracy and stability of the task processing result are ensured.

[0239] The above is a schematic solution of a task processing device in this embodiment. It should be noted that the technical solution of this task processing device and the technical solution of the above task processing method belong to the same concept. For the details not described in detail in the technical solution of the task processing device, reference can be made to the description of the technical solution of the above task processing method.

[0240] Corresponding to the above embodiment of the information extraction model training method, this specification also provides an embodiment of an information extraction model training device. Figure 11 FIG. shows a schematic structural diagram of an information extraction model training device provided by an embodiment of this specification. As Figure 11 shown, this device is applied to a cloud-side device and includes:

[0241] The second acquisition module 1102 is configured to acquire a sample set, where the sample set includes sample texts and sample task description information of multiple sample tasks, the sample texts carry sample matrix tags, and the sample matrix tags are obtained based on the sample processing results of the sample texts;

[0242] The second construction module 1104 is configured to, for any sample task, extract sample extraction category information from the sample task description information according to the sample task type of the sample task, and construct sample prompt information according to the sample extraction category information and the preset prompt format corresponding to the sample task type;

[0243] The second input module 1106 is configured to splice the sample prompt information and the sample text, and input the spliced sample information into the information extraction model to obtain a predicted information extraction matrix, where the predicted information extraction matrix represents the correspondence between the sample text and the sample prompt information;

[0244] The adjustment module 1108 is configured to adjust the model parameters of the information extraction model according to the sample matrix tag and the predicted information extraction matrix to obtain a trained information extraction model.

[0245] Applying the solution of the embodiment of this specification, adjusting the model parameters of the information extraction model according to the sample matrix tag and the predicted information extraction matrix to obtain a trained information extraction model. By continuously adjusting the model parameters of the information extraction model, the finally obtained information extraction model can be made more accurate.

[0246] The above is a schematic solution of an information extraction model training device in this embodiment. It should be noted that the technical solution of this information extraction model training device and the technical solution of the above information extraction model training method belong to the same concept. For the details not described in the technical solution of the information extraction model training device, reference can be made to the description of the technical solution of the above information extraction model training method.

[0247] Corresponding to the above embodiment of the classification task processing method, this specification also provides an embodiment of a classification task processing device. Figure 12 The structural schematic diagram of a classification task processing device provided by an embodiment of this specification is shown. As Figure 12 shown, the device includes:

[0248] The third acquisition module 1202 is configured to acquire the text to be processed and task description information of the target classification task;

[0249] The third construction module 1204 is configured to extract the category information to be extracted from the task description information according to the task type of the target classification task, and construct target prompt information according to the category information to be extracted and the preset prompt format corresponding to the task type;

[0250] A third input module 1206, configured to splice the target prompt information and the text to be processed, and input the spliced information into the information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the correspondence between the text to be processed and the target prompt information;

[0251] A second generation module 1208, configured to determine the classification result of the target classification task according to the information extraction matrix.

[0252] Applying the solution of the embodiment of this specification, by constructing the target prompt information based on the task type, the target prompt information can be used to prompt the information extraction model that the task is a classification task, so that the information extraction model can process the classification task, improving the scalability of the information extraction model and the generality of task processing. Moreover, since the classification result is derived from the task description information and the text to be processed, the accuracy and stability of the classification result are ensured.

[0253] The above is a schematic solution of a classification task processing device in this embodiment. It should be noted that the technical solution of this classification task processing device and the technical solution of the above classification task processing method belong to the same concept. For the details not described in the technical solution of the classification task processing device, reference can be made to the description of the technical solution of the above classification task processing method.

[0254] Figure 13 The structural block diagram of a computing device provided by an embodiment of this specification is shown. The components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 through a bus 1330, and a database 1350 is used to store data.

[0255] The computing device 1300 also includes an access device 1340, which enables the computing device 1300 to communicate via one or more networks 1360. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Networks (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0256] In one embodiment of the present specification, the above components of the computing device 1300 and Figure 13 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 13 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.

[0257] The computing device 1300 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1300 can also be a mobile or stationary server.

[0258] Wherein, the processor 1320 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above task processing method or information extraction model training method or classification task processing method are implemented.

[0259] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above task processing method, information extraction model training method, and classification task processing method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the above task processing method, information extraction model training method, or classification task processing method.

[0260] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above task processing method, information extraction model training method, or classification task processing method are implemented.

[0261] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above task processing method, information extraction model training method, and classification task processing method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the above task processing method, information extraction model training method, or classification task processing method.

[0262] An embodiment of this specification also provides a computer program. When the computer program is executed on a computer, the computer is made to execute the steps of the above task processing method, information extraction model training method, or classification task processing method.

[0263] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solutions of the above task processing method, information extraction model training method, and classification task processing method belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the descriptions of the above task processing method, information extraction model training method, or classification task processing method.

[0264] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.

[0265] The computer instructions include computer program code, which may be in the form of source code, object code, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0266] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.

[0267] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0268] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A task processing method, comprising: obtaining the text to be processed and task description information of a target task; extracting category information to be extracted from the task description information according to the task type of the target task, and constructing target prompt information according to the category information to be extracted and a preset prompt format corresponding to the task type; concatenating the target prompt information and the text to be processed, and inputting the concatenated information into an information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the correspondence between the text to be processed and the target prompt information; determining a task processing result of the target task according to the information extraction matrix.

2. The method according to claim 1, wherein the target task comprises a multi-level task; the extracting category information to be extracted from the task description information according to the task type of the target task comprises: extracting category information to be extracted at the current level from the task description information according to the task type of the multi-level task, where when the current level is not the first level, the category information to be extracted at the current level is obtained based on the task processing result of the previous level of the current level; the constructing target prompt information according to the category information to be extracted and a preset prompt format corresponding to the task type comprises: constructing target prompt information according to the category information to be extracted, the preset prompt format corresponding to the task type, and the task processing result of the previous level of the current level.

3. The method according to claim 1, after constructing the target prompt information according to the category information to be extracted and a preset prompt format corresponding to the task type, further comprises: when the length of the target prompt information is greater than a preset prompt length threshold, splitting the target prompt information to obtain a plurality of target sub-prompt information, where the lengths of the plurality of target sub-prompt information are less than or equal to the preset prompt length threshold; the concatenating the target prompt information and the text to be processed, and inputting the concatenated information into an information extraction model to obtain an information extraction matrix comprises: extracting first target sub-prompt information from the plurality of target sub-prompt information, where the first target sub-prompt information is the target sub-prompt information that has not been concatenated with the text to be processed among the plurality of target sub-prompt information; concatenating the first target sub-prompt information and the text to be processed, and inputting the concatenated information into the information extraction model until there is no target sub-prompt information that has not been concatenated with the text to be processed among the plurality of target sub-prompt information, to obtain an information extraction matrix.

4. The method according to claim 1, before concatenating the target prompt information and the text to be processed, and inputting the concatenated information into an information extraction model to obtain an information extraction matrix, further comprises: when the length of the text to be processed is greater than a preset text length threshold, splitting the text to be processed to obtain a plurality of sub-texts to be processed, where the lengths of the plurality of sub-texts to be processed are less than or equal to the preset text length threshold; Splicing the target prompt information and the text to be processed, and inputting the spliced information into an information extraction model to obtain an information extraction matrix, including: Extracting a first text to be processed from the multiple texts to be processed, where the first text to be processed is a text to be processed that has not been spliced with the target prompt information among the multiple texts to be processed; Splicing the first text to be processed and the target prompt information, and inputting the spliced information into the information extraction model until there is no text to be processed that has not been spliced with the target prompt information among the multiple texts to be processed, to obtain an information extraction matrix.

5. According to the method described in claim 1, the target task includes an extraction task; Constructing the target prompt information according to the preset prompt format corresponding to the information to be extracted category and the task type, including: Obtaining the preset prompt format corresponding to the extraction task, where the preset prompt format includes a category prompt symbol; Splicing the information to be extracted category and the category prompt symbol to obtain the target prompt information.

6. According to the method described in claim 1, the target task includes a classification task; Constructing the target prompt information according to the preset prompt format corresponding to the information to be extracted category and the task type, including: Obtaining the preset prompt format corresponding to the classification task, where the preset prompt format includes a category prompt symbol and a classification identifier; Splicing the information to be extracted category, the category prompt symbol and the classification identifier to obtain the target prompt information.

7. According to the method described in claim 1, the information extraction model includes an encoding unit, a prompt separation unit and a feature processing unit; Inputting the spliced information into the information extraction model to obtain an information extraction matrix, including: Inputting the spliced information into the information extraction model, and through the prompt separation unit, separating the position of the target prompt information to obtain a prompt separation matrix; Through the encoding unit, encoding the prompt separation matrix and the spliced information to obtain an encoded feature matrix; Through the feature processing unit, performing a fusion process on the prompt separation matrix and the encoded feature matrix to obtain an information extraction matrix.

8. Before splicing the target prompt information and the text to be processed, and inputting the spliced information into the information extraction model to obtain an information extraction matrix, it also includes: Obtaining a sample set, where the sample set includes sample texts and sample task description information of multiple sample tasks, the sample texts carry sample matrix labels, and the sample matrix labels are obtained based on the sample processing results of the sample texts; For any sample task, extracting sample extraction category information from the sample task description information according to the sample task type of the sample task, and constructing sample prompt information according to the sample extraction category information and the preset prompt format corresponding to the sample task type; Concatenate the sample prompt information and the sample text, and input the concatenated sample information into the information extraction model to obtain a predicted information extraction matrix, where the predicted information extraction matrix represents the correspondence between the sample text and the sample prompt information; Adjust the model parameters of the information extraction model according to the sample matrix label and the predicted information extraction matrix to obtain a trained information extraction model.

9. The method according to claim 1, wherein determining the task processing result of the target task according to the information extraction matrix, comprises: Construct a text matrix according to the text to be processed and the target prompt information; Extract the task processing result of the target task from the text matrix according to the information extraction matrix.

10. The method according to claim 1, wherein determining the task processing result of the target task according to the information extraction matrix, comprises: Convert the information extraction matrix to obtain a target information extraction matrix, where the target information extraction matrix represents the information extraction probability of the text to be processed; Determine the task processing result of the target task according to the target information extraction matrix.

11. An information extraction model training method applied to a cloud-side device, comprises: Obtain a sample set, where the sample set includes the sample texts and sample task description information of multiple sample tasks, the sample texts carry sample matrix labels, and the sample matrix labels are obtained based on the sample processing results of the sample texts; For any sample task, extract sample extraction category information from the sample task description information according to the sample task type of the sample task, and construct sample prompt information according to the sample extraction category information and a preset prompt format corresponding to the sample task type; Concatenate the sample prompt information and the sample text, and input the concatenated sample information into the information extraction model to obtain a predicted information extraction matrix, where the predicted information extraction matrix represents the correspondence between the sample text and the sample prompt information; Adjust the model parameters of the information extraction model according to the sample matrix label and the predicted information extraction matrix to obtain a trained information extraction model.

12. A classification task processing method, comprises: Obtain the text to be processed and task description information of the target classification task; Extract the category information to be extracted from the task description information according to the task type of the target classification task, and construct target prompt information according to the category information to be extracted and a preset prompt format corresponding to the task type; Concatenate the target prompt information and the text to be processed, and input the concatenated information into the information extraction model to obtain an information extraction matrix, where the information extraction matrix represents the correspondence between the text to be processed and the target prompt information; Determine the classification result of the target classification task according to the information extraction matrix.

13. A computing device, comprises: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10, or claim 11, or claim 12 are implemented.

14. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10, or claim 11, or claim 12.