Task processing method and apparatus, question answering processing method and apparatus in target domain, domain task model test method and apparatus, and computing device, computer-readable storage medium and computer program product
Patent Information
- Application Number
- PCT/CN2025/076840
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-04
- Filing Date
- 2025-02-11
- Publication Date
- 2025-10-02
AI Technical Summary
In existing technologies, after fine-tuning a large model for a specific field, the evaluation method is one-sided and cannot fully assess the model's chain thinking ability in the specific field, resulting in poor iteration and tuning effects.
By obtaining the task data of the target task, dividing it into thinking subtasks and decision-making subtasks, gradually inputting it into the domain task model, and performing chain thinking, the model's thinking analysis and task processing capabilities are improved.
The accuracy of target task results has been improved. By adjusting the model through closed-loop training, the model's chain thinking ability and task processing effect in professional fields have been improved.
Smart Images

Figure CN2025076840_02102025_PF_FP_ABST
Abstract
Description
Task processing, question-answering processing in a target domain, domain task model testing method and device, computing device, computer-readable storage medium, and computer program product
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on March 4, 2024, with application number 202410245362.0, and invention name “Task processing, question and answer processing under target domain, domain task model testing method and device, computing equipment, computer-readable storage medium, and computer program product”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of this specification relate to the field of computer technology, and in particular to a method and apparatus for task processing, question-and-answer processing in a target domain, and domain task model testing, a computing device, a computer-readable storage medium, and a computer program product. Background Art
[0003] In practical applications, after training a large model, the large model can also be fine-tuned for specific fields or scenarios, so that the large model has the ability to reason about knowledge in professional fields.
[0004] However, currently, large models that have undergone fine-tuning are often evaluated based on model perplexity or their knowledge reasoning capabilities based on specific datasets. These evaluation methods can easily lead to one-sided evaluation results and fail to fully assess the model's reasoning capabilities in specialized fields. Summary of the Invention
[0005] In view of this, embodiments of this specification provide a task processing method. One or more embodiments of this specification also relate to a question-and-answer processing method in a target domain, a domain task model testing method, a task processing device, a question-and-answer processing device in a target domain, a domain task model testing device, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0006] According to a first aspect of an embodiment of this specification, a task processing method is provided, including:
[0007] Acquiring task data of a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask;
[0008] Input the task data into the domain task model, perform the thinking subtask, and obtain the target thinking result of the domain task model for the task data;
[0009] The target thinking result is input into the domain task model, the decision subtask is executed, and the target task result corresponding to the target task is obtained. The domain task model is trained based on the test results of the thinking subtask and the decision subtask.
[0010] According to a second aspect of the embodiments of this specification, a question-answering processing method in a target domain is provided, including:
[0011] Receive problem information in the target area sent by front-end users;
[0012] Input the problem information into the domain model of the target domain, execute the problem thinking subtask, and obtain the first thinking result of the domain model for the problem information;
[0013] Input the first thinking result into the domain model, execute the component decision subtask, and obtain the component information of the target component;
[0014] Based on the component information, the target component is called to process the problem information and obtain the component output information;
[0015] Input the component output information into the domain model, execute the component output thinking subtask, and obtain the second thinking result of the domain model for the component output information;
[0016] Input the second thinking result into the domain model, execute the question-answering decision subtask, and obtain the answer information, wherein the domain model is trained based on the test results of the question-thinking subtask, component decision subtask, component output thinking subtask, and question-answering decision subtask;
[0017] Feedback the answer information to the front-end user.
[0018] According to a third aspect of the embodiments of this specification, a domain task model testing method is provided, including:
[0019] Obtain a test set and a domain task model, where the test set includes test pairs, each test pair includes domain test data of the target domain, and label thinking results and label task results corresponding to the domain test data. The domain task model is pre-trained based on sample data of the target domain.
[0020] Input the domain test data into the domain task model, perform the thinking subtask, obtain the predicted thinking results, and generate thinking test indicators based on the predicted thinking results and the labeled thinking results;
[0021] Input the prediction thinking results into the domain task model, execute the decision subtask, obtain the prediction task results, and generate decision test indicators based on the prediction task results and labeling task results;
[0022] Based on thinking test indicators and decision test indicators, the test results of the domain task model are obtained.
[0023] According to a fourth aspect of the embodiments of this specification, there is provided a task processing device, including:
[0024] A first acquisition module is configured to acquire task data of a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision subtask;
[0025] The first input module is configured to input task data into the domain task model, perform a thinking subtask, and obtain a target thinking result of the domain task model for the task data;
[0026] The first execution module is configured to input the target thinking result into the domain task model, execute the decision subtask, and obtain the target task result corresponding to the target task, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask.
[0027] According to a fifth aspect of the embodiments of this specification, a question-answering processing apparatus in a target domain is provided, comprising:
[0028] A receiving module is configured to receive problem information in a target field sent by a front-end user;
[0029] The second input module is configured to input the problem information into the domain model of the target domain, perform the problem thinking subtask, and obtain the first thinking result of the domain model for the problem information;
[0030] A second execution module is configured to input the first thinking result into the domain model, execute the component decision subtask, and obtain component information of the target component;
[0031] The calling module is configured to call the target component based on the component information to process the problem information and obtain the component output information;
[0032] The third input module is configured to input the component output information into the domain model, execute the component output thinking subtask, and obtain the second thinking result of the domain model for the component output information;
[0033] a third execution module configured to input the second thinking result into a domain model, execute the question-answering decision subtask, and obtain answer information, wherein the domain model is trained based on the test results of the question-thinking subtask, the component decision subtask, the component output thinking subtask, and the question-answering decision subtask;
[0034] The feedback module is configured to feed back answer information to the front-end user.
[0035] According to a sixth aspect of the embodiments of this specification, a domain task model testing device is provided, including:
[0036] A second acquisition module is configured to acquire a test set and a domain task model, wherein the test set includes test pairs, the test pairs include domain test data of the target domain, and labeled thinking results and labeled task results corresponding to the domain test data, and the domain task model is pre-trained based on sample data of the target domain;
[0037] a fourth input module configured to input the domain test data into the domain task model, execute the thinking subtask, obtain the predicted thinking result, and generate the thinking test indicator according to the predicted thinking result and the labeled thinking result;
[0038] The fourth execution module is configured to input the prediction thinking results into the domain task model, execute the decision subtask, obtain the prediction task results, and generate decision test indicators based on the prediction task results and the labeling task results;
[0039] The result generation module is configured to obtain the test results of the domain task model based on the thinking test indicators and the decision test indicators.
[0040] According to a seventh aspect of the embodiments of this specification, a computing device is provided, including:
[0041] memory and processor;
[0042] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above method are implemented.
[0043] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and the steps of the above method are implemented when the instructions are executed by a processor.
[0044] According to a ninth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0045] One embodiment of the present specification implements the acquisition of task data for a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask; inputting the task data into a domain task model, executing the thinking subtask, and obtaining the target thinking result of the domain task model for the task data; inputting the target thinking result into the domain task model, executing the decision subtask, and obtaining the target task result corresponding to the target task, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask. By inputting the task data into the domain task model and executing the thinking subtask, the target thinking result of the domain task model for the task data can be obtained; by inputting the target thinking result into the domain task model again and executing the decision subtask, the target task result corresponding to the target task can be obtained; through gradual interaction with the model, the model's thinking and analysis capabilities and task processing capabilities for the target task can be improved based on the model's chain thinking, thereby improving the accuracy of the target task result. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] FIG1 is an architecture diagram of a task processing system provided by one embodiment of this specification;
[0047] FIG2 is a flowchart of a task processing method provided by one embodiment of this specification;
[0048] FIG3 is a flowchart of a question-answering processing method in a target domain provided by one embodiment of this specification;
[0049] FIG4 is a flowchart of a domain task model testing method provided by one embodiment of this specification;
[0050] FIG5 is a schematic diagram of a processing process of a task processing method provided by an embodiment of this specification;
[0051] FIG6 is a schematic diagram of a process for generating domain evaluation data of a task processing method provided by one embodiment of this specification;
[0052] FIG7 is a schematic diagram of a model evaluation process of a task processing method provided by one embodiment of this specification;
[0053] FIG8 is a schematic diagram of an evaluation result analysis process of a task processing method provided in one embodiment of this specification;
[0054] FIG9 is a schematic diagram of the structure of a task processing device provided by one embodiment of this specification;
[0055] FIG10 is a schematic diagram of the structure of a question-answering processing device in a target domain provided by one embodiment of this specification;
[0056] FIG11 is a schematic diagram of the structure of a domain task model testing device provided by one embodiment of this specification;
[0057] FIG12 is a structural block diagram of a computing device provided in one embodiment of this specification. DETAILED DESCRIPTION
[0058] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0059] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0060] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0061] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0062] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. It is pre-trained on a large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.
[0063] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0064] First, the terms involved in one or more embodiments of this specification are explained.
[0065] CEVAL: is a comprehensive Chinese-based model evaluation suite that includes questions from multiple professional fields and at various difficulty levels for assessing LLM capabilities.
[0066] CMMUL: is a comprehensive Chinese assessment dataset that contains questions from multiple professional fields and at various difficulty levels. It is used to assess LLM knowledge and reasoning ability in the Chinese language and context.
[0067] Large Language Model (LLM): A large language model (LLM) is a natural language processing model based on neural networks. It has strong understanding, high cognitive and generalization capabilities. It can handle a variety of natural language tasks, including text classification, text generation, and machine translation.
[0068] Domain model: In this specification, it refers to a large model that has been pre-trained with domain knowledge or fine-tuned with instructions.
[0069] Chain of Thought (COT): This capability emerges when a large model reaches a certain number of parameters. By gradually interacting with the model and providing prompts for its judgment and execution, it can guide the model's thinking and execution.
[0070] Fine-tuning is a technique used in deep learning to apply a pre-trained model to a specific task or domain. The basic idea of fine-tuning is to take a pre-trained model that has been trained on a large amount of data and then continue training it on a small amount of specific data, hoping to achieve good results.
[0071] GPT (Generative Pre-training Transformer): is a pre-training generative model based on the Transformer architecture, which has been widely used in speech recognition, machine translation, language generation and other fields.
[0072] Perplexity: is a metric used to measure the predictive ability of a language model. It measures its performance by measuring the uncertainty of the model's predicted probability distribution for a given sequence.
[0073] BLEU (Bilingual Evaluation Understudy): is a reference text-driven evaluation metric used to evaluate the similarity between generated text and reference text.
[0074] ROUGE (Recall-Oriented Understudy for Gisting Evaluation): This is a recall-based metric that compares the word-level and sentence-level similarity between a generated summary and a reference summary. It can be used to assess the quality of tasks such as text summarization, machine translation, and text generation. Its score ranges from 0 to 1, with scores closer to 1 indicating a high degree of similarity between the generated summary and the reference summary.
[0075] Precision: Accuracy measures the proportion of samples predicted by the model as positive that are actually positive. A higher accuracy indicates a lower false positive rate. Accuracy = True Positives / (True Positives + False Positives).
[0076] Recall: Recall measures the proportion of samples that the model correctly identifies as positive out of all true positive examples. A higher recall indicates a lower rate of missed detections. Recall = True Positives / (True Positives + False Negatives).
[0077] F1 score: The F1 score is a comprehensive evaluation metric for precision and recall. It is the harmonic mean of precision and recall. A higher F1 score indicates better overall model performance. F1 = 2 * (Precision * Recall) / (Precision + Recall).
[0078] In practical applications, due to a lack of specialized domain training data and differing interpretations of specific terms within and outside the domain, large, general-purpose models that haven't been fine-tuned for specialized domains struggle to demonstrate excellent chained thinking capabilities within specialized domains. With the rise of LLMs and the discovery of their chained thinking capabilities, it's become possible to apply domain-tuned LLMs to specific scenarios, allowing the models to perform thinking and decision-making to solve problems in a variety of specialized scenarios.
[0079] However, after pre-training and fine-tuning the model's instructions in specialized fields, the knowledge reasoning capabilities of large models are often evaluated based on model perplexity or specific datasets (such as question-answer pairs in specialized fields). These evaluation methods can only assess individual indicators and fail to take into account the overall process of the model's chain thinking. This can easily lead to one-sided evaluation results and an inability to fully and accurately assess the model's chain thinking capabilities in specialized fields. This makes it difficult to effectively iterate and tune the model based on the evaluation results, resulting in poor performance on the target tasks in specialized fields.
[0080] Based on this, an embodiment of the present specification provides a task processing method, which obtains task data of a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision subtask; inputs the task data into a domain task model, executes the thinking subtask, and obtains the target thinking result of the domain task model for the task data; inputs the target thinking result into the domain task model, executes the decision subtask, and obtains the target task result corresponding to the target task, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask. By inputting the task data into the domain task model and executing the thinking subtask, the target thinking result of the domain task model for the task data can be obtained. By inputting the target thinking result into the domain task model again and executing the decision subtask, the target task result corresponding to the target task can be obtained. Through gradual interaction with the model, the model's thinking and analysis capabilities and task processing capabilities for the target task can be improved based on the model's chain thinking, thereby improving the accuracy of the target task result.
[0081] In this specification, a task processing method is provided. This specification also involves a question and answer processing method in a target domain, a domain task model testing method, a task processing device, a question and answer processing device in a target domain, a domain task model testing device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.
[0082] 1 , which shows an architecture diagram of a task processing system according to an embodiment of the present disclosure, specifically, the task processing system 100 includes a client 102 and a server 104 , wherein the server 104 includes a domain task model 1042 .
[0083] The client 102 is used to send the task data of the target task to the server 104 .
[0084] Server 104: used to obtain task data of a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask; input the task data into the domain task model, execute the thinking subtask, and obtain the target thinking result of the domain task model for the task data; input the target thinking result into the domain task model, execute the decision subtask, and obtain the target task result corresponding to the target task, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask.
[0085] The client 102 is also used to receive the target task result returned by the server 104.
[0086] In practical applications, the task processing system may include multiple clients 102 and a server 104. Each of the multiple clients 102 may establish communication connections with the server 104. In the task processing scenario, the server 104 is used to obtain task data for a target task sent by each client 102, process the target task, and return the target task result corresponding to the target task to each client 102.
[0087] The client 102 and the server 104 are connected via a network. The network provides a medium for the communication link between the client 102 and the server 104. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 102 may need to be encoded, transcoded, compressed, and other processing before being released to the server 104.
[0088] The client 102 can be deployed in an electronic device and rely on the device or certain apps in the device to run. For example, the electronic device may have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, tablet computer, personal computer, or other end-side device. Various other types of applications can also be configured in the electronic device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0089] The server 104 may include servers that provide various services, such as servers that provide communication services, servers that support background training for models, and servers that process data sent by the client 102. It should be noted that the server 104 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server that integrates a blockchain. The server can also be a cloud server (cloud-side device) that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0090] It is worth noting that the task processing methods provided in the embodiments of this specification are generally executed by the server 104. However, in other embodiments of this specification, the client 102 may also have similar functions to the server and thus execute the task processing methods provided in the embodiments of this specification. In other embodiments, the task processing methods provided in the embodiments of this specification may also be executed jointly by the client 102 and the server.
[0091] In actual applications, a front-end user can send the task data of a target task to the server 104 via the client 102. The server 104 receives the task data, inputs it into the domain task model 1042, and obtains the target thinking result for the task data, which is output by the domain task model 1042 after executing the thinking subtask. The server 104 then inputs the target thinking result into the domain task model 1042 again, and obtains the target task result corresponding to the target task, which is output by the domain task model 1042 after executing the decision subtask. The server 104 returns the target task result to the client 102.
[0092] Specifically, the domain task model 1042 can be applied to a variety of application scenarios across various domains, such as institutional service scenarios like sending emails, booking hotels, searching and arranging travel itineraries, or question-and-answer scenarios. Furthermore, the domain task model 1042 can complete the processing of target tasks through chained thinking, leveraging the model's inherent analytical and decision-making capabilities.
[0093] For example, the target task may be to answer a math multiple-choice question, the task data may be the title information of the multiple-choice question, and the target task result may be the options corresponding to the multiple-choice question.
[0094] In the embodiments of this specification, the client sends the task data of the target task to the server, which receives the task data and inputs it into the domain task model. The domain task model obtains the target thinking result for the task data, which is output by the thinking subtask. The target thinking result is then input into the domain task model to obtain the target task result corresponding to the target task output by the decision subtask. The target task result is then returned to the client. By gradually interacting with the model and based on model chain thinking, the model's thinking and analysis capabilities for the target task and its task processing capabilities can be improved, thereby improving the accuracy of the target task results.
[0095] Referring to FIG. 2 , FIG. 2 shows a flowchart of a task processing method provided according to an embodiment of this specification, which specifically includes the following steps.
[0096] Step 202: Obtain task data of a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask.
[0097] In practical applications, the task data of the target task can be obtained, and the target task can be processed based on the task data.
[0098] Specifically, the target task can be understood as a task corresponding to any scenario in the target professional field. For example, the professional field can be the humanities field, the natural field, the scientific field, etc., or it can be a more subdivided field, such as Chinese, mathematics, English, traditional Chinese medicine, computer, engineering and architecture, etc. It should be noted that the above is only an example of the division of professional fields, which can actually be determined according to the needs of the specific application process, and this manual does not impose any restrictions on this.
[0099] Specifically, each professional field may also include one or more application scenarios. For example, the computer field may include software development, hardware development, database maintenance, project requirements analysis, artificial intelligence, and other different scenarios.
[0100] Specifically, task data can be understood as task data corresponding to the target task. Task data can include task description information of the target task, task information of the task to be processed itself, prompt information of the target task, and so on. The target task can include at least one task stage, and each task stage can include a thinking subtask and a decision-making subtask. Furthermore, a task stage can correspond to a chain thinking process of the model, that is, a task stage includes multiple interaction processes with the model. Through step-by-step interaction with the model, the model can complete thinking and decision-making on the target task. Among them, the thinking subtask can be understood as a subtask in which the model performs thinking analysis on the input data and outputs the thinking analysis results; the decision-making subtask can be understood as a subtask in which the model performs decision-making on the thinking analysis results and outputs the target task results.
[0101] Exemplarily, when the target task includes a task stage, the task data corresponding to the target task may be task information related to the task to be processed, and the target task result corresponding to the target task may be the task execution result of the task to be processed.
[0102] For example, when the target task includes two task stages, the target task may include a tool identification task stage and a question-answering decision-making task stage. In the tool identification task stage, the task data corresponding to the target task may be a question raised by the user regarding the target professional field, and the target task result corresponding to the target task may be the tool required to solve the problem and the input parameters of the tool. In the question-answering decision-making task stage, the task data corresponding to the target task may be the tool processing result obtained based on the above tool processing, and the target task result corresponding to the target task may be the answer to the question.
[0103] Step 204: Input the task data into the domain task model, execute the thinking subtask, and obtain the target thinking result of the domain task model for the task data.
[0104] In actual applications, after obtaining the task data corresponding to the target task, the task data can be input into the domain task model, and the thinking subtasks can be executed to obtain the target thinking results of the domain task model for the task data.
[0105] Specifically, the domain task model can be understood as a large model that performs domain knowledge pre-training or instruction fine-tuning for the target domain. It can also be understood as a general large model that can handle tasks in multiple domains. A thinking subtask can be understood as a subtask in which the model analyzes and thinks about the task data. The target thinking result can be understood as the data output by the domain task model after analyzing and thinking about the task data. The target thinking result can be output as a textual representation.
[0106] Step 206: Input the target thinking result into the domain task model, execute the decision subtask, and obtain the target task result corresponding to the target task, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask.
[0107] In practical applications, based on the target thinking results, the target thinking results can be further input into the domain task model, and the decision subtasks can be executed to obtain the target task results corresponding to the target task.
[0108] Specifically, the decision-making subtask can be understood as a subtask in which the model makes a problem result decision for the target task based on the target thinking result. The target task result can be understood as the task processing result corresponding to the target task. In the case where the target task is a question, the target task result can be the answer information corresponding to the question; in the case where the target task is a task to be processed in a certain institutional service scenario, the target task result can be the feedback information after the service is completed, or the corresponding feedback information when the service is not completed. The test result of the thinking subtask can be understood as the evaluation result obtained by evaluating the task execution of the thinking subtask executed by the domain task model. The test result of the decision subtask can be understood as the evaluation result obtained by evaluating the task execution of the decision subtask executed by the domain task model.
[0109] Applying the embodiments of this specification, task data of a target task is obtained, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask; the task data is input into a domain task model, the thinking subtask is executed, and the target thinking result of the domain task model for the task data is obtained; the target thinking result is input into the domain task model, the decision subtask is executed, and the target task result corresponding to the target task is obtained, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask. By inputting the task data into the domain task model and executing the thinking subtask, the target thinking result of the domain task model for the task data can be obtained, and by inputting the target thinking result into the domain task model again and executing the decision subtask, the target task result corresponding to the target task can be obtained; through gradual interaction with the model, the model's thinking and analysis capabilities and task processing capabilities for the target task can be improved based on the model's chain thinking, thereby improving the accuracy of the target task result.
[0110] Furthermore, in order to improve the domain task model's ability to learn and reason about professional domain knowledge, thereby improving the model's task processing capabilities, each time the domain task model is trained, the domain task model's chain thinking ability in the professional domain can be evaluated, and the training of the domain task model can be adjusted based on the evaluation results. Among them, the test results of the thinking subtask can reflect the model's thinking and analysis capabilities during the chain thinking process, and the test results of the decision-making subtask can reflect the model's problem parameter identification capabilities, that is, decision-making capabilities, during the chain thinking process. Training the domain task model based on the test results of the thinking subtask and the decision-making subtask can break through the boundary between model evaluation and model training, forming a closed loop of model training-effect evaluation-training adjustment, thereby achieving effective iteration and tuning of the domain task model and improving the chain thinking capabilities of the domain task model in the professional domain.
[0111] Based on this, in an optional embodiment of this specification, before inputting the task data into the domain task model, the following steps S2002-S2008 may also be included:
[0112] S2002: Obtain a test set and a domain task model, wherein the test set includes test pairs, the test pairs include domain test data of the target domain, and label thinking results and label task results corresponding to the domain test data, and the domain task model is pre-trained based on sample data of the target domain.
[0113] In practical applications, before inputting task data into the domain task model to process the target task and obtain the target task results, you can also obtain a test set and domain task model. Based on the test set, you can evaluate the domain task model's chain thinking ability in the professional field and adjust the domain task model based on the evaluation results. Then, you can input the task data into the adjusted domain task model. Based on the domain task model adjusted according to the evaluation results, you can think about and make decisions about the target task and obtain the target task results.
[0114] Specifically, the domain task model is pre-trained based on sample data from the target domain. This sample data can include training data corresponding to various scenarios within the target domain, specifically question-answer pairs. Furthermore, the domain task model can be a generative model obtained by pre-training a general large model or fine-tuning its metrics based on sample data from the target domain.
[0115] In an optional implementation of the present specification, the sample data may be obtained based on currently available public datasets (such as CEVAL, CMMUL, etc.), or may be obtained based on paid third-party datasets.
[0116] In another optional implementation of the present specification, the sample data may also be obtained by manually labeling the task data in the target domain.
[0117] Specifically, a test set may include one or more test pairs, each of which may include domain test data from the target domain, as well as labeled thinking results and labeled task results corresponding to the domain test data. The target domain can be understood as the domain corresponding to the target task. Domain test data can be understood as test data within the target domain, which can correspond to different specific scenarios within the target domain. Domain test data can be understood as test data used to assess the domain task model's chained thinking ability within the target domain, and may include task information, task description data, task prompts, and so on. For example, domain test data can be questions from any scenario within the target domain, such as multiple-choice questions or application questions; or task information corresponding to performing a specific service, such as sending an email, booking a hotel, or searching and arranging a travel itinerary. Labeled thinking results can be understood as the thinking results annotated on the domain test data. Labeled thinking results can be used to represent the expected output of the domain task model's thinking and analysis of the domain test data. Labeled thinking results can include thinking results on the task data and thinking results on the tool processing results. Exemplarily, when the target task includes at least two task stages, the label thinking result corresponding to the first task stage may be a thinking result on the task data, and the label thinking result corresponding to the second task stage may be a thinking result on the tool processing result.
[0118] In an optional implementation of the present specification, the label thinking result can be obtained based on manual labeling of the domain test data.
[0119] In another optional implementation of the present specification, the tag thinking result may also be obtained based on semantic generalization of the labeled tag thinking result.
[0120] Specifically, the labeling task result can be used to represent the expected output task result of the domain task model's execution of the decision subtask. The labeling task result can include tool parameter information or task processing results. Tool parameter information can include tool identification information and the input parameter information corresponding to the tool. For example, when the target task includes at least two task stages, the labeling task result corresponding to the first task stage can be the tool parameter information, and the labeling task result corresponding to the second task stage can be the task processing result.
[0121] In an optional implementation of the present specification, the labeling task results can be obtained based on currently available public datasets (such as CEVAL, CMMUL, etc.), or based on paid third-party datasets.
[0122] In another optional implementation of the present specification, the labeling task result may also be obtained based on manual labeling of the domain test data.
[0123] In another optional implementation of the present specification, the labeling task result may also be obtained by enumerating and generalizing the labeled labeling task results.
[0124] It should be noted that the target domain sample data can include training data for pre-training and fine-tuning the large model in the target domain, as well as test data for evaluating the chained thinking ability of the pre-trained domain task model. The specific sample data used as training data and test data can be determined based on actual application needs. The target domain sample data is labeled data.
[0125] Optionally, in one embodiment of this specification, obtaining a test set may include the following steps:
[0126] Acquire a sample set, wherein the sample set includes a first sample pair, and the first sample pair includes domain sample data of a target domain, a label thinking result corresponding to the domain sample data, and a label task result;
[0127] Performing generalization processing on the first sample pair to generate multiple second sample pairs;
[0128] A test set is obtained according to the first sample pair and the second sample pair.
[0129] Specifically, the sample set can be understood as a data set consisting of sample data in the target domain, wherein the sample set includes a first sample pair, which includes domain sample data in the target domain, label thinking results corresponding to the domain sample data, and label task results.
[0130] Optionally, the domain sample data and corresponding labeling task results in the first sample pair can be obtained from existing public datasets or manually annotated. The domain sample data and corresponding labeling task results can be in the form of question-answer pairs. The labeling thinking results in the first sample pair can be obtained by manually annotating the domain sample data.
[0131] Specifically, the second sample pairs can be understood as sample pairs in the target domain obtained by generalizing the first sample pairs. The number of the second sample pairs can be greater than the number of the first sample pairs.
[0132] In actual implementation, after obtaining the second sample pair, the second sample pair can be added to the sample set to enrich the sample data in the target domain. Furthermore, the proportion of sample data used in the training phase and the test phase can be determined based on the actual application situation to obtain the test set.
[0133] By applying this embodiment, by obtaining a sample set, generalizing the first sample pair, generating multiple second sample pairs, and obtaining a test set based on the first sample pairs and the second sample pairs, it is possible to generalize and obtain rich sample data based on a small amount of labeled data, thereby improving the richness of semantics and parameters and providing better data support for model training and evaluation.
[0134] Optionally, in one embodiment of the present specification, performing generalization processing on the first sample pair to generate multiple second sample pairs may include the following steps:
[0135] Input the domain sample data and the label thinking result in the first sample pair into the language processing model, perform the synonymous rewriting task, and obtain the target domain sample data that is synonymous with the domain sample data, and the target label thinking result that is synonymous with the label thinking result;
[0136] Based on the target domain sample data, the target label thinking results and the label task results, a second sample pair is generated.
[0137] Specifically, the synonym rewriting task can be understood as a task of giving appropriate prompt words and using the semantic understanding and text generation capabilities of the language processing model to rewrite the sentence synonymously. The language processing model can include a general large model.
[0138] Optionally, the synonym rewriting task can be performed not only through the model, but also through traditional word order rewriting, synonym replacement and other methods.
[0139] Specifically, the target domain sample data can be understood as sample data that is synonymous with the domain sample data after synonymous rewriting. The target label thinking result can be understood as a thinking result that is synonymous with the label thinking result after synonymous rewriting.
[0140] By applying this embodiment, by performing the synonym rewriting task on the domain sample data and label reflection results in the first sample pair, it is possible to generalize and enhance a small amount of labeled data, thereby obtaining a larger amount of sample pairs, thereby providing better data support for model training and evaluation. Furthermore, by performing the synonym rewriting task on the language processing model, the efficiency of synonym rewriting can be improved, and more generalized data can be obtained.
[0141] Optionally, before generating the second sample pair based on the target domain sample data, target label thinking results and label task results, an enumeration replacement task can also be performed to enumerate and replace parameters of the label task results in the first sample pair, thereby obtaining more label task results.
[0142] It should be noted that while performing the synonym rewriting task can generate target domain sample data and target label thinking results that are semantically consistent with the domain sample data and label thinking results, we cannot rule out the possibility that the newly generated text may contain missing, tampered, or added parameters in specific domain scenarios.
[0143] Based on this, in an optional embodiment of the present specification, after inputting the domain sample data and the label thinking result in the first sample pair into the language processing model, performing the synonymous rewriting task, and obtaining the target domain sample data synonymous with the domain sample data and the target label thinking result synonymous with the label thinking result, the following may also be included:
[0144] Perform parameter consistency identification tasks, perform parameter consistency identification on parameters in domain sample data and target domain sample data, and perform parameter consistency identification on parameters in label thinking results and target label thinking results.
[0145] Optionally, the parameter consistency recognition task can be performed by GPT (generative language model), or NLP (deep learning model), or traditional statistical learning model.
[0146] By applying this embodiment, by performing the parameter consistency identification task, the problem of parameter inconsistency in the synonym generalization results can be discovered in a timely manner and handled accordingly, thereby improving the parameter stability before and after synonym generalization, thereby avoiding negative impacts on the stability of the domain task model.
[0147] Optionally, in one embodiment of the present specification, generating a second sample pair based on the target domain sample data, the target label thinking result, and the label task result may include the following steps:
[0148] Identify whether the semantic parameters of the target domain sample data and the target label thinking results are consistent;
[0149] When the semantic parameters are consistent, a second sample pair is generated based on the target domain sample data, the target label thinking results and the label task results.
[0150] In practical applications, since the semantic parameter information between any two target domain sample data and target label thinking results obtained through synonymous rewriting may not be consistent, it is possible to identify whether the semantic parameters of the target domain sample data and the target label thinking results are consistent. If the semantic parameters are consistent, a second sample pair is constructed based on the target domain sample data, the target label thinking results, and the labeling task results, thereby improving the accuracy of the second sample pair construction.
[0151] It should be noted that although synonym generalization can greatly improve the richness of parameters and semantics, if the sample data constructed by the first sample pair and the second sample pair are directly applied to the training or evaluation stage of the domain task model without evaluation, the distribution information of the sample data may be inappropriate, resulting in insufficient recognition ability of the model in specific problem scenarios, or overfitting to the training data of specific scenarios.
[0152] Based on this, in an optional embodiment of the present specification, obtaining a test set according to the first sample pair and the second sample pair may include the following steps:
[0153] Add the first sample pair and the second sample pair to the test set;
[0154] Determine the distribution information of each sample pair in the test set;
[0155] According to the distribution information, the distribution of samples in the test set is adjusted.
[0156] In an optional implementation of the present specification, both the first sample pair and the second sample pair may be added to the test set.
[0157] In another optional embodiment of the present specification, the first sample pair and the second sample pair can be added to the test set and the training set according to a preset ratio. It should be noted that the preset ratio can be determined according to the needs of the actual application and is not limited in this specification.
[0158] Specifically, the distribution information may include task scenario distribution information, tool type distribution information, and tool parameter distribution information. The task scenario distribution information may be used to characterize the distribution of domain sample data in each scenario of the target domain. For example, in the mathematics domain, 10 domain sample data are included, and the domain sample data include 3 multiple-choice questions and 7 word problems. The scenario distribution information may be: the data distribution ratio of multiple-choice question scenarios to word problem scenarios is 3:7. The tool type distribution information may be used to characterize the frequency distribution of tool use in each target scenario. The tool parameter distribution information may be used to characterize the frequency distribution of parameter enumeration corresponding to tool calls in each target scenario.
[0159] Optionally, adjusting the distribution of sample pairs in the test set according to the distribution information may include at least one of the following three steps:
[0160] When the distribution of task scenario information is uneven, adjust the distribution of domain sample data in each scenario.
[0161] In the case of uneven distribution of tool type information, adjust the distribution of labeling task results in each scenario.
[0162] When tool parameter distribution information is uneven, adjust the distribution of labeling task results in each scenario.
[0163] Alternatively, the uniformity of the task scenario distribution can be determined by counting the number of domain sample data in each task scenario. The uniformity of the tool type distribution can be determined by counting the number of tools of different types used in each task scenario. The uniformity of the tool type distribution can be determined by counting the number of occurrences of parameter enumerations for each tool in each task scenario.
[0164] Optionally, adjusting the distribution of domain sample data in each scenario can be achieved by generalizing the domain sample data in scenarios with less distribution. Adjusting the distribution of labeling task results in each scenario can be achieved by generalizing the tool information or tool parameters in scenarios with less distribution. Specifically, this can be achieved by increasing the number of times the tool type or tool parameter is enumerated.
[0165] By applying this embodiment, the first sample pair and the second sample pair are added to the test set; the distribution information of each sample pair in the test set is determined; and the distribution of the sample pairs in the test set is adjusted according to the distribution information, so that the domain sample data and label task results in each specific task scenario are evenly distributed, thereby ensuring that the evaluation data is evenly distributed in each specific task scenario, avoiding the domain task model's insufficient recognition ability for specific scenarios, ensuring that the frequency distribution of the tools used in each scenario is evenly distributed, avoiding the evaluation data being too concentrated in the evaluation of individual simple or identical tools, resulting in the inability to reasonably evaluate the domain task model's true tool calling capability, and ensuring that the recognition frequency distribution of the enumeration values of each parameter in the evaluation data when a specific tool is called is evenly distributed, avoiding the domain task model only being able to recognize individual specific parameters and having a low accuracy rate in recognizing other parameters.
[0166] In practical applications, for task scenarios in different professional fields or target areas, it is also possible to aggregate and evaluate the distribution of task scenario distribution information, tool type distribution information, and tool parameter distribution information. Optionally, big data computing solutions such as Hadoop and Spark can be used to implement aggregate calculations of various distribution information in the case of large data volumes.
[0167] S2004: Input the domain test data into the domain task model, execute the thinking subtask, obtain the predicted thinking result, and generate the thinking test indicator based on the predicted thinking result and the labeled thinking result.
[0168] Specifically, the predicted thinking result can be understood as the thinking result output by the domain task model after analyzing the domain test data. This thinking result can be output in the form of text. The predicted thinking result can include the thinking result of the task data or the thinking result of the tool processing result.
[0169] Thinking test indicators can include perplexity, BLEU index and ROUGE index.
[0170] Perplexity is a metric used to measure the predictive power of a language model. It evaluates the model's analytical ability by measuring the uncertainty of the probability distribution of the predicted thought outcomes and the label thought outcomes. A lower perplexity value indicates that the model is able to more accurately predict the label thought outcomes.
[0171] For example, given a test set W = w1, w2, ...wn, the perplexity is calculated as shown in the following formula (1):
[0172] The BLEU metric is a reference text-driven evaluation metric that quantifies similarity by comparing n-gram matches between predicted and labeled thought results. Its score ranges from 0 to 1, with scores closer to 1 indicating a high degree of similarity between the predicted and labeled thought results. The basic principles of BLEU scoring are as follows:
[0173] (1) Calculate the number of matches of each n-gram (n consecutive words) in the predicted thinking result in the label thinking result.
[0174] (2) Calculate the number of occurrences of each n-gram in the predicted thinking results.
[0175] (3) Calculate the number of occurrences of each n-gram in the label thinking results.
[0176] (4) Calculate the BLEU score based on the above data, taking into account the number of n-gram occurrences in the predicted thought results and the label thought results, as well as the text length of the predicted thought results.
[0177] The ROUGE metric is a recall-based evaluation metric that evaluates the generated summary text by comparing the word-level and sentence-level similarities between the generated summary text and the reference summary text. It can be used to evaluate the quality of tasks such as text summarization, machine translation, and text generation. Its score ranges from 0 to 1, with a score close to 1 indicating a high degree of similarity between the generated summary text and the reference summary text. The basic principle of ROUGE score calculation is as follows:
[0178] (1) Calculate the number of n-gram matches between the generated summary text and the reference summary text.
[0179] (2) Calculate the number of n-gram occurrences in the generated summary text and the reference summary text.
[0180] (3) Calculate the ROUGE score based on the above data, taking into account the number of matches and the number of occurrences.
[0181] By generating thinking test indicators of different dimensions based on predicted thinking results and labeled thinking results, it is possible to evaluate the thinking and analysis ability of the domain task model in the chain thinking process in the target domain based on thinking test indicators of different dimensions, thereby improving the accuracy and effectiveness of the evaluation of thinking and analysis ability.
[0182] S2006: Input the prediction thinking results into the domain task model, execute the decision subtask, obtain the prediction task results, and generate decision test indicators based on the prediction task results and labeling task results.
[0183] Specifically, the predicted task result can be understood as the task result output by the domain task model based on the predicted thinking results. The predicted task result can include tool parameters or task processing results, where tool parameters can include tool representation information and tool input information.
[0184] Decision test indicators can include precision, recall and F1 value.
[0185] By generating decision test indicators of different dimensions based on the prediction task results and labeling task results, it is possible to evaluate the problem parameter recognition ability and decision-making ability of the domain task model in the chain thinking process in the target domain based on decision test indicators of different dimensions, thereby improving the accuracy and effectiveness of the evaluation of problem parameter recognition ability and decision-making ability.
[0186] S2008: Training domain task models based on thinking test indicators and decision-making test indicators.
[0187] Optionally, in one embodiment of the present specification, training a domain task model based on thinking test indicators and decision-making test indicators may include the following steps:
[0188] Obtain the rate of change of the test indicators of the domain task model under multiple rounds of pre-training;
[0189] Determine the target number of training rounds for the domain task model based on the rate of change;
[0190] According to the distribution information of each test pair in the test set, the decision test indicators are aggregated, and the training sample set is determined based on the aggregation processing results;
[0191] Use the training sample set to train the domain task model according to the target number of training rounds.
[0192] In practical applications, the rate of change of the thinking test indicator after each round of pre-training of the domain task model can be calculated, and the target number of training rounds for the domain task model can be determined based on the rate of change. Specifically, the first-order and second-order rate of change analysis can be performed on the indicators under different numbers of training rounds. The first-order rate of change can be approximated as the difference in the thinking test indicator after each round of model training, and the second-order rate of change is the difference in the first-order rate of change of the thinking test indicator after each round of model training.
[0193] In practical applications, the decision test indicators can be aggregated based on the distribution information of each test pair in the test set, and the training sample set can be adjusted based on the aggregation results. The training sample set can then be used to train the domain task model according to the target number of training rounds.
[0194] Optionally, determining the target number of training rounds for the domain task model based on the rate of change, and determining the training sample set based on the aggregate processing results can be achieved through expert experience, expert models, and rule engines. Among them, expert experience can be used to evaluate the model size and amount of training data in specific fields and specific scenarios. It can decide how much training data to give based on the number of parameters to be trained in the model. It can determine the number of parameters to be increased and the amount of training data to be increased according to a preset ratio. The expert model can evaluate the more consistent and common indicators in each training scenario. The rule engine can make a more logically complex combination of expert experience and the rules in the expert model to form hard rules, and evaluate the thinking test indicators and decision test indicators of the domain task model obtained from each training based on the hard rules.
[0195] For example, the expert model can be used to evaluate the rate of change of an indicator. For example, under normal circumstances, the first-order change difference of the thinking test indicator should decrease with each round, and the first-order change rate should approach 0 with each round. If the indicator change rate does not meet the requirements, the training data set needs to be adjusted accordingly.
[0196] By applying this embodiment, by calculating the rate of change of the thinking test index after each round of training, the number of training rounds of the model can be adjusted, and a model corresponding to the appropriate number of training rounds can be selected to achieve a balance between training data fitting and domain knowledge generalization. The training sample set is adjusted according to the aggregation processing results. The domain task model is trained according to the target number of training rounds using the training sample set, and the model training process can be optimized and adjusted based on the evaluation results of the multi-dimensional aggregation index to improve the training effect of the domain task model. Based on expert experience, expert model and rule engine, the adjustment of the parameters and data sets of the domain task model can be achieved, and the conditions for the model to stop training can be accurately obtained, so as to realize automatic optimization of model training under given rules and knowledge, and achieve better model training effect.
[0197] Optionally, in one embodiment of the present specification, after the domain task model is trained using the training sample set according to the target number of training rounds, the following steps may also be included:
[0198] Obtain the target prediction task result generated by the domain task model based on the target domain test data, where the target domain test data is the domain test data of any test pair in the test set;
[0199] If the target prediction task result is an abnormal result, any test pair containing target domain test data is deleted from the test set.
[0200] Optionally, if abnormal cases are found after multiple rounds of training, the test pairs with abnormal results can be deleted from the evaluation data. Alternatively, if the abnormal cases are unambiguous, the number of repetitions of abnormal cases in the training set can be increased. The specific method for handling abnormal results can be determined based on the actual application.
[0201] By applying this embodiment, when the target prediction task result is an abnormal result, any test pair containing target domain test data is deleted from the test set, so that abnormal cases can be analyzed to avoid the impact of abnormal cases on the model chain thinking ability evaluation results.
[0202] One embodiment of the present specification implements the acquisition of task data for a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask; inputting the task data into a domain task model, executing the thinking subtask, and obtaining the target thinking result of the domain task model for the task data; inputting the target thinking result into the domain task model, executing the decision subtask, and obtaining the target task result corresponding to the target task, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask. By inputting the task data into the domain task model and executing the thinking subtask, the target thinking result of the domain task model for the task data can be obtained; by inputting the target thinking result into the domain task model again and executing the decision subtask, the target task result corresponding to the target task can be obtained; through gradual interaction with the model, the model's thinking and analysis capabilities and task processing capabilities for the target task can be improved based on the model's chain thinking, thereby improving the accuracy of the target task result.
[0203] Referring to FIG3 , FIG3 shows a flowchart of a question-answering processing method in a target domain provided according to an embodiment of this specification, which specifically includes the following steps.
[0204] Step 302: Receive problem information in the target field sent by the front-end user.
[0205] Step 304: Input the problem information into the domain model of the target domain, execute the problem thinking subtask, and obtain the first thinking result of the domain model for the problem information.
[0206] Step 306: Input the first thinking result into the domain model, execute the component decision subtask, and obtain component information of the target component.
[0207] Step 308: Based on the component information, call the target component to process the problem information and obtain component output information.
[0208] Step 310: Input the component output information into the domain model, execute the component output thinking subtask, and obtain the second thinking result of the domain model for the component output information.
[0209] Step 312: Input the second thinking result into the domain model, execute the question-answering decision subtask, and obtain answer information, wherein the domain model is trained based on the test results of the question thinking subtask, component decision subtask, component output thinking subtask, and question-answering decision subtask.
[0210] Step 314: Feedback the answer information to the front-end user.
[0211] Specifically, the question thinking subtask can be understood as a subtask of thinking and analyzing the question information. The component decision subtask can be understood as a subtask of determining the component information of the target component based on the first thinking result. The component output thinking subtask can be understood as a subtask of thinking and analyzing the component output result. The question-answer decision subtask can be understood as a subtask of outputting the answer information corresponding to the question information based on the second thinking result. Among them, the first thinking result and the second thinking result can be in the form of text representation. The target component can be an API or a tool for processing specific scenario problems. Component information may include component identification information and component parameter information.
[0212] In actual applications, after receiving question information about the target domain sent by the front-end user, the question information can be input into the domain model of the target domain, and the question-thinking subtask can be executed to obtain the domain model's first thinking result for the question information. Furthermore, the first thinking result can be input into the domain model, and the component-decision subtask can be executed to obtain the component information of the target component. Based on the component information obtained, the target component can be called through an external framework and the component parameters can be input so that the target component processes the question information and obtains the component output information. The component output information can then be input into the domain model, and the component-output-thinking subtask can be executed to obtain the domain model's second thinking result for the component output information. The second thinking result can be input into the domain model, and the question-answering-decision-making subtask can be executed to obtain the answer information. The answer information is then fed back to the front-end user. The question-thinking subtask, component-decision-making subtask, component-output-thinking subtask, and question-answering-decision-making subtask constitute a complete process for the domain model to chain-think about question information. For each process, the chain-thinking ability can be evaluated using multi-dimensional indicators.
[0213] One embodiment of this specification implements receiving question information of a target domain sent by a front-end user; inputting the question information into the domain model of the target domain, executing the question thinking subtask, and obtaining the first thinking result of the domain model for the question information; inputting the first thinking result into the domain model, executing the component decision subtask, and obtaining the component information of the target component; based on the component information, calling the target component to process the question information and obtaining the component output information; inputting the component output information into the domain model, executing the component output thinking subtask, and obtaining the second thinking result of the domain model for the component output information; inputting the second thinking result into the domain model, executing the question-answering decision subtask, and obtaining the answer information, wherein the domain model is trained based on the test results of the question thinking subtask, the component decision subtask, the component output thinking subtask, and the question-answering decision subtask; and feeding back the answer information to the front-end user. Through gradual interaction with the model, the model's thinking and analysis capabilities and task processing capabilities for the target task can be improved based on the model's chain thinking, thereby improving the accuracy of the target task results.
[0214] Referring to FIG4 , FIG4 shows a flowchart of a domain task model testing method provided according to an embodiment of this specification, which specifically includes the following steps.
[0215] Step 402: Obtain a test set and a domain task model, wherein the test set includes test pairs, the test pairs include domain test data of the target domain, and label thinking results and label task results corresponding to the domain test data, and the domain task model is pre-trained based on sample data of the target domain.
[0216] Step 404: Input the domain test data into the domain task model, execute the thinking subtask, obtain the predicted thinking result, and generate the thinking test indicator based on the predicted thinking result and the labeled thinking result.
[0217] Step 406: Input the prediction thinking results into the domain task model, execute the decision subtask, obtain the prediction task results, and generate decision test indicators based on the prediction task results and the labeling task results.
[0218] Step 408: Based on the thinking test indicators and the decision test indicators, obtain the test results of the domain task model.
[0219] It should be noted that the implementation of steps 402 to 408 is the same as the implementation of steps S2002 to S2008 described above, and will not be described in detail in this embodiment of the specification.
[0220] One embodiment of the present specification implements obtaining a test set and a domain task model, wherein the test set includes a test pair, the test pair includes domain test data of the target domain, and labeled thinking results and labeled task results corresponding to the domain test data, and the domain task model is pre-trained based on sample data of the target domain; the domain test data is input into the domain task model, a thinking subtask is executed, a predicted thinking result is obtained, and a thinking test indicator is generated based on the predicted thinking result and the labeled thinking result; the predicted thinking result is input into the domain task model, a decision subtask is executed, a predicted task result is obtained, and a decision test indicator is generated based on the predicted task result and the labeled task result; the test result of the domain task model is obtained based on the thinking test indicator and the decision test indicator. By obtaining the test result of the domain task model based on the thinking test indicator and the decision test indicator, each step in the chain thinking process of the domain task model can be evaluated accordingly based on the evaluation results of the thinking subtask and the evaluation results of the decision subtask, thereby improving the comprehensiveness and accuracy of the evaluation of the chain thinking ability of the model in the professional field.
[0221] The following is a further explanation of the task processing method provided in this specification, with reference to Figures 5 to 8, taking the application of the task processing method provided in this specification in the detection of model chain thinking ability in the target domain as an example. Among them, Figure 5 shows a schematic diagram of the processing process of a task processing method provided in one embodiment of this specification. Figure 6 shows a schematic diagram of the domain evaluation data generation process of a task processing method provided in one embodiment of this specification. Figure 7 shows a schematic diagram of the model evaluation process of a task processing method provided in one embodiment of this specification. Figure 8 shows a schematic diagram of the evaluation result analysis process of a task processing method provided in one embodiment of this specification. The following specifically explains Figures 5 to 8.
[0222] As shown in Figure 5, for the four steps of chain thinking (model thinking, tool parameter identification, tool result understanding and result return) of the domain task model in the process of executing the target task, the embodiment of this specification uses three main parts: domain evaluation data generation, model chain thinking ability detection and detection result aggregation analysis, to conduct different types of multi-dimensional indicator evaluation for each of the above four steps, thereby achieving a comprehensive, accurate and effective evaluation of the chain thinking ability of the domain task model in the professional field. Among them, model thinking is that the domain task model thinks and analyzes the questions asked by the user and generates model thinking statements; tool parameter identification is that the domain task model outputs the tool to be called and the input parameters of the tool based on the model thinking statement. On the basis of calling the tool through the external framework and inputting the input parameters, the tool result output by the tool can be obtained. Furthermore, tool result understanding is that the domain task model thinks and analyzes the tool result and outputs a thinking statement; result return is the answer corresponding to the question output by the domain task model based on the thinking result of the tool result.
[0223] Among them, the domain evaluation data generation part can realize the generation of domain test data by supplementing manually annotated data and generalizing semantic parameters. Through data distribution evaluation, the problem scenario distribution, tool type distribution and tool parameter distribution of the domain test data are adjusted to obtain a uniformly distributed evaluation data set. And based on the evaluation data set, the domain task model that has completed pre-training in the target domain is tested. The model chain thinking ability detection part can adopt the link customized evaluation method for the above four steps, and evaluate the scenario reasoning ability index and parameter identification accuracy index for thinking and analysis ability and problem parameter identification ability respectively. The detection result aggregation analysis part can obtain training optimization suggestions based on expert experience, expert model and rule engine through indicator dimension aggregation and abnormal case analysis. According to the training optimization suggestions, the training stage of the domain task model can be feedback optimized, and the chain thinking ability evaluation of the next cycle can be carried out based on the re-trained domain task model.
[0224] As shown in Figure 6, semantic parameter generalization can be performed on a small amount of manually annotated data (a sample pair can include multiple test data such as user questions, model thinking, parameter identification, and tool results) to obtain rich annotated data. Semantic parameter generalization can include synonym generalization through large models, parameter consistency verification, and parameter enumeration generalization sampling. Based on the obtained semantically and parameter-richer dataset, data distribution assessment can be performed, including the distribution of problem scenarios, tool types, and tool parameters.
[0225] As shown in Figure 7, the model chaining process specifically includes four steps: model thinking, tool parameter identification, tool result understanding, and result return. The evaluation of model thinking and tool result understanding is a test of thinking and analytical ability, and metrics such as Perplexity, BLEU, and ROUGE can be calculated based on test data. Tool parameter identification and result return are part of the problem parameter identification evaluation, and metrics such as Precision, Recall, and F1 can be calculated based on test data.
[0226] As shown in Figure 8, the results of the thinking and analytical ability assessment can be analyzed for rate of change, while the results of the problem parameter identification assessment can be analyzed for exception cases and dimension aggregation. Based on these analysis, model training optimization recommendations can be obtained using the rule engine, expert experience, and expert models. These recommendations include adjusting the training argument, generalizing the indicator enumeration, and adjusting the scenario distribution.
[0227] In the domain evaluation data generation stage, the embodiments of this specification provide a solution for generating sufficient generalized evaluation data based on a small amount of manually annotated data through the basic text generation and generalization capabilities of a general large model, and help the model trainer understand the feature distribution and change trends in the data set through data distribution evaluation, thereby ensuring the data quality of the model evaluation data set. In the model chain thinking ability detection stage, a specific solution for the domain model chain thinking ability detection is provided, and customized evaluation is performed on each step of the chain thinking link. The various steps of the chain thinking link can be mainly divided into two categories: thinking analysis ability evaluation and parameter identification ability evaluation. These two categories can rely on different indicators and evaluation systems respectively, and the chain thinking ability of the model can be evaluated in multiple dimensions through composite indicators under multiple categories. In the detection result aggregation and analysis stage, an analysis and evaluation scheme for the domain model chain thinking ability detection results is given. According to the thinking and analysis ability evaluation results and the parameter identification ability evaluation results, the aggregation result analysis of indicators in different dimensions and abnormal case analysis can be performed. In combination with the rule engine, expert model and expert experience, guidance suggestions for model training are given, which opens the boundary between model evaluation and model training, forms a closed loop of model training-effect evaluation-training adjustment, realizes automatic optimization of model training under given rules and knowledge, and achieves better model training effect.
[0228] Corresponding to the above method embodiment, this specification also provides an embodiment of a task processing device. FIG9 shows a schematic diagram of the structure of a task processing device provided in one embodiment of this specification. As shown in FIG9 , the device includes:
[0229] The first acquisition module 902 is configured to acquire task data of a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask.
[0230] The first input module 904 is configured to input task data into the domain task model, execute the thinking subtask, and obtain the target thinking result of the domain task model for the task data.
[0231] The first execution module 906 is configured to input the target thinking result into the domain task model, execute the decision subtask, and obtain the target task result corresponding to the target task, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask.
[0232] Optionally, the task processing device further includes a training module configured to:
[0233] Obtain a test set and a domain task model, where the test set includes test pairs, each test pair includes domain test data of the target domain, and label thinking results and label task results corresponding to the domain test data. The domain task model is pre-trained based on sample data of the target domain.
[0234] Input the domain test data into the domain task model, perform the thinking subtask, obtain the predicted thinking results, and generate thinking test indicators based on the predicted thinking results and the labeled thinking results;
[0235] Input the prediction thinking results into the domain task model, execute the decision subtask, obtain the prediction task results, and generate decision test indicators based on the prediction task results and labeling task results;
[0236] Train domain task models based on thinking test indicators and decision-making test indicators.
[0237] Optionally, the training module is further configured to:
[0238] Acquire a sample set, wherein the sample set includes a first sample pair, and the first sample pair includes domain sample data of a target domain, a label thinking result corresponding to the domain sample data, and a label task result;
[0239] Performing generalization processing on the first sample pair to generate multiple second sample pairs;
[0240] A test set is obtained according to the first sample pair and the second sample pair.
[0241] Optionally, the training module is further configured to:
[0242] Input the domain sample data and the label thinking result in the first sample pair into the language processing model, perform the synonymous rewriting task, and obtain the target domain sample data that is synonymous with the domain sample data, and the target label thinking result that is synonymous with the label thinking result;
[0243] Based on the target domain sample data, the target label thinking results and the label task results, a second sample pair is generated.
[0244] Optionally, the training module is further configured to:
[0245] Identify whether the semantic parameters of the target domain sample data and the target label thinking results are consistent;
[0246] When the semantic parameters are consistent, a second sample pair is generated based on the target domain sample data, the target label thinking results and the label task results.
[0247] Optionally, the training module is further configured to:
[0248] Add the first sample pair and the second sample pair to the test set;
[0249] Determine the distribution information of each sample pair in the test set;
[0250] According to the distribution information, the distribution of samples in the test set is adjusted.
[0251] Optionally, the training module is further configured to:
[0252] Obtain the rate of change of the test indicators of the domain task model under multiple rounds of pre-training;
[0253] Determine the target number of training rounds for the domain task model based on the rate of change;
[0254] According to the distribution information of each test pair in the test set, the decision test indicators are aggregated, and the training sample set is determined based on the aggregation processing results;
[0255] Use the training sample set to train the domain task model according to the target number of training rounds.
[0256] Optionally, the training module is further configured to:
[0257] Obtain the target prediction task result generated by the domain task model based on the target domain test data, where the target domain test data is the domain test data of any test pair in the test set;
[0258] If the target prediction task result is an abnormal result, any test pair containing target domain test data is deleted from the test set.
[0259] One embodiment of the present specification implements the acquisition of task data for a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask; inputting the task data into a domain task model, executing the thinking subtask, and obtaining the target thinking result of the domain task model for the task data; inputting the target thinking result into the domain task model, executing the decision subtask, and obtaining the target task result corresponding to the target task, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask. By inputting the task data into the domain task model and executing the thinking subtask, the target thinking result of the domain task model for the task data can be obtained; by inputting the target thinking result into the domain task model again and executing the decision subtask, the target task result corresponding to the target task can be obtained; through gradual interaction with the model, the model's thinking and analysis capabilities and task processing capabilities for the target task can be improved based on the model's chain thinking, thereby improving the accuracy of the target task result.
[0260] The above is a schematic scheme of a task processing device of this embodiment. It should be noted that the technical scheme of the task processing device and the technical scheme of the task processing method described above are of the same concept. For details not described in detail in the technical scheme of the task processing device, please refer to the description of the technical scheme of the task processing method described above.
[0261] Corresponding to the above method embodiments, this specification also provides an embodiment of a question-answering processing device in a target domain. FIG10 shows a schematic diagram of the structure of a question-answering processing device in a target domain provided by one embodiment of this specification. As shown in FIG10 , the device includes:
[0262] Receiving module 1002: configured to receive problem information in the target field sent by the front-end user.
[0263] The second input module 1004 is configured to input the problem information into the domain model of the target domain, execute the problem thinking subtask, and obtain the first thinking result of the domain model for the problem information.
[0264] The second execution module 1006 is configured to input the first thinking result into the domain model, execute the component decision subtask, and obtain component information of the target component.
[0265] The calling module 1008 is configured to call the target component based on the component information to process the problem information and obtain component output information.
[0266] The third input module 1010 is configured to input the component output information into the domain model, execute the component output thinking subtask, and obtain the second thinking result of the domain model for the component output information.
[0267] The third execution module 1012 is configured to input the second thinking result into the domain model, execute the question-answering decision subtask, and obtain answer information, wherein the domain model is trained based on the test results of the question thinking subtask, the component decision subtask, the component output thinking subtask, and the question-answering decision subtask.
[0268] Feedback module 1014: configured to feed back answer information to the front-end user.
[0269] One embodiment of this specification implements receiving question information of a target domain sent by a front-end user; inputting the question information into the domain model of the target domain, executing the question thinking subtask, and obtaining the first thinking result of the domain model for the question information; inputting the first thinking result into the domain model, executing the component decision subtask, and obtaining the component information of the target component; based on the component information, calling the target component to process the question information and obtaining the component output information; inputting the component output information into the domain model, executing the component output thinking subtask, and obtaining the second thinking result of the domain model for the component output information; inputting the second thinking result into the domain model, executing the question-answering decision subtask, and obtaining the answer information, wherein the domain model is trained based on the test results of the question thinking subtask, the component decision subtask, the component output thinking subtask, and the question-answering decision subtask; and feeding back the answer information to the front-end user. Through gradual interaction with the model, the model's thinking and analysis capabilities and task processing capabilities for the target task can be improved based on the model's chain thinking, thereby improving the accuracy of the target task results.
[0270] The above is a schematic scheme of a question-and-answer processing device for a target domain of this embodiment. It should be noted that the technical scheme of the question-and-answer processing device for this target domain and the technical scheme of the question-and-answer processing method for the target domain are based on the same concept. For details not described in detail in the technical scheme of the question-and-answer processing device for the target domain, please refer to the description of the technical scheme of the question-and-answer processing method for the target domain.
[0271] Corresponding to the above method embodiment, this specification also provides an embodiment of a domain task model testing device. Figure 11 shows a schematic diagram of the structure of a domain task model testing device provided by one embodiment of this specification. As shown in Figure 11, the device includes:
[0272] The second acquisition module 1102 is configured to obtain a test set and a domain task model, wherein the test set includes a test pair, the test pair includes domain test data of the target domain, and label thinking results and label task results corresponding to the domain test data, and the domain task model is pre-trained based on sample data of the target domain.
[0273] The fourth input module 1104 is configured to input the domain test data into the domain task model, execute the thinking subtask, obtain the predicted thinking result, and generate the thinking test indicator according to the predicted thinking result and the labeled thinking result.
[0274] The fourth execution module 1106 is configured to input the prediction thinking result into the domain task model, execute the decision subtask, obtain the prediction task result, and generate a decision test indicator based on the prediction task result and the labeling task result.
[0275] Result generation module 1108: configured to obtain the test result of the domain task model based on the thinking test indicator and the decision test indicator.
[0276] One embodiment of the present specification implements obtaining a test set and a domain task model, wherein the test set includes a test pair, the test pair includes domain test data of the target domain, and labeled thinking results and labeled task results corresponding to the domain test data, and the domain task model is pre-trained based on sample data of the target domain; the domain test data is input into the domain task model, a thinking subtask is executed, a predicted thinking result is obtained, and a thinking test indicator is generated based on the predicted thinking result and the labeled thinking result; the predicted thinking result is input into the domain task model, a decision subtask is executed, a predicted task result is obtained, and a decision test indicator is generated based on the predicted task result and the labeled task result; the test result of the domain task model is obtained based on the thinking test indicator and the decision test indicator. By obtaining the test result of the domain task model based on the thinking test indicator and the decision test indicator, each step in the chain thinking process of the domain task model can be evaluated accordingly based on the evaluation results of the thinking subtask and the evaluation results of the decision subtask, thereby improving the comprehensiveness and accuracy of the evaluation of the chain thinking ability of the model in the professional field.
[0277] The above is a schematic scheme of a domain task model testing device of this embodiment. It should be noted that the technical scheme of the domain task model testing device and the technical scheme of the aforementioned domain task model testing method are based on the same concept. For details not described in detail in the technical scheme of the domain task model testing device, please refer to the description of the technical scheme of the aforementioned domain task model testing method.
[0278] Figure 12 shows a block diagram of a computing device 1200 according to one embodiment of this specification. Components of the computing device 1200 include, but are not limited to, a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 via a bus 1230, and a database 1250 is used to store data.
[0279] The computing device 1200 also includes an access device 1240 that enables the computing device 1200 to communicate via one or more networks 1260. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 1240 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0280] In one embodiment of the present specification, the aforementioned components of the computing device 1200 and other components not shown in FIG12 may also be connected to each other, for example, via a bus. It should be understood that the block diagram of the computing device structure shown in FIG12 is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0281] Computing device 1200 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1200 may also be a mobile or stationary server.
[0282] The processor 1220 is configured to execute the following computer-executable instructions, which implement the steps of the above method when executed by the processor.
[0283] The above is a schematic solution of a computing device of this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above method.
[0284] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above method when executed by a processor.
[0285] The above is a schematic solution of a computer-readable storage medium of this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the above method belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the above method.
[0286] An embodiment of the present specification further provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0287] The above is an illustrative solution of a computer program product of this embodiment. It should be noted that the technical solution of the computer program product and the technical solution of the above method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above method.
[0288] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0289] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0290] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0291] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0292] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A task processing method, comprising: Acquiring task data of a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask; Inputting the task data into a domain task model, executing the thinking subtask, and obtaining a target thinking result of the domain task model for the task data; The target thinking result is input into the domain task model, the decision subtask is executed, and the target task result corresponding to the target task is obtained, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask.
2. The task processing method according to claim 1, before inputting the task data into the domain task model, further comprising: Obtaining a test set and the domain task model, wherein the test set includes test pairs, the test pairs include domain test data of the target domain, and labeled thinking results and labeled task results corresponding to the domain test data, and the domain task model is pre-trained based on sample data of the target domain; Inputting the domain test data into the domain task model, executing the thinking subtask, obtaining a predicted thinking result, and generating a thinking test indicator based on the predicted thinking result and the labeled thinking result; Input the prediction thinking result into the domain task model, execute the decision subtask, obtain the prediction task result, and generate a decision test indicator based on the prediction task result and the label task result; The domain task model is trained based on the thinking test indicator and the decision-making test indicator.
3. The task processing method according to claim 2, wherein obtaining a test set comprises: Acquire a sample set, wherein the sample set includes a first sample pair, the first sample pair including domain sample data of a target domain, a label thinking result corresponding to the domain sample data, and a label task result; performing generalization processing on the first sample pairs to generate a plurality of second sample pairs; A test set is obtained according to the first sample pair and the second sample pair.
4. The task processing method according to claim 3, wherein the generalizing the first sample pair to generate a plurality of second sample pairs comprises: Inputting the domain sample data and the label thinking result in the first sample pair into a language processing model, performing a synonymous rewriting task, and obtaining target domain sample data that is synonymous with the domain sample data, and a target label thinking result that is synonymous with the label thinking result; A second sample pair is generated based on the target domain sample data, the target label thinking result and the label task result.
5. The task processing method according to claim 4, wherein generating a second sample pair based on the target domain sample data, the target label thinking result, and the label task result comprises: Identify whether the semantic parameters of the target domain sample data and the target label thinking result are consistent; When the semantic parameters are consistent, a second sample pair is generated based on the target domain sample data, the target label thinking result and the label task result.
6. The task processing method according to claim 3, wherein obtaining a test set based on the first sample pair and the second sample pair comprises: Adding the first sample pair and the second sample pair to a test set; Determining distribution information of each sample pair in the test set; According to the distribution information, the distribution of the samples in the test set is adjusted.
7. The task processing method according to claim 2, wherein training the domain task model based on the thinking test indicator and the decision test indicator comprises: Obtaining the rate of change of the thinking test indicator of the domain task model under multiple rounds of pre-training; Determining a target number of training rounds for the domain task model based on the rate of change; Aggregating the decision test indicators according to the distribution information of each test pair in the test set, and determining the training sample set according to the aggregation result; The domain task model is trained using the training sample set according to the target number of training rounds.
8. The task processing method according to claim 7, further comprising, after training the domain task model using the training sample set according to the target number of training rounds: Obtaining a target prediction task result generated by the domain task model based on target domain test data, wherein the target domain test data is the domain test data of any test pair in the test set; In the case where the target prediction task result is an abnormal result, any test pair containing the target domain test data is deleted from the test set.
9. A method for question answering in a target domain, comprising: Receive problem information in the target area sent by front-end users; Inputting the problem information into the domain model of the target domain, executing the problem thinking subtask, and obtaining a first thinking result of the domain model for the problem information; Inputting the first thinking result into the domain model, executing the component decision subtask, and obtaining component information of the target component; Based on the component information, calling the target component to process the problem information and obtain component output information; Inputting the component output information into the domain model, executing the component output thinking subtask, and obtaining a second thinking result of the domain model for the component output information; Inputting the second thinking result into the domain model, executing the question-answering decision subtask, and obtaining answer information, wherein the domain model is trained based on the test results of the question-thinking subtask, the component decision subtask, the component output thinking subtask, and the question-answering decision subtask; Feedback the answer information to the front-end user.
10. A domain task model testing method, comprising: Obtaining a test set and a domain task model, wherein the test set includes test pairs, the test pairs include domain test data of the target domain, and labeled thinking results and labeled task results corresponding to the domain test data, and the domain task model is pre-trained based on sample data of the target domain; Inputting the domain test data into the domain task model, executing the thinking subtask, obtaining a predicted thinking result, and generating a thinking test indicator based on the predicted thinking result and the labeled thinking result; Input the prediction thinking result into the domain task model, execute the decision subtask, obtain the prediction task result, and generate a decision test indicator based on the prediction task result and the label task result; Based on the thinking test indicator and the decision-making test indicator, a test result of the domain task model is obtained.
11. A task processing device comprising: A first acquisition module is configured to acquire task data of a target task, wherein the target task includes at least one task stage, and the task stage includes a thinking subtask and a decision-making subtask; A first input module is configured to input the task data into a domain task model, execute the thinking subtask, and obtain a target thinking result of the domain task model for the task data; The first execution module is configured to input the target thinking result into the domain task model, execute the decision subtask, and obtain the target task result corresponding to the target task, wherein the domain task model is trained based on the test results of the thinking subtask and the decision subtask.
12. The task processing device according to claim 11, further comprising a training module configured to: Get the test set and the domain task model, where: The test set includes a test pair, each of which includes domain test data of a target domain, and a label thinking result and a label task result corresponding to the domain test data, wherein the domain task model is pre-trained based on sample data of the target domain; Inputting the domain test data into the domain task model, executing the thinking subtask, obtaining a predicted thinking result, and generating a thinking test indicator based on the predicted thinking result and the labeled thinking result; Input the prediction thinking result into the domain task model, execute the decision subtask, obtain the prediction task result, and generate a decision test indicator based on the prediction task result and the label task result; The domain task model is trained based on the thinking test indicator and the decision-making test indicator.
13. The task processing device according to claim 12, wherein the training module is further configured to: Get a sample set, where The sample set includes a first sample pair, wherein the first sample pair includes domain sample data of a target domain, a label thinking result and a label task result corresponding to the domain sample data; performing generalization processing on the first sample pairs to generate a plurality of second sample pairs; A test set is obtained according to the first sample pair and the second sample pair.
14. The task processing device according to claim 13, wherein the training module is further configured to: Inputting the domain sample data and the label thinking result in the first sample pair into a language processing model, performing a synonymous rewriting task, and obtaining target domain sample data that is synonymous with the domain sample data, and a target label thinking result that is synonymous with the label thinking result; A second sample pair is generated based on the target domain sample data, the target label thinking result and the label task result.
15. The task processing device according to claim 14, wherein the training module is further configured to: Identify whether the semantic parameters of the target domain sample data and the target label thinking result are consistent; When the semantic parameters are consistent, a second sample pair is generated based on the target domain sample data, the target label thinking result and the label task result.
16. The task processing device according to claim 13, wherein the training module is further configured to: Adding the first sample pair and the second sample pair to a test set; Determining distribution information of each sample pair in the test set; According to the distribution information, the distribution of the samples in the test set is adjusted.
17. The task processing device according to claim 12, wherein the training module is further configured to: Obtaining the rate of change of the thinking test indicator of the domain task model under multiple rounds of pre-training; Determining a target number of training rounds for the domain task model based on the rate of change; Aggregating the decision test indicators according to the distribution information of each test pair in the test set, and determining the training sample set according to the aggregation result; The domain task model is trained using the training sample set according to the target number of training rounds.
18. The task processing device according to claim 17, wherein the training module is further configured to: Obtain the target prediction task result generated by the domain task model based on the target domain test data, where: The target domain test data is the domain test data of any test pair in the test set; In the case where the target prediction task result is an abnormal result, any test pair containing the target domain test data is deleted from the test set.
19. A question-answering processing device in a target domain, comprising: A receiving module is configured to receive problem information in a target field sent by a front-end user; A second input module is configured to input the problem information into the domain model of the target domain, execute the problem thinking subtask, and obtain a first thinking result of the domain model for the problem information; a second execution module, configured to input the first thinking result into the domain model, execute a component decision subtask, and obtain component information of a target component; A calling module is configured to call the target component based on the component information to process the problem information and obtain component output information; a third input module, configured to input the component output information into the domain model, execute the component output thinking subtask, and obtain a second thinking result of the domain model for the component output information; a third execution module, configured to input the second thinking result into the domain model, execute the question-answering decision subtask, and obtain answer information, wherein the domain model is trained based on the test results of the question-thinking subtask, the component decision subtask, the component output thinking subtask, and the question-answering decision subtask; The feedback module is configured to feed back the answer information to the front-end user.
20. A domain task model testing device, comprising: a second acquisition module configured to acquire a test set and a domain task model, wherein the test set includes test pairs, the test pairs include domain test data of the target domain, and labeled thinking results and labeled task results corresponding to the domain test data, and the domain task model is pre-trained based on sample data of the target domain; a fourth input module configured to input the domain test data into the domain task model, execute the thinking subtask, obtain a predicted thinking result, and generate a thinking test indicator based on the predicted thinking result and the labeled thinking result; a fourth execution module, configured to input the prediction thinking result into the domain task model, execute the decision subtask, obtain the prediction task result, and generate a decision test indicator based on the prediction task result and the labeling task result; The result generation module is configured to obtain the test result of the domain task model based on the thinking test indicator and the decision test indicator.
21. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 10 are implemented.
22. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
23. A computer program product comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.