Task processing method, information processing method based on task processing model, and task platform
By iteratively executing multi-level tasks in the task processing model and using the labeled task results to train the model, the problems of consistency between stages and reward signal design in multi-stage tasks are solved, and the efficiency and accuracy of task processing are improved.
Patent Information
- Application Number
- CN202510727406.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-06-03
AI Technical Summary
In model training of multi-stage tasks, stage-by-stage training leads to weakened consistency and coherence between stages, while the reward signal design is difficult to set effectively during overall training, affecting the task execution effect.
By obtaining the pending task data of the target task, inputting it into the task processing model to execute the initial level task, obtaining the intermediate results, and iteratively executing other level tasks until the target task results are obtained, the model is trained using the label task results to ensure the dependencies of tasks at each level and the consistency of the final results.
It improves the efficiency of task processing and the quality of results, simplifies the reward signal design, and ensures the coordination between stages and the accuracy of task execution.
Smart Images

Figure CN120234126B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular to a task processing method, an information processing method based on a task processing model, and a task platform. Background Art
[0002] In the field of automated task processing in artificial intelligence, tasks are typically broken down into multiple, sequential stages, with the output of each stage crucial for subsequent steps. In reinforcement learning, in particular, these multi-stage tasks require models to understand not only the operations of individual stages but also the dependencies between them to achieve appropriate execution.
[0003] However, when training models for multi-stage tasks, stage-by-stage training can lead to a loss of consistency and coherence between stages, as individual training fails to capture the global objective. Attempting to train the model as a whole presents the challenge of effectively setting reward signals, particularly when ensuring that dependencies between stages are accurately reflected. This makes it challenging to achieve both efficient training and coordinated learning across stages. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a task processing method. One or more embodiments of this specification also relate to a task processing model training method, an information processing method based on a task processing model, a task platform, a task processing device, a task processing model training device, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0005] According to a first aspect of an embodiment of this specification, a task processing method is provided, including:
[0006] Obtaining pending task data of a target task, wherein the target task includes hierarchical tasks of at least two levels;
[0007] Inputting the task data to be processed into a task processing model to execute an initial level task among at least two levels of hierarchical tasks to obtain an intermediate result, wherein the task processing model is trained based on at least two sample level tasks and a labeling task result, wherein the labeling task result is information that labels expected results after executing the at least two sample level tasks;
[0008] The intermediate results are input into the task processing model, and the tasks of at least two levels of hierarchical tasks except the initial level tasks are iteratively executed until the task processing results of the target task are obtained.
[0009] According to a second aspect of an embodiment of this specification, a task processing model training method is provided, comprising:
[0010] Acquire sample task data of a sample task and a label task result corresponding to the sample task, wherein the sample task includes at least two levels of sample-level tasks;
[0011] Input the sample task data into the initial processing model to execute the initial sample level task and obtain the sample intermediate result;
[0012] Inputting the sample intermediate result into the initial processing model, iteratively executing the other sample-level tasks except the initial sample-level task in at least two levels of sample-level tasks until a predicted processing result of the sample task is obtained;
[0013] Based on the prediction processing results and the label task results, the initial processing model is trained to obtain the task processing model.
[0014] According to a third aspect of the embodiments of this specification, there is provided an information processing method based on a task processing model, which is applied to a task platform, including:
[0015] Receiving a task generation request sent by a terminal device, wherein the task generation request includes request information;
[0016] Based on the request information, a task processing model is obtained, wherein the task processing model is used to execute an initial level task in at least two levels of hierarchical tasks based on the to-be-processed task data of the target task, obtain an intermediate result, and iteratively execute other level tasks in at least two levels of hierarchical tasks except the initial level task based on the intermediate result until a task processing result of the target task is obtained, wherein the target task includes at least two levels of hierarchical tasks, and the task processing model is trained based on at least two sample level tasks and the label task result;
[0017] Based on the task processing model, task information is generated, wherein the task information is used by the terminal device to execute the target task.
[0018] According to a fourth aspect of the embodiments of this specification, there is provided a task platform, comprising a request interface and a response unit;
[0019] A request interface, configured to receive a task generation request sent by a terminal device, wherein the task generation request includes request information;
[0020] A response unit is used to obtain a task processing model based on the request information, wherein the task processing model is used to execute an initial level task in at least two levels of hierarchical tasks based on the to-be-processed task data of the target task, obtain an intermediate result, and iteratively execute other level tasks except the initial level task in at least two levels of hierarchical tasks based on the intermediate result until the task processing result of the target task is obtained, wherein the target task includes at least two levels of hierarchical tasks, and the task processing model is trained based on at least two sample level tasks and label task results.
[0021] According to a fifth aspect of the embodiments of this specification, there is provided a task processing device, including:
[0022] A first acquisition module is configured to acquire to-be-processed task data of a target task, wherein the target task includes hierarchical tasks of at least two levels;
[0023] a first task execution module configured to input the task data to be processed into a task processing model to execute an initial level task of at least two levels of hierarchical tasks to obtain an intermediate result, wherein the task processing model is trained based on at least two sample level tasks and a labeling task result, wherein the labeling task result is information that labels expected results after executing the at least two sample level tasks;
[0024] The second task execution module is configured to input the intermediate result into the task processing model, and iteratively execute the tasks of at least two levels of hierarchical tasks except the initial level task until the task processing result of the target task is obtained.
[0025] According to a sixth aspect of the embodiments of this specification, a task processing model training device is provided, comprising:
[0026] A second acquisition module is configured to acquire sample task data of a sample task and a label task result corresponding to the sample task, wherein the sample task includes sample-level tasks of at least two levels;
[0027] A third task execution module is configured to input the sample task data into the initial processing model to execute the initial sample level task and obtain a sample intermediate result;
[0028] a fourth task execution module, configured to input the sample intermediate result into the initial processing model, and iteratively execute the other sample-level tasks except the initial sample-level task in the at least two levels of sample-level tasks until a predicted processing result of the sample task is obtained;
[0029] The training module is configured to train the initial processing model based on the prediction processing results and the label task results to obtain the task processing model.
[0030] According to a seventh aspect of the embodiments of this specification, a computing device is provided, including:
[0031] memory and processor;
[0032] Among them, the memory is used to store computer programs / instructions, and the processor is used to execute computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the above-mentioned task processing method, task processing model training method, and information processing method based on the task processing model are implemented.
[0033] According to the eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned task processing method, task processing model training method, and information processing method based on the task processing model.
[0034] According to the ninth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned task processing method, task processing model training method, and information processing method based on the task processing model.
[0035] One embodiment of the present specification realizes the following steps: obtaining the pending task data of the target task, wherein the target task includes at least two levels of hierarchical tasks; inputting the pending task data into the task processing model to execute the initial hierarchical task in the at least two levels of hierarchical tasks to obtain an intermediate result, wherein the task processing model is trained based on at least two sample hierarchical tasks and labeling task results, and the labeling task results are information for labeling the expected results after executing at least two sample hierarchical tasks; inputting the intermediate result into the task processing model, and iteratively executing the other hierarchical tasks except the initial hierarchical task in the at least two levels of hierarchical tasks until the task processing result of the target task is obtained. By using the final labeling task result to reversely adjust and optimize the hierarchical tasks of at least two sample levels, it is ensured that each sample level task is considered correct only when the entire task process achieves the expected goal. This enables the task processing model to not only learn the operations within a single level, but also understand and grasp the dependencies between levels, thereby achieving overall optimization of the target task. At the same time, since the learning and adjustment at each stage are oriented towards the final results, the design requirements of complex reward signals are effectively simplified, and the task processing efficiency and result quality are significantly improved while maintaining coordination between stages. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 It is a task processing flow chart of a two-stage task;
[0037] Figure 2It is a structural interaction diagram of a rule-based reward model;
[0038] Figure 3 This is an application architecture diagram of a target task provided by an embodiment of this specification;
[0039] Figure 4 This is a flowchart of a task processing method provided by one embodiment of this specification;
[0040] Figure 5 This is a flowchart of a task processing model training method provided by one embodiment of this specification;
[0041] Figure 6 This is a schematic diagram of a processing process of a task processing method provided by an embodiment of this specification;
[0042] Figure 7 This is a schematic diagram of the structure of an end-to-end task consistency verification reward model including an illusion penalty mechanism provided by an embodiment of this specification;
[0043] Figure 8 This is a flowchart of an information processing method based on a task processing model provided by one embodiment of this specification;
[0044] Figure 9 This is a schematic diagram of the structure of a task platform provided by one embodiment of this specification;
[0045] Figure 10 This is a structural diagram of a task processing device provided by one embodiment of this specification;
[0046] Figure 11 This is a structural diagram of a task processing model training device provided by one embodiment of this specification;
[0047] Figure 12 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0048] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0049] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0050] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0051] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0052] In one or more embodiments of this specification, a large model refers to a deep learning model with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a foundation model. It is pre-trained on a large amount of unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large language model (LLM) and a multi-modal pre-training model.
[0053] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image description (IC, Image Caption), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0054] First, the terms involved in one or more embodiments of this specification are explained.
[0055] Reinforcement Learning (RL): A machine learning method that learns appropriate policies based on reward signals through interaction with the environment to maximize cumulative rewards.
[0056] GRPO: Group Relative Policy Optimization, an advanced reinforcement learning algorithm based on policy optimization, improves training performance through group relative policy optimization technology.
[0057] Multi-StageTask: A task is divided into multiple sequential stages, and the execution of subsequent stages depends on the output of previous stages.
[0058] Multi-Agent System (MAS): A system composed of multiple agents that can work together or independently perform tasks to achieve complex goals.
[0059] Reward Model (RM): In reinforcement learning, it is used to evaluate the quality of actions or strategies and guide the training process to optimize strategies to obtain higher rewards.
[0060] Rule-based Reward Model (RM): A reward model that scores based on predefined rules, typically using a binary score (0-1) to evaluate the correctness of the output.
[0061] Task consistency validation: Ensures that the outputs of each stage in a multi-stage task are consistent and correct as a whole to ensure the accuracy of the final result.
[0062] Policy Optimization: Adjusting and improving an agent’s policy to perform better on a given task.
[0063] In today's rapidly evolving technological landscape, artificial intelligence (AI), particularly reinforcement learning (RL), is finding an increasingly broad range of applications. From gaming and entertainment to optimizing complex task processes, RL provides a powerful tool for automating a wide range of tasks. Especially when dealing with tasks requiring continuous decision-making, RL optimizes objectives by continuously interacting with the environment and adjusting strategies. However, as application areas expand and task complexity increases, effectively training models to adapt to these complex scenarios becomes a pressing challenge.
[0064] One particularly challenging aspect is the handling of multi-stage tasks. In these tasks, the output of each stage not only determines the success of the current stage, but also has a direct impact on subsequent stages. Therefore, coordination and consistency between stages are crucial. Although traditional stage-separated training methods can independently optimize the performance of each stage, the lack of a global perspective often leads to inconsistencies between stages, affecting the overall performance of the task. Although end-to-end training methods aim to optimize the entire task chain, they face huge challenges in designing effective reward signals suitable for complex tasks, which increases the difficulty of training and may limit its practical application.
[0065] This problem is particularly prominent when taking tasks in multi-agent systems (MAS) as an example. Figure 1 As shown, Figure 1 This is a flowchart of the task processing process for a two-stage task. MAS tasks are typically divided into two main phases: Phase 1, Task Planning, and Phase 2, API Selection. First, in the task planning phase, the system develops a detailed execution plan based on the input task processing data and the capabilities of the sub-agents, extracting entities from the sub-agent identifiers, such as entity a, entity b, and so on. Second, in the API selection phase, based on the plan generated in the first phase, a specific API is selected for task execution. For example, the agent corresponding to entity a is called, and the agent corresponding to entity b is called, to obtain the output of each agent.
[0066] For reinforcement learning training of multi-stage tasks, the Rule-based Reward Model (RM) is a commonly used reward evaluation method. Figure 2 As shown, Figure 2This is a schematic diagram of the structural interaction of a rule-based reward model. Specifically, this model integrates a policy model into the model. During the first stage of task planning, the user's question is fed into the policy model to generate a Stage 1 answer. The sub-agent name is then extracted based on the Stage 1 answer. By extracting the sub-agent name and performing a rule-based judgment against a predefined list of sub-agents, a binary correctness score (0 or 1) is assigned to the output. If the sub-agent's selection and task planning conform to the rules, a score of 1 is awarded; otherwise, a score of 0 is awarded. However, each stage is scored separately. While this allows for independent performance optimization, the lack of a global perspective often leads to inconsistencies between stages, impacting the overall task execution.
[0067] To address the above technical issues, this specification provides a task processing method. One or more embodiments of this specification also relate to a task processing model training method, an information processing method based on a task processing model, a task platform, a task processing device, a task processing model training device, a computing device, a computer-readable storage medium, and a computer program product, each of which is described in detail in the following embodiments.
[0068] Considering the huge number of model parameters of large models and the limited computing resources of mobile terminals, the task processing method provided in the embodiment of the present application can be applied to Figure 3 The application scenarios shown are not limited to this. Figure 3 This is an application architecture diagram of a target task provided by an embodiment of this specification. Figure 3 In the illustrated application scenario, the large model is deployed on a server 10. Server 10 can be connected to one or more client devices 20 via a local area network (LAN), a wide area network (WAN), the Internet, or other types of data networks. Client devices 20 herein include, but are not limited to, smartphones, tablet computers, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. Client devices 20 can interact with users via a graphical user interface (GUI) to access the large model and implement the methods provided in the embodiments of this specification.
[0069] In an embodiment of the present specification, a system composed of a client device and a server can perform the following steps: the client device generates task data to be processed and sends it to the server. The server executes to obtain task data to be processed of the target task, wherein the target task includes at least two levels of hierarchical tasks; the task data to be processed is input into the task processing model to execute the initial hierarchical task in at least two levels of hierarchical tasks to obtain an intermediate result, wherein the task processing model is trained based on at least two sample hierarchical tasks and label task results, and the label task result is information for labeling the expected result after executing at least two sample hierarchical tasks; the intermediate result is input into the task processing model, and the hierarchical tasks other than the initial hierarchical task in at least two levels of hierarchical tasks are iteratively executed until the task processing result of the target task is obtained. It should be noted that, when the operating resources of the client device can meet the deployment and operating conditions of the large model, the embodiment of the present application can be carried out in the client device.
[0070] See also Figure 4 , Figure 4 This is a flowchart of a task processing method provided by an embodiment of this specification, which specifically includes the following steps.
[0071] Step 402: Obtain to-be-processed task data of a target task, wherein the target task includes hierarchical tasks of at least two levels.
[0072] Task data is the input information required before executing the target task. It typically includes a detailed description of the target task, a hierarchical task breakdown, and relevant contextual information. For example, in a multi-agent system, the target task might be to complete a complex user request, while task data might include a description of the user's question, an introduction to the sub-agent's functionality, and the constraints associated with task execution.
[0073] Hierarchical tasks are the multiple phases of the target task. Each hierarchical task corresponds to an execution phase of the target task, and dependencies exist between hierarchical tasks. For example, in a MAS task, hierarchical tasks can be divided into two phases: task planning and API selection. The former is responsible for developing the overall plan, while the latter is responsible for selecting specific tools or interfaces based on the plan.
[0074] In practical applications, obtaining pending task data for a target task first requires clarifying the overall objective of the target task and its underlying hierarchical task structure. One option is to extract the target task's description and hierarchical task division from a task management system. Another option is to dynamically generate the definition of the target task and its hierarchical tasks through user input. Additionally, pending data for similar tasks can be loaded from a historical task database for reference. Dependencies between hierarchical tasks can be modeled by analyzing task logic or invoking a rules engine to ensure coherence between tasks at each level.
[0075] In the embodiments of this specification, by clarifying the target task and its hierarchical task structure, it is possible to effectively support the phased processing requirements of subsequent task execution, while ensuring that the dependencies between hierarchical tasks are clearly visible, providing a basis for optimizing the task chain.
[0076] For example, in a multi-agent system, the target task is to help users complete a cross-platform data query and integration task. The system first obtains the description information of the target task from the user input. For example, the user's query requirement is "to count the sales data of an e-commerce platform in the past week and generate a trend chart." Then, the system divides the target task into two levels of tasks: the task planning stage and the API selection stage. In the task planning stage, the system needs to clarify which sub-agents will participate in the task, such as the data query agent and the data visualization agent; in the API selection stage, the system needs to determine the specific API to be used, such as the data interface provided by the e-commerce platform and the interface of the chart generation tool. This information constitutes the task data to be processed and lays the foundation for the subsequent task execution.
[0077] Step 404: Input the task data to be processed into the task processing model to execute the initial level task in at least two levels of hierarchical tasks to obtain an intermediate result, wherein the task processing model is trained based on at least two sample level tasks and label task results, and the label task result is information for labeling the expected results after executing at least two sample level tasks.
[0078] The task processing model is an algorithm model trained based on multiple sample-level tasks and corresponding label task results, and is used to predict or execute specific task processes.
[0079] Sample-level tasks are representative task instances used when training task processing models. Each sample-level task represents a sub-stage of the sample task, and there are dependencies between tasks at each level. These level tasks work together to complete the entire target task, and their outputs serve as input to subsequent level tasks or as part of the final result.
[0080] Labeled task results are annotations of the expected comprehensive results after executing a series of sample-level tasks. They not only reflect the completion of a single-level task, but also reflect the overall quality of the task completed by combining multiple levels of tasks. Labeled task results provide a critical supervisory signal for training task processing models, enabling the model to learn how to map from input task data to the correct multi-level task results.
[0081] In practical applications, task data is input into a task processing model to execute the initial level of tasks in at least two hierarchical tasks. First, the input task data must be properly preprocessed to ensure that its format meets the model's requirements. The system can directly feed the raw task data into the task processing model, or extract key features based on the task requirements before inputting them into the model. Key features refer to attribute information that is highly relevant to the target task and typically include user behavior characteristics, object attribute characteristics, and contextual characteristics. For example, in an intelligent customer service system, if a user requests "Help me find a nearby restaurant with a rating higher than 4 points," the system can extract multiple key features from the request: "geographic location" (based on the user's device location), "rating requirement" (≥4 points), and "catering type" (if explicitly specified by the user, such as "Sichuan cuisine"). Once the model receives the input data, it begins executing the initial level of tasks, generating intermediate results through a series of complex computational processes. For the task processing model in this step, one option is to build a model using deep learning methods, while another option is to use a rule-based expert system. Regardless of the approach, the goal is to use the trained model to analyze the input data and output meaningful intermediate results, providing a basis for subsequent layer-by-layer tasks. After multiple iterations of layer-by-layer tasks, a comprehensive result set is ultimately generated. This result set integrates the outputs of all layer-by-layer tasks to form the labeling task result.
[0082] In the embodiments of this specification, by inputting the task data to be processed into the task processing model trained based on multiple sample-level tasks and label task results to execute the initial-level task, an intermediate result containing contributions from multiple-level tasks can be effectively obtained. This not only improves the automation level of task execution, but also enhances the ability to understand complex task structures, while ensuring the accuracy and completeness of the final result.
[0083] For example, in a multi-agent system (MAS), the goal is to create a detailed user travel plan. The task data to be processed includes information such as the user's preferences, budget constraints, and schedule. This data is input into a task processing model trained on multiple sample hierarchical tasks (such as destination selection, itinerary planning, and accommodation booking) and the corresponding labeled task results (i.e., the complete travel plan). The model first performs the initial hierarchical task—destination selection. By analyzing the user's provided preference information, it identifies several possible destination options and forms a preliminary list of recommendations. The model then sequentially performs subsequent hierarchical tasks, such as itinerary planning and accommodation booking. Each step builds on the output of the previous step to gradually refine and improve the travel plan. Ultimately, the outputs of all hierarchical tasks are combined into a complete travel plan, which is the labeled task result. This fully reflects the collaborative work of multiple hierarchical tasks.
[0084] Furthermore, at least two levels of tasks include a first-level task and a second-level task subsequent to the first-level task; inputting the task data to be processed into the task processing model to execute the initial level task among the at least two levels of hierarchical tasks to obtain an intermediate result, including: inputting the task data to be processed into the task processing model to execute the first level task to obtain an intermediate result; inputting the intermediate result into the task processing model, iteratively executing other level tasks among the at least two levels of hierarchical tasks except the initial level task until the task processing result of the target task is obtained, including: inputting the intermediate result into the task processing model to execute the second level task to obtain the task processing result.
[0085] In actual applications, the task data to be processed is input into the task processing model to execute the initial level task in at least two levels of hierarchical tasks, that is, the first level task, to obtain an intermediate result. Specifically, the system first preprocesses the input data appropriately to ensure that its format meets the requirements of the model, and then sends this data to the task processing model to execute the first level task. This process generates preliminary intermediate results through a series of calculations. For example, in a task on an online education platform, the first level task may be to analyze the learning interests of students. By processing data such as students' browsing history and learning time, a preliminary analysis report on the students' areas of interest is generated as an intermediate result.
[0086] Next, the system takes the above intermediate results as input and feeds them into the task processing model again, iteratively executing other level tasks in at least two levels until the task processing result of the target task is obtained. What is emphasized here is that the output of each level task is strictly used as the input of the next level task, forming a multi-stage task processing process. Taking the previous online education platform as an example, the second-level task can be to recommend corresponding course content based on the student's interest areas derived from the first-level task. Therefore, the system takes the interest analysis report generated by the first-level task as input, further executes the second-level task, and finally generates a personalized course recommendation list for the student as the task processing result.
[0087] In the embodiments of this specification, by inputting the task data to be processed into the task processing model in sequence, the first-level task is executed first and then the subsequent-level tasks are iteratively executed based on its output. This strict input-output relationship between levels ensures that the tasks of each stage can be analyzed and processed more deeply based on the results of the previous stage, thereby improving the accuracy and effectiveness of task processing and making the final target task results more in line with actual needs.
[0088] For example, an online education platform hopes to provide users with personalized learning path planning services. The task data to be processed includes the user's basic information, learning history, and preference settings. First, these data are input into the task processing model to perform the first-level task - learning interest analysis, identify the subject areas of interest to the user, and generate an interest analysis report as an intermediate result. Subsequently, this report is used as the input for the second-level task, namely course recommendation, which selects the most suitable courses from the many courses on the platform and recommends them to the user based on the user's interest areas. In this process, each level of task is based on the output of the previous level of task, ensuring the consistency and accuracy of the entire task chain, and ultimately providing users with a highly customized learning plan. This processing approach not only improves the user experience, but also increases user engagement and satisfaction.
[0089] Furthermore, the first-level task is a task planning subtask, and the second-level task is an interface calling subtask; the data of the task to be processed is input into the task processing model to execute the first-level task and obtain an intermediate result, including: inputting the data of the task to be processed into the task processing model to execute the task planning subtask and obtain the identifier of the interface to be called; inputting the intermediate result into the task processing model to execute the second-level task and obtain the task processing result, including: inputting the identifier of the interface to be called into the task processing model to execute the interface calling subtask and obtain the task processing result output by the interface to be called.
[0090] The Task Planning subtask is a first-level task. Its purpose is to analyze and decompose the task execution path based on the task data to be processed, and determine the specific interface call sequence or call strategy required to complete the target task. The intermediate results output by the Task Planning subtask include key information for subsequent operations, such as interface identifiers, execution sequence, and input parameters.
[0091] The interface call subtask is a second-level task. Based on the interface identification information output by the task planning subtask, it actually initiates the interface call request and receives the data returned by the interface as the task processing result. This subtask typically involves communication with external systems, services, or agents to ensure the execution of the task logic.
[0092] In actual applications, the first-level task is executed by inputting the task data to the task processing model. Specifically, the task description or structured data entered by the user is fed into the task processing model. The task planning module in the model analyzes the task intent and identifies the required external interface resources. For example, in an intelligent customer service scenario, if a user asks "Please check the weather conditions for the next three days," the system will identify the need to call the "Weather Query API" through the task planning subtask and extract the parameter "Time Range = 3 Days."
[0093] Subsequently, the intermediate results generated by the task planning subtask are fed into the task processing model to execute the second-level task. Specifically, this involves using the interface identifier (e.g., weather query API name) and related parameters (e.g., city name, time period) obtained in the previous step to trigger the actual execution of the interface call subtask. The system constructs a request message that meets the requirements and sends it to the target interface. After waiting for a response, it obtains structured data, such as the daily maximum and minimum temperatures, and precipitation forecasts for the next three days. This data is then compiled into the final task processing result and returned to the user.
[0094] For the task planning subtask in the step, one optional method is to adopt a rule-based reasoning mechanism, and another optional method is to use a sequence generation model for multi-step decision-making; for the interface call subtask, one optional method is to call the RESTful API through the HTTP protocol, and another optional method is to call the internal microservice interface through the remote procedure call RPC method.
[0095] In the embodiments of this specification, by dividing tasks into two clear levels: task planning and interface calling, a layer-by-layer transformation from abstract understanding to specific execution is achieved, which improves the structural and scalability of the task processing flow and facilitates the introduction of training and optimization strategies for different levels.
[0096] For example, in a multi-agent system (MAS), a user makes a request: "Book me a high-speed rail ticket to Beijing." The system first inputs this request as pending task data into the task processing model to execute the task planning subtask. It then identifies the interfaces to be called as the "High-Speed Rail Ticket Remaining Query Interface" and the "Ticket Purchase Interface," and extracts the parameters "Departure = Shanghai," "Destination = Beijing," and "Date = Current Date + 1 Day." The system then inputs these interface identifiers and parameter information into the task processing model, executes the interface call subtask, and sequentially initiates requests to the remaining ticket interface and the ticket purchase interface to obtain a list of available trains, complete the ticket order, and ultimately return the order confirmation information to the user as the task processing result. This entire process embodies the complete chain from task planning to determine interface dependencies and interface calls to implement functional implementation.
[0097] Step 406: Input the intermediate result into the task processing model, and iteratively execute the tasks of at least two levels of hierarchical tasks except the initial level task until the task processing result of the target task is obtained.
[0098] During this process, the intermediate results can be used to execute tasks at other levels separately, or the output of the previous level can be used as the input of the next level to gradually refine and improve the task results.
[0099] For example, when processing complex tasks, the initial-level task may output some preliminary information, such as the destination selection in the user's travel plan. This information is then used as input for the itinerary planning-level task, which is further refined into specific activity arrangements. The results of itinerary planning can then serve as input for the accommodation booking-level task to help determine a suitable place to stay. At the same time, other intermediate results may be directly used to execute other independent-level tasks. For example, the budget allocation task can directly use the results of destination selection and itinerary planning to develop a detailed spending plan without having to rely on the output of the previous-level task in a strict order.
[0100] For example, the destination selection results are used as input to develop a detailed activity schedule. Simultaneously, the budget allocation task can run independently, leveraging the destination selection and itinerary planning results to create a detailed spending plan. Finally, the outputs of all tasks are combined into a complete travel plan, which is the result of the labeling task. This demonstrates how tasks at different levels can collaborate, either through sequential iteration or parallel processing.
[0101] Furthermore, after obtaining the task processing result of the target task, it also includes: sending the task processing result to the front end; receiving correction information returned by the front end for the task processing result; and training the task processing model based on the correction information.
[0102] The front end is the interface or platform for interacting with users. It is responsible for displaying task processing results and receiving user feedback and correction information.
[0103] Correction information is feedback provided by users or operators based on task processing results. It may include direct modifications to task processing results, additional explanations, or pointing out errors. This feedback is crucial to improving system performance and accuracy.
[0104] In real-world applications, after obtaining the task processing results for the target task, the system first sends these results to the front-end for display. For example, in a multi-agent system (MAS), after the system completes processing a user's request to "find nearby restaurants that offer vegetarian options," it sends the task processing results, including a list of eligible restaurants, to the user's mobile app for viewing. The system then waits for corrections to the task processing results from the front-end. For example, if a user discovers that a restaurant does not actually offer vegetarian options, they can submit this correction to the system through the front-end interface.
[0105] Once the correction information is received, the system will use it to train the task processing model. Specifically, the system can add the correction information as a new sample to the existing training dataset, or directly use the correction information to adjust model parameters to optimize its performance. For example, in the above example, the system can use the user feedback to update the knowledge base about restaurant service content and retrain the relevant entity extraction or interface call models, so that similar queries in the future can obtain more accurate task processing results.
[0106] In the embodiments of this specification, by sending the task processing results to the front end for display and further training the task processing model based on the correction information received from the front end, not only the transparency of the system and user participation are improved, but also the learning ability and adaptability of the model are enhanced, which helps to continuously improve the system performance and make it meet user needs more accurately.
[0107] For example, in a MAS environment of an online education recommendation system, the target task is to recommend courses that suit students' interests and learning levels. After completing the preliminary course recommendations, the system sends these recommendation results to the students' personal learning portal for display. If students feel that some recommendations do not meet their interests or ability levels, they can submit correction information through the learning portal, such as marking inappropriate courses or suggesting more suitable course types. After the system collects this correction information, it uses it to train the task processing model, especially for optimizing the interest analysis part in the task planning subtask and the course matching algorithm in the interface call subtask. In this way, as more correction information is accumulated and applied, the system can continuously learn and adjust the recommendation strategy to better serve students' learning needs.
[0108] Furthermore, sending the task processing results to the front end includes: determining at least two levels of task results corresponding to at least two levels of hierarchical tasks, wherein at least two levels of task results include intermediate results; generating a task processing process report based on the hierarchical task results and the task processing results; and sending the task processing process report to the front end.
[0109] The task processing report is a document that summarizes and explains the results of tasks at all levels during the entire task execution process. It helps users understand how the tasks at each stage are completed step by step.
[0110] In actual applications, the system generates a task processing report based on the results of these hierarchical tasks and the final task processing results. This report not only includes the specific results of each hierarchical task, but also uses different identification methods (such as highlighting, bolding, color changes, etc.) to differentiate the results of different levels or stages, so that users can quickly identify and understand the content of each part. For example, in the MAS environment of an online education recommendation system, assuming that the first-level task is student interest analysis and the second-level task is course matching, then:
[0111] First-level task results: The results of the student interest area analysis will be specially marked, such as highlighted with a blue background and bold font, "[interest area] students mainly show a strong interest in mathematics and physics."
[0112] Second-level task results: The course list recommended based on interest analysis will be highlighted in green, "[Course Recommendation] Based on your interests, we recommend the following courses for you: 1. Introduction to Advanced Calculus; 2. Foundations of Classical Mechanics."
[0113] Next, the system sends the task processing report containing the detailed task processing process to the front end. Specifically, this report will list in detail the entire process from inputting the task data to be processed, through entity extraction, interface calls and other steps to obtaining the final task processing results. For the online education recommendation system in the above example, the report may be described as follows: "First, the system receives the student's personal information and learning history as the task data to be processed. Then, it performs the first-level task - interest analysis, and identifies that the student's main areas of interest are mathematics and physics (shown here with a blue background and bold font). Subsequently, this result is used to perform the second-level task - course matching, and screen out several advanced courses suitable for the student (highlighted in green here)."
[0114] In this embodiment, detailed task processing reports are generated and sent to the front-end, improving transparency within the task processing process. This allows users to clearly see the results of each level of tasks and their role in the overall task chain. This differentiated identification method makes key information readily available, enhancing the user experience while also making it easier for users to provide feedback or make corrections to specific areas.
[0115] For example, in a multi-agent system (MAS), the objective is to provide users with personalized health advice. The task data to be processed includes the user's health status, medical history, and lifestyle preferences. The system first performs the first-level task—health assessment. Results such as "[Health Assessment] You are at slight risk for high blood pressure" are displayed in bold orange text. Based on this result, the system then performs the second-level task—personalized recommendation formulation. Results such as "[Recommendation] We recommend reducing salt intake and increasing daily exercise" are highlighted in green text. Finally, the system sends a report on the entire task processing process to the front-end, ensuring that users can clearly see each step and its results, from health assessment to personalized advice, allowing them to better understand and accept the provided health advice.
[0116] Furthermore, the task processing model is trained by the following method: obtaining sample task data of the sample task and the label task result corresponding to the sample task, wherein the sample task includes at least two levels of sample-level tasks; inputting the sample task data into the initial processing model to execute the initial sample-level task to obtain a sample intermediate result; inputting the sample intermediate result into the initial processing model, iteratively executing other sample-level tasks except the initial sample-level task in at least two levels of sample-level tasks until a predicted processing result of the sample task is obtained; training the initial processing model based on the predicted processing result and the label task result to obtain a task processing model.
[0117] Sample task data refers to a dataset of task instances used to train task processing models. It includes a detailed description of the task and various information involved in its execution. Labeled task results are information that annotates the expected results after executing the sample task. They provide a supervisory signal for the model to learn how to map inputs to correct outputs.
[0118] Sample-level tasks are the different execution phases or steps within a sample task. Each sample-level task represents a sub-phase of the target task, and there are dependencies between tasks at each level. For example, in an intelligent customer service system, the first-level task might be to identify the core topic of a user's question, while the second-level task might be to provide specific answers or suggestions based on the identified topic.
[0119] In practical applications, sample task data and corresponding labeling task results are obtained for a sample task, where the sample task consists of at least two levels of sample-level tasks. First, the sample task data is fed into an initial processing model to execute the initial sample-level task, obtaining sample intermediate results. These sample intermediate results are then used as input and fed back into the initial processing model to iteratively execute additional sample-level tasks until a predicted processing result for the sample task is obtained. It is noteworthy that this approach does not directly evaluate the accuracy of the output of the first stage. Instead, it uses the output of the first stage to execute the task of the second stage and indirectly judges the quality of the first stage output by evaluating the correctness of the execution result of the second stage task. For example, in the context of an intelligent customer service system, even if the first stage accurately identifies the core topic of a user's question, if this identification does not help the second stage provide an appropriate answer, the first stage is considered to have performed poorly. Based on the difference between the predicted processing results and the labeling task results, the initial processing model is trained to optimize its performance, ultimately obtaining a task processing model. This process not only verifies the accuracy of the planning but also promotes the continuous optimization of the planning output to adapt to the second stage task, thereby improving the performance of the entire task chain. This reward model requires that the first stage must not only be correct but also adapt to the second stage tasks, so that it can be scored and optimized at a deeper level, ensuring that the overall link effect can be improved by training only the first stage tasks.
[0120] In the examples of this specification, the task processing model constructed using this method can effectively ensure that the output of each stage of the task is not only accurate at its own level, but also effectively supports the correct execution of tasks in subsequent stages. This not only improves the overall performance of the model, but also promotes better collaboration between tasks at all levels, making the entire task processing process more efficient.
[0121] For example, in a multi-agent system (MAS), the goal is to provide a user with a personalized health plan. First, the system obtains sample task data, including the user's health status, medical history, and lifestyle preferences, and assigns corresponding labeled task results—the ideal health plan. The system then feeds this data into an initial processing model to perform the first-level task—health assessment—and arrive at a preliminary conclusion, such as "[Health Assessment] You are at slight risk for hypertension." The system then uses this conclusion to proceed to the second-level task—personalized recommendation formulation, offering specific recommendations, such as "[Recommendation] We recommend reducing salt intake and increasing daily exercise." Throughout this process, the system does not directly evaluate the output of the first-level task, but instead focuses on whether the second-level task can provide appropriate recommendations based on the first-level results. This approach indirectly assesses the quality of the first-level task and adjusts model parameters accordingly, optimizing the first-stage output to better support the second-stage task, thereby improving the overall quality of the health plan. Ultimately, based on multiple iterations of training, the system generates an efficient health recommendation model, significantly improving the user experience.
[0122] Furthermore, based on the prediction processing results and the labeling task results, the initial processing model is trained, including: in response to the selection operation of the front end, determining the sample-level tasks to be trained in at least two levels of sample-level tasks; determining the parameters to be adjusted corresponding to the sample-level tasks to be trained in the initial processing model; based on the prediction processing results and the labeling task results, training the parameters to be adjusted in the initial processing model.
[0123] Parameters to be adjusted refer to the set of parameters in the initial processing model that need to be adjusted and optimized based on the prediction processing results and labeling task results. These parameters directly affect the model's ability to understand and process input data.
[0124] In practical applications, first, in response to the front-end selection operation, the sample-level task to be trained is determined from at least two levels of sample-level tasks. For example, in an intelligent customer service system, assuming that user feedback indicates that the second-level task (such as providing specific answers or suggestions) is not ideal, the system selects this level as the sample-level task to be trained. Next, the parameters to be adjusted corresponding to the sample-level task to be trained are determined in the initial processing model. Continuing with the above example, if the second-level task involves an information retrieval process based on the output of the first-level task, the relevant parameters may include weights, thresholds, etc. in the information retrieval algorithm.
[0125] Based on the predicted processing results and the labeling task results, the parameters to be adjusted in the initial processing model are trained. This process focuses on optimizing the performance of tasks at a specific level by comparing the differences between the predicted results and the actual results. For example, when it is found that the second-level task fails to accurately provide the specific answer required by the user, the system will focus on adjusting the parameters related to the task at that level to improve its accuracy. This method allows the system to optimize only for specific level tasks that perform poorly, rather than retraining the entire model, thereby improving efficiency and reducing unnecessary computational overhead. For the training method in the step, one optional way is to use the gradient descent method to update the model parameters, and another optional way is to use a reinforcement learning strategy to adjust the parameter values. Regardless of which method is used, the goal is to make the task output of the level to be trained closer to the labeling task result, thereby improving the performance of the overall task link.
[0126] In the examples of this specification, by responding to front-end selections, accurately locating the hierarchical tasks that need improvement and adjusting the relevant parameters accordingly, the performance of these hierarchical tasks can be effectively improved without requiring large-scale adjustments to the entire model. This not only improves training efficiency but also ensures that the model can quickly adapt to new requirements and changes, enhancing the flexibility and adaptability of the system.
[0127] For example, in a multi-agent system (MAS), the objective is to provide users with personalized health advice. The system first obtains sample task data, including the user's health status, medical history, and lifestyle preferences, and sets corresponding labeled task results—the ideal health advice plan. After one iteration of training, the system discovered that while the first-level task—health assessment—accurately identified the user's health risks (e.g., "[Health Assessment] You are at a slight risk of hypertension"), the second-level task—personalized recommendation—failed to provide effective advice (e.g., "[Recommendation] We recommend reducing salt intake and increasing daily exercise," even though the user had already taken such measures). Therefore, in response to front-end operations, the system selected the second-level task as the sample task for training and determined the parameters to be adjusted for this task in the initial processing model, such as the priority weight in the recommendation algorithm. Subsequently, based on the discrepancy between the predicted processing results and the labeled task results, the system adjusted these parameters, enabling the second-level task to more accurately provide personalized recommendations tailored to the user's actual situation based on the first-level results. In this way, the system not only improves the performance of the second-level tasks, but also indirectly verifies the effectiveness of the first-level tasks, achieving the goal of significantly improving the overall task processing effect by only training some levels of tasks.
[0128] Furthermore, the task processing model training method also includes: obtaining a target output result of a target sample level task and a set output result corresponding to the target sample level task, wherein the target sample level task is at least one of the sample level tasks of at least two levels; determining an illusion loss of the target sample level task based on the target output result and the set output result, wherein the illusion loss is used to characterize whether there is an illusion output in the target output result that is not included in the set output result; training an initial processing model based on the prediction processing result and the label task result to obtain a task processing model, including: determining a prediction loss based on the prediction processing result and the label task result; training an initial processing model based on the prediction loss and the illusion loss to obtain a task processing model.
[0129] Target sample-level tasks refer to one or more specific tasks selected from at least two sample-level tasks to further optimize model performance. The target output is the actual output after executing these level tasks, while the target output is the expected correct output based on the labeling task results.
[0130] Hallucination loss is a metric used to assess whether the target output contains additional information or erroneous information that is not contained in the target output (the so-called "hallucination output"). This loss is mainly used to characterize the degree of deviation between the model output and the expected output, especially the additional information that should not be present.
[0131] In practical applications, the target output result of the target sample level task and its corresponding set output result are first obtained. For example, in an intelligent customer service system, if the first-level task is to identify the core topic of the user's question and the second-level task is to provide an answer based on the topic, the second level can be selected as the target sample level task, and its actual output and expected output can be obtained respectively. Then, based on the target output result and the set output result, the hallucination loss of the target sample level task is determined. This process aims to check and quantify any additional or erroneous information in the output to ensure that the results generated by the model strictly meet expectations. If hallucination output is detected, it indicates that there may be serious errors in the task at that level, and it needs to be adjusted even if the prediction loss is low.
[0132] Next, based on the prediction processing results and the labeling task results, the prediction loss is determined. The prediction loss reflects the gap between the overall output of the model and the ideal result. However, in the embodiments of this specification, it is particularly emphasized that when hallucinations occur, the layer where the hallucinations occur needs to be trained regardless of the prediction loss. This is because the illusion loss has a more direct impact on user experience and trust than the prediction loss, so it needs to be solved first. For example, if the second-level task provides information that does not meet user needs, targeted parameter adjustments must be made to this layer to eliminate unnecessary hallucination outputs. Ultimately, the training process of the initial processing model is guided by the prediction loss and the hallucination loss to obtain an optimized task processing model.
[0133] In the examples presented herein, by introducing the concept of hallucination loss and assigning it a higher priority, we can effectively reduce unnecessary or erroneous information in the model output, significantly improving user experience and satisfaction. This approach not only focuses on the overall accuracy of the model but also places particular emphasis on the specific performance of each layer of the task, ensuring that each step accurately and flawlessly completes its designated task.
[0134] For example, in a multi-agent system (MAS), the goal is to provide users with personalized health advice. The system first obtains sample task data, including the user's health status, medical history, and lifestyle preferences, and sets corresponding labeled task results, namely, the ideal health advice plan. After one iteration of training, the system discovered that while the first-level task—health assessment—accurately identified the user's health risks (e.g., "[Health Assessment] You are at a slight risk of hypertension"), the second-level task—personalized advice formulation—included information that was inconsistent with the user's actual situation (e.g., "[Recommendation] Recommends increasing salt intake," when the user should actually reduce it). Therefore, the system calculated the hallucination loss for the second-level task and found it high, indicating significant hallucination output. Although the prediction loss may be relatively low at this point, the presence of hallucination loss forces the system to focus on training the second-level task, adjusting relevant parameters until the hallucination loss is minimized. In this way, the system not only eliminates misleading advice but also improves the quality of the overall health advice plan, strengthening user trust in the system. Throughout the process, the system keeps phantom loss top of mind, ensuring that all recommendations are based on accurate and reliable data.
[0135] One embodiment of the present specification implements the following steps: obtaining the pending task data of the target task, wherein the target task includes at least two levels of hierarchical tasks; inputting the pending task data into the task processing model to execute the initial hierarchical task in the at least two levels of hierarchical tasks to obtain an intermediate result, wherein the task processing model is trained based on at least two sample hierarchical tasks and labeling task results, and the labeling task results are information that labels the expected results after executing at least two sample hierarchical tasks; inputting the intermediate result into the task processing model, and iteratively executing the other hierarchical tasks except the initial hierarchical task in the at least two levels of hierarchical tasks until the task processing result of the target task is obtained. By using the final labeling task result to reversely adjust and optimize the hierarchical tasks of at least two sample levels, it is ensured that each sample level is considered correct only when the entire task process achieves the expected goal. This enables the task processing model to not only learn the operations within a single level, but also understand and grasp the dependencies between levels, thereby achieving overall optimization of the target task. At the same time, since the learning and adjustment at each stage are oriented towards the final results, the design requirements of complex reward signals are effectively simplified, and the task processing efficiency and result quality are significantly improved while maintaining coordination between stages.
[0136] See also Figure 5 , Figure 5 This is a flowchart of a task processing model training method provided by an embodiment of this specification, which specifically includes the following steps.
[0137] Step 502: Obtain sample task data of the sample task and the label task result corresponding to the sample task, wherein the sample task includes at least two levels of sample-level tasks, and the label task result is information for labeling the expected results after executing at least two sample-level tasks.
[0138] Step 504: Input the sample task data into the initial processing model to execute the initial sample-level task among at least two levels of sample-level tasks to obtain a sample intermediate result.
[0139] Step 506: Input the sample intermediate result into the initial processing model, and iteratively execute the other sample-level tasks except the initial sample-level task in at least two levels of sample-level tasks until the predicted processing result of the sample task is obtained.
[0140] Step 508: Based on the prediction processing results and the label task results, train the initial processing model to obtain the task processing model.
[0141] It is understandable that steps 502-508 are described with reference to the training process of the aforementioned task processing model and will not be repeated here.
[0142] One embodiment of the present specification implements the following steps: obtaining sample task data and label task results corresponding to a sample task, wherein the sample task includes at least two levels of sample-level tasks; inputting the sample task data into an initial processing model to execute the initial sample-level task and obtain a sample intermediate result; inputting the sample intermediate result into the initial processing model and iteratively executing the sample-level tasks other than the initial sample-level task in the at least two levels of sample-level tasks until a predicted processing result for the sample task is obtained; and training the initial processing model based on the predicted processing result and the label task result to obtain a task processing model. By using the final label task result to reversely adjust and optimize the hierarchical tasks of at least two sample levels, it is ensured that each sample level is considered correct only when the entire task process achieves the expected goal. This enables the task processing model to not only learn the operations within a single level, but also understand and grasp the dependencies between levels, thereby achieving overall optimization of the target task. At the same time, since the learning and adjustment at each stage are guided by the final result, the design requirements of complex reward signals are effectively simplified, and the task processing efficiency and result quality are significantly improved while maintaining coordination between stages.
[0143] The following combined Figure 6 , taking the application of the task processing method provided in this specification in two task stages as an example, the task processing method is further explained. Figure 6 This is a schematic diagram of a processing process of a task processing method provided by an embodiment of this specification, which specifically includes the following steps.
[0144] Step 602: Execute the phase 1 task according to the question and the strategy model to obtain the phase 1 answer.
[0145] Step 604: Perform hallucination scoring on the stage 1 answers.
[0146] Step 606: Execute the phase 2 task according to the phase 1 answer and the policy model to obtain the phase 2 answer.
[0147] Step 608: Score the correctness of the answers in stage 2.
[0148] Step 610: Train a strategy model based on the hallucination score and the correctness score.
[0149] For details, see Figure 7 As shown, Figure 7This is a schematic diagram of the structure of an end-to-end task consistency verification reward model with a hallucination penalty mechanism, provided in one embodiment of this specification. First, the Phase 1 task is executed based on the question and the policy model to generate a Phase 1 answer. The Phase 1 answer is then hallucinated and scored to assess whether any sub-agent names are fabricated. The Phase 2 task is then executed using the Phase 1 answer and the policy model to obtain a Phase 2 answer. The Phase 2 answer is then scored for correctness to determine whether it meets the expected goal. Finally, the policy model is trained and optimized by combining the hallucination and correctness scoring results. This entire process ensures the consistency and accuracy of the outputs at each stage through a multi-stage evaluation and feedback mechanism.
[0150] Through steps 602-610, the performance of the entire task chain was effectively improved. The dual evaluation of hallucination scoring and correctness scoring not only prevented the model from fabricating sub-agent names in the first phase, but also ensured that the output of the first phase effectively supported the correct execution of the second phase. This approach not only verified the accuracy of the plan but also promoted the continuous optimization of the planning output, thereby achieving deep optimization and improved consistency across the entire task chain, significantly improving the efficiency and quality of task processing.
[0151] See also Figure 8 , Figure 8 This is a flowchart of an information processing method based on a task processing model provided by an embodiment of this specification, which is applied to a task platform and specifically includes the following steps 802-806.
[0152] Step 802: Receive a task generation request sent by a terminal device, wherein the task generation request includes request information.
[0153] Step 804: Based on the request information, a task processing model is obtained, wherein the task processing model is used to execute the initial level task in at least two levels of hierarchical tasks based on the to-be-processed task data of the target task, obtain an intermediate result, and iteratively execute the other level tasks except the initial level task in at least two levels of hierarchical tasks based on the intermediate result until the task processing result of the target task is obtained, wherein the target task includes at least two levels of hierarchical tasks, and the task processing model is trained based on at least two sample level tasks and label task results, and the label task result is information for labeling the expected results after executing at least two sample level tasks.
[0154] Step 806: Generate task information based on the task processing model, wherein the task information is used for the terminal device to execute the target task.
[0155] It should be noted that based on the model request, the corresponding task processing model is determined from at least one task processing model. One optional method is: based on the model request, the corresponding task processing model is searched from at least one task processing model included in the model library; another optional method is: based on the model request, the task processing model is obtained by training; and another optional method is: based on the model request, the task processing model is constructed, which is not limited here.
[0156] For example, based on the scene identification of the target scene, you can first search for at least one pre-trained task processing model from the model library, then based on the model specification parameters, filter out a task processing model of corresponding size from at least one task processing model, and then train the task processing model of corresponding size based on the scene input data of the target scene to obtain a task processing model that meets user needs.
[0157] At least one task processing model is obtained by training according to the training method of the task processing model. Figure 5 The embodiments of the specification are based on the same inventive concept. For specific methods, please refer to the above-mentioned model training content and will not be repeated here.
[0158] In an optional implementation of this embodiment, the model request includes a scene identifier of the target scene; and based on the model request, determining a corresponding task processing model from at least one task processing model includes:
[0159] Based on the scenario identifier of the target scenario, a task processing model adapted to the target scenario is searched from a model library, wherein the model library stores at least one task processing model adapted to different processing scenarios.
[0160] It should be noted that the model library is a database for storing and managing various pre-trained deep learning models. Multiple task processing models adapted to different processing scenarios cover different application scenarios and needs. The model library allows users to select appropriate models according to their needs, or directly use the models for task processing through API calls.
[0161] The multiple task processing models adapted to different processing scenarios are stored in the model library and are specifically designed for different processing scenarios. Each model is optimized for a specific application environment. Each task processing model is trained using the task processing model training method described above. For example, based on the target scenario identifier "image generation," a task processing model adapted for the image generation scenario can be searched from the model library.
[0162] In the embodiments of this specification, based on scenario requirements, the task processing model suitable for the scenario is accurately found through scenario identification, so that the target processing result is more accurate and fits the scenario, thereby improving user experience and task processing quality.
[0163] As an example, the task platform can provide task processing models for a variety of scenarios, such as image generation scenarios, and can provide corresponding task processing models based on the model request sent by the terminal device. Since the task processing model is selected from at least one task processing model trained based on the training method of the above-mentioned task processing model, the task processing model has been precisely trained and can realize customized image generation of single objects and multiple objects.
[0164] In an optional implementation of this embodiment, the model request includes scene input data of the target scene; and based on the model request, determining a corresponding task processing model from at least one task processing model includes:
[0165] determining an initial task processing model adapted to the target scenario from at least one task processing model;
[0166] Based on the scene input data of the target scene, the initial task processing model is trained to obtain a task processing model.
[0167] In actual implementation, the model request may include scene input data of the target scene, and the task processing model is a task processing model suitable for the target scene.
[0168] For example, a general task processing model is a basic task processing model that is trained to be adaptable to different processing scenarios, but is not optimized for any specific scenario. For example, a general task processing model is trained based on scene input data from an image generation scenario to obtain a task processing model that is adapted to the image generation scenario.
[0169] In the embodiments of this specification, based on scenario requirements, the general task processing model is further trained through scenario input data to obtain a task processing model adapted to the scenario, so that the target processing results are more accurate and fit the scenario, thereby improving the user experience and task processing quality.
[0170] In an optional implementation of this embodiment, the model request includes model specification parameters; and based on the model request, determining a corresponding task processing model from at least one task processing model includes:
[0171] Based on the model specification parameters, a corresponding task processing model is searched from a model library, wherein the model library stores a plurality of task processing models with different model specification parameters.
[0172] The model specification parameter may be a model size. For example, based on the model size of 32 GB, a task processing model of a corresponding size is searched from the model library.
[0173] In the embodiments of this specification, based on the model specification requirements, the corresponding task processing model is accurately found through the model specification parameters, which ensures the efficient and stable operation of the task processing model and improves the user experience.
[0174] In an optional implementation of this embodiment, after determining a corresponding task processing model from at least one task processing model based on the model request, the method further includes:
[0175] Deploy a task processing model and build a processing interface based on the task processing model to enable the terminal device to schedule the task processing model to execute tasks.
[0176] It should be noted that the processing interface is an interactive programming interface for the terminal device scheduling task processing model, typically provided as an API. Through the processing interface, users can input task data for the target task, such as at least one object, a reference image, and text prompt information, and effectively control the model's output, such as the generated target image.
[0177] In actual implementation, one option for deploying the task processing model is to deploy it on the task platform's distributed system. For example, the task processing model can be deployed on the task platform's distributed system and, based on the task processing model, a processing interface can be built and provided to the terminal device, allowing the terminal device to schedule the task processing model to execute the target task in the processing scenario.
[0178] In the embodiments of this specification, efficient terminal calling is achieved, task processing is optimized, and task processing quality and response speed are improved.
[0179] The information processing method based on the task processing model provided in the embodiments of this specification is applied to obtain the task processing model according to user needs, realize personalized model service, provide users with an efficient, flexible and easy-to-use model service method, and improve user experience.
[0180] Corresponding to the above method embodiment, this specification also provides a task platform embodiment, Figure 9 FIG1 shows a schematic diagram of the structure of a task platform provided by an embodiment of this specification. Figure 9 As shown, the task platform 900 includes: a request interface 902 and a response unit 904;
[0181] The request interface 902 is used to receive a task generation request sent by a terminal device, wherein the task generation request includes request information;
[0182] Response unit 904 is used to obtain a task processing model based on the request information, wherein the task processing model is used to execute the initial level task in at least two levels of hierarchical tasks based on the to-be-processed task data of the target task, obtain an intermediate result, and iteratively execute other level tasks except the initial level task in at least two levels of hierarchical tasks based on the intermediate result until the task processing result of the target task is obtained, wherein the target task includes at least two levels of hierarchical tasks, and the task processing model is trained based on at least two sample level tasks and label task results, and the label task result is information for labeling the expected results after executing at least two sample level tasks.
[0183] Optionally, the task platform further includes a processing interface, and the processing interface is constructed based on the task processing model;
[0184] Processing interface, used for terminal devices to schedule and execute tasks.
[0185] Optionally, the model request includes a scene identifier of the target scene;
[0186] The response unit 904 is further configured to:
[0187] Based on the scenario identifier of the target scenario, a task processing model adapted to the target scenario is searched from a model library, wherein the model library stores at least one task processing model adapted to different processing scenarios.
[0188] Optionally, the model request includes scene input data of the target scene;
[0189] The response unit 904 is further configured to:
[0190] determining an initial task processing model adapted to the target scenario from at least one task processing model;
[0191] Based on the scene input data of the target scene, the initial task processing model is trained to obtain a task processing model.
[0192] Optionally, the model request includes model specification parameters;
[0193] The response unit 904 is further configured to:
[0194] Based on the model specification parameters, a corresponding task processing model is searched from a model library, wherein the model library stores a plurality of task processing models with different model specification parameters.
[0195] Optionally, the task platform further includes a deployment module configured to:
[0196] Deploy a task processing model and build a processing interface based on the task processing model to enable the terminal device to schedule the task processing model to execute tasks.
[0197] In the embodiments of this specification, the task platform adapts to user needs to obtain task processing models, realizes personalized model services, provides users with an efficient, flexible and easy-to-use model service platform, and improves user experience.
[0198] The above is a schematic scheme of a task platform of this embodiment. It should be noted that the technical scheme of this task platform and the technical scheme of the information processing method based on the task processing model described above are based on the same concept. For details not described in detail in the technical scheme of the task platform, please refer to the description of the technical scheme of the information processing method based on the task processing model described above.
[0199] Corresponding to the above method embodiment, this specification also provides a task processing device embodiment, Figure 10 This is a structural diagram of a task processing device provided by an embodiment of this specification. Figure 10 As shown, the device includes:
[0200] A first acquisition module 1002 is configured to acquire to-be-processed task data of a target task, wherein the target task includes hierarchical tasks of at least two levels;
[0201] The first task execution module 1004 is configured to input the task data to be processed into the task processing model to execute the initial level task of the at least two levels of hierarchical tasks, and obtain an intermediate result, wherein the task processing model is trained based on the at least two sample level tasks and the labeling task result, and the labeling task result is information that labels the expected results after executing the at least two sample level tasks;
[0202] The second task execution module 1006 is configured to input the intermediate result into the task processing model, and iteratively execute the tasks of at least two levels except the initial level task until the task processing result of the target task is obtained.
[0203] Optionally, at least two levels of tasks include a first-level task and a second-level task subsequent to the first-level task; accordingly, the first task execution module 1004 is further configured to input the task data to be processed into the task processing model to execute the first-level task and obtain an intermediate result; accordingly, the second task execution module 1006 is further configured to input the intermediate result into the task processing model to execute the second-level task and obtain a task processing result.
[0204] Optionally, the first-level task is a task planning subtask, and the second-level task is an interface calling subtask; accordingly, the first task execution module 1004 is further configured to input the task data to be processed into the task processing model to execute the task planning subtask, and obtain the identifier of the interface to be called; accordingly, the second task execution module 1006 is further configured to input the identifier of the interface to be called into the task processing model to execute the interface calling subtask, and obtain the task processing result output by the interface to be called.
[0205] Optionally, the task processing device further includes an interactive adjustment module configured to send the task processing result to the front end; receive correction information returned by the front end for the task processing result; and train the task processing model based on the correction information.
[0206] Optionally, the adjustment module is further configured to determine at least two levels of task results corresponding to at least two levels of hierarchical tasks, wherein at least two levels of task results include intermediate results; generate a task processing process report based on the hierarchical task results and the task processing results; and send the task processing process report to the front end.
[0207] Optionally, the task processing device also includes a training module, which is configured to obtain sample task data of a sample task and a label task result corresponding to the sample task, wherein the sample task includes at least two levels of sample-level tasks; input the sample task data into the initial processing model to execute the initial sample-level task to obtain a sample intermediate result; input the sample intermediate result into the initial processing model, and iteratively execute other sample-level tasks except the initial sample-level task in at least two levels of sample-level tasks until a predicted processing result of the sample task is obtained; and train the initial processing model based on the predicted processing result and the label task result to obtain a task processing model.
[0208] Optionally, the training module is further configured to determine the sample-level tasks to be trained in at least two levels of sample-level tasks in response to the selection operation of the front end; determine the parameters to be adjusted corresponding to the sample-level tasks to be trained in the initial processing model; and train the parameters to be adjusted in the initial processing model based on the prediction processing results and the label task results.
[0209] Optionally, the training module is further configured to obtain a target output result of a target sample level task and a set output result corresponding to the target sample level task, wherein the target sample level task is at least one of the sample level tasks of at least two levels; based on the target output result and the set output result, determine the hallucination loss of the target sample level task, wherein the hallucination loss is used to characterize whether there is a hallucination output in the target output result that is not included in the set output result; based on the prediction processing result and the label task result, determine the prediction loss; based on the prediction loss and the hallucination loss, train the initial processing model to obtain the task processing model.
[0210] Applied to this task processing device, the final labeling task results are used to reversely adjust and optimize the hierarchical tasks of at least two sample levels, ensuring that each sample level is considered correct only when the entire task process achieves the expected goal. This allows the task processing model to not only learn the operations within a single level, but also understand and grasp the dependencies between levels, thereby achieving overall optimization of the target task. At the same time, because the learning and adjustment at each stage are guided by the final result, it effectively simplifies the design requirements of complex reward signals, significantly improving task processing efficiency and result quality while maintaining coordination between stages.
[0211] The above is a schematic scheme of a task processing device of this embodiment. It should be noted that the technical scheme of the task processing device and the technical scheme of the task processing method described above are of the same concept. For details not described in detail in the technical scheme of the task processing device, please refer to the description of the technical scheme of the task processing method described above.
[0212] Corresponding to the above method embodiment, this specification also provides an embodiment of a task processing model training device, Figure 11 This is a structural diagram of a task processing model training device provided by an embodiment of this specification. Figure 11 As shown, the device includes:
[0213] The second acquisition module 1102 is configured to acquire sample task data of a sample task and a label task result corresponding to the sample task, wherein the sample task includes sample-level tasks of at least two levels;
[0214] The third task execution module 1104 is configured to input the sample task data into the initial processing model to execute the initial sample level task and obtain a sample intermediate result;
[0215] The fourth task execution module 1106 is configured to input the sample intermediate result into the initial processing model, and iteratively execute the sample-level tasks except the initial sample-level task in the at least two levels of sample-level tasks until a predicted processing result of the sample task is obtained;
[0216] The training module 1108 is configured to train the initial processing model based on the prediction processing results and the label task results to obtain the task processing model.
[0217] Applied to this task processing model training device, the final labeled task results are used to reversely adjust and optimize hierarchical tasks at at least two sample levels, ensuring that each sample level is considered correct only when the entire task process achieves the intended goal. This enables the task processing model to not only learn the operations within a single level, but also understand and grasp the dependencies between levels, thereby achieving overall optimization of the target task. At the same time, because the learning and adjustment at each stage are guided by the final result, the design requirements of complex reward signals are effectively simplified, significantly improving task processing efficiency and result quality while maintaining coordination between stages.
[0218] The above is a schematic diagram of a task processing model training device according to this embodiment. It should be noted that the technical solution of the task processing model training device and the technical solution of the task processing model training method described above are based on the same concept. For details not described in detail in the technical solution of the task processing model training device, please refer to the description of the technical solution of the task processing model training method described above.
[0219] Figure 12 The following is a block diagram of a computing device 1200 according to one embodiment of the present disclosure. Components of the computing device 1200 include, but are not limited to, a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 via a bus 1230, and a database 1250 is used to store data.
[0220] Computing device 1200 also includes an access device 1240 that enables computing device 1200 to communicate via one or more networks 1260. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 1240 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0221] In one embodiment of the present specification, the above components of the computing device 1200 and Figure 12 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 12 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0222] Computing device 1200 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1200 can also be a mobile or stationary server.
[0223] Among them, the memory 1210 is used to store computer programs / instructions, and the processor 1220 is used to execute the computer programs / instructions stored in the memory 1210. When the computer program / instructions are executed by the processor, the steps of the above-mentioned task processing method or task processing model training method or information processing method based on the task processing model are implemented.
[0224] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the aforementioned task processing method, task processing model training method, or information processing method based on a task processing model are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned task processing method, task processing model training method, or information processing method based on a task processing model.
[0225] An embodiment of the present specification also provides a computer-readable storage medium storing a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned task processing method or task processing model training method or information processing method based on the task processing model.
[0226] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of this storage medium and the technical scheme of the above-mentioned task processing method, task processing model training method, or information processing method based on the task processing model are based on the same concept. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical scheme of the above-mentioned task processing method, task processing model training method, or information processing method based on the task processing model.
[0227] An embodiment of this specification also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned task processing method or task processing model training method or information processing method based on the task processing model.
[0228] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of this computer program product and the technical scheme of the task processing method, task processing model training method, or information processing method based on the task processing model described above are based on the same concept. For details not described in detail in the technical scheme of the computer program product, please refer to the description of the technical scheme of the task processing method, task processing model training method, or information processing method based on the task processing model described above.
[0229] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0230] The computer program / instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0231] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0232] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0233] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A task processing method, comprising: Get the pending task data of the target task, where: The target task includes at least two levels of hierarchical tasks; Inputting the task data to be processed into a task processing model to execute an initial level task among the at least two levels of level tasks to obtain an intermediate result, wherein the task processing model is obtained by training the sample-level tasks of at least two sample levels based on the prediction loss and the hallucination loss, the training priority of the hallucination loss is higher than the prediction loss, the prediction loss is determined based on the labeling task result and the prediction processing results of the sample-level tasks of the at least two sample levels, the hallucination loss is determined based on the target output result and the set output result of each sample-level task, the labeling task result is information for labeling the expected results after executing the at least two sample-level tasks, and the set output result is the output of each sample level expected by the labeling task result; The intermediate result is input into the task processing model, and the other level tasks except the initial level task in the at least two levels of level tasks are iteratively executed until the task processing result of the target task is obtained.
2. The method according to claim 1, wherein the at least two level tasks include a first level task and a second level task subsequent to the first level task; The step of inputting the task data to be processed into the task processing model to execute the initial level task and obtain the intermediate result includes: Inputting the to-be-processed task data into the task processing model to execute the first-level task and obtain the intermediate result; Inputting the intermediate result into the task processing model, and iteratively executing the tasks of the at least two levels except the initial level task until the task processing result of the target task is obtained, includes: The intermediate result is input into the task processing model to execute the second-level task and obtain a task processing result.
3. The method according to claim 2, wherein the first-level task is a task planning subtask, and the second-level task is an interface call subtask; The step of inputting the to-be-processed task data into the task processing model to execute the first-level task and obtain the intermediate result includes: Inputting the pending task data into the task processing model to execute the task planning subtask and obtain an identifier of the interface to be called, wherein the task planning subtask is used to determine the interface to be called required to execute the pending task; Inputting the intermediate result into the task processing model to execute the second-level task to obtain the task processing result includes: The identifier of the interface to be called is input into the task processing model to execute the interface calling subtask, and the task processing result output by the interface to be called is obtained.
4. The method according to claim 1, further comprising, after obtaining the task processing result of the target task: Determining at least two hierarchical task results corresponding to the hierarchical tasks of the at least two hierarchical levels, wherein the at least two hierarchical task results include the intermediate results; Generate a task processing process report based on the hierarchical task results and the task processing results; Send the task processing progress report to the front end.
5. The method according to claim 1, wherein the task processing model is trained by the following method: Obtain sample task data of a sample task and a label task result corresponding to the sample task, wherein: The sample tasks include at least two levels of sample-level tasks; Inputting the sample task data into the initial processing model to execute the initial sample level task and obtain a sample intermediate result; Inputting the sample intermediate result into the initial processing model, and iteratively executing the other sample-level tasks except the initial sample-level task among the at least two levels of sample-level tasks until a predicted processing result of the sample task is obtained; Based on the prediction processing result and the label task result, the initial processing model is trained to obtain a task processing model.
6. The method according to claim 5, wherein training the initial processing model based on the prediction processing result and the labeling task result comprises: In response to a selection operation of the front end, determining a sample-level task to be trained among the sample-level tasks of the at least two levels; Determining, in the initial processing model, parameters to be adjusted corresponding to the sample-level task to be trained; Based on the prediction processing result and the labeling task result, the parameters to be adjusted in the initial processing model are trained.
7. The method according to claim 5, further comprising: Obtaining a target output result of a target sample-level task and a set output result corresponding to the target sample-level task, wherein the target sample-level task is at least one of the sample-level tasks of the at least two levels; Determining a hallucination loss of the target sample-level task based on the target output result and the set output result, wherein the hallucination loss is used to characterize whether there is a hallucination output in the target output result that is not included in the set output result; The step of training the initial processing model based on the prediction processing result and the label task result to obtain a task processing model includes: Determining a prediction loss based on the prediction processing result and the labeling task result; Based on the prediction loss and the hallucination loss, the initial processing model is trained to obtain a task processing model.
8. A task processing model training method, comprising: Obtaining sample task data of a sample task and a labeling task result corresponding to the sample task, wherein the sample task includes at least two levels of sample-level tasks, and the labeling task result is information that labels expected results after executing the at least two sample-level tasks; Inputting the sample task data into an initial processing model to execute an initial sample-level task among the at least two levels of sample-level tasks to obtain a sample intermediate result; Inputting the sample intermediate result into the initial processing model, and iteratively executing the other sample-level tasks except the initial sample-level task among the at least two levels of sample-level tasks until a predicted processing result of the sample task is obtained; A prediction loss is determined based on the prediction processing result and the label task result, an illusion loss is determined based on the target output result of each sample level task and the set output result, and the sample level tasks of the at least two sample levels are trained based on the prediction loss and the illusion loss to obtain a task processing model, wherein the training priority of the illusion loss is higher than the prediction loss, and the set output result is the output of each sample level expected by the label task result.
9. An information processing method based on a task processing model, applied to a task platform, comprising: Receiving a task generation request sent by a terminal device, wherein the task generation request includes request information; Based on the request information, a task processing model is obtained, wherein the task processing model is used to execute an initial level task in at least two levels of hierarchical tasks based on the to-be-processed task data of the target task, obtain an intermediate result, and iteratively execute other level tasks in at least two levels of hierarchical tasks except the initial level task based on the intermediate result until the task processing result of the target task is obtained, wherein the target task includes the at least two levels of hierarchical tasks, the task processing model is obtained by training sample-level tasks of at least two sample levels based on prediction loss and hallucination loss, the training priority of the hallucination loss is higher than the prediction loss, the prediction loss is determined based on the labeling task result and the prediction processing results of the sample-level tasks of the at least two sample levels, the hallucination loss is determined based on the target output result and the set output result of each sample-level task, the labeling task result is information for labeling the expected result after executing the at least two sample-level tasks, and the set output result is the output of each sample level expected by the labeling task result; Based on the task processing model, task information is generated, wherein the task information is used by the terminal device to perform the target task.
10. The method according to claim 9, wherein the request information includes a task scenario identifier of the target task, or a task model identifier; The acquiring of the task processing model based on the request information includes: Based on the task scenario identifier, determining a target scenario template from a plurality of preset scenario templates, and based on the target scenario template, searching for a task processing model from a model library, wherein the model library stores a plurality of task processing models; or, Based on the task model identifier, a task processing model is searched from the model library.
11. The method according to claim 9, wherein the request information includes sample task data of a sample task; The acquiring of the target task model based on the request information includes: Based on the sample task data, an initial processing model corresponding to the sample task is trained to obtain a trained task processing model.
12. A task platform comprising a request interface and a response unit; The request interface is used to receive a task generation request sent by a terminal device, wherein: The task generation request includes request information; The response unit is used to obtain a task processing model based on the request information, wherein the task processing model is used to execute an initial level task in at least two levels of level tasks based on the target task to be processed task data, obtain an intermediate result, and iteratively execute other level tasks in at least two levels of level tasks except the initial level task based on the intermediate result until the task processing result of the target task is obtained, wherein the target task includes the at least two levels of level tasks, the task processing model is obtained by training the sample level tasks of at least two sample levels based on the prediction loss and the hallucination loss, the training priority of the hallucination loss is higher than the prediction loss, the prediction loss is determined based on the label task result and the prediction processing results of the sample level tasks of the at least two sample levels, the hallucination loss is determined based on the target output result and the set output result of each sample level task, the label task result is information for labeling the expected result after executing the at least two sample level tasks, and the set output result is the output of each sample level expected by the label task result.
13. The task platform according to claim 12, further comprising a model library, wherein the model library stores a plurality of task processing models; the request information comprises a task scenario identifier of the target task, or a task model identifier; The response unit is configured to determine a target scenario template from a plurality of preset scenario templates based on the task scenario identifier, and search a task processing model from a model library based on the target scenario template; or, Based on the task model identifier, a task processing model is searched from the model library.
14. A computing device comprising: memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 11.
16. A computer program product comprising a computer program / instruction, which implements the steps of the method according to any one of claims 1 to 11 when executed by a processor.
Citation Information
Patent Citations
Multi-stage task processing method and device based on intelligent Agent model and medium
CN118656196A
Financial question and answer method, system and equipment based on multi-agent interaction and medium
CN119539095A