Autonomous task processing method and device, storage medium and electronic equipment

By performing step-by-step behavior quality evaluation on the task execution action sequence, fine-grained action supervision and evaluation information are generated, and based on this information, the inference process of the agent service object is optimized, and the problem of lack of process supervision and relying on manual annotation in the existing technology is solved, and efficient learning and reasoning of the agent is achieved.

CN120144236APending Publication Date: 2025-06-13ALIPAY (HANGZHOU) INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510180434.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing technology lacks process supervision in task autonomous processing, relies on manual annotation and difficulty in achieving inference time expansion, resulting in insulting the learning and reasoning of agents.

Method used

By performing step-by-step behavior quality evaluation on the task execution action sequence, fine-grained action supervision and evaluation information is generated, and the reasoning process of the agent service object is optimized based on this information.

Benefits of technology

It realizes precise positioning of the task execution process, improves the learning and reasoning efficiency of the agent, reduces the need for manual labeling, reduces processing costs, and enhances the processing capabilities of complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144236A_ABST
    Figure CN120144236A_ABST
Patent Text Reader

Abstract

The invention discloses a task autonomous processing method and device, a storage medium and electronic equipment, and the method comprises the steps: determining a transaction task of an agent service object corresponding to a transaction processing large model, based on the transaction task, an agent service object is adopted to call the transaction processing large model to perform task autonomous reasoning processing to obtain a task execution action sequence, the task execution action sequence comprises at least one task execution action, and gradual behavior quality evaluation processing for the transaction task is performed on the task execution action. And obtaining action supervision evaluation information of each task execution action, and performing task autonomous reasoning optimization processing on the agent service object based on the action supervision evaluation information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and particularly to a method, apparatus, storage medium, and electronic device for autonomous task processing. Background Art

[0002] With the development of multi-modal large language models (MLLMs) and the increasing complexity of transaction services provided by electronic devices, autonomous transaction service interactions have become increasingly important. Users hope that intelligent agent service objects driven by multi-modal large language models can implement input transaction tasks, and then autonomously process the transaction tasks based on the intelligent agent service objects, which can effectively achieve transaction service navigation and interaction with transaction services, while providing convenient and easy-to-use intelligent interaction services. Summary of the Invention

[0003] This specification provides a method, apparatus, storage medium, and electronic device for autonomous task processing. The technical solutions are as follows:

[0004] In a first aspect, this specification provides a method for autonomous task processing. The method includes:

[0005] Determine the transaction task of the intelligent agent service object corresponding to the transaction processing large model;

[0006] Based on the transaction task, use the intelligent agent service object to call the transaction processing large model to perform autonomous task reasoning processing to obtain a task execution action sequence, where the task execution action sequence includes at least one task execution action;

[0007] Perform step-by-step behavior quality evaluation processing on the task execution actions for the transaction task to obtain action supervision evaluation information for each task execution action;

[0008] Based on the action supervision evaluation information, perform autonomous task reasoning optimization processing on the intelligent agent service object.

[0009] In a second aspect, this specification provides an apparatus for autonomous task processing. The apparatus includes:

[0010] A task determination module for determining the transaction task of the intelligent agent service object corresponding to the transaction processing large model;

[0011] An autonomous reasoning module for using the intelligent agent service object to call the transaction processing large model to perform autonomous task reasoning processing based on the transaction task to obtain a task execution action sequence, where the task execution action sequence includes at least one task execution action;

[0012] A quality evaluation module for performing step-by-step behavior quality evaluation processing on the task execution actions for the transaction task to obtain action supervision evaluation information for each task execution action;

[0013] An inference optimization module for performing task autonomous inference optimization processing on the intelligent agent service object based on the action supervision evaluation information.

[0014] In a third aspect, this specification provides a computer storage medium storing at least one instruction, and the instruction is adapted to be loaded and executed by a processor to perform the method steps of one or more embodiments of this specification.

[0015] In a fourth aspect, this specification provides a computer program product storing at least one instruction, and the instruction is adapted to be loaded and executed by a processor to perform the method steps of one or more embodiments of this specification.

[0016] In a fifth aspect, this specification provides an electronic device, which may include: a processor and a memory; wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the method steps of one or more embodiments of this specification.

[0017] The beneficial effects brought by the technical solutions provided in some embodiments of this specification at least include:

[0018] In one or more embodiments of this specification, by determining the transaction task of the intelligent agent service object corresponding to the transaction processing large model, using the intelligent agent service object to call the transaction processing large model for task autonomous inference processing based on the transaction task to obtain a task execution action sequence, and by performing step-by-step behavior quality evaluation on the task execution action sequence during the task autonomous inference process, action supervision evaluation information with a higher granularity is generated, thereby replacing the single supervision signal based on the task completion result. It effectively solves the limitations of lack of process supervision, dependence on manual annotation, and difficulty in realizing inference time expansion, can accurately locate the execution details during the task execution process, improve the learning and inference efficiency of the intelligent agent, and also can reduce the manual annotation requirements, realize the automatic generation of supervision signals during the task autonomous processing process, reduce the processing cost, enhance the processing ability of complex tasks by gradually evaluating and optimizing task inference, and thus significantly improve the task autonomous processing performance and application value of the multi-modal virtual intelligent agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions in this specification or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0020] Figure 1 is a schematic diagram of the scenario of a task autonomous processing system provided by this specification;

[0021] Figure 2 is a schematic flowchart of a task autonomous processing method provided by this specification;

[0022] Figure 3 is a schematic diagram of a scenario for comparing solutions provided by this specification;

[0023] Figure 4 is a schematic flowchart of a step-by-step behavior quality evaluation process provided by this specification;

[0024] Figure 5 is a schematic diagram of a scenario for multi-dimensional step-by-step evaluation provided by this specification;

[0025] Figure 6 is a schematic flowchart of a task autonomous reasoning optimization provided by this specification;

[0026] Figure 7 is a schematic flowchart of the model training process of a step-by-step multi-dimensional evaluation reward model provided by this specification;

[0027] Figure 8 is a schematic diagram of the structure of a task autonomous processing device provided by this specification;

[0028] Figure 9 is a schematic diagram of the structure of an electronic device provided by this specification. Detailed implementation manners

[0029] The following will clearly and completely describe the technical solutions in this specification in conjunction with the accompanying drawings in this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by this specification.

[0030] In the description of this specification, it should be understood that terms such as "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In the description of this specification, it should be noted that unless otherwise clearly specified and defined, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally also include steps or units not listed, or may optionally also include other steps or units inherent to these processes, methods, products, or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood according to specific circumstances. In addition, in the description of this specification, unless otherwise stated, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0031] In the related art, with the development of multi-modal large language models (MLLMs), general virtual agents (GVAs) have shown great potential in autonomous task execution. However, the current training and task autonomous processing evaluation paradigms have significant limitations, including the lack of process supervision and reliance on human-annotated trajectory data, which will affect the task autonomous processing effect in practical applications. General virtual agents GVAs driven by multi-modal large language models (MLLMs) can process multi-modal inputs to navigate digital environments, perform transaction tasks, and generate manipulation interfaces or provide responsive outputs. Currently, the training and actual reasoning of virtual agents use the task completion result as the main supervision signal for evaluating the quality of behavior. This task autonomous processing paradigm based on task completion results has significant limitations:

[0032] 1) Lack of multi-dimensional fine-grained process supervision: Existing methods usually focus on the global success or final state of the task, ignoring the intermediate steps during execution. This neglect makes it difficult to locate problems in failed trajectories or identify errors in successful trajectories during transaction task autonomous processing, resulting in an inefficient learning and reasoning process.

[0033] 2) Reliance on human-annotated trajectories with reward signals: Domain experts need to carefully annotate trajectories containing dozens of steps to ensure that each step has an accurate result-based reward signal for training GVAs. In addition, obtaining step-by-step fine-grained process reward signals makes the process labor-intensive, time-consuming, and difficult to implement on a large scale.

[0034] 3) Difficulty of Inference Time Expansion: The latest research shows that expanding the inference time can significantly enhance the agent performance. However, relying on result-based training and a large number of manual annotations limits the ability to handle complex tasks by gradually evaluating and selecting the best actions.

[0035] Based on this, to address all or part of the above limitations, one or more of the task autonomous processing methods described in this specification are involved. The autonomous task processing method performs step-by-step behavior quality evaluation on the task execution action sequence in the task autonomous inference processing to obtain the action supervision evaluation information for each task execution action, and then uses the action supervision evaluation information to replace the supervision signal of the intelligent agent service object GVAS in the related technology, and relies on the action supervision evaluation information to perform task autonomous inference optimization processing on the intelligent agent service object.

[0036] The following will describe this specification in detail with specific embodiments;

[0037] Please refer to Figure 1 , which is a schematic diagram of the scenario of a task autonomous processing system provided in this specification. As Figure 1 shown, the task autonomous processing system can at least include a client cluster and a service platform 100.

[0038] The client cluster can include at least one client. As Figure 1 shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2,..., client n corresponding to user n, where n is an integer greater than 0.

[0039] Each client in the client cluster can be an electronic device with communication functions, and this electronic device includes but is not limited to: wearable devices, handheld devices, personal computers, tablet computers, vehicle-mounted devices, smartphones, computing devices, or other processing devices connected to a wireless modem, etc. In different networks, the electronic device can be called different names, for example: user equipment, access terminal, user unit, user station, mobile station, mobile platform, remote station, remote terminal, mobile device, user terminal, electronic device, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), electronic device in a 5G network or future evolved network, etc.

[0040] The service platform 100 may be a separate server device, such as a rack-mounted, blade, tower, or cabinet-style server device, or a hardware device with strong computing capabilities such as a workstation or a mainframe computer; it may also be a server cluster composed of multiple servers. The servers in the service cluster may be composed symmetrically, where each server is functionally equivalent and has an equivalent status in the transaction link, and each server can provide services independently. The independent provision of services can be understood as not requiring the assistance of another server.

[0041] In one or more embodiments of the present specification, the service platform 100 can establish a communication connection with at least one client in the client cluster, and complete the interaction of data during the autonomous task processing process based on this communication connection, such as online transaction data interaction. For example, the service platform 100 can implement autonomous task processing for the transaction tasks of the client based on the autonomous task processing method described in this specification.

[0042] It should be noted that the service platform 100 and at least one client in the client cluster establish a communication connection through a network for interactive communication. Among them, the network can be a wireless network or a wired network. The wireless network includes, but is not limited to, a cellular network, a wireless local area network, an infrared network, or a Bluetooth network. The wired network includes, but is not limited to, an Ethernet, a universal serial bus (USB), or a controller area network. In one or more embodiments of the specification, technologies and / or formats including HyperText Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data (such as a target compressed package) exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0043] The embodiments of the task autonomous processing system provided in this specification and the task autonomous processing method in one or more embodiments belong to the same concept. The execution subject corresponding to the task autonomous processing method involved in one or more embodiments of the specification can be the above service platform 100; the execution subject corresponding to the task autonomous processing method involved in one or more embodiments of the specification can also be the electronic device corresponding to the client, which is specifically determined based on the actual application environment. For the embodiments of the task autonomous processing system, the implementation process can be seen in the following method embodiments in detail, which will not be elaborated here.

[0044] Based on Figure 1 the scene schematic diagram shown below, the task autonomous processing method provided in one or more embodiments of this specification will be introduced in detail.

[0045] Please refer to Figure 2 , which is a flowchart of a task autonomous processing method provided in one or more embodiments of this specification. This method can be implemented depending on a computer program and can run on a task autonomous processing device based on the von Neumann architecture. This computer program can be integrated in an application or run as an independent tool application. The task autonomous processing device can be an electronic device.

[0046] Specifically, the task autonomous processing method includes:

[0047] S102: Determine the transaction tasks of the intelligent agent service object corresponding to the transaction processing large model;

[0048] In this specification, the intelligent agent service object (Agent service object, also known as GVAs) is a computer program with perception, reasoning, and decision-making relying on an electronic device. The intelligent agent service object is the service carrier object of the intelligent agent Agent, and relies on the transaction processing large model to drive the task autonomous processing function for the user to provide transaction tasks for the intelligent agent service object;

[0049] For example, the intelligent agent service object can be a service control that carries the intelligent agent service object. The user can trigger the service control corresponding to the intelligent agent service object and input the target task for the target transaction in the intelligent agent service object interface corresponding to the service control;

[0050] The intelligent agent service object can be applied to application scenarios such as artificial intelligence, robotics, virtual reality, games, automatic control, information retrieval, recommendation systems, and natural language processing.

[0051] The transaction tasks can be user tasks in transaction scenarios such as application transactions, Web automation transactions, robotic process automation transactions, and intelligent virtual assistant transactions;

[0052] Taking transaction tasks as application transaction examples, the application transaction can be a system-based application or a third-party application, and the type of transaction tasks is not limited. Transaction tasks can provide services for users. With the popularization of electronic devices and the increasing complexity of transaction services, in one or more embodiments of this specification, for the convenience of users, based on the intelligent agent service object, it is possible to realize the transaction tasks input by the user, autonomously interact with the target transaction service indicated by the transaction task, and perform task navigation, so as to complete the autonomous processing of the transaction task and save the user's operation path for completing the target task, and improve the task processing efficiency.

[0053] A transaction task can be understood as a task request based on a certain transaction. A transaction task can be understood as the target task request information input by the user for the target transaction, that is, it refers to the request made by the user to the intelligent agent service object, asking the intelligent agent service object to directly perform the autonomous task processing of the transaction task relying on the transaction processing large model;

[0054] For example, the transaction task can be "delete all items with calendar activities enabled in the calendar application", can be "help me order an Americano latte in the application", can be "help me call an express car from place A to place B", etc., and there is no limitation on this;

[0055] In this specification, the task autonomous processing method for the large model scenario can be applied to the intelligent agent service object, that is, the user can interact with the intelligent agent service object, so that the intelligent agent service object can accurately and efficiently perform "autonomous operation of the target transaction" for the user through the transaction processing large model to complete the transaction task input by the user.

[0056] Among them, the intelligent agent service object can be associated with multiple services of the transaction. Each service can have a corresponding service function depending on the transaction. The number of service functions of the transaction can cover multiple different types of applications. The services of multiple different types of applications can run jointly to autonomously complete the transaction task on a certain application service, and realize the transformation of the human-computer interaction based on the transaction task into the interaction between the intelligent agent and the application.

[0057] S104: Based on the transaction task, use the intelligent agent service object to call the transaction processing large model to perform task autonomous reasoning processing to obtain a task execution action sequence, and the task execution action sequence includes at least one task execution action;

[0058] A task execution action refers to the specific execution steps or operation instructions inferred and generated by the transaction processing large model. Each action is a component in the process of completing the transaction task.

[0059] Schematically, the agent serves the object to output a transaction task to the transaction processing large model, triggering the transaction processing large model to perform autonomous reasoning on the task. The transaction processing large model autonomously reasons about the solution to the task and generates a task execution action sequence. The task execution action sequence consists of one or more task execution actions, and the task execution action sequence determines the specific path for solving the transaction task;

[0060] For example, assume the transaction task is: customer refund processing. The action sequence obtained through reasoning can be as follows:

[0061] Action 1: Verify customer information

[0062] Action 2: Check purchase records

[0063] Action 3: Initiate the refund process

[0064] S106: Perform step-by-step behavior quality evaluation processing on the task execution actions for the transaction task to obtain action supervision evaluation information for each task execution action;

[0065] Step-by-step behavior quality evaluation: For each task execution action, evaluate its quality in the actual transaction task execution process step by step from multiple dimensions, such as: correctness: whether the action conforms to the task goal and logic, efficiency: whether the action execution is fast and concise, adaptability: whether the action conforms to the context of the transaction task, etc.;

[0066] Action supervision evaluation information: It is the quality evaluation result of each task execution action. The action supervision evaluation information can include detailed information about the execution situation, helpfulness or helpfulness, odds of success, efficiency, task relevance, and coherence, etc.

[0067] Schematically, a multi-dimensional process supervision system is preset to evaluate the quality of each task execution action of the agent serving the object. Through the multi-dimensional process supervision system, step-by-step behavior quality evaluation processing is performed on the task execution actions for the transaction task to obtain action supervision evaluation information for each task execution action. This action supervision evaluation information provides fine-grained supervision information and can be used for reasoning optimization in the proxy training stage of the agent serving the object and the reasoning extension stage in the actual application process.

[0068] Optionally, in some embodiments, the multi-dimensional process supervision system introduces multiple key process supervision dimensions such as helpfulness, odds of success, efficiency, task relevance, and coherence, etc. These dimensions provide a comprehensive evaluation of the action quality of each task execution action. As Figure 3 shownFigure 3 It is a schematic diagram of a scenario for comparing solutions, which Figure 3 illustrates the differences between the traditional solution and the present solution. For the same transaction task, regarding the agent's action trajectory for the agent service object, the traditional solution uses a coarse-grained comparison based on the result as the supervision signal, while the present solution uses the action supervision and evaluation information of step-by-step multi-dimensional evaluation based on the process of the agent's action trajectory as the supervision signal. Compared with the traditional solution, its supervision signal has a higher fine-grained level, can provide a comprehensive evaluation of the quality of each action, and realizes autonomous optimization based on the process.

[0069] S108: Perform task autonomous reasoning and optimization processing on the agent service object based on the action supervision and evaluation information.

[0070] Schematically, the action supervision and evaluation information can provide quality feedback on each task execution action (such as success rate, error reason, improvement suggestions, etc.). According to the supervision and evaluation information, it can assist the transaction processing large model to optimize the reasoning logic of the transaction task to optimize the task execution process. After the task autonomous reasoning and optimization processing, it can reduce errors in action execution, improve the efficiency and adaptability of task completion, and enhance the overall performance of the agent service.

[0071] Optionally, adjusting the reasoning logic of the transaction processing large model includes, but is not limited to, the following methods:

[0072] Action generation: Optimize the task decomposition strategy to ensure the generation of a more reasonable action sequence.

[0073] Action order: Reorder the priorities of action execution to reduce context conflicts.

[0074] Condition handling: Increase the adaptability to special scenarios (such as data loss or permission issues).

[0075] In one or more embodiments of the present specification, by determining the transaction task of the transaction processing large model corresponding to the agent service object, based on the transaction task, the agent service object is used to call the transaction processing large model to perform task autonomous reasoning processing to obtain a task execution action sequence. By performing step-by-step behavior quality evaluation on the task execution action sequence during the task autonomous reasoning process, action supervision and evaluation information with a higher fine-grained level is generated, thereby replacing the single supervision signal based on the task completion result. It effectively solves the limitations of lack of process supervision, dependence on manual annotation, and difficulty in realizing reasoning time extension, can accurately locate the execution details during the task execution process, improve the learning and reasoning efficiency of the agent, and, in addition, can reduce the demand for manual annotation, realize the automatic generation of supervision signals during the task autonomous processing process, reduce the processing cost, enhance the processing ability of complex tasks by gradually evaluating and optimizing task reasoning, and thus significantly improve the task autonomous processing performance and application value of the multi-modal virtual agent.

[0076] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a step-by-step behavior quality evaluation process proposed in one or more embodiments of this specification. To specifically perform the step-by-step behavior quality evaluation process for the task execution actions with respect to the transaction task and obtain the action supervision and evaluation information for each task execution action, the following embodiments may be referred to:

[0077] S202: Construct a Monte Carlo tree for the transaction task based on the task execution action sequence, and perform an execution trajectory simulation process for each task execution action based on the Monte Carlo tree to obtain the number of inference steps and the basic reward for reaching the simulated task goal;

[0078] In this specification, based on the improvement of the Monte Carlo algorithm, its random simulation decision-making process is combined into the autonomous task inference and evaluation scenario. By introducing a multi-dimensional process supervision system (i.e., performing S204), the optimal task execution strategy is searched for. This method can also be called the MCTS-P algorithm; optionally, in some embodiments, the advantages of the Monte Carlo tree search (MCTS) based on the MCTS-P algorithm, such as its scalability and the balance between efficient exploration and exploitation, are used to optimize task execution in combination with the multi-dimensional process supervision system, comprehensively evaluate each task execution action of the virtual agent (intelligent agent service object), and thus generate the action supervision and evaluation information for each task execution action.

[0079] In this specification, a Monte Carlo tree for the transaction task is constructed based on the task execution action sequence;

[0080] In MCTS-P, the action supervision and evaluation information calculated by the multi-dimensional process supervision system is used as the basis for node selection and backpropagation. Specifically, the action supervision and evaluation information is used as the value v of each node S in the search tree i,j in MCTS-P. The tree structure in MCTS-P is similar to that of the traditional MCTS. Each node S i,j stores the action a i,j , the visit count n i,j , and the value v i,j . i,j

[0081] Use the action supervision and evaluation information for training data annotation, and automatically perform data annotation to collect the action supervision and evaluation information as a process supervision signal. The steps include:

[0082] 1) Calculate the minimum number of steps: For each node S in the Monte Carlo tree T q i,j ​​, calculate the minimum number of steps M required to achieve the simulation task goal (such as the optimal action, finally completing the transaction task), and the minimum number of steps M is used as the number of reasoning steps to meet the simulation task goal.

[0083] 2) Trajectory simulation in the expansion stage: In the expansion stage, the algorithm simulates N possible results of each step of the task execution action to determine the basic reward r i .

[0084] 3) Automatic calculation: According to M and r i, calculate three automatically generated dimensions: the helpfulness parameter, the success probability parameter, and the efficiency parameter.

[0085] 4) Task relevance and coherence evaluation: Use a multimodal large language model (such as GPT-4) to evaluate the task relevance parameter (Task Relevance) and coherence parameter (Coherence) of each step.

[0086] 5) Prune unfinished branches: Prune all incomplete branches (those that cannot finally complete the transaction task). Optionally, a preset benchmark evaluation method can be used to verify the correctness of the remaining paths.

[0087] Furthermore, the current task execution step is denoted as S i , and Monte Carlo tree search is performed through the Monte Carlo tree as follows:

[0088] First, select the optimal node S according to the current state i for expansion;

[0089] Second, generate new child nodes on the selected node to represent possible next operations. At this time, it is the expansion process for S i , and N possible results of each step of the task execution action S i are obtained, representing possible next operations (which can be called trajectory simulation operations. The jth trajectory simulation operation of the ith task execution step S i is denoted as S i,j , and in the Monte Carlo tree T q is a node);

[0090] Third, starting from the expanded node S i,j , simulate the execution process of the transaction task (which can be the transaction task in the sample training stage or the transaction task in the actual application stage) until the task is completed or fails. The completion of the transaction task can be regarded as achieving the simulation task goal. At this time, record the minimum number of steps M required to achieve the simulation task goal (such as the optimal action, finally completing the transaction task) as the number of reasoning steps to meet the simulation task goal, and then calculate the basic reward r i;

[0091] Calculate three automatically generated dimensions: the helpfulness parameter, the success probability parameter, and the efficiency parameter based on M and r i. Use a multi-modal large language model (such as GPT-4) to evaluate the task relevance parameter and coherence parameter of each step, so as to generate action supervision evaluation information and simulation results simulating success or failure states;

[0092] Finally, transmit the simulation results back to the root node of the Monte Carlo tree T q to update the evaluation value v of each node i,j ; The evaluation value can be updated based on the action supervision evaluation information.

[0093] Among them, for step S i Through trajectory simulation until the termination condition is met (such as completing the transaction task or reaching the maximum step length), the basic reward r i is defined as follows

[0094]

[0095] Among them, the a i,j represents the jth trajectory simulation action of the ith task execution action, the a* represents the standard simulation action, and the standard simulation action can usually be preset or select the optimal trajectory simulation action from the current N standard simulation actions (which can be regarded as the next final action corresponding to the ith task execution action), and A is the set of trajectory simulation actions;

[0096] Through the above steps, the collected task execution action trajectories are used as multi-dimensional evaluation sample data for finally training and evaluating the multi-step multi-dimensional evaluation reward model.

[0097] S204: Determine the helpfulness parameter, success probability parameter, and efficiency parameter of each task execution action based on the number of inference steps and the basic reward. Use the first multi-modal large model to perform a relevance analysis evaluation of the task execution action for the transaction task to obtain the task relevance parameter corresponding to each task execution action, and use the first multi-modal large model to determine the task relevance parameter and coherence parameter of the task execution action;

[0098] Helpfulness Parameter (H): The helpfulness parameter quantifies the positive or negative contribution of a step to task completion and assigns a corresponding quantitative value. This dimension aims to evaluate the impact of each step on the overall task trajectory. Steps that facilitate task completion (which can be regarded as task execution actions) are considered helpful, while steps that impede task progress are assigned negative values. The helpfulness parameter is particularly useful for comparing steps in different trajectories and distinguishing different degrees of contribution to success or failure. For example, among two successful trajectories, the trajectory with fewer steps will have a higher helpfulness parameter due to its more efficient and direct impact.

[0099] Odds of Success: Measures the probability that a particular step will lead to the successful completion of a task. This dimension is crucial for identifying steps that are more likely to result in a successful outcome. Steps with higher values are more likely to succeed, while steps with lower values are less likely to succeed. Conversely, incorrect steps will lead to failure. The odds of success are calculated by evaluating the proportion of successful paths in the simulation results of a given step.

[0100] Efficiency Parameter (E): Efficiency measures the operational efficiency of a step in terms of resource consumption (such as computational time or the number of operations). Assuming that fewer steps generally indicate higher efficiency because shorter paths are usually associated with lower resource usage. Steps that reduce the total number of steps required to complete a task are considered efficient steps.

[0101] Task Relevance Parameter (TR) evaluates whether a step is directly relevant to the current task. Some steps may be insignificant (e.g., performing an invalid operation), while others may be critical steps (e.g., clicking an important button). These differences cannot be captured by automated calculations. This dimension is evaluated through multimodal large language models (MLLMs) because task relevance often depends on context and semantic understanding. The task relevance parameter is a binary value: relevant (1) or irrelevant (0).

[0102] Coherence parameter (Coherence, abbreviated as C): Coherence evaluates the logical relationship and process consistency between consecutive steps. Some operations, although relevant to the task, efficient, and potentially successful, may lack coherence with the previous step. For example, in a browser operation task, after performing the operation "query the latest game results and record them in a note", opening the browser and note functions simultaneously may disrupt coherence. Poor coherence may lead to a logical break between steps, affecting task success. Compared with directly searching for game results through the browser, steps lacking coherence may seem less efficient. Coherence is evaluated by multi-modal large language models (MLLMs) and is a binary classification dimension with values {0, 1}.

[0103] Schematic, please refer to Figure 5 , Figure 5 is a schematic diagram of a multi-dimensional step-by-step evaluation scenario. Each agent step of the intelligent object service involves 5 task elements: instruction, last step, next step, final state, and step count. The evaluation parameters of the above 5 dimensions correspond to these. Specifically, the helpfulness parameter H (final state), success probability OS (next step), efficiency parameter E (step count), task relevance parameter TR (instruction), and coherence C (last step). In Figure 5 the step S2 is evaluated through these dimensions. This task requires 3 steps: S1, S2, and S 3,2 , and the H value of all steps is 1 / 3. In the subsequent steps of S2, S 3,2 and S3, 3 is correct, and the resulting OS value is 2 / 3. The E value is 1 / 3, and the calculated value is 2 - 1 / 3. Since S2 is aligned with S1 and the instruction, the TR and C values are 1.

[0104] In a feasible implementation manner, performing the execution trajectory simulation process for each of the tasks based on the Monte Carlo tree to obtain the number of inference steps and the basic reward for reaching the simulated task goal can refer to the following implementation manner:

[0105] S3002: Perform node search and expansion on each execution action tree node corresponding to the task execution action in the Monte Carlo tree to obtain multiple trajectory simulation action nodes of the execution action tree node. The trajectory simulation action node represents the inference expansion action after the task execution action;

[0106] S3004: Based on the multiple trajectory simulation action nodes, perform task execution trajectory simulation processing from the execution action tree node to obtain a task execution simulation result;

[0107] S3006: Determine the target task execution simulation result that meets the simulation task objective from the task execution simulation results, and determine the number of inference steps and the basic reward based on the target task execution simulation result.

[0108] Use action supervision evaluation information for training data annotation, and automatically perform data annotation to collect action supervision evaluation information as a process supervision signal. The steps include:

[0109] 1) Calculate the minimum number of steps: For each node S q in the Monte Carlo tree T i,j , calculate the minimum number of steps M required to reach the simulation task objective (such as the optimal action, the final completion of the transaction task), and the minimum number of steps M is used as the number of inference steps to meet the simulation task objective.

[0110] 2) Expand the trajectory simulation in the expansion stage: In the expansion stage, the algorithm simulates N possible results of each step of the task execution action to determine the basic reward r i .

[0111] 3) Automatic calculation: According to M and r i, calculate three automatically generated dimensions: the helpfulness parameter, the success probability parameter, and the efficiency parameter.

[0112] 4) Task relevance and coherence evaluation: Use a multi-modal large language model (such as GPT-4) to evaluate the task relevance parameter (Task Relevance) and coherence parameter (Coherence) of each step.

[0113] 5) Prune incomplete branches: Prune all incomplete branches (those that cannot finally complete the transaction task). Optionally, a preset benchmark evaluation method can be used to verify the correctness of the remaining paths.

[0114] Furthermore, the current task execution step is denoted as S i , and Monte Carlo tree search is performed through the Monte Carlo tree as follows:

[0115] First, select the optimal node S i according to the current state for expansion;

[0116] Secondly, generate new child nodes on the selected node to represent possible next operations. At this time, expansion processing is performed for the execution of S i , and N possible results of each step of the task execution action S i are obtained, representing possible next operations (which can be called trajectory simulation operations or inference expansion actions. The jth trajectory simulation operation of the ith task execution step S i is denoted as S i,j , and in the Monte Carlo tree T q is a node);

[0117] Next, starting from the expanded node S i,j begin, simulate the execution process of a transaction task (which can be a transaction task in the sample training phase or a transaction task in the actual application phase) until the task is completed or fails. Completion of the transaction task can be regarded as achieving the simulation task goal. At this time, record the minimum number of steps M required to achieve the simulation task goal (such as the optimal action, final completion of the transaction task) as the number of inference steps to meet the simulation task goal, and then calculate the basic reward r i;

[0118] According to M and ri, calculate three automatically generated dimensions: the helpfulness parameter, the success probability parameter, and the efficiency parameter. Use a multi-modal large language model (such as GPT-4) to evaluate the task relevance parameter (Task Relevance) and coherence parameter (Coherence) of each step, so as to generate action supervision evaluation information and the simulation result of the simulation success or failure state;

[0119] Finally, transmit the simulation result back to the root node of the Monte Carlo tree T q to update the evaluation value v of each node i,j ; The evaluation value can be updated based on the action supervision evaluation information.

[0120] Among them, for step S i through trajectory simulation until the termination condition is met (such as completing the transaction task or reaching the maximum step length), the basic reward r i is defined as follows

[0121]

[0122] where the a i,j represents the jth trajectory simulation action of the ith task execution action, the a* represents the standard simulation action, and the standard simulation action can usually be preset or select the optimal trajectory simulation action from the current N standard simulation actions (which can be regarded as the next final action corresponding to the ith task execution action), and A is the set of trajectory simulation actions;

[0123] Through the above steps, the collected task execution action trajectories are used as multi-dimensional evaluation sample data for finally training and evaluating the multi-dimensional evaluation reward model.

[0124] S206: Obtain the action supervision evaluation information of each task execution action based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter, and the coherence parameter;

[0125] S208: Perform multi-dimensional parameter weighted calculation based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter, and the coherence parameter to obtain the action comprehensive score of each task execution action, and calculate the mean value of the step scores based on all the action comprehensive scores to obtain the mean value of the trajectory action scores. Obtain the action supervision and evaluation information of each task execution action based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter, the action comprehensive score, and the mean value of the trajectory action scores.

[0126] Schematically, perform multi-dimensional parameter weighted calculation based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter, and the coherence parameter to obtain the action comprehensive score of each task execution action;

[0127] Schematically, calculate the mean value of the step scores based on all the action comprehensive scores to obtain the mean value of the trajectory action scores;

[0128] Furthermore, the step-by-step multi-dimensional evaluation reward model corresponding to the intelligent agent service object can be trained based on the action supervision and evaluation information;

[0129] In one or more embodiments of this specification, the following technical effects can be achieved by the above method:

[0130] 1. Improve the learning and reasoning ability of GVA: By providing fine-grained process supervision signals, this solution can significantly improve the learning and reasoning ability of GVA in complex tasks, making it show higher autonomy and efficiency in various digital environments.

[0131] 2. Reduce the training cost: Automated data collection and annotation greatly reduce the need for manual annotation, reduce the training cost, and improve the efficiency of data generation at the same time.

[0132] 3. Applicability across platforms and multiple tasks: This solution supports cross-platform and multi-task evaluation and can be widely applied to various actual scenarios, such as intelligent office, automated process processing, etc.

[0133] 4. Promote the development of AI Agent technology: This solution provides a new training and evaluation paradigm for the field of AI Agent, promoting the technological progress and commercial application in this field.

[0134] In an alternative embodiment, the step-by-step multi-dimensional evaluation reward model corresponding to the intelligent agent service object can be trained based on the action supervision and evaluation information. Refer to the following embodiments:

[0135] S4002: Generate multi-dimensional evaluation sample data based on the action supervision and evaluation information, the transaction task, and the task execution action sequence;

[0136] In some embodiments, to overcome the problems of high cost and low efficiency of manual annotation, an improvement is made based on the Monte Carlo algorithm and its random simulation decision-making process is combined into the autonomous task inference and evaluation scenario. By introducing a multi-dimensional process supervision system (i.e., performing S204), the optimal task execution strategy is found. This method can also be called the MCTS-P algorithm. Using the MCTS-P algorithm, the action supervision and evaluation information for each task execution step is obtained step by step in multiple dimensions, and every operation of the virtual agent is comprehensively evaluated. Then, based on the action supervision and evaluation information, the transaction task, and the task execution action sequence, multi-dimensional evaluation sample data can be generated to automatically generate multi-dimensional annotation data.

[0137] S4004: Determine the initial reward model of the intelligent agent's service object, and use the multi-dimensional evaluation sample data to train the initial reward model to obtain a step-by-step multi-dimensional evaluation reward model after model training;

[0138] The reward model (RMs) is an important tool for guiding the action quality evaluation of the virtual agent (intelligent agent's service object). In related technologies, the result reward model (ORMs) is often used. The result reward model focuses on the final success of the task, and this paradigm has limitations. Based on this, one or more embodiments of this specification involve a brand-new virtual agent reward model - the step-by-step multi-dimensional evaluation reward model. By providing feedback on intermediate steps through the step-by-step multi-dimensional evaluation reward model, the agent performance can be evaluated in complex reasoning tasks. In some embodiments, the MCTS-P algorithm based on Monte Carlo tree search is shown to automate data collection, achieving remarkable results. On this basis, a step-by-step, multi-dimensional system - the step-by-step multi-dimensional evaluation reward model (Step-wise Multidimensional Generalist Reward ModelSimilar, Similar) is introduced. Using MCTS to collect fine-grained annotations provides a robust reward model for the intelligent agent's service object GVAs.

[0139] For details, please refer to the subsequent step interpretations.

[0140] In a feasible implementation manner, the multi-dimensional process supervision system is detailed as follows. Specifically, to execute and determine the helpfulness parameter, success probability parameter, and efficiency parameter of each task execution action based on the number of reasoning steps and the basic reward, the following method can be referred to:

[0141] 1) Helpfulness parameter: Obtain the cumulative execution action value corresponding to the task execution action, the number of reasoning steps, the execution action number, and the basic reward. Calculate the helpfulness parameter of each task execution action using the first calculation formula, and update the cumulative execution action value using the second calculation formula; the first calculation formula satisfies the following formula:

[0142]

[0143] where, the H i represents the helpfulness parameter, the AC i-1 represents the cumulative execution action value of the (i - 1)-th step, the M represents the number of reasoning steps, the i represents the execution action number, and the value range of i is 1 - M, the r i represents the basic reward of the task execution action S i , and usually can be a basic reward vector or a basic reward value;

[0144] The second calculation formula satisfies the following formula:

[0145]

[0146] where, the AC i represents the cumulative execution action value of the i-th step;

[0147] For the second calculation formula, when the step index i = 0 (i.e., the initial state of the task), the cumulative action number ACi is initialized to 0. When the step index i is greater than 0, for the i-th step, the cumulative action number AC i is the sum of the cumulative action number AC i-1 of the previous step and the helpfulness H i of the current step, but not less than 0.

[0148] The cumulative execution action value is a dynamic variable used to gradually record the effective contributions of all steps in the task execution process. It accumulatively changes with the increase of each step H i , but is not less than 0, indicating that the effectiveness of the action is always non-negative. And Hi is the positive or negative contribution of the current step to the task:

[0149] If H i > 0, it means that the current step is effective and the cumulative action number ACi increases.

[0150] If H i < 0, it means that there is a negative impact on the current step and the cumulative action number decreases.

[0151] 2) Success probability parameter: Obtain the trajectory simulation action, standard simulation action, and total number of trajectory simulation actions corresponding to the task execution action, and calculate the success probability parameter of each task execution action using the third calculation formula; the third calculation formula satisfies the following formula:

[0152]

[0153] where the OS i represents the success probability parameter, the N represents the total number of trajectory simulation actions, the j is the number of the trajectory simulation action, and the a i,j represents the jth trajectory simulation action of the ith task execution action, the a* represents the standard simulation action, and the II() is an indicator function, which is used to compare whether the trajectory simulation action matches the standard simulation action;

[0154] The Indicator Function is a mathematical function used to determine whether a certain condition holds. In the success probability formula, the condition in the parentheses of the indicator function is used to compare whether the trajectory simulation action matches the standard simulation action. Usually, if the condition holds, the value of the indicator function is 1, otherwise it is 0; in the subsequent reward model of the virtual agent, the indicator function is used to measure the correctness of the agent's action, so as to provide a quantitative index for the reward model through the success probability parameter OS. This function is crucial in guiding the virtual agent to learn and optimize the task path.

[0155] 3) Efficiency parameter: Determine the remaining action length of the historical task and the remaining action length of the current task corresponding to the task execution action, and calculate the efficiency parameter of each task execution action using the fourth calculation formula; the fourth calculation formula satisfies the following formula:

[0156]

[0157] where the E i represents the efficiency parameter of the ith task execution action, the Len i-1 represents the remaining action length of the historical task, which can be understood as the remaining task length of the (i - 1)th step (the expected total number of steps from the current step to the task goal)

[0158] The Len i represents the remaining action length of the current task, the remaining task length of step i (the expected total number of steps from step i to the task goal)

[0159] The remaining action length of the current task satisfies the following fifth calculation formula:

[0160]

[0161] Among them, the Len(S i,j ) represents the remaining action length of the task corresponding to each of the trajectory simulation actions.

[0162] Meaning of efficiency parameters: Evaluating the quality of operations: Helping to quantify the advancement effect of each operation on the task goal. Optimizing the reward model: Using efficiency metrics in the reward model to distinguish between efficient and inefficient steps and guiding the virtual agent to optimize action selection. Improving task planning: Improving task decomposition and planning methods by analyzing inefficient steps.

[0163] In a feasible implementation, when performing the task of determining the task relevance parameter and the coherence parameter of the task execution action by using the first multimodal large model, the following method can be referred to:

[0164] In this specification, the first multimodal large model can be the same multimodal large model as the transaction processing large model, or it can be a different multimodal large model, which is specifically set based on the actual application environment and is not specifically limited here;

[0165] 1) Based on a preset task relevance evaluation prompt, use the first multimodal large model to perform a relevance analysis and evaluation of the task execution action for the transaction task to obtain the task relevance parameter corresponding to each task execution action;

[0166] Preset task relevance evaluation prompt: A prompt specifically designed in advance (Prompt) used to guide the multimodal large model (MLLM) to analyze the relevance of task execution actions based on the task goal. The prompt usually includes the goal information, context of the transaction task, and the relevance evaluation requirements for the action.

[0167] Optionally, create a prompt template for guiding the MLLM to perform relevance evaluation to generate a preset task relevance evaluation prompt. The prompt content includes, but is not limited to: the goal description of the transaction task, the current task execution action, and the requirement for the model to judge the relevance of the action to the task goal.

[0168] Furthermore, input the evaluation prompt into the first multimodal large model (such as GPT-4 or a similar multimodal model) to let the model analyze the relevance between the task execution action and the transaction task goal. The analysis result returned by the multimodal large model includes a relevance score. For example, the task relevance score of the current action is 0.95, indicating a high degree of relevance.

[0169] Exemplarily, a preset task relevance evaluation prompt for task relevance can be referred to as follows:

[0170] [Task Relevance]

[0171] 1. Meaning: Indicates whether the agent's operation is related to achieving the instruction.

[0172] 2. Design Motivation: Some operation steps may prevent the task from being completed, but they are related to the task (for example, we need to ask the agent to take notes, and the agent takes notes, which is related to the task, but the content of the notes recorded is incorrect, indicating that this is an incorrect step). Some operation steps may be meaningless, but they can still lead to the completion of the task (for example, clicking on a blank screen without generating any response, which is not related to the task, but the subsequent actions of the agent can still lead to the success of the task). Therefore, a metric is needed to determine whether the current operation step is related to the task.

[0173] 3. Range of the Mapped Value: {0, 1}. The larger the value, the greater the correlation between the step and the task. That is, 1 indicates a higher task correlation of the operation, while 0 indicates that the operation has little correlation with the task.

[0174] 2) Using the first multi-modal large model based on the preset action coherence evaluation prompt, perform action coherence evaluation processing on the task execution actions for the transaction task to obtain the coherence parameter corresponding to each task execution action.

[0175] Preset Action Coherence Evaluation Prompt: A prompt specifically designed in advance, used to guide the MLLM to evaluate the logical rationality and consistency of the current action in the entire task execution action sequence.

[0176] Optionally, create a prompt template to guide the model to evaluate the logical coherence of the current action. The prompt content includes, but is not limited to: the description of the current transaction task, the context of the executed action sequence, the current action to be evaluated, and require the model to judge the logical coherence of the current action with the context.

[0177] Furthermore, input the prompt into the first multi-modal large model to analyze whether the current action conforms to the task execution logic. The parsing result returned by the multi-modal large model includes a coherence score. For example, the coherence score of the current action is 0.85, indicating relatively coherent with the context.

[0178] Exemplarily, a preset action coherence evaluation prompt for the coherence parameter (Coherence) can be as follows:

[0179] [Coherence]

[0180] 1. Meaning: It represents the compactness and coherence between the current step and the previous step.

[0181] 2 Design Motivation: Although some operations are task-related, they are inefficient and may likely lead to success, but lack consistency with the previous step. For example, the task is "query the game results of the Lakers and record them in a note. For example, the agent operations are as follows: a. Open the browser; b. Open the note; c. Create a new note; d. Search for the Lakers' game; e. Query the game results; f. Record the game results in your note. It can be found that operation a and the operation with black consistency, directly searching for the competition results after opening the browser instead of opening Note simultaneously is more in line with human preferences.

[0182] 3. Mapped Value Range: {0, 1}. The larger the value, the greater the coherence of the step size, that is, 1 indicates that the operation has greater consistency and conforms to human preferences, and 0 indicates lower consistency of the operation.

[0183] Optionally, please refer to Figure 6 , Figure 6 is a schematic flow diagram of task autonomous reasoning optimization. Perform the task autonomous reasoning optimization process on the intelligent agent service object based on the action supervision and evaluation information. Specifically, please refer to the following embodiments:

[0184] S302: Train the step-by-step multi-dimensional evaluation reward model corresponding to the intelligent agent service object based on the action supervision and evaluation information;

[0185] In an optional embodiment, when performing the training of the step-by-step multi-dimensional evaluation reward model corresponding to the intelligent agent service object based on the action supervision and evaluation information, the following embodiments can be referred to:

[0186] S4002: Generate multi-dimensional evaluation sample data based on the action supervision and evaluation information, the transaction task, and the task execution action sequence;

[0187] According to some embodiments, in the process of automated multi-dimensional evaluation sample data, the MCTS-P algorithm is used to collect annotated step-by-step data from multiple platforms, including execution environments such as Web, Android, Linux, and Windows;

[0188] Exemplarily, the following shows a process of collecting multi-dimensional evaluation sample data. The process of collecting multi-dimensional evaluation sample data can refer to the step interpretations of S202 - S208, S3002 - S3006. The method shown below realizes the collection of multi-dimensional evaluation sample data D. Using the MCTS-P algorithm, annotated step-by-step data is collected from multiple platforms, including Web, Android, Linux, and Windows. The following pseudocode describes the data collection process, including the calculation of five-dimensional scores, the pruning of incomplete branches, and the verification of trajectory correctness. The specific example is as follows:

[0189] Steps 1 - 3 Initialization:

[0190] Step 1: Input task q and the set of target platforms P, and each platform needs to be processed independently.

[0191] Step 2: The output target is a labeled multi-dimensional evaluation sample data - dataset D, which should at least contain the input, result, and evaluation label of the task.

[0192] Step 3: Initialize an empty dataset D to store the generated labeled data.

[0193] Steps 4 - 6: Platform-level loop:

[0194] Step 4: Traverse all platforms (such as Web, Android, etc.) and construct task data for each platform separately.

[0195] Step 5: Initialize an MCTS-P search tree T q for the current platform p, representing the search and simulation of task q on platform p.

[0196] Step 6: Process each node S i,j in the search tree.

[0197] Steps 7 - 9: Simulation and reward calculation:

[0198] Step 7: Calculate the minimum number of steps M required to reach the correct answer from node Si,j. This provides a benchmark path length for the algorithm (the optimality of the reference path), and generally, the shorter the path, the greater the potential of the current node.

[0199] Step 8: Simulate N trajectories starting from node Si,j and calculate their basic rewards ri. The simulation tries different task execution paths randomly or heuristically. The reward calculates the score based on the effect and quality of the path.

[0200] Step 9: Calculate the following three-dimensional metrics using the formula:

[0201] H i The accuracy or relevance of the result (Heuristic Score).

[0202] OS i The coverage of the path (Outcome Similarity Score).

[0203] E i The efficiency of the execution process (Efficiency Score).

[0204] Steps 10 - 11: Large model generation and verification:

[0205] Step 10: Use a large model (such as GPT-4) to evaluate the following two-dimensional metrics:

[0206] TR i : Task Relevance Score. It represents the semantic relevance between the result and the task instruction qqq.

[0207] C i : Consistency Score. It represents the compatibility of the result with the platform context (such as platform specifications, formats, etc.).

[0208] Step 11: Determine whether the current node Si,j can generate a complete trajectory:

[0209] If the node Si,j generates a complete task execution path, proceed to subsequent processing.

[0210] Result annotation and storage for Steps 12 - 15:

[0211] Step 12: Verify the correctness of the result generated by the current node Si,j, specifically using an evaluation method for the platform.

[0212] Step 13: Add the annotation data of the current node to the dataset D, including: information of the node Si,j, scores of the five evaluation dimensions, etc.;

[0213] Steps 14 - 15: Complete the processing of the current node.

[0214] Platform - level and global processing for Steps 16 - 18:

[0215] Step 16: After completing the processing of the current platform p, proceed to the next platform.

[0216] Step 17: Further prune the search tree Tq, remove the branches that do not generate complete trajectories, and retain the valid paths.

[0217] Step 18: Finally, return the generated annotation dataset D.

[0218] In this specification, the task execution path is gradually evaluated and annotated through multi - dimensional indicators, combined with platform characteristics and large - model reasoning, to generate a high - quality annotation dataset D. Its features are: applicable to the automatic annotation of multi - platform task instructions. Combining the capabilities of large models with heuristic evaluation to optimize the evaluation efficiency. The multi - dimensional indicators comprehensively cover the accuracy, efficiency, relevance, and environmental adaptability of the results.

[0219] S4004: Determine the initial reward model for the intelligent agent's service object, and use the multi - dimensional evaluation sample data to train the initial reward model to obtain the gradually multi - dimensional evaluation reward model after model training;

[0220] Reward models (RMs) are important tools for guiding the quality assessment of the actions of virtual agents (agent service objects). In related technologies, outcome reward models (ORMs) are often used. Outcome reward models focus on the ultimate success of tasks, but this paradigm has limitations. Based on this, one or more embodiments of this specification involve a brand-new virtual agent reward model - the step-wise multidimensional evaluation reward model. By providing feedback on intermediate steps through the step-wise multidimensional evaluation reward model, the agent performance can be evaluated in complex reasoning tasks. In some embodiments, the MCTS-P algorithm based on Monte Carlo tree search is demonstrated to automate data collection, achieving remarkable results. On this basis, a step-wise and multi-dimensional system - the step-wise multidimensional generalist reward model (Similar) is introduced. MCTS is used to collect fine-grained annotations to generate multi-dimensional evaluation samples, providing a robust reward model for the agent service object GVAs.

[0221] The initial reward model can be an untrained initial model created based on a machine learning model, containing the basic reward evaluation structure for task execution actions. After training with multi-dimensional sample data, the step-wise multidimensional evaluation reward model is obtained. The step-wise multidimensional evaluation reward model can perform precise multi-dimensional evaluations on each task execution action according to the action supervision evaluation information.

[0222] Initializing the reward model: The input is multi-dimensional evaluation sample data. Output during the model training phase: Action supervision evaluation information containing the evaluation scores for each step. In some embodiments, scalarized reward information may also be included.

[0223] Illustratively, multi-dimensional sample data is used as the training set to optimize the reward model parameters so that it can accurately predict the reward information and action supervision evaluation information for each action until the initialized reward model meets the preset model end training condition, and the step-wise multidimensional evaluation reward model is obtained.

[0224] In this specification, the training of the step-wise multidimensional evaluation reward model has achieved the following beneficial effects:

[0225] Fine-grained reward signal generation: Based on the action supervision evaluation information and multi-dimensional sample data, the step-wise reward evaluation is more detailed, avoiding the limitations of traditional single reward signals.

[0226] Dynamically weighing multi-dimensional parameters: The model can dynamically evaluate the reward value of task execution actions according to multi-dimensional parameters such as helpfulness, success probability, efficiency, relevance, and coherence, providing more accurate guidance for the reasoning optimization of complex tasks.

[0227] Improving training efficiency: The automatically generated multi-dimensional evaluation sample data reduces the need for manual intervention and significantly improves the training efficiency of the reward model.

[0228] Enhancing agent adaptability: The trained step-by-step multi-dimensional evaluation reward model is applicable to various task scenarios, capable of supporting dynamic reasoning and optimized decision-making for complex tasks, and enhancing the practical application value of the agent's service object.

[0229] S304: Use the step-by-step multi-dimensional evaluation reward model to perform task autonomous reasoning optimization processing on the agent's service object.

[0230] After the step-by-step multi-dimensional evaluation reward model completes model training, it can assist the agent's service object in providing action supervision and evaluation information for each target task execution action inferred, serving as a guiding signal for the agent during the reasoning stage.

[0231] The action supervision and evaluation information provided for the agent's service object during the actual application stage includes, but is not limited to, helpfulness parameters, success probability parameters, efficiency parameters, task relevance parameters, coherence parameters, etc. These action supervision and evaluation information provide step-by-step optimization feedback signals for the agent, guiding it to select better actions during the reasoning process.

[0232] In an alternative implementation, when performing the task autonomous reasoning optimization processing on the agent's service object using the step-by-step multi-dimensional evaluation reward model, the following method can be adopted:

[0233] S2: In response to the user's target transaction task for the agent's service object, based on the target transaction task, use the agent's service object to call the transaction processing large model to perform task autonomous reasoning processing to obtain a target task execution action sequence, where the target task execution action sequence includes at least one target task execution action;

[0234] It can be understood that one or more target task execution actions in the target task execution action sequence can be adjusted based on the action supervision and evaluation information during the execution action optimization process;

[0235] S4: Use the target task execution action as the task execution action, and use the step-by-step multi-dimensional evaluation reward model to perform step-by-step behavior quality evaluation processing on the task execution action for the transaction task to obtain the action supervision and evaluation information for each task execution action;

[0236] S6: Based on the action supervision and evaluation information, drive the agent's service object to perform execution action optimization processing on the target task execution action sequence.

[0237] Calculation of action supervision and evaluation information for target task execution actions:

[0238] The intelligent agent's service object performs action S in each target task i During the inference process of (it can also be combined with the step-by-step multi-dimensional evaluation reward model), a set of task execution actions A = {a i-1 corresponding to a certain task execution action will be generated, which contains multiple possible actions (also called the actual task execution action S 1 determined currently) and the next trajectory simulation actions), a 2 ...a n}, a 1 、a 2 ...a n are trajectory simulation actions;

[0239] The step-by-step multi-dimensional evaluation reward model will generate action supervision evaluation information for each action a j regarded as the current step S to be evaluated: H(S), OS(S), E(S), TR(S), C(S);

[0240] Furthermore, based on the action supervision evaluation information, the multi-dimensional reward information v(S) of each action a j will be generated: v(S) = w h H(S) + w os OS(S) + w e E(S) + w tr TR(S) + w c C(S);

[0241] Among them, w h 、w os 、w e 、w tr 、w c are the importance weights of each dimension, which are set according to the task requirements.

[0242] Furthermore, the multi-dimensional reward information v(S) of each action a j selects the optimal trajectory simulation action from the set of task execution actions A = {a 1 、a 2 ...a n} as the target task execution action S i , after each target task execution action S i is executed by the intelligent agent's service object, the multi-dimensional evaluation values of the remaining target task execution actions are updated by collecting real-time environmental feedback information (the current task status and task feedback information), and the remaining action sequence A' corresponding to the target task execution action sequence will continue to be regenerated or corrected to determine the next target task execution action to be executed until the target transaction task is completed.

[0243] In a specific embodiment, in combination with the above explanations, during the inference process of the intelligent agent's service object for each target task execution action S i a task execution action set A = {a i-1 that contains multiple possible actions (also referred to as the next trajectory simulation actions of the currently determined actual task execution action S 1 ), a 2 ...a n} can be generated by jointly using the step-by-step multi-dimensional evaluation reward model during the inference process of the intelligent agent's service object for each target task execution action S 1 、a 2 ...a n are trajectory simulation actions, and the optimal task execution action is determined from the task execution action set A as the target task execution action for the intelligent agent's service object to execute. Please refer to the following example process. The following is a process for selecting the target task execution action, and the following is an example for interpretation:

[0244] Step 1: Initialize the root node S0 of the MCTS-P search tree Tq, representing the initial state s0. The purpose is to construct the basic structure of the MCTS-P search tree Tq. At this time, the root node of the search tree has been initialized and is ready to enter the main loop. It can be based on the currently determined target task execution step action in the target task execution step action sequence or directly based on the target task execution step action sequence to initialize the root node S 0 .

[0245] Steps 2 - 6: The search process (main loop) of the MCTS-P search tree

[0246] Step 2: Set the calculation budget target. The operation indicates entering the main loop and continuously performing search steps within the calculation budget. The purpose is to limit the computing resources for the search, such as time or the number of iterations.

[0247] Step 3: By calling the function TreePolicy(S 0) select the currently to-be-inferred node S i , the intelligent agent's service object jointly uses the step-by-step multi-dimensional evaluation reward model and selects a strategy to find an unfully expanded node Si along the tree path starting from the root node;

[0248] If the current node has been fully expanded, use the BestChild function to select the next node.

[0249] If the node is not fully expanded, use the Expand function to add new child nodes.

[0250] Step 4: Call the function DefaultPolicy(sSi) to evaluate the node S i , starting from the state s(S i)Start progressive random simulation until the termination state is reached. Return node S i The simulated reward value Δ of i is obtained based on the action supervision evaluation information.

[0251] Step 5: Call the function: Backup(Si,Δ) to backpropagate the reward.

[0252] Backpropagate the simulated reward value Δ from node Si to all parent nodes on the path of the tree, and update the statistical values (such as the number of visits and cumulative rewards) of the nodes along the way.

[0253] Step 6: End the loop; continue the next iteration until the calculation budget is exhausted.

[0254] Furthermore, Steps 7 - 10 are processed using the TreePolicy(S) function

[0255] Step 7: TreePolicy is defined to select a node S in the search tree until an unexpanded non - terminal node is found.

[0256] Step 8: Determine whether the termination state is reached. If the current node S is in a non - termination state (not reaching the leaf node), continue to select downward.

[0257] Step 9: Check whether the node is fully expanded. If node S still has unadded child nodes, it means it has not been fully expanded.

[0258] Step 10: Expand the tree call: Expand(S), add a new child node to node S, and return this new node. Each new child node corresponds to a trajectory simulation action, and the trajectory simulation action of node S is denoted as A = {a 1 、a 2 ...a n}.

[0259] Furthermore, Steps 15 - 18 indicate the end of TreePolicy

[0260] Step 15: Return the selected node. At this time, TreePolicy returns a non - terminal and unexpanded node S.

[0261] Step 16: End of the TreePolicy function;

[0262] Steps 19 - 23: Are processed using the BestChild(S, C) function

[0263] Step 19: The BestChild function is defined to select the child node S′ with the highest score among all child nodes of node S;

[0264] Step 21: Scoring formula v(S):

[0265] Reference may be made to v(S) = whH(S) + wosOS(S) + weE(S) + wtrTR(S) + wcC(S);

[0266] Action supervision and evaluation information H(S), O(S), E(S), TR(S), C(S): As heuristic functions representing different aspects, each action a is generated based on the action supervision and evaluation information j The multi-dimensional reward information v(S) is used to evaluate the performance of the simulated action corresponding to the current node in the task execution step. It is the weighted comprehensive scoring result of multiple dimensions; based on the multi-dimensional reward information v(S), the scores of each child node are calculated, and the child node with the highest score is selected to determine the optimal trajectory simulation action, realizing the selection of the optimal trajectory simulation action from the set of task execution step actions A = {a j of the multi-dimensional reward information v(S) 1 and a 2 ... a n} as the target task execution step action S i

[0267] Step 22: Return the corresponding child node;

[0268] The specific implementation of the best child node selection strategy return result in the tree search process combines the value evaluation and exploration mechanism of the node, and is used to select the best child node from the child nodes of the current node S and further return.

[0269] Among them, in Step 22, S′ represents: a child node of the current node S. children: the set of all child nodes of the current node S, that is, the above set A, v(S′) represents: the comprehensive value of the child node S′ (calculated by the v(S) formula in Step 21), and n(S′) represents: the number of times the child node S′ is visited.

[0270] n(S) represents: the number of times the parent node S is visited.

[0271] c: exploration parameter, controlling the balance between exploration and exploitation (usually a positive number)

[0272] In Step 22, v(S′) / n(S′) represents the average value of the child node S′, and the child node with a high average value is preferentially selected.

[0273] In Step 22 represents the preference for the under-explored child nodes, encouraging exploration by increasing the confidence interval,

[0274] Calculate the scores of each child node: For each child node S′ of S, calculate:

[0275]

[0276] Then in the 23rd step, compare the scores of each child node and select the child node S* with the highest score, return the action of the selected child node, and output the corresponding action a, that is, the decision of the best child node.

[0277] Optionally, please refer to Figure 7 , Figure 7 is a schematic diagram of the model training process of a step-by-step multi-dimensional evaluation reward model. To execute the determination of the initial reward model of the intelligent agent service object and perform model training on the initial reward model using the multi-dimensional evaluation sample data to obtain the step-by-step multi-dimensional evaluation reward model after model training, the following method can be referred to:

[0278] S5002: Obtain a pre-trained second multi-modal large model and a gating network, determine a backbone feature extraction network based on the second multi-modal large model, and create an initial reward model for the intelligent agent service object based on the backbone feature extraction network and the gating network;

[0279] Optionally, the second multi-modal large model can be the same as or different from the first multi-modal large model or the transaction processing large model.

[0280] Second multi-modal large model: A pre-trained large-scale multi-modal model used to extract multi-modal features (such as text, images, audio, etc.) from input data. This model can be GPT, CLIP, or other pre-trained models that support multi-modal input.

[0281] Backbone feature extraction network: A feature extraction sub-network generated by the second multi-modal large model, specifically used to capture multi-modal information in transaction tasks. Its main function is to encode task context and action data into high-dimensional semantic feature representations.

[0282] Gating Network: Used to dynamically adjust the weights of the reward model on multi-dimensional supervision information (such as helpfulness, relevance, efficiency, etc.). Through context-aware weight allocation, it ensures that the model can adapt to task requirements in different scenarios.

[0283] Initial reward model: A reward prediction model composed of at least the backbone feature extraction network and the gating network, used to provide a basic framework before training.

[0284] Schematically, load the second multi-modal large model to initialize the feature extraction ability, and then construct a backbone feature extraction network: extract a feature extraction network adapted to a specific task from the second multi-modal large model (such as freezing some weights or fine-tuning specific layers), and load a gating network: introduce a gating network to dynamically generate weight allocation in combination with the task context for evaluating multi-dimensional parameters of task execution actions. It is at least composed of the backbone feature extraction network and the gating network to form an initial reward prediction model, which can input multi-dimensional evaluation sample data and output action reward values.

[0285] S5004: Use the multi-dimensional evaluation sample data to train the initial reward model. During the model training process, perform multi-dimensional supervised evaluation information regression training on the initial reward model and dynamically adjust the weights of the initial reward model to obtain a step-by-step multi-dimensional evaluation reward model after the model training ends.

[0286] Multi-dimensional supervised evaluation information regression training: According to the supervision information in the multi-dimensional evaluation sample data (such as helpfulness, relevance, efficiency, etc.), perform regression optimization on the reward model so that it can accurately predict the multi-dimensional reward scores of each action; train a regression layer for five-dimensional score prediction. For each input sequence x⊕y (where x represents the prompt and y represents the response), extract the last hidden state h∈R d, which has d-dimensional features, and map it to the multi-dimensional reward score through a linear regression layer W. The model is optimized using the regression loss.

[0287] Dynamic weight adjustment: During the training process, use the gating network to dynamically adjust the attention degree of the reward model to different evaluation dimensions to achieve task-specific weight adaptation. Train a gating network gating to dynamically balance the scores of five dimensions and solve the multi-objective optimization problem. Introduce a gating network gφ that perceives the prompt, using a shallow multi-layer perceptron (MLP), which dynamically adjusts the focus of the model according to the input prompt x. The gating network calculates non-negative coefficients for multiple reward dimensions. These coefficients are derived from the last hidden state corresponding to the prompt x and are normalized through the softmax function.

[0288] Step-by-step multi-dimensional evaluation reward model: The trained reward model can comprehensively consider multi-dimensional evaluation information and generate accurate reward values for each task execution action.

[0289] Optionally, the backbone feature extraction network includes a backbone feature extractor and a linear regression layer. To perform the multi-dimensional supervised evaluation information regression training on the initial reward model, the following method can be referred to:

[0290] Step B2: Based on the backbone feature extractor, perform hidden state extraction on the data input sequence corresponding to the multi-dimensional evaluation sample data using the sixth calculation formula to obtain hidden state features. Based on the linear regression layer, perform linear mapping on the hidden state features using the seventh calculation formula to obtain a predicted action supervision evaluation reward vector;

[0291] Schematically, the pre-trained MLLM fθ is the feature extractor of the entire system, responsible for extracting semantically rich multi-modal features from the input (x⊕y). It can be referred to the sixth calculation formula to satisfy the following formula:

[0292]

[0293] where, the h represents the hidden state features, the f θ represents the hidden state extraction operation of the backbone feature extractor, the x represents the predicted input task sample in the data input sequence, the y represents the candidate action response item information generated based on the action supervision evaluation information; (x⊕y) is the concatenation of the task condition x and the candidate response y. h is the high-dimensional hidden feature extracted by the backbone feature extractor of the MLLM, representing the semantic relationship between the context and the candidate response.

[0294] Schematically, the regression layer is immediately followed by the feature extractor, and converts the high-dimensional feature h into a scalar reward value through a linear mapping (weight matrix W). It can be referred to the seventh calculation formula to satisfy the following formula:

[0295]

[0296] where, the represents the predicted action supervision evaluation reward vector, the W represents the regression weight parameter; for the extended relationship, the regression layer does not belong to the original parameters of the MLLM in the pre-training stage, but during the training of the reward model, it is used as an extended module of the feature extractor and is co-trained and optimized with f θ together.

[0297] Step B4: Determine the true action supervision evaluation reward vector based on the multi-dimensional evaluation sample data;

[0298] The true action supervision evaluation reward vector is used to measure the performance of the candidate response y under the given task condition x. This is a key variable that plays a core role in the training of the reward model. The true action supervision evaluation reward vector can be sourced from:

[0299] Manual annotation: The evaluator scores according to the quality of the candidate response.

[0300] Automatically generated: Score the response through automated rules or heuristic methods.

[0301] Preference pair: A reward signal extracted from human preferences (preferably a comparison of acceptance and rejection).

[0302] In the context, r is mainly used to guide the reward model to learn how to predict the quality of candidate responses.

[0303] Step B6: Based on the predicted action supervised evaluation reward vector and the true action supervised evaluation reward vector, obtain the optimized regression loss using the eighth calculation formula, and use the optimized regression loss to adjust the backbone network parameters and regression weight parameters of the backbone feature extraction network;

[0304] The eighth calculation formula satisfies the following formula:

[0305]

[0306] Among them, the L RG represents the optimized regression loss, and the r represents the true action supervised evaluation reward vector. L RG The goal is to train the reward model to accurately predict the quality of candidate responses by minimizing the error between the predicted reward value and the true reward value.

[0307] Optionally, in the two-stage preference modeling, r may indirectly affect the preference scores (R chosen and R rejected ) of y chosen and y rejected . The goal of two-stage training is to dynamically adjust the weights of the reward model through a gating network to adapt to different task conditions x and the specific requirements of candidate responses. The core method is to model through preference pairs (chosen / rejected pairs), optimize the preference scores (R chosen ,R rejected ), and update the gating network parameters through the Bradley-Terry (BT) loss function L BT .

[0308] Furthermore, to perform the dynamic weight adjustment of the initial reward model, the following method can be referred to:

[0309] Input preference data, traverse the training dataset D, and take out a set of preference data each time:

[0310] x: Task condition.

[0311] y chosen : The candidate response selected in the preference pair.

[0312] y rejected : The candidate response rejected in the preference pair.

[0313] Step C2: Obtain the hidden state features, and perform weighted calculation processing on the hidden state features through the gating network using the ninth calculation formula to obtain a multi-dimensional dynamic weight vector;

[0314] The ninth calculation formula satisfies the following formula:

[0315] ω = g φ (f θ (x))

[0316] where, ω represents the multi-dimensional dynamic weight vector, f θ (x) represents the hidden state features extracted by the backbone feature extractor, and g φ () represents the weighted calculation processing operation of the gating network; The gating network generates a dynamic weight w for calculating the preference score according to the features of the task condition x, so that the reward model can adapt to the requirements of different tasks.

[0317] Step C4: Determine candidate action response item information based on the multi-dimensional evaluation sample data, where the candidate action response item information includes action preferred response item information and action rejection response item information;

[0318] Step C6: Determine the preferred response feature representation corresponding to the action preferred response item information and the rejection response feature representation corresponding to the action rejection response item information through the gating network;

[0319] Step C8: Determine the action preferred response preference parameter based on the multi-dimensional dynamic weight vector and the preferred response feature representation through the gating network using the tenth calculation formula, and determine the action rejection response preference parameter based on the multi-dimensional dynamic weight vector and the rejection response feature representation using the eleventh calculation formula;

[0320] The tenth calculation formula satisfies the following formula:

[0321] R chosen = ω T T chosen

[0322] where, R chosen represents the action preferred response preference parameter, and T chosen represents the preferred response feature representation, which is the feature representation obtained through the feature extractor; The preference score R chosen represents the adaptability of the preferred response under the current task condition x. The higher the score, the more the preferred response meets the task requirements.

[0323] The eleventh calculation formula satisfies the following formula:

[0324] R rejected = ωT T rejected

[0325] Among them, the R rejected represents the action rejection response preference parameter, and the T rejected represents the rejection response feature representation, which is the feature representation obtained by the feature extractor; the preference score R rejected represents the adaptability of the rejection response under the current task condition x. The lower the score, the less the rejection response meets the task requirements.

[0326] Step C10: Calculate the preference alignment loss based on the preferred response preference parameter and the action rejection response preference parameter, and adjust the parameters of the gating network based on the preference alignment loss;

[0327] The twelfth calculation formula satisfies the following formula:

[0328]

[0329] Among them, the L BT represents the preference alignment loss. The optimization goal is to maximize the score difference between R chosen and R rejected so that the model can correctly distinguish the preferred and rejected candidate responses, and then update the parameters φ of the gating network. By optimizing L BT , the gating network gφ learns to generate more accurate dynamic weights w, improving the ability of the reward model to distinguish between preferred and rejected responses.

[0330] Optionally, after completing all model training processes, return the trained step-by-step multi-dimensional evaluation reward model, which can be represented by the following model calculation formula:

[0331]

[0332] Among them, the feature extractor corresponds to f θ and provides multi-modal features.

[0333] The gating network corresponds to g φ The dynamically generated weight w.

[0334] R represents the prediction of the step-by-step multi-dimensional evaluation reward model for the reward information and the action supervision evaluation information.

[0335] In one or more embodiments of this specification, the step-by-step multi-dimensional evaluation reward model Similar has made significant progress in the training and inference of general virtual agents. It can provide fine-grained, multi-dimensional feedback, addressing the key limitations of current reward modeling methods and making learning more efficient and effective. The impact of this model is not limited to reinforcement learning training but also extends to the scale and data cleaning during inference, making it a versatile and practical solution applicable to a wide range of applications.

[0336] The following will be combined with Figure 8 to introduce the task autonomous processing device provided in this specification in detail. It should be noted that Figure 8 the task autonomous processing device shown is used to execute the method of the embodiment shown in this specification. For the sake of convenience of description, only the parts related to this specification are shown. For those specific technical details not disclosed, please refer to the embodiment shown in this specification Figures 1 to 7 Figures 1 to 7 shown.

[0337] Please refer to Figure 8 , which shows the structural schematic diagram of the task autonomous processing device of this specification. The task autonomous processing device 1 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the task autonomous processing device 1 includes a task determination module 11, a task autonomous inference module 12, and a task autonomous inference optimization module 13, specifically used for:

[0338] The task determination module 11 is used to determine the transaction tasks of the intelligent agent service object corresponding to the transaction processing large model;

[0339] The task autonomous inference module 12 is used to perform task autonomous inference processing on the transaction tasks by invoking the transaction processing large model using the intelligent agent service object to obtain a task execution action sequence, and the task execution action sequence includes at least one task execution action;

[0340] The quality evaluation module 13 is used to perform step-by-step behavior quality evaluation processing on the task execution actions for the transaction tasks to obtain action supervision evaluation information for each task execution action;

[0341] The inference optimization module 14 is used to perform task autonomous inference optimization processing on the intelligent agent service object based on the action supervision evaluation information.

[0342] Optionally, the step-by-step behavior quality evaluation processing on the task execution actions for the transaction tasks to obtain action supervision evaluation information for each task execution action includes:

[0343] ​Construct a Monte Carlo tree for the transaction task based on the action sequence for task execution, and perform simulation processing on the execution trajectory of each task execution action based on the Monte Carlo tree to obtain the number of inference steps and the basic reward for reaching the simulated task goal.

[0344] Determine the helpfulness parameter, success probability parameter, and efficiency parameter of each task execution action based on the number of inference steps and the basic reward, and use the first multi-modal large model to determine the task relevance parameter and coherence parameter of the task execution action.

[0345] Obtain the action supervision evaluation information of each task execution action based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter, and the coherence parameter; or, perform multi-dimensional parameter weighted calculation based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter, and the coherence parameter to obtain the action comprehensive score of each task execution action, and calculate the mean value of the step scores based on all the action comprehensive scores to obtain the mean value of the trajectory action scores. Obtain the action supervision evaluation information of each task execution action based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter, the action comprehensive score, and the mean value of the trajectory action scores.

[0346] Optionally, the determining the helpfulness parameter, success probability parameter, and efficiency parameter of each task execution action based on the number of inference steps and the basic reward includes:

[0347] Obtain the cumulative execution action value, the number of inference steps, the execution action number, and the basic reward corresponding to the task execution action, calculate the helpfulness parameter of each task execution action using the first calculation formula, and update the cumulative execution action value using the second calculation formula; the first calculation formula satisfies the following formula:

[0348]

[0349] where the i H i-1 represents the helpfulness parameter, the i AC i represents the cumulative execution action value at the (i - 1)-th step, the

[0350] M represents the number of inference steps, the

[0351]

[0352] i represents the execution action number, and thei represents the cumulative execution action value of the i-th step;

[0353] and,

[0354] Obtain the trajectory simulation action, standard simulation action, and total number of trajectory simulation actions corresponding to the task execution action, and calculate the success probability parameter of each task execution action using a third calculation formula; the third calculation formula satisfies the following formula:

[0355]

[0356] wherein, the OS i represents the success probability parameter, the N represents the total number of trajectory simulation actions, the j is the number of the trajectory simulation action, and the a i,j represents the j-th trajectory simulation action of the i-th task execution action, the a* represents the standard simulation action, and the II() is an indicator function, and the indicator function is used to compare whether the trajectory simulation action matches the standard simulation action;

[0357] and,

[0358] Determine the historical task remaining action length and the current task remaining action length corresponding to the task execution action, and calculate the efficiency parameter of each task execution action using a fourth calculation formula; the fourth calculation formula satisfies the following formula:

[0359]

[0360] wherein, the E i represents the efficiency parameter of the i-th task execution action, the Len i-1 represents the historical task remaining action length, and the Len i represents the current task remaining action length; the current task remaining action length satisfies the following fifth calculation formula:

[0361]

[0362] wherein, the Len(S i,j) represents the task remaining action length corresponding to each trajectory simulation action.

[0363] Optionally, the determining the task relevance parameter and coherence parameter of the task execution action by using the first multimodal large model includes:

[0364] Performing a relevance analysis and evaluation of the task execution action for the transaction task by using the first multimodal large model based on a preset task relevance evaluation prompt to obtain the task relevance parameter corresponding to each task execution action;

[0365] Based on the preset action coherence evaluation prompt, use the first multi-modal large model to perform action coherence evaluation processing on the task execution actions for the transaction task to obtain the coherence parameters corresponding to each task execution action.

[0366] Optionally, the task autonomous reasoning optimization processing for the intelligent agent service object based on the action supervision evaluation information includes:

[0367] Train the step-by-step multi-dimensional evaluation reward model corresponding to the intelligent agent service object based on the action supervision evaluation information;

[0368] Use the step-by-step multi-dimensional evaluation reward model to perform task autonomous reasoning optimization processing on the intelligent agent service object.

[0369] Optionally, the execution trajectory simulation processing for each task execution action based on the Monte Carlo tree to obtain the number of reasoning steps and the basic reward for reaching the simulated task goal includes:

[0370] Perform node search and expansion on each execution action tree node corresponding to the task execution action in the Monte Carlo tree to obtain multiple trajectory simulation action nodes of the execution action tree node, where the trajectory simulation action node represents the reasoning expansion action after the task execution action;

[0371] Based on the multiple trajectory simulation action nodes, perform task execution trajectory simulation processing from the execution action tree node to obtain a task execution simulation result;

[0372] Determine the target task execution simulation result that meets the simulated task goal from the task execution simulation result, and determine the number of reasoning steps and the basic reward based on the target task execution simulation result.

[0373] Optionally, the training of the step-by-step multi-dimensional evaluation reward model corresponding to the intelligent agent service object based on the action supervision evaluation information includes:

[0374] Generate multi-dimensional evaluation sample data based on the action supervision evaluation information, the transaction task, and the task execution action sequence;

[0375] Determine the initial reward model of the intelligent agent service object, and use the multi-dimensional evaluation sample data to perform model training on the initial reward model to obtain the step-by-step multi-dimensional evaluation reward model after model training.

[0376] Optionally, the determination of the initial reward model of the intelligent agent service object, the use of the multi-dimensional evaluation sample data to perform model training on the initial reward model, and the obtaining of the step-by-step multi-dimensional evaluation reward model after model training include:

[0377] Obtain a pre-trained second multi-modal large model and a gating network, determine a backbone feature extraction network based on the second multi-modal large model, and create an initial reward model for the intelligent agent service object based on the backbone feature extraction network and the gating network;

[0378] Use the multi-dimensional evaluation sample data to train the initial reward model. During the model training process, perform multi-dimensional supervised evaluation information regression training on the initial reward model and perform dynamic weight adjustment on the initial reward model to obtain a step-by-step multi-dimensional evaluation reward model after the model training ends.

[0379] Optionally, the backbone feature extraction network includes a backbone feature extractor and a linear regression layer. The multi-dimensional supervised evaluation information regression training on the initial reward model includes:

[0380] Based on the backbone feature extractor, use the sixth calculation formula to extract the hidden state features from the data input sequence corresponding to the multi-dimensional evaluation sample data, and use the seventh calculation formula to perform a linear mapping on the hidden state features based on the linear regression layer to obtain a predicted action supervised evaluation reward vector;

[0381] Determine a true action supervised evaluation reward vector based on the multi-dimensional evaluation sample data;

[0382] Based on the predicted action supervised evaluation reward vector and the true action supervised evaluation reward vector, use the eighth calculation formula to obtain an optimized regression loss, and use the optimized regression loss to adjust the backbone network parameters and regression weight parameters of the backbone feature extraction network;

[0383] The sixth calculation formula satisfies the following formula:

[0384]

[0385] where, the h represents the hidden state feature, the f θ represents the hidden state extraction operation of the backbone feature extractor, the x represents the predicted input task sample in the data input sequence, and the y represents the candidate action response item information generated based on the action supervised evaluation information;

[0386] The seventh calculation formula satisfies the following formula:

[0387]

[0388] where, the represents the predicted action supervised evaluation reward vector, and the W represents the regression weight parameter;

[0389] The eighth calculation formula satisfies the following formula:

[0390]

[0391] Wherein, the L RG represents the optimized regression loss, and the r represents the true action supervision evaluation reward vector.

[0392] Optionally, the dynamic weight adjustment of the initial reward model includes:

[0393] Obtain the hidden state features, and perform weight calculation processing on the hidden state features through the gated network using the ninth calculation formula to obtain a multi-dimensional dynamic weight vector;

[0394] Determine candidate action response item information based on the multi-dimensional evaluation sample data, where the candidate action response item information includes action preferred response item information and action rejection response item information;

[0395] Determine the preferred response feature representation corresponding to the action preferred response item information and the rejection response feature representation corresponding to the action rejection response item information through the gated network;

[0396] Determine the action preferred response preference parameter through the gated network based on the multi-dimensional dynamic weight vector and the preferred response feature representation using the tenth calculation formula, and determine the action rejection response preference parameter based on the multi-dimensional dynamic weight vector and the rejection response feature representation using the eleventh calculation formula;

[0397] Calculate the preference alignment loss based on the preferred response preference parameter and the action rejection response preference parameter, and adjust the parameters of the gated network based on the preference alignment loss;

[0398] The ninth calculation formula satisfies the following formula:

[0399] ω = g φ (f θ (x))

[0400] Wherein, the ω represents the multi-dimensional dynamic weight vector, the f θ (x) represents the hidden state features extracted by the backbone feature extractor, and the g φ () represents the weight calculation processing operation of the gated network;

[0401] The tenth calculation formula satisfies the following formula:

[0402] R chosen = ω T T chosen

[0403] Among them, the R chosen represents the preferred response preference parameter of the action, and the T chosen represents the preferred response feature representation;

[0404] The eleventh calculation formula satisfies the following formula:

[0405] R rejected = ω T T rejected

[0406] Among them, the R rejected represents the action rejection response preference parameter of the action, and the T rejected represents the rejection response feature representation;

[0407] The twelfth calculation formula satisfies the following formula:

[0408]

[0409] Among them, the L BT represents the preference alignment loss.

[0410] Optionally, using the step-by-step multi-dimensional evaluation reward model to perform task autonomous reasoning optimization processing on the intelligent agent service object includes:

[0411] In response to the user's target transaction task for the intelligent agent service object, based on the target transaction task, the intelligent agent service object is used to call the transaction processing large model to perform task autonomous reasoning processing to obtain a target task execution action sequence, and the target task execution action sequence includes at least one target task execution action;

[0412] Taking the target task execution action as the task execution action, and using the step-by-step multi-dimensional evaluation reward model to perform step-by-step behavior quality evaluation processing on the task execution action for the transaction task, to obtain the action supervision evaluation information of each task execution action;

[0413] Based on the action supervision evaluation information, driving the intelligent agent service object to perform execution action optimization processing on the target task execution action sequence.

[0414] It should be noted that when the task autonomous processing device provided in the above embodiments executes the task autonomous processing method, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the task autonomous processing device provided in the above embodiments and the embodiments of the task autonomous processing method belong to the same concept. For the implementation process, please refer to the method embodiments, which will not be elaborated here.

[0415] The serial numbers in this specification are only for description and do not represent the superiority or inferiority of the embodiments.

[0416] In one or more embodiments of this specification, by determining the transaction tasks of the intelligent agent service object corresponding to the transaction processing large model, based on the transaction tasks, the intelligent agent service object is used to call the transaction processing large model to perform task autonomous reasoning processing to obtain a task execution action sequence. By gradually evaluating the behavior quality of the task execution action sequence during the task autonomous reasoning process, action supervision and evaluation information with a higher degree of granularity is generated, thereby replacing the single supervision signal based on the task completion result. It effectively solves the limitations of lack of process supervision, dependence on manual annotation, and difficulty in realizing reasoning time extension, can accurately locate the execution details during the task execution process, improve the learning and reasoning efficiency of the intelligent agent, and, in addition, can also reduce the demand for manual annotation, realize the automatic generation of supervision signals during the task autonomous processing process, reduce the processing cost, enhance the processing ability of complex tasks by gradually evaluating and optimizing task reasoning, and thus significantly improve the task autonomous processing performance and application value of the multi-modal virtual intelligent agent.

[0417] This specification also provides a computer storage medium, which can store multiple instructions. The instructions are suitable for being loaded and executed by a processor to perform the task autonomous processing method as described in the above Figures 1 to 7 illustrated embodiments. The specific execution process can refer to the specific description of the Figures 1 to 7 illustrated embodiments and will not be elaborated here.

[0418] This specification also provides a computer program product, which stores at least one instruction. The at least one instruction is loaded and executed by the processor to perform the task autonomous processing method as described in the above Figures 1 to 7 illustrated embodiments. The specific execution process can refer to the specific description of the Figures 1 to 7 illustrated embodiments and will not be elaborated here.

[0419] Please refer to Figure 9, which is a block diagram of the structure of an electronic device provided by an embodiment of this specification. The electronic device in this specification may include one or more of the following components: a processor 1010, a memory 1020, an input device 1030, an output device 1040, and a bus 1050. The processor 1010, the memory 1020, the input device 1030, and the output device 1040 may be connected through the bus 1050.

[0420] The processor 1010 may include one or more processing cores. The processor 1010 uses various interfaces and lines to connect various parts within the entire electronic device. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1020, and by calling data stored in the memory 1020, it executes various functions of the electronic device and processes data. Optionally, the processor 1010 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1010 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 1010 and may be implemented separately through a communication chip.

[0421] The memory 1020 may include random access memory (RAM) and may also include read-only memory (ROM). Optionally, the memory 1020 includes a non-transitory computer-readable storage medium. The memory 1020 can be used to store instructions, programs, code, code sets, or instruction sets.

[0422] Among them, the input device 1030 is used to receive input instructions or data. The input device 1030 includes, but is not limited to, a keyboard, a mouse, a camera, a microphone, or a touch device. The output device 1040 is used to output instructions or data. The output device 1040 includes, but is not limited to, a display device, a speaker, etc. In the embodiments of this specification, the input device 1030 may be a temperature sensor for obtaining the operating temperature of the electronic device. The output device 1040 may be a speaker for outputting an audio signal.

[0423] In addition, those skilled in the art can understand that the structure of the electronic device shown in the above drawings does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements. For example, the electronic device further includes components such as a radio frequency circuit, an input unit, a sensor, an audio circuit, a wireless fidelity (WIFI) module, a power supply, a Bluetooth module, etc., which will not be elaborated here.

[0424] In the embodiments of this specification, the execution subject of each step may be the electronic device introduced above. Optionally, the execution subject of each step is the operating system of the electronic device. The operating system may be the Android system, the IOS system, or other operating systems, and the embodiments of this specification do not limit this.

[0425] In Figure 9 the electronic device, the processor 1010 may be used to call the program stored in the memory 1020 and execute to implement the task autonomous processing method as described in each method embodiment of this specification.

[0426] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.

[0427] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the transaction tasks and user information involved in this specification are all obtained under sufficient authorization.

[0428] The above-disclosed is only the preferred embodiment of this specification. Of course, it cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification still fall within the scope covered by this specification.

Claims

1. A method for autonomous task processing, the method comprising: Determine the transaction tasks of the intelligent agent service object corresponding to the transaction processing model; Based on the transaction task, the agent service object is used to call the transaction processing large model to perform task autonomous reasoning processing to obtain a task execution action sequence, wherein the task execution action sequence includes at least one task execution action; Performing step-by-step behavior quality evaluation processing on the task execution action for the transaction task to obtain action supervision evaluation information of each task execution action; Based on the action supervision and evaluation information, task autonomous reasoning and optimization processing is performed on the intelligent agent service object.

2. According to the method of claim 1, the step of performing a step-by-step behavior quality evaluation process on the task execution action to obtain action supervision and evaluation information of each task execution action comprises: Constructing a Monte Carlo tree for the transaction task based on the task execution action sequence, and performing execution trajectory simulation processing for each task execution action based on the Monte Carlo tree to obtain the number of reasoning steps and basic rewards for achieving the simulation task goal; Determine a helpfulness parameter, a success probability parameter, and an efficiency parameter of each of the task execution actions based on the number of reasoning steps and the basic reward, and determine a task relevance parameter and a coherence parameter of the task execution actions using a first multimodal large model; Obtaining action supervision and evaluation information of each of the task execution actions based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter, and the coherence parameter; Alternatively, a multi-dimensional parameter weighted calculation is performed based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter and the continuity parameter to obtain a comprehensive action score for each task execution action, and a step score average is calculated based on all the action comprehensive scores to obtain a trajectory action score average, and action supervision and evaluation information for each task execution action is obtained based on the helpfulness parameter, the success probability parameter, the efficiency parameter, the task relevance parameter, the action comprehensive score and the trajectory action score.

3. The method according to claim 2, wherein determining the helpfulness parameter, success probability parameter and efficiency parameter of each task execution action based on the number of reasoning steps and the basic reward comprises: Obtain the cumulative execution action value, the number of reasoning steps, the execution action number and the basic reward corresponding to the task execution action, use the first calculation formula to calculate the helpfulness parameter of each task execution action, and use the second calculation formula to update the cumulative execution action value; the first calculation formula satisfies the following formula: Among them, the H i represents the helpfulness parameter, the AC i-1 represents the cumulative execution action value of the i-1th step, M represents the number of reasoning steps, i represents the execution action number, and r i Indicates the task execution action S i The basic reward of The second calculation formula satisfies the following formula: Among them, the AC i represents the cumulative execution action value of the i-th step; as well as, The trajectory simulation action, the standard simulation action and the total number of trajectory simulation actions corresponding to the task execution action are obtained, and the success probability parameter of each task execution action is calculated using a third calculation formula; the third calculation formula satisfies the following formula: Among them, the OS i represents the success probability parameter, N represents the total number of trajectory simulation actions, j is the number of the trajectory simulation action, and a i,j represents the jth trajectory simulation action of the i-th task execution action, the a* represents the standard simulation action, and the II() is an indicator function, which is used to compare whether the trajectory simulation action matches the standard simulation action; as well as, Determine the remaining action length of the historical task and the remaining action length of the current task corresponding to the task execution action, and use the fourth calculation formula to calculate the efficiency parameter of each task execution action; the fourth calculation formula satisfies the following formula: Among them, the E i represents the efficiency parameter of the i-th task execution action, the Len i-1 Indicates the remaining action length of the historical task, the Len i Indicates the remaining action length of the current task; the remaining action length of the current task satisfies the following fifth calculation formula: Wherein, the Len(S i,j ) represents the remaining action length of the task corresponding to each of the trajectory simulation actions.

4. The method according to claim 2, wherein the step of using the first multimodal macro model to determine the task relevance parameter and the coherence parameter of the task execution action comprises: Based on the preset task relevance evaluation prompt, the first multimodal large model is used to perform relevance analysis and evaluation on the task execution action for the transaction task to obtain a task relevance parameter corresponding to each task execution action; Based on the preset action consistency evaluation prompt, the first multimodal large model is used to perform action consistency evaluation processing on the task execution action for the transaction task to obtain the consistency parameter corresponding to each task execution action.

5. According to the method of claim 1, the step of performing task autonomous reasoning optimization processing on the intelligent agent service object based on the action supervision evaluation information comprises: Based on the action supervision evaluation information, a step-by-step multi-dimensional evaluation reward model corresponding to the agent service object is trained; The step-by-step multi-dimensional evaluation and reward model is used to perform task autonomous reasoning optimization processing on the intelligent agent service object.

6. The method according to claim 2 or 5, wherein the execution trajectory simulation processing for each task execution action is performed based on the Monte Carlo tree to obtain the number of reasoning steps and basic rewards for achieving the simulated task goal, comprises: Performing node search and expansion on each execution action tree node corresponding to the task execution action in the Monte Carlo tree to obtain a plurality of trajectory simulation action nodes of the execution action tree node, wherein the trajectory simulation action node represents an inference expansion action after the task execution action; Based on the plurality of trajectory simulation action nodes, task execution trajectory simulation processing is performed from the execution action tree nodes to obtain a task execution simulation result; A target task execution simulation result that meets the simulation task goal is determined from the task execution simulation results, and the number of reasoning steps and the basic reward are determined based on the target task execution simulation result.

7. The method according to claim 5, wherein the step-by-step multi-dimensional evaluation reward model corresponding to the agent service object is trained based on the action supervision evaluation information, comprising: Generate multi-dimensional evaluation sample data based on the action supervision evaluation information, the transaction task and the task execution action sequence; An initial reward model for the agent service object is determined, and the multi-dimensional evaluation sample data is used to perform model training on the initial reward model to obtain a step-by-step multi-dimensional evaluation reward model after model training.

8. According to the method of claim 7, the determining of the initial reward model of the agent service object, using the multi-dimensional evaluation sample data to perform model training on the initial reward model to obtain a step-by-step multi-dimensional evaluation reward model after model training, comprises: Acquire a pre-trained second multimodal large model and a gating network, determine a backbone feature extraction network based on the second multimodal large model, and create an initial reward model for an intelligent agent service object based on the backbone feature extraction network and the gating network; The multidimensional evaluation sample data is used to perform model training on the initial reward model. During the model training process, the initial reward model is subjected to multidimensional supervisory evaluation information regression training and the initial reward model is dynamically weighted to obtain a step-by-step multidimensional evaluation reward model after the model training is completed.

9. The method according to claim 8, wherein the backbone feature extraction network comprises a backbone feature extractor and a linear regression layer, and the performing multi-dimensional supervised evaluation information regression training on the initial reward model comprises: Based on the backbone feature extractor, the data input sequence corresponding to the multi-dimensional evaluation sample data is subjected to hidden state extraction using the sixth calculation formula to obtain hidden state features, and based on the linear regression layer, the hidden state features are subjected to linear mapping using the seventh calculation formula to obtain a predicted action supervision evaluation reward vector; Determining a real action supervision evaluation reward vector based on the multi-dimensional evaluation sample data; Based on the predicted action supervision evaluation reward vector and the real action supervision evaluation reward vector, an optimized regression loss is obtained by using the eighth calculation formula, and the backbone network parameters and regression weight parameters of the backbone feature extraction network are adjusted by using the optimized regression loss; The sixth calculation formula satisfies the following formula: Wherein, h represents the hidden state feature, and f θ represents a hidden state extraction operation of the backbone feature extractor, the x represents a prediction input task sample in the data input sequence, and the y represents candidate action response item information generated based on the action supervision evaluation information; The seventh calculation formula satisfies the following formula: Among them, the represents the predicted action supervision evaluation reward vector, and W represents the regression weight parameter; The eighth calculation formula satisfies the following formula: Among them, the L RG Represents the optimized regression loss, and r represents the real action supervision evaluation reward vector.

10. The method according to claim 9, wherein the dynamically adjusting the weight of the initial reward model comprises: Acquire the hidden state feature, and perform weight calculation processing on the hidden state feature by using the gating network using the ninth calculation formula to obtain a multi-dimensional dynamic weight vector; Determine candidate action response item information based on the multi-dimensional evaluation sample data, wherein the candidate action response item information includes action preferred response item information and action rejected response item information; Determine, by the gating network, a preferred response feature representation corresponding to the action preferred response item information and a rejection response feature representation corresponding to the action rejection response item information; Determine the action preferred response preference parameter by using the gating network based on the multidimensional dynamic weight vector and the preferred response feature representation using a tenth calculation formula, and determine the action rejection response preference parameter by using an eleventh calculation formula based on the multidimensional dynamic weight vector and the rejection response feature representation; Calculate the preference alignment loss using the twelfth calculation formula based on the preferred response preference parameter and the action rejection response preference parameter, and adjust the parameters of the gating network based on the preference alignment loss; The ninth calculation formula satisfies the following formula: ω=g φ (f θ (x)) Wherein, ω represents the multi-dimensional dynamic weight vector, and f θ (x) represents the hidden state feature extracted by the backbone feature extractor, and the g φ () represents the weight calculation processing operation of the gating network; The tenth calculation formula satisfies the following formula: R chosen =ω T T chosen Among them, the R chosen represents the preferred response preference parameter of the action, the T chosen indicating said preferred response characteristic representation; The eleventh calculation formula satisfies the following formula: R rejected =ω T T rejected Among them, the R rejected represents the action rejection response preference parameter, the T rejected indicating said rejection response characteristic representation; The twelfth calculation formula satisfies the following formula: Among them, the L BT represents the preference alignment loss.

11. The method according to claim 5, wherein the step-by-step multi-dimensional evaluation and reward model is used to perform task autonomous reasoning optimization processing on the intelligent agent service object, comprising: In response to a target transaction task of the user for the agent service object, the agent service object is used to call the transaction processing big model to perform task autonomous reasoning based on the target transaction task to obtain a target task execution action sequence, wherein the target task execution action sequence includes at least one target task execution action; The target task execution action is used as the task execution action, and the step-by-step multi-dimensional evaluation reward model is used to perform the step-by-step behavior quality evaluation process on the task execution action for the transaction task, so as to obtain the action supervision evaluation information of each task execution action; Based on the action supervision and evaluation information, the intelligent agent service object is driven to perform action optimization processing on the target task execution action sequence.

12. A task autonomous processing device, the device comprising: A task determination module is used to determine the transaction tasks of the intelligent agent service object corresponding to the transaction processing large model; An autonomous reasoning module is used to use the agent service object to call the transaction processing large model to perform task autonomous reasoning processing based on the transaction task to obtain a task execution action sequence, wherein the task execution action sequence includes at least one task execution action; A quality evaluation module is used to perform a step-by-step behavior quality evaluation process on the task execution action for the transaction task, and obtain action supervision and evaluation information of each task execution action; The reasoning optimization module is used to perform task autonomous reasoning optimization processing on the intelligent agent service object based on the action supervision evaluation information.

13. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 11.

14. A computer program product, the computer program product storing at least one instruction, wherein the at least one instruction is loaded by a processor and executes the method steps according to any one of claims 1 to 11.

15. An electronic device, comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps as claimed in any one of claims 1 to 11.