Multi-agent collaboration method based on large language model

By constructing a multi-agent coordinated resource allocation model and an improved particle swarm algorithm, the problems of uncertainty accumulation and unbalanced resource allocation in the multi-agent system are solved, and efficient, stable and high-quality task allocation under complex tasks are achieved.

CN120338035AActive Publication Date: 2025-07-18北京长河数智科技有限责任公司 +1

Patent Information

Application Number
CN202510827689.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing multi-agent collaborative system has problems of uncertainty accumulation and uneven resource allocation in complex tasks, resulting in reduced results reliability and unstable system performance.

Method used

A multi-agent coordinated resource allocation model is constructed, task allocation problems are transformed into multi-objective optimization problems, and an improved particle swarm algorithm is introduced, and uncertainty is evaluated through multiple random sampling, and a multi-agent cross-verification mechanism is adopted to optimize the task allocation strategy.

Benefits of technology

The inference quality and efficiency of multi-agent systems under complex tasks are improved, the stability of results and the balanced utilization of resources are ensured, and the changes in the dynamic environment are adapted to dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338035A_ABST
    Figure CN120338035A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent cooperation method based on a large language model, and relates to the field of agent cooperation control, and the method comprises the steps: constructing a multi-agent system; establishing a multi-agent coordination resource allocation model; a task allocation problem in the resource allocation model is converted into a multi-objective optimization problem, and optimization objectives comprise the minimum task completion time and the optimal resource utilization rate; and solving a multi-objective optimization problem by using an improved particle swarm optimization (PSO) to obtain an optimal cooperation scheme among the master control agent, the task creation agent, the intention arrangement agent and the at least one professional agent. The improved particle swarm optimization (PSO) evaluates the uncertainty of each candidate solution through multiple times of random sampling, and when the uncertainty exceeds a threshold value, cross validation is performed through a multi-agent voting mechanism. For uncertainty accumulation of a multi-agent system under a complex cognitive task, the efficiency and stability of collaborative decision and task scheduling of the multi-agent system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of collaborative control of agents, and particularly to a multi-agent collaboration method based on large language models. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models (LLMs) have become the core technology driving intelligent applications due to their powerful natural language understanding and generation capabilities. To address the limitations of a single large language model in handling complex tasks, multi-agent collaborative systems have emerged. By organizing agents with different functions, multi-agent systems form a "division of labor and cooperation" working mode, which can handle more complex cognitive tasks and significantly expand the application boundaries of large language models.

[0003] Such systems have been widely applied in fields such as intelligent finance, medical diagnosis, scientific research, and complex decision-making support. Taking intelligent financial risk assessment as an example, the system needs to simultaneously analyze multi-dimensional information such as enterprise financial statements, industry risks, credit history, and market prospects, and make accurate risk judgments by integrating various factors. This exceeds the processing capacity of a single model and must rely on multi-agent collaboration to complete.

[0004] However, the existing large language model multi-agent collaboration technologies have obvious deficiencies, mainly reflected in the following aspects: On the one hand, there is uncertainty in the reasoning process of large language models itself. When multiple agents collaborate serially according to a specific process, the uncertainty of the previous agent will be transmitted and amplified to the subsequent agents, resulting in a significant decrease in the reliability of the final result. For example, in financial risk assessment, if the financial analysis agent has a deviation in the interpretation of certain financial indicators, this deviation will affect the judgment of the subsequent industry risk assessment agent, and may ultimately lead to incorrect credit decisions. On the other hand, traditional multi-agent systems usually use methods such as fixed rule allocation, simple polling mechanism, or static priority queue for task allocation and resource scheduling. These methods lack environmental adaptability and cannot dynamically adjust the resource allocation strategy according to task complexity and system load. In high-concurrency scenarios, it is easy to cause uneven resource utilization, with some agents overloaded while others are idle, affecting the overall system performance.

[0005] For example, the related technology CN202211459220.1 discloses a multi-agent collaborative task allocation method based on an improved particle swarm algorithm with non-dominated sorting, including: establishing an objective benefit model for multi-agent task allocation in combination with job environment information; establishing a loss cost model for multi-agent task allocation in combination with job environment information; establishing a multi-agent collaborative task allocation model based on multi-objective functions and combining the constraint conditions in the process of agents executing tasks; using an improved particle swarm algorithm based on non-dominated sorting to solve the obtained model to obtain a Pareto solution set; and obtaining a Pareto optimal solution through the maximum distance method based on the Pareto solution set. However, this solution assumes that the job environment is static and determined, and does not consider the impact of environmental parameter fluctuations on task allocation. In practical applications, parameters such as task processing time, resource consumption, and agent availability have random fluctuations, resulting in the "optimal solution" generated by it may be unstable or even completely ineffective under environmental perturbations. Summary of the Invention

[0006] Aiming at the problem of uncertainty accumulation in multi-agent systems under complex cognitive tasks, the present application provides a multi-agent collaborative method based on a large language model, which improves the traditional particle swarm algorithm by introducing uncertainty perception, and breaks through the limitations of traditional fixed task allocation strategies through a multi-agent cross-validation mechanism.

[0007] The present application provides a multi-agent collaborative method based on a large language model, including: constructing a multi-agent system, the multi-agent architecture includes a master agent, a task creation agent, an intent orchestration agent, and at least one professional agent; the master agent is responsible for overall coordination and decision-making; the task creation agent is responsible for task decomposition; the intent orchestration agent is responsible for execution process planning; the professional agent is responsible for target domain task processing.

[0008] Establish a multi-agent coordinated resource allocation model for managing the computing resources and task processing time of the multi-agent system; the resource allocation model includes a resource supply matrix, a task demand vector, and resource allocation constraint conditions, including: establishing a resource supply matrix R, representing the types and quantities of computing resources provided by each agent, where, represents the supply volume of the i-th agent for the j-th type of resource; establish a task demand vector D, representing the demand for computing resources and processing time requirements of each task, where represents the demand for the j-th type of resource of the k-th task; set resource allocation constraint conditions C, including execution time constraints; establish a task allocation matrix X, where, represents that the k-th task is assigned to the i-th agent, represents not assigned; construct a multi-agent coordinated resource allocation model , where the task assignment problem is defined as: determining the optimal task assignment matrix X while satisfying the constraint condition C.

[0009] The task assignment problem in the resource allocation model is transformed into a multi-objective optimization problem. The optimization objectives include minimizing the task completion time and the optimal resource utilization rate, including: defining the task completion time objective function , which represents the total completion time after all tasks are assigned according to the task assignment matrix X. , where represents the completion time of the k-th task, which depends on the processing capacity and current load of the assigned agent; defining the resource utilization rate objective function , which represents the balance degree of system resource utilization after assignment according to the task assignment matrix X. , where represents the resource load rate of the i-th agent. represents the standard deviation function. The smaller the standard deviation, the more balanced the resource utilization; constructing the objective function , where and respectively represent the weight coefficients of the task completion time and the resource utilization rate; setting the total resource constraint , which means that the total demand of all tasks assigned to each agent for various resources shall not exceed the supply of the corresponding resources of the corresponding agent; setting the execution time constraint. For task k assigned to agent i, its completion time shall not exceed the preset task deadline , that is ;

[0010] The task assignment problem is expressed as a multi-objective optimization problem: ; Constraint conditions:

[0011] (1) ;

[0012] (2) ;

[0013] (3) ;

[0014] (4) ;

[0015] Among them, X is a 0-1 integer matrix, representing the assignment relationship from tasks to agents.

[0016] In particular, the task assignment in traditional multi-agent systems often relies on rule engines or heuristic algorithms. In this application, by introducing objective functions and constraint conditions, the multi-agent cooperation problem becomes a solvable optimization problem.

[0017] Using the improved Particle Swarm Optimization (PSO) algorithm to solve the multi-objective optimization problem and obtain the optimal cooperation plan, including: initializing the particle swarm, where each particle represents a candidate task assignment matrix X, the number of particles is N, and the number of iterations is G; calculating the objective function value F(X) of each particle;

[0018] Evaluating the uncertainty index of each particle through multiple random samplings , including: generating m perturbations of environmental parameters for each particle X, where the environmental parameter perturbations include random changes in task processing time, resource consumption, and agent availability; each perturbation is actually a simulation of a possible state of the real environment, constructing a "digital twin" of the solution to predict its behavior in actual deployment. Specifically, for task processing time perturbation: it simulates the processing time fluctuations in the real environment due to data complexity and quality changes, usually perturbed within the range of [-15%, +25%]. For resource consumption perturbation: it simulates the changes in resource requirements for different task instances, usually perturbed within the range of [-10%, +20%]. For agent availability perturbation: it simulates the situation where an agent is temporarily unavailable with a certain probability (usually 5%), which reflects possible model updates, maintenance, or failure scenarios in the actual system.

[0019] Under each perturbation condition, calculate the objective function value ; Introduce a time decay factor λ to increase the influence of the most recent sampling results on uncertainty assessment, and calculate the uncertainty index : , where 0 < λ < 1, and σ is the standard deviation function; specifically, the environmental perturbation is essentially a non-stationary process, that is, the statistical characteristics of the system change over time, so the most recent observations have higher predictive value for the current state. By weight, making the influence of the most recent sampling significantly higher than that of early sampling.

[0020] When the uncertainty index of the particle is greater than the preset threshold τ, the master agent broadcasts an evaluation request to at least one professional agent; each professional agent evaluates the candidate solution X and calculates the local score , and the calculation formula is: , where represents the priority of the kth task, represents the processing ability coefficient of the ith professional agent for the kth task, represents the execution efficiency of the ith professional agent for the kth task; set the credibility weights of each professional agent : , where represents the number of times the professional agent's historical evaluations were correct, Denote the total number of evaluations involved; the evaluation results of each professional agent are integrated using a weighted voting mechanism to obtain a comprehensive score ; According to the comprehensive score the objective function value is corrected, and the calculation formula is: , where α is the correction coefficient, and its value range is [0, 2], and S(X) ∈ [0, 1].

[0021] For each particle, compare the corrected objective function value of the current position X with the objective function value of the corresponding individual historical optimal position ; Calculate the criterion according to the corrected objective function value and the uncertainty index; ;

[0022] , where β is the uncertainty sensitivity coefficient, and its value range is [0, 1], and are the uncertainty indexes of the current position and the individual historical optimal position respectively; when , update the individual optimal position , otherwise keep it unchanged; after each iteration, select the position with the minimum objective function value from all the individual optimal positions of the particles as the new global optimal position ; Set the global optimal position update counter , when the global optimal position has not changed for consecutive iterations, perform local random perturbation on to avoid falling into local optimality: randomly select r% of the elements in the matrix and reverse their values, that is, for the selected element (k, i), execute .

[0023] Specifically, traditional PSO requires strict monotonic improvement, that is , while this application allows solutions with slightly worse objective function values but significantly reduced uncertainty in some cases. This non-monotonic improvement strategy enables the algorithm to jump out of local optimality and improve the global search ability. In addition, at the individual level in this application, non-monotonic updates are allowed through uncertainty trade-offs; at the global level, large-scale jumps are provided through random perturbations. This multi-level escape strategy significantly improves the performance of the algorithm in complex multi-modal optimization problems.

[0024] Update the velocity and position of the particle;

[0025] ;

[0026] ; where, represents the velocity vector of the particle at time t, represents the position vector of the particle at time t, is the inertia weight, controlling the influence of the velocity at the previous moment, is the acceleration coefficient of the individual optimal position, controlling the degree to which the particle moves towards its own historical optimal position, is the acceleration coefficient of the global optimal position, controlling the degree to which the particle moves towards the global optimal position of the population, is the uncertainty correction coefficient, is a random number within the interval [0, 1], is the position of the particle with the lowest uncertainty in the current population. represents the position of the particle with the lowest uncertainty index μ(X) in the current population.

[0027] Perform constraint processing on the updated particle positions to meet the set constraint conditions (1) to (4);

[0028] Repeat the iteration until the preset number of iterations G or the convergence condition, and output the global optimal solution as the final task assignment matrix X.

[0029] Specifically, traditional PSO adopts a dual-source guidance mechanism (individual optimal and global optimal), while this solution introduces a third guidance source (the position with the lowest uncertainty), realizing a fundamental transformation from "performance-driven" to "performance-stability dual-driven". This three-source guidance mechanism significantly expands the search ability of the algorithm, enabling it to optimize simultaneously in two dimensions of the objective function value and stability.

[0030] Compared with the prior art, the advantages of this application are as follows:

[0031] In a multi-agent system under complex cognitive tasks, the uncertainty of the reasoning results of each agent will accumulate and amplify during the serial collaboration process. Traditional methods such as fixed rule assignment, simple polling mechanism, or static priority queue adopt fixed task assignment strategies, which have defects such as lack of environmental adaptability, inability to perceive reasoning uncertainty, unbalanced resource utilization, and easy significant decline in reasoning quality in a dynamic environment;

[0032] In this application, by constructing a multi-agent coordinated resource allocation model and an accurate task-resource requirement mapping, the task allocation problem is transformed into a multi-objective optimization problem that takes into account both time efficiency and resource balance. An uncertainty-aware improved particle swarm optimization algorithm is introduced. The stability of the solution is evaluated through multiple samplings with time decay. When the uncertainty exceeds the threshold, a multi-agent cross-validation mechanism is triggered, and an update strategy that comprehensively considers the objective function value and uncertainty is adopted, realizing the active perception and intervention of uncertainty in a dynamic environment, and improving the inference quality and efficiency. Brief Description of the Drawings

[0033] This application will be further described in the form of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:

[0034] Figure 1 is an exemplary flowchart of a multi-agent collaboration method based on a large language model shown in some embodiments of this application;

[0035] Figure 2 is an exemplary flowchart of constructing a task allocation problem shown in some embodiments of this application;

[0036] Figure 3 is an exemplary flowchart of a method for obtaining the global optimal solution shown in some embodiments of this application. Detailed Description of the Specific Embodiments

[0037] The methods and systems provided in the embodiments of this application will be described in detail below with reference to the drawings.

[0038] As Figure 1 shown, a multi-agent system is constructed. The multi-agent architecture includes a master agent, a task creation agent, an intent orchestration agent, and at least one professional agent. A multi-agent coordinated resource allocation model is established to manage the computing resources and task processing time of the multi-agent system. The resource allocation model includes a resource supply matrix, a task demand vector, and resource allocation constraints. The task allocation problem in the resource allocation model is transformed into a multi-objective optimization problem, and the optimization objectives include minimizing the task completion time and optimizing the resource utilization rate. An improved particle swarm optimization algorithm (PSO) is used to solve the multi-objective optimization problem to obtain the optimal collaboration plan among the master agent, the task creation agent, the intent orchestration agent, and at least one professional agent. The improved particle swarm optimization algorithm (PSO) evaluates the uncertainty of each candidate solution through multiple random samplings. When the uncertainty exceeds the threshold, cross-validation is performed through a multi-agent voting mechanism.

[0039] In this embodiment, a multi-agent system for enterprise financing application risk assessment is constructed. Based on large language model technology, the system includes four types of core agents: a master agent, a task creation agent, an intent orchestration agent, and multiple professional agents.

[0040] The master agent uses a large language model with a parameter scale of 175B and has global decision-making and coordination capabilities. It is responsible for the overall control of the system, the integration and output of the final risk assessment results. Its main responsibilities include: receiving the enterprise financing application and confirming the start of the assessment task; generating a final risk assessment report when all assessment steps are completed. Monitoring the system resource usage in real time, including the computing resource occupancy rate of each agent, the task queue depth, and the processing delay. When detecting abnormal situations (such as too high uncertainty in a certain assessment result), triggering a multi-agent voting mechanism for cross-verification. Conducting a quality assessment of the output results of each agent to ensure the accuracy and consistency of the risk assessment process. The master agent adopts a global state representation method based on the attention mechanism, encoding the state information and output results of each agent into a unified vector representation to achieve the perception of the system's global state. When processing a large number of concurrent tasks, the master agent uses a hierarchical scheduling strategy to allocate applications with different risk levels and complexities to the corresponding processing queues.

[0041] The task creation agent uses a large language model with a parameter scale of 70B and focuses on task parsing and decomposition. Its main responsibilities include: parsing the financing application materials submitted by the enterprise (including documents such as financial statements, business plans, and market analyses), and extracting key information. Decomposing the risk assessment task into multiple subtasks, such as financial condition analysis, industry risk assessment, credit history review, fraud detection, and compliance review. Assigning priority weights to each subtask according to factors such as the financing amount, enterprise scale, and industry type. Standardizing and structuring the extracted data to generate a data format suitable for each professional agent to process. In this embodiment, the task creation agent adopts multi-modal information processing capabilities and can process application materials in text, table, and image forms simultaneously.

[0042] The intention orchestration agent adopts a large language model with a parameter scale of 35B and is responsible for planning the task execution process. Its main responsibilities include: designing the optimal task execution process according to the task type and dependencies, determining which tasks can be executed in parallel and which tasks need to be processed serially. Dynamically adjusting the subsequent execution plan based on the intermediate results during the execution process. Providing resource allocation suggestions to the main control agent according to the task complexity and urgency. Monitoring the task execution progress to ensure that the tasks are completed as planned. The intention orchestration agent uses a directed acyclic graph (DAG) to represent task dependencies and calculates the optimal execution path through a dynamic programming algorithm. In this embodiment, the standard process of financing risk assessment includes 5 main stages and 18 sub-tasks, forming a complex task network.

[0043] The system includes multiple professional agents, and each agent is responsible for risk assessment tasks in a specific field: Financial analysis agent: Adopts a professional model in the financial field with a parameter scale of 13B, focusing on enterprise financial data analysis, including solvency, profitability, operating ability, and development ability assessment. Industry risk assessment agent: Adopts an industry analysis model with a parameter scale of 20B, focusing on evaluating industry prospects, market competition situation, and macroeconomic impact. Credit history agent: Adopts a model with a parameter scale of 7B, responsible for analyzing the enterprise's historical credit records, repayment behaviors, and credit events. Fraud detection agent: Adopts a model with a parameter scale of 15B, focusing on identifying potential financial fraud behaviors and abnormal patterns. Compliance review agent: Adopts a model with a parameter scale of 10B, responsible for evaluating the compliance risks of enterprise business activities and financing applications. Each professional agent has been specifically fine-tuned with data in the relevant field and is equipped with specific reasoning plugins to enhance its professional capabilities. For example, the financial analysis agent integrates a financial ratio calculation tool and an industry standard comparison database; the fraud detection agent is equipped with an anomaly detection algorithm and a historical fraud case database.

[0044] As Figure 2 shown, in the intelligent financial risk assessment system, resource allocation is crucial and directly affects the system processing efficiency and assessment quality. In this embodiment, there are a total of 8 agents in the system (1 main control agent, 1 task creation agent, 1 intention orchestration agent, and 5 professional agents), and the resource types include three categories: computing resources (GPU computing power), memory resources, and bandwidth resources. Therefore, the resource supply matrix R is an 8×3 matrix, with units of GPU cores, GB, and Mbps respectively:

[0045] R = [[32, 128, 1000], / / Main control agent;

[0046] [16, 64, 800], / / Task creation agent;

[0047] [8, 32, 500], / / Intention orchestration agent;

[0048] [16, 48, 300], / / Financial analysis agent;

[0049] [24, 56, 300], / / Industry risk assessment agent;

[0050] [4, 16, 200], / / Credit history agent;

[0051] [12, 32, 250], / / Fraud detection agent;

[0052] [8, 24, 200] / / Compliance review agent].

[0053] For example, R[0, 0] = 32 indicates that the master agent provides computing resources of 32 GPU cores, and R[3, 1] = 48 indicates that the financial analysis agent provides 48 GB of memory resources. The system is deployed on the private cloud platform of a financial institution, with a total configuration of 128 GPU cores, 400 GB of memory, and 3550 Mbps of bandwidth, which are allocated according to the functional complexity and processing requirements of each agent. Resource allocation also takes into account the typical loads of various tasks in historical data. For example, industry risk assessment usually requires processing a large amount of external market data, so higher computing resources are allocated.

[0054] In the scenario of enterprise financing risk assessment, a typical assessment task consists of 18 subtasks, and each subtask has different resource requirements. The task requirement vector D is an 18×3 matrix, which also corresponds to computing resources, memory resources, and bandwidth resources:

[0055] D = [[4, 8, 100], / / Enterprise basic information extraction;

[0056] [6, 12, 150], / / Financial statement data extraction;

[0057] [8, 16, 200], / / Financial ratio calculation and analysis;

[0058] [10, 24, 120], / / Cash flow status assessment;

[0059] [12, 32, 150], / / Debt repayment ability analysis;

[0060] [14, 28, 160], / / Profitability analysis;

[0061] [8, 16, 180], / / Industry development trend analysis;

[0062] [16, 36, 220], / / Market competition situation assessment;

[0063] [12, 24, 180], / / Macroeconomic impact analysis;

[0064] [3, 8, 120], / / Credit history record retrieval;

[0065] [5, 12, 100], / / Historical repayment behavior analysis;

[0066] [4, 10, 80], / / Credit event impact assessment;

[0067] [10, 20, 150], / / Financial data consistency check;

[0068] [12, 28, 120], / / Abnormal transaction pattern recognition;

[0069] [8, 16, 100], / / Financial indicator anomaly detection;

[0070] [6, 12, 80], / / Industry compliance requirement review;

[0071] [8, 20, 100], / / Anti-money laundering compliance assessment;

[0072] [4, 8, 60] / / Internal control risk assessment].

[0073] For example, D[2, 0] = 8 means that the "Financial ratio calculation and analysis" task requires 8 GPU cores of computing resources, and D[7, 1] = 36 means that the "Market competition situation assessment" task requires 36 GB of memory resources. The task requirements are determined by analyzing historical execution data. The system records the resource consumption of 2,500 financing applications processed in the past six months, and establishes a mapping relationship between task complexity and resource requirements through regression analysis. For computing resource requirements, the computing volume and parallel processing capabilities required for model inference are considered; for memory requirements, the model size and storage requirements for intermediate data processing results are considered; for bandwidth requirements, the requirements for external data access and internal communication are considered.

[0074] For the special requirements of the financial risk assessment scenario, the following constraint conditions are set: Execution time constraint: Different assessment time limits are set according to the financing amount. For financing applications below 10 million yuan, the system needs to complete the assessment within 2 hours; applications from 10 million yuan to 50 million yuan need to be completed within 4 hours; applications above 50 million yuan need to be completed within 8 hours. Priority constraint: Applications in high-risk industries (such as industries with large fluctuations like real estate and new energy) are processed first. A table of industry risk coefficients is set up in the system, which includes the risk levels of 42 industries. Quality assurance constraint: For large-scale financing applications (above 50 million yuan), it is required that at least 3 professional agents cross-verify the same key indicators to ensure the reliability of the assessment results. Resource utilization efficiency constraint: The average resource utilization rate of any agent should not be lower than 60% or higher than 90% to avoid resource waste and system bottlenecks.

[0075] The task allocation matrix X is an 18×8 0-1 matrix, representing the allocation relationship of 18 subtasks to 8 agents. For example, X[2, 3]=1 means that the 3rd task "Calculation and analysis of financial ratios" is assigned to the 4th agent (financial analysis agent). During the resource allocation process, the system needs to determine the optimal task allocation matrix X to minimize the overall assessment time and balance the resource utilization rates of each agent as much as possible. In addition, the task allocation also needs to meet the professional matching requirements. For example, financial tasks should be preferentially assigned to financial analysis agents to make full use of their professional capabilities.

[0076] Integrate the above elements to construct a multi-agent coordinated resource allocation model . The core of this model is to determine the optimal task allocation matrix X to achieve the optimal efficiency and quality of the financing risk assessment while satisfying the constraint conditions C.

[0077] As Figure 3 shown, in the intelligent financial risk assessment system, the task allocation problem is transformed into a multi-objective optimization problem considering the assessment time and resource utilization rate. The task completion time objective function T(X) is defined as the total completion time after all tasks are allocated according to the allocation matrix X, and the critical path method is used for calculation: , where, represents the completion time of the kth task, which is determined by the following factors: Basic processing time: Through historical data analysis, a basic processing time matrix B of each task on different agents is established, where, represents the benchmark time (unit: seconds) for the kth task to be processed by the ith agent. For example, the financial ratio analysis task takes 120 seconds to be processed by the financial analysis agent, but may take 180 seconds or longer to be processed by other non-professional agents. Load impact factor: The current load of the agent will affect the task processing time. Define the load impact function When the agent load L exceeds 60%, the processing time begins to increase significantly. Task complexity adjustment: Adjust the basic processing time according to the complexity of the financing application (considering enterprise scale, industry type, financing amount, etc.). Define the complexity coefficient with a range of [0.8, 1.5]. Task dependencies: There are dependencies between many tasks. For example, "financial data extraction" must be completed before "financial ratio analysis". The system uses a directed acyclic graph to represent task dependencies and considers the completion time of previous tasks when calculating the completion time.

[0078] Taking the above factors into account, the formula for calculating the completion time of the k-th task is: where represents the set of completion times of all previous tasks of the k-th task. During implementation, the system dynamically adjusts the basic processing time and complexity coefficient according to the characteristics of the financing application. For example, for large enterprises with complex group structures, the complexity coefficient of financial analysis tasks may reach 1.4; while for small enterprises with a high degree of standardization, the complexity coefficient may only be 0.9.

[0079] Resource utilization objective function represents the balance degree of system resource utilization and is defined as: where represents the resource load rate of the i-th agent, and σ represents the standard deviation function. The smaller the standard deviation, the more balanced the load of each agent, and the higher the value of the resource utilization objective function.

[0080] Resource load rate of agent i The calculation considers the comprehensive usage of three types of resources (computing resources, memory resources, and bandwidth resources): where: (computing resource load rate); (memory resource load rate); (bandwidth resource load rate); In the financial risk assessment scenario, uneven load distribution of agents may cause delays in processing some high-value assessment tasks, affecting the overall risk control effect. By optimizing the resource utilization balance, the system can avoid the situation where some agents are overloaded while others are idle, and improve the overall processing capacity.

[0081] The comprehensive objective function F(X) combines two objectives: task completion time and resource utilization: where and represent the weight coefficients of task completion time and resource utilization respectively. In this embodiment, the setting of the weight coefficient is related to the urgency of the financing application: for applications marked as "urgent", ​ Give priority to the completion time; for regular applications, balance time and resources; for batch-processed applications, pay more attention to resource utilization efficiency.

[0082] Express the task allocation problem as a multi-objective optimization problem: Constraints:

[0083] (1) , (each task allocation variable is 0 or 1);

[0084] (2) ; (each task must be assigned to one and only one agent);

[0085] (3) ; (total agent resource constraint);

[0086] (4) ; (task completion time constraint).

[0087] To solve the above multi-objective optimization problem, in the intelligent financial risk assessment system, initialize a particle swarm containing N = 50 particles, and each particle represents a candidate task allocation matrix X (a 0-1 matrix of 18×8). The initialization process considers professional matching degree to make the initial particles reasonable: 60% of the particles are initialized based on the professional matching rule, that is, they tend to assign tasks to agents with high professional matching degree. For example, financial analysis tasks are preferentially assigned to financial analysis agents. 30% of the particles are initialized completely randomly to maintain population diversity. 10% of the particles are initialized based on the historical optimal solution, using the system's past successful task allocation schemes. The system maintains a database containing 500 historical high-quality solutions, indexed by financing type, enterprise scale, and industry classification. For each initialized particle, the system performs a constraint repair process to ensure that all particles meet the basic constraint conditions, especially constraints (1) and (2). The repair process adopts a greedy strategy, giving priority to resource utilization efficiency and professional matching degree. The number of iterations is set to G = 200, but if the improvement of the objective function value does not exceed 0.1% for 30 consecutive iterations, the iteration is terminated early.

[0088] In the financial risk assessment scenario, there are various uncertainties in the task execution environment, such as data quality fluctuations, external system response time changes, model inference performance fluctuations, etc. To evaluate the stability of each candidate solution in the uncertain environment, the system generates m = 20 environmental parameter perturbations for each particle X:

[0089] Apply random perturbations to the basic processing time matrix B, with the perturbation range being [-15%, +25%], to reflect the time fluctuations in actual execution. The time fluctuations in financial data processing mainly stem from changes in data quality and complexity. For example, the processing time of non-standard financial statements may increase by 25% compared to standard statements.

[0090] Apply random perturbations to the task demand vector D, with the perturbation range being [-10%, +20%], to reflect the uncertainty of task resource requirements. For example, when the number of external data sources to be processed increases, the memory consumption may exceed expectations.

[0091] Simulate the situation where an agent is temporarily unavailable with a probability of 5%. In this case, its tasks need to be reallocated to other agents. This reflects the possible system maintenance, model updates, or hardware failures that may occur during actual operation.

[0092] Under each perturbation condition, recalculate the objective function value . The calculation of the perturbed objective function takes into account the additional costs and delays of task reallocation. Introduce a time decay factor λ = 0.9 to calculate the uncertainty index : ; among them, the newer sampling results (the larger the serial number) have a greater impact on the uncertainty assessment, reflecting the sensitivity of the system to the recent state. This time decay mechanism is particularly suitable for the financial risk assessment scenario because recent market fluctuations and policy changes usually have a greater impact on risk assessment.

[0093] Set the uncertainty threshold τ = 0.15. When the uncertainty index of the particle , trigger the multi-agent voting mechanism. In the financial risk assessment system, high uncertainty may mean that in the case of large market fluctuations or complex corporate financial conditions, a simple task allocation scheme may lead to unstable risk assessment results. When a candidate solution X with high uncertainty is detected, the master agent broadcasts an evaluation request to 5 professional agents, including the detailed task allocation scheme and historical execution statistics of the candidate solution X.

[0094] Each professional agent evaluates the candidate solution from its own professional perspective and calculates the local score : ; where: is the priority of the k-th task, with the range [1, 10]. The higher the financing amount and the higher the risk level of the task, the higher the priority; is the processing ability coefficient of the i-th professional agent for the k-th task, with the range [0.6, 1.2], reflecting the professional matching degree; is the execution efficiency of the i-th professional agent for the k-th task, with the range [0.7, 1.3], reflecting the historical execution speed. For example, the processing ability coefficient of the financial analysis agent for financial tasks The processing ability coefficient for fraud detection tasks reflects the expertise differences among professional agents.

[0095] Set the credibility weight of professional agents: ; among which, represents the number of times the professional agent's historical evaluation is correct, represents the total number of evaluations participated. The system records the difference between each evaluation result and the actual execution result. When the difference is less than 10%, it is determined as a correct evaluation. In the initial stage, the credibility weights of each agent are similar; as the system runs and historical data accumulates, more accurate agents obtain higher weights. For example, after running for three months, the historical accuracy rate of the financial analysis agent is 92% (460 / 500), and its weight W = 0.92; while the historical accuracy rate of the compliance review agent is 78% (312 / 400), and its weight W = 0.78.

[0096] Adopt a weighted voting mechanism to integrate the evaluation results of each professional agent: ; according to the comprehensive score correct the objective function value: ; among which, is the correction coefficient, S(X) ∈ [0, 1]. When , it indicates that the professional agent believes that this solution is better than the average level, and the corrected objective function value becomes smaller (for a minimization problem, this means an improvement in the solution quality); when , it indicates that the professional agent believes that this solution is not good, and the corrected objective function value becomes larger.

[0097] Based on the corrected objective function value , update the individual best position and the global best position of each particle: For each particle, calculate the criterion : ; among which, β = 0.3 is the uncertainty sensitivity coefficient. This criterion takes into account both the objective function value and uncertainty. When the objective function values of two solutions are close, it tends to select the solution with lower uncertainty.

[0098] When , update the individual best position . After each iteration ends, select the position with the minimum objective function value from all the individual best positions of the particles as the new global best position . Set the global best position update counter , when the global best position has not changed for Ks = 20 consecutive iterations, perform local random perturbation on : Randomly select For the elements with r = 10% in the matrix, their values are inverted. This perturbation mechanism helps the algorithm to jump out of local optima, which is particularly important in complex scenarios such as financial risk assessment.

[0099] Update the velocity and position of the particle:

[0100] ; where: ω = 0.7 is the inertia weight; is the acceleration coefficient for the individual optimal position; is the acceleration coefficient for the global optimal position; is the uncertainty correction coefficient; is a random number in the interval [0, 1]; is the position of the particle with the lowest uncertainty in the current population. The innovation of this velocity update formula lies in the introduction of the third uncertainty correction term, which makes the particle tend to move to the area with lower uncertainty during the update process. For applications such as financial risk assessment that require stability, this improvement is particularly important.

[0101] After the particle position is updated, the elements of the X matrix are often continuous values, which need to be converted into 0-1 integers and ensure that each task is only assigned to one agent:

[0102] The system adopts an improved S-shaped mapping function to map the continuous value into a binary value: ; when random(0, 1) < probability, , in other cases . This probabilistic mapping maintains the randomness in the search process and helps the algorithm to jump out of local optima. For example, when , the probability of conversion to 1 is about 0.82, rather than being deterministically set to 1.

[0103] Handling of constraints (1) and (2): The integrity and uniqueness of task assignment. After binarization, the system checks the assignment of each task, and three situations may occur: Multiple assignments: Task k is assigned to multiple agents ; The system calculates the comprehensive score according to the professional matching degree matrix E and the current load situation: . Retain the assignment with the highest score and set the others to 0. For example, the task "Financial Ratio Calculation and Analysis" is assigned to both the Financial Analysis Agent (matching degree 0.95, load 0.75) and the Intent Orchestration Agent (matching degree 0.45, load 0.55). Calculate the scores: Financial Analysis Agent: 0.7 * 0.95 + 0.3 * (1 - 0.75) = 0.665 + 0.075 = 0.74; Intent Orchestration Agent: 0.7 * 0.45 + 0.3 * (1 - 0.55) = 0.315 + 0.135 = 0.45; The system retains the assignment of the Financial Analysis Agent (X[2, 3] = 1) and sets the assignment of the Intent Orchestration Agent to 0 (X[2, 2] = 0).

[0104] No assignment: Task k is not assigned to any agent ( ), and the system also selects the best agent based on the comprehensive score: ; . For example, the task "Credit History Retrieval" is not assigned. The system calculates the comprehensive scores of each agent and selects the Credit History Agent with the highest score (X[9, 5] = 1).

[0105] Correct assignment: Task k is exactly assigned to one agent ( ), no adjustment is needed, and the current assignment is maintained.

[0106] Processing of Constraint (3): Total resource constraint. After processing Constraints (1) and (2), the system checks whether the resource usage of each agent exceeds its supply:

[0107] Calculate the total demand of each agent i for various resources j: ; Check whether there is resource overrun: . For agents with resource overrun, the system adopts a greedy adjustment algorithm. For example, the computing resources of the Fraud Detection Agent are overrun (Demand[6, 0] = 14 > R[6, 0] = 12). The system attempts to migrate its lowest-priority task "Financial Indicator Anomaly Detection" (Priority = 5.8) to the Financial Analysis Agent (E[14, 3] = 0.65 > E_min

[14] = 0.6) to make the resource usage meet the constraint. Resource adjustment may trigger a chain reaction. The migration of tasks from one agent may cause resource overrun in another agent. The system adopts an iterative adjustment strategy until the resource usage of all agents meets the constraint or reaches the maximum number of iterations (usually set to 5 times).

[0108] Processing of Constraint (4): Execution time constraint. After processing the resource constraint, the system checks whether the task completion time meets the deadline requirement:

[0109] Calculate the completion time of each task according to the task dependency graph and the current assignment Check whether there is a task whose completion time exceeds the deadline: For time-out tasks, high-priority tasks can suspend the execution of low-priority tasks. For example, when the "abnormal transaction pattern recognition" task is expected to time out, the system raises its priority from 7 to 9, enabling it to preempt other tasks that the agent is processing, etc.

[0110] The present application and its implementation manners are schematically described above. The description is not restrictive. Without departing from the spirit or basic features of the present application, the present application can be implemented in other specific forms. What is shown in the drawings is only one of the implementation manners of the present application, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and, without departing from the purpose of this creation, design similar structural manners and embodiments to this technical solution without creative efforts, they shall fall within the protection scope of the present application. In addition, the term "including" does not exclude other elements or steps, and the term "a" before an element does not exclude including "a plurality of" such elements. The terms such as first and second are used to indicate names and do not represent any specific order.

Claims

1. A multi-agent collaboration method based on large language models, characterized in that, Including: Construct a multi-agent system, where the multi-agent architecture includes a master agent, a task creation agent, an intention orchestration agent, and at least one professional agent; Establish a multi-agent coordinated resource allocation model for managing the computing resources and task processing time of the multi-agent system; the resource allocation model includes a resource supply matrix, a task demand vector, and resource allocation constraints; Convert the task allocation problem in the resource allocation model into a multi-objective optimization problem, where the optimization objectives include minimizing the task completion time and achieving optimal resource utilization; Use the improved Particle Swarm Optimization (PSO) algorithm to solve the multi-objective optimization problem and obtain the optimal collaboration plan among the master agent, the task creation agent, the intention orchestration agent, and at least one professional agent; the improved PSO algorithm evaluates the uncertainty of each candidate solution through multiple random samplings. When the uncertainty exceeds the threshold, cross-validation is performed through a multi-agent voting mechanism.

2. The multi-agent collaboration method based on a large language model according to claim 1, wherein: The master agent is responsible for overall coordination and decision-making; The task creation agent is responsible for task decomposition; The intention orchestration agent is responsible for execution process planning; The professional agent is responsible for handling tasks in the target domain.

3. The multi-agent collaboration method based on a large language model according to claim 2, wherein: Establish a multi-agent coordinated resource allocation model, including: Construct a resource supply matrix R, which represents the types and quantities of computing resources provided by each agent. Among them, represents the supply quantity of the j-th type of resource by the i-th agent; Establish a task demand vector D, representing the demand for computing resources and processing time requirements of each task, where D[k, j] represents the demand of the kth task for the jth type of resource; Set resource allocation constraints C, including execution time constraints; Construct a task assignment matrix X, where, represents that the k-th task is assigned to the i-th agent, represents non-assignment; Construct a multi-agent coordinated resource allocation model , where the task allocation problem is defined as: determining the optimal task allocation matrix X under the condition of satisfying the constraint condition C.

4. The multi-agent collaboration method based on a large language model according to claim 3, wherein: Convert the task allocation problem in the resource allocation model into a multi-objective optimization problem, including: Define the task completion time objective function \(T(X)\), which represents the total completion time after all tasks are assigned according to the task assignment matrix \(X\). , where represents the completion time of the \(k\)-th task, which depends on the processing capacity and current load of the assigned agent. Define the resource utilization objective function , representing the system resource utilization balance degree after allocation according to the task allocation matrix X, , where represents the resource load rate of the i-th agent, represents the standard deviation function, and the smaller the standard deviation, the more balanced the resource utilization; Construct the objective function , where and represent the weight coefficients of the task completion time and resource utilization rate, respectively; Set the total resource constraint, which means that the total demand of all tasks assigned to each agent for various resources shall not exceed the supply of the corresponding resources of the corresponding agent; Set an execution time constraint. For task k assigned to agent i, its completion time shall not exceed the preset task deadline , that is ; Express the task assignment problem as a multi-objective optimization problem: ; Constraint conditions: (1) ; (2) ; (3) ; (4) ; where X is a 0-1 integer matrix representing the allocation relationship from tasks to agents.

5. The multi-agent collaboration method based on a large language model according to claim 4, wherein: Use the improved Particle Swarm Optimization (PSO) algorithm to solve the multi-objective optimization problem and obtain the optimal collaboration plan, including: Initialize the particle swarm, where each particle represents a candidate task allocation matrix X, the number of particles is N, and the number of iterations is G; Calculate the objective function value F(X) of each particle; Evaluate the uncertainty index of each particle through multiple random samplings ; Set the uncertainty threshold , when the uncertainty index of the particle , trigger the multi-agent voting mechanism. The master agent organizes multiple professional agents to evaluate the corresponding solution and correct the objective function value according to the voting results; Update the individual best position of each particle according to the corrected objective function value and the global best position ; Update the velocity and position of the particles; Perform constraint processing on the updated particle positions to meet the set constraints (1) to (4); Iterate repeatedly until the preset number of iterations G or the convergence condition is reached, and output the global optimal solution As the final task assignment matrix X.

6. The multi-agent collaboration method based on a large language model according to claim 5, wherein: Evaluate the uncertainty index of each particle through multiple random samplings , including: Generate m perturbations of environmental parameters for each particle X, and the environmental parameter perturbations include random changes in task processing time, resource consumption, and agent availability; Under each perturbation condition, calculate the objective function value ; Introduce a time decay factor , to increase the impact of the most recent sampling results on the uncertainty assessment and calculate the uncertainty index : , where \(0 < \lambda < 1\) and \(\sigma\) is the standard deviation function.

7. The multi-agent collaboration method based on a large language model according to claim 5, wherein: When the uncertainty index of the particle is reached, a multi-agent voting mechanism is triggered, including: When the uncertainty index of the particle is greater than the preset threshold τ, the master agent broadcasts an evaluation request to at least one professional agent; Each professional agent evaluates the candidate solution X and calculates the local score , and the calculation formula is: , where represents the priority of the k-th task, represents the processing ability coefficient of the i-th professional agent for the k-th task, represents the execution efficiency of the i-th professional agent for the k-th task; Set the credibility weights of each professional agent : , where represents the number of times the professional agent's historical evaluation is correct, represents the total number of evaluations participated; The evaluation results of each professional agent are integrated by using a weighted voting mechanism to obtain a comprehensive score ; The objective function value is corrected according to the comprehensive score S(X), and the calculation formula is: , where α is the correction coefficient, and its value range is [0, 2], and S(X) ∈ [0, 1].​ 8. The multi-agent collaboration method based on a large language model according to claim 5, wherein: Update the individual optimal position of each particle according to the corrected objective function value and the global optimal position , including: For each particle, the corrected objective function value at the current position X is compared with the objective function value at the corresponding individual historical best position ; According to the corrected objective function value and the uncertainty index, calculate the criterion ; When update the individual optimal position otherwise, keep it unchanged; After each iteration, select the position with the minimum objective function value from the individual best positions of all particles as the new global best position ; Set the global optimal position update counter , when the global optimal position has not changed for consecutive iterations, perform local random perturbation on to avoid falling into local optimum: randomly select r% of the elements in the matrix and invert their values, that is, for the selected element (k, i), execute .

9. The multi-agent collaboration method based on a large language model according to claim 8, wherein: According to the corrected objective function value and the uncertainty index, calculate the criterion , using the following formula: , where β is the uncertainty sensitivity coefficient, and its value range is [0, 1], and are the uncertainty indicators of the current position and the individual historical optimal position respectively.

10. The multi-agent collaboration method based on a large language model according to claim 5, wherein: The velocity and position of the particle are updated through the following formula: ; ; Among them, represents the velocity vector of the particle at time t, represents the position vector of the particle at time t, ω is the inertia weight, controlling the influence of the velocity at the previous moment, is the acceleration coefficient of the individual optimal position, controlling the degree to which the particle moves towards its own historical optimal position, is the acceleration coefficient of the global optimal position, controlling the degree to which the particle moves towards the global optimal position of the population, is the uncertainty correction coefficient, is a random number within the interval [0, 1], is the position of the particle with the lowest uncertainty in the current population.

Citation Information

Patent Citations

  • Multi-agent cooperative task allocation method of improved particle swarm algorithm based on non-dominated sorting

    CN115809547A

  • Unmanned cluster cooperation strategy reconstruction method and device based on two-layer scheduling

    CN115857558A

  • Digital project management method and system based on artificial intelligence

    CN119151259A

  • Unmanned aerial vehicle-based river hydrological sampling inspection method and system

    CN119151387A

  • AU2020103709A4

Cited By

  • Structured data calculation result optimization method and system

    CN120724010A

  • Large model material formula design and evaluation method and system based on multi-agent collaboration and storage medium

    CN120878009A

  • Multi-agent collaboration-based large model material formula design and evaluation method and system and storage medium

    CN120878009B

  • Enterprise financing compliance auditing method and device and storage medium

    CN121052714A

  • Enterprise financing capability assessment method and device, and storage medium

    CN121073258A