High-concurrency task reasoning method and system driven by multi-agent non-cooperative game

Through the high-concurrency task inference method driven by multi-agent non-cooperative game, the decision-making delay problem of centralized scheduling architecture in high concurrency scenarios is solved, efficient and fair distribution of resources and stable operation of the system are achieved, and task execution efficiency and response speed are improved.

CN120256073BActive Publication Date: 2025-09-05PANDA ELECTRONICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510744894.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-05
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Centralized scheduling architecture is easily a performance bottleneck in high concurrent task scenarios, resulting in high decision-making delays and difficult to meet real-time requirements, especially in dynamic environments, task changes and resource competition intensify, affecting system efficiency and stability.

Method used

The high-concurrent task inference method driven by multi-agent non-cooperative game is adopted. By decomposing heterogeneous tasks into standardized task units, building a task auction mechanism for the non-cooperative game model based on dynamic priority adjustment and resource demand prediction models, a non-cooperative game model is constructed, resource allocation is combined with sharded asynchronous protocols, and malicious behavior is suppressed through the credit pledge mechanism to achieve efficient and fair distribution of resources and stable system operation.

Benefits of technology

It significantly improves resource utilization efficiency and system response speed, ensures priority execution of high-value tasks, reduces communication overhead, improves the stability and flexibility of the system in a complex competitive environment, and is suitable for industrial scenarios with large fluctuations in resource demand.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256073B_ABST
    Figure CN120256073B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-concurrency task reasoning method and system driven by multi-agent non-cooperative game. The method includes decomposing heterogeneous tasks into standardized task units and quantitatively modeling the computing resource requirements of the tasks; constructing a task bidding mechanism for the multi-agent non-cooperative game model, adopting anti-strategy bidding rules and sharding asynchronous protocol to allocate tasks, and constraining the resource declaration behavior of the agents through a credit pledge mechanism; monitoring the real-time resource utilization of the system, dynamically adjusting the task priority based on the historical status of task preemption, triggering a resource soft preemption strategy and establishing a compensation queue; adopting an adaptive model splitting strategy to allocate reasoning tasks to terminal devices and edge computing nodes, achieving collaborative execution through data compression transmission, and rescheduling tasks based on feedback from execution results; the present invention can significantly improve task execution efficiency, resource utilization and system stability in high-concurrency scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of industrial automation technology, and in particular relates to a high-concurrency task reasoning method and system driven by multi-agent non-cooperative game. Background Art

[0002] In recent years, the rapid development of the Industrial Internet and intelligent manufacturing has driven profound changes in production models. Highly concurrent tasks, such as multi-robot collaborative operations in smart factories, large-scale equipment monitoring, and real-time data analysis, have become increasingly commonplace in the industrial sector. These scenarios place higher demands on system collaboration efficiency, resource utilization, and real-time responsiveness, driving the widespread application of technologies such as distributed computing and edge intelligence.

[0003] Currently, a typical solution employs a centralized scheduling architecture, where a central controller manages the task allocation and resource scheduling of a multi-agent system. In this approach, the central controller distributes tasks to edge nodes or local execution units using static priorities or fixed allocation strategies based on a global task queue and resource status information. Edge nodes primarily handle data collection and preliminary processing, while complex decision-making and optimization tasks are handled by the central server.

[0004] However, the most significant drawback of a centralized scheduling architecture is high decision latency. As the number of tasks and data size increase, the central controller can easily become a performance bottleneck, significantly increasing task response times and making it difficult to meet the real-time requirements of high-concurrency scenarios. Especially in dynamic environments, frequent task changes and resource competition further exacerbate scheduling delays, impacting the overall efficiency and stability of the system. Summary of the Invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a multi-agent non-cooperative game-driven high-concurrency task reasoning method that can improve decision-making efficiency, reduce response time, and enhance system stability; on the other hand, to provide a multi-agent non-cooperative game-driven high-concurrency task reasoning system.

[0006] Technical solution: The high-concurrency task reasoning method described in the present invention includes the following steps:

[0007] (1) Heterogeneous tasks are decomposed into standardized task units. Based on the dynamic priority adjustment mechanism, resource demand prediction model and virtual resource unit (VRU) normalization method, the computing resource requirements of the tasks are quantitatively modeled. This effectively solves the resource allocation problem caused by task heterogeneity in industrial scenarios, realizes the unified quantitative expression of diversified computing requirements, provides an accurate quantitative basis for subsequent resource scheduling, and significantly improves the accuracy and comparability of resource description.

[0008] (2) Based on standardized task units and quantitative modeling results, a task bidding mechanism for a multi-agent non-cooperative game model is constructed. Anti-strategy bidding rules and a sharded asynchronous protocol are used for task allocation. The resource declaration behavior of agents is constrained by a credit pledge mechanism, which effectively suppresses false reporting of resources and malicious bidding behaviors. Combined with the sharded asynchronous protocol, the communication overhead in high-concurrency scenarios is reduced, achieving efficient and fair allocation of resources and ensuring the stable operation of the system in a complex competitive environment.

[0009] (3) Based on the task allocation results, the system monitors the real-time resource utilization rate, dynamically adjusts the task priority based on the task preemption history, triggers the resource soft preemption strategy and establishes a compensation queue. This not only ensures the priority execution of high-value tasks, but also maintains the fairness of the system through the compensation queue mechanism, significantly improves the resource utilization efficiency, and effectively responds to the resource needs of sudden high-priority tasks;

[0010] (4) Based on task priority, an adaptive model splitting strategy is adopted to allocate inference tasks to terminal devices and edge computing nodes, and collaborative execution is achieved through data compression transmission. Tasks are rescheduled based on feedback from execution results, ensuring the reliability of task execution and building an efficient and collaborative edge computing architecture, which greatly improves the system response speed and execution quality.

[0011] Preferably, the decomposition of heterogeneous tasks into standardized task units in step 1 includes:

[0012] According to the task type identifier T type and input data features D input , through the predefined operator resource mapping function Generate the computing power requirement R of the subtask compute , where T type Including convolution calculation, matrix operation and stream processing, D input Including data scale, dimensionality and batch size;

[0013] Decompose hierarchical tasks into serialized subtasks based on the computational graph topology. The computing power requirement of each subtask is determined by the type of the corresponding computation operator and the characteristics of the input data.

[0014] Decompose parallel tasks into a set of stateless subtasks based on data partitioning rules and hardware parallelism constraints, ensuring that the total resource requirements of the subtasks do not exceed the upper limit of the node computing power capacity.

[0015] Through the standardized definition of task type identifiers and input data characteristics, combined with predefined operator resource mapping functions, accurate resource demand quantification for diverse industrial tasks such as convolution calculations, matrix operations, and stream processing is achieved; the dual-track processing mechanism of computational graph topology decomposition and data sharding rules is adopted to ensure the orderly execution of hierarchical tasks and achieve efficient resource allocation for parallel tasks. Through strict control of the total amount of sub-task resource requirements, the risk of node computing power overload is effectively avoided, providing a precise and reliable task decomposition basis for subsequent resource scheduling.

[0016] Preferably, the dynamic priority adjustment mechanism described in step 1 is implemented in the following manner:

[0017] The initial priority S0 is defined by the user service level agreement SLA, and the decay rate λ is negatively correlated with the system real-time load;

[0018] When the task waiting time t delay Reaching delay tolerance D max When the preset ratio threshold is reached, the priority is smoothly enhanced through the Sigmoid function, and the enhancement amplitude is proportional to the computing power requirement R compute Positive correlation;

[0019] The calculation formula of the priority S(t) is:

[0020] Where, is the weight coefficient of the delay-sensitive compensation mechanism. The design of the compensation function P is related to the task computing power requirement R compute and delay tolerance D max Strong correlation;

[0021] The calculation formula of the compensation function P is:

[0022] .

[0023] Adaptive adjustment of task priority is achieved through the initial priority defined by the service level agreement (SLA) and the dynamic decay rate negatively correlated with the system load; combined with the delay-sensitive compensation mechanism based on the Sigmoid function, the priority is intelligently increased when the task approaches the maximum tolerable delay, and the compensation amplitude is positively correlated with the task computing power requirement, effectively ensuring the timely scheduling of tasks with high computing requirements. At the same time, fine control of priority adjustment is achieved through the weight coefficient, significantly improving the system's processing capabilities for time-sensitive tasks and the rationality of resource allocation.

[0024] Preferably, the resource demand forecasting in step 1 includes:

[0025] Based on computing power requirement R computeAs the benchmark value, combined with the dynamic correction value ΔR output by the time series prediction network, the final computing power requirement is generated ;

[0026] The input of the time series prediction network includes the historical resource utilization sequence {R compute} and node real-time status code E node , where E node Including GPU utilization, memory usage, and network bandwidth fluctuations.

[0027] By combining static computing power benchmark values ​​and dynamic timing corrections, accurate predictions of task resource requirements are achieved: the timing prediction network based on historical resource utilization sequences and real-time node status (including GPU utilization, memory occupancy, and network bandwidth fluctuations) can dynamically correct the initial computing power assessment value, effectively capturing the short-term fluctuation characteristics of resource demand, significantly improving the accuracy of resource demand predictions, providing a reliable quantitative basis for subsequent resource allocation decisions, and solving the problem that traditional static prediction methods are difficult to adapt to dynamic load changes.

[0028] Preferably, the normalization processing of the virtual resource unit VRU in step 1 includes:

[0029] The multi-dimensional resource demand is mapped into a single virtual resource unit through the normalized weight coefficient ω, where ω is dynamically calculated by the global resource competition intensity;

[0030] The global resource competition intensity is defined as the ratio of total resource demand to available supply;

[0031] The node dynamic capacity Calculated as total capacity minus allocated capacity.

[0032] Multi-dimensional heterogeneous resources (such as computing power, memory, and bandwidth) are uniformly quantified into a single comparable indicator through the dynamic weight coefficient ω, where ω is dynamically adjusted based on the real-time resource supply and demand ratio (competition intensity), so that the resource bottleneck dimension automatically obtains a higher weight. Combined with the real-time calculation of the node's dynamic capacity (total capacity minus the allocated amount), it realizes the intelligent unified measurement and dynamic balance of multi-dimensional resources, effectively solves the problem of incomparability of heterogeneous resources, provides a fair and accurate resource quantification basis for multi-agent games, and significantly improves the rationality of resource allocation in complex industrial scenarios.

[0033] Preferably, step 2 includes:

[0034] Receive bidding requests B submitted by each agent i , each bidding request B i Contains virtual resource requirements (VRUs) i , dynamic priority S(t) i and delay tolerance D maxi;

[0035] Press S(t) i / VRU i The winning task set A is selected in descending order of the ratio to ensure that a near-optimal solution is obtained in polynomial time:

[0036] ; Where T is the set of all bidding tasks, and the result A is the set of winning tasks;

[0037] Calculate the actual payment price p for the winning task i i :

[0038] Where, is the set of winning bids when task i is not involved; is the set of winning bids when task i participates;

[0039] Introducing a credit pledge mechanism, the legal bid of the intelligent entity must meet Credit i Dynamically correlated with historical task completion rates, resource utilization rates, and timeout times;

[0040] Divide the global resource pool into multiple sub-markets T1, T2, ..., T K , each sub-market runs the auction process independently, resource capacity constraints satisfy , the winning bid set A in each sub-market k Merge into global allocation result A through distributed consensus protocol:

[0041] .

[0042] By receiving bidding requests containing virtual resource requirements, dynamic priorities and delay tolerance, S(t) is adopted i / VRU i A near-optimal solution selection strategy based on ratio sorting achieves efficient task allocation in polynomial time; a real-bid incentive mechanism is established by calculating the actual payment price of the winning task (based on the marginal contribution difference), and a credit pledge mechanism dynamically linked to task completion performance is combined to effectively suppress malicious bidding behavior; resource pool sharding and distributed consistency protocols are used to decompose the global auction into multiple parallel sub-market competitions, significantly reducing the communication and computing overhead in high-concurrency scenarios while ensuring the fairness of resource allocation, achieving a dual improvement in the efficiency and stability of resource allocation in complex industrial environments.

[0043] Preferably, step 3 includes:

[0044] Based on the deviation ratio between the actual resource occupancy ARO and the virtual resource requirement VRU of each task in the winning task set A, the adjusted priority P(t) is dynamically calculated:

[0045] Where, N preempt is the number of times the task has been preempted in history, λ is the global control factor, and the resource efficiency deviation term Measures the rationality of task resource usage. The greater the deviation, the more significant the priority decay. Preemption penalty Malicious competition can be suppressed by accumulating the number of preemptions;

[0046] When a high priority task T is detected high ∈A cannot be started due to insufficient resources, and there is a low-priority task T low ∈A satisfies When , the resource preemption operation is triggered, and the threshold reuse λ dynamically adjusts the priority difference:

[0047] And the system load ≥80%

[0048] Reclaim low-priority tasks proportionally The amount of resources recovered is determined by VRU low Together with λ, the preempted task is allowed to continue running in degraded mode:

[0049] Tasks that will be preempted Join the compensation queue, its compensation priority is dynamically improved through the time-dependent exponential compensation function, and the compensation intensity is determined by λ and VRU low Automatic adjustment:

[0050] Where, t wait The waiting time of the task in the compensation queue

[0051] Set a cooldown period mechanism to limit the maximum number of preemptions of the same task within a unit of time. The cooldown period expression is:

[0052] ; In the formula, the λ value can be dynamically adjusted through system strategy.

[0053] By real-time monitoring of the deviation ratio between the actual resource usage of tasks and the reported demand, combined with the historical preemption times, the task priority (P(t)) is intelligently adjusted, which not only punishes unreasonable resource usage but also suppresses malicious competition; when the system load exceeds the threshold and there is a significant priority difference, a soft preemption strategy based on proportional recovery is triggered, so that the preempted task continues to run in a degraded mode, while its subsequent execution rights are guaranteed through the compensation queue and exponential compensation function; combined with a dynamically adjustable cooling period mechanism, it effectively prevents system shocks caused by frequent preemption, achieves a balance between stability and flexibility of resource scheduling under high load, and significantly improves the response timeliness of high-value tasks in complex industrial scenarios.

[0054] Preferably, step 4 includes:

[0055] The model split point k is dynamically selected based on the computing power of the end-side device and the real-time load of the edge node, where the end-side performs the first k layers of inference to generate intermediate feature data F k ;

[0056] For the intermediate feature data F k Compression coding is performed, and the compression rate η is based on the bandwidth requirement R bw and the maximum tolerable delay of the task Dynamic adjustment to meet transmission time :

[0057]

[0058] Calculate the confidence C(Y) of the inference result through the built-in confidence estimation module of the model;

[0059] When C(Y) is lower than the confidence threshold C thr When , a re-reasoning request is triggered and the task priority P(t) is increased to , where the confidence threshold Cthr is dynamically set according to the task reliability requirement, and γ is the priority gain coefficient, which is dynamically configured by the system to control the impact of confidence on task priority;

[0060] Submit the re-inference task to the task auction mechanism in step 2, and its virtual resource requirement is set to VRU retry =VRU⋅(1+ω), where ω is the redundancy coefficient, and the priority inherits the updated P new (t), ensure that high-priority tasks quickly obtain resources;

[0061] When C(Y) is higher than the confidence threshold C thr When the edge node receives the intermediate feature data F compressed with compression rate η k The obtained compression features The remaining layers are then inferred to generate the final result Y and return it to the client.

[0062] Through the dynamic model splitting strategy (adaptive selection of the splitting point k based on the device computing power and edge load) and intelligent compression transmission (the compression rate η is optimized in real time according to the bandwidth demand and latency requirements), efficient collaborative execution of computing tasks between terminals and edge nodes is achieved; combined with the confidence feedback mechanism, a re-inference process with priority enhancement is automatically triggered for low-confidence results (the redundancy coefficient ω is increased according to resource requirements), and at the same time, the dynamically adjusted priority gain coefficient γ is used to ensure that critical tasks quickly obtain resources. This significantly reduces end-to-end latency while ensuring inference accuracy, providing a highly reliable, low-latency distributed intelligent computing solution for industrial Internet scenarios.

[0063] The high-concurrency task reasoning system of the present invention includes:

[0064] High-performance central control module, including a high-performance computing server cluster, for centralized scheduling of global tasks and complex decision-making calculations;

[0065] A strong real-time edge computing module, including an edge server with a built-in AI accelerator, for real-time collection and pre-processing of local data;

[0066] The business agent module includes market agent, procurement agent, warehouse agent, process agent, planning agent, production agent and emergency agent, which are used to complete specific order business;

[0067] The communication module, including 5G base stations and industrial switches, is used to provide high-speed, low-latency network connections between devices.

[0068] Global intelligent scheduling is achieved through a high-performance central control module, and combined with the localized processing capabilities of a strong real-time edge computing module, a "center-edge" collaborative computing architecture is constructed; a cluster of specialized business intelligence entities realizes closed-loop management of the entire order process, and combined with the high-speed and low-latency transmission characteristics of the 5G industrial communication network, a complete solution covering task scheduling, resource allocation, production execution and exception response is formed, which significantly improves the task processing efficiency, system response speed and business continuity guarantee capabilities in complex industrial scenarios.

[0069] A computer device, characterized in that it includes a memory and a processor, wherein the memory stores a computer program that can be loaded and executed by the processor for the multi-agent non-cooperative game-driven high-concurrency task reasoning method.

[0070] A computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the multi-agent non-cooperative game-driven high-concurrency task reasoning method is implemented.

[0071] Beneficial effects: Compared with the existing technology, the present invention has the following significant advantages: 1. It can ensure the robustness of the system, significantly improve the task execution efficiency, resource utilization and system stability in high concurrency scenarios, and provide complex industrial scenarios with intelligent solutions that are both real-time and reliable; 2. Based on the normalized processing of virtual resource units and the sharded asynchronous auction protocol, the system can perceive the resource status in real time and dynamically adjust the allocation strategy, greatly shortening the resource allocation response time, which is particularly suitable for industrial scenarios with large fluctuations in resource demand, ensuring that critical tasks obtain the required resources first; 3. Through adaptive model splitting and dynamic compression transmission technology, the computing power of terminal devices is organically combined with edge computing resources, significantly reducing transmission delays. At the same time, the reliability of task execution is guaranteed through the rescheduling mechanism of confidence feedback, effectively solving the bottleneck problem of insufficient edge computing capabilities of traditional systems; 4. The closed-loop design from task modeling, auction allocation, dynamic preemption to collaborative execution realizes intelligent control of the entire production process, greatly improving the abnormal response speed and production scheduling flexibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 Schematic diagram of the method flow of the present invention;

[0073] Figure 2 Schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0074] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0075] The embodiment of the present invention provides a high-concurrency task reasoning method driven by a multi-agent non-cooperative game, which is implemented based on a high-concurrency task reasoning system driven by a multi-agent non-cooperative game. Figure 1 As shown, the following steps are included:

[0076] Step 1: Multi-agent task modeling and resource requirement quantification

[0077] In a heterogeneous computing environment, the scheduling of multi-agent tasks needs to be based on a unified resource quantification model. This step uses a structured modeling method to abstract diverse tasks (including real-time reasoning, stream processing, batch computing and other tasks in the embodiment of the present invention) into decomposable and measurable standardized units to solve the resource allocation problem caused by task heterogeneity. The core of task modeling lies in decomposition and quantification, which is divided into the following three steps: First, task decomposition is performed, and the task is broken down into multiple subtasks according to the task type and input data characteristics. The resource requirements of each subtask are generated through predefined rules, and it is ensured that the total demand does not exceed the computing power limit of the node; secondly, dynamic priority adjustment is performed, and the priority of the task will change over time; finally, resource requirements are predicted and allocated, and the multi-dimensional resource requirements are mapped into a single virtual resource unit (VRU) based on the normalization method, providing computable input parameters for non-cooperative games. The following are the specific steps:

[0078] First, in the task decomposition phase, define the task type identifier T type (In the embodiment of the present invention, convolution calculation, matrix operation, and stream processing are included) and input data feature D input (In the embodiment of the present invention, including data scale, dimension, batch size), through the predefined operator resource mapping function Generate the computing power requirements for each subtask. Hierarchical tasks are decomposed into serialized subtasks according to the computational graph topology. The computing power requirements of each subtask are determined by the type of corresponding operator and the characteristics of the input data. Parallel tasks are decomposed into a set of stateless subtasks based on data sharding rules and hardware parallelism constraints, ensuring that the total resource requirements of the subtasks do not exceed the node computing power capacity limit.

[0079]

[0080] Memory requirements in task characteristics R mem and bandwidth requirements R bw , which is strongly related to the inherent characteristics of the task, for different task type identifiers T type Takes a specific value.

[0081] Secondly, the dynamic priority S(t) is calculated jointly by the time decay function and the delay sensitivity compensation mechanism: the initial priority S0 is defined by the user service level agreement (SLA), the decay rate λ is negatively correlated with the real-time load of the system; the task waiting time t delay The compensation function P is triggered. When a task approaches the maximum tolerable delay, the priority is increased through a nonlinear enhancement mechanism to prevent high-delay-sensitive tasks from timing out due to resource competition failure.

[0082] Where, is the weight coefficient of the delay-sensitive compensation mechanism. The design of the compensation function P is related to the task computing power requirement Rcompute and delay tolerance D max Strong correlation:

[0083]

[0084] When the task waiting time t delay When the preset ratio threshold is exceeded (in the embodiment of the present invention, D max 80% of the total), the priority is smoothly enhanced through the Sigmoid function, and the enhancement amplitude is related to the computing power requirement R compute Positive correlation ensures that high-computing tasks receive higher scheduling weights in emergency situations.

[0085] The resource demand prediction model is based on computing power demand R compute As the benchmark value, combined with the dynamic correction value ΔR output by the time series prediction network, the final computing power requirement R is generated. final The input of the prediction network is the historical resource utilization sequence {R compute} and node real-time status code E node (In the embodiment of the present invention, this includes GPU utilization, memory occupancy, and network bandwidth fluctuations). Time series modeling is used to predict short-term fluctuations in computing power demand. The dynamic correction term ΔR is used to compensate for the deviation between the static baseline value and the actual load, thereby improving the accuracy of resource demand prediction.

[0086]

[0087] Finally, through virtual resource units Normalize heterogeneous resource requirements to solve the problem of incomparability of multi-dimensional resources (computing power, memory, bandwidth). The normalization weight coefficient ω is dynamically calculated based on the global resource competition intensity (competition intensity is defined as the ratio of total resource demand to available supply), so that the resource bottleneck dimension receives a higher weight in the allocation. The current available resource amount is the total capacity minus the allocated amount, ensuring that resource allocation decisions are based on the real-time available resource status.

[0088]

[0089] Step 2: Non-cooperative auction mechanism

[0090] In order to achieve efficient resource allocation in high-concurrency scenarios, a non-cooperative bidding mechanism is designed in this step. It suppresses the false reporting of resource requirements or malicious bidding behaviors of intelligent agents through anti-strategy rules, while optimizing computational complexity and communication overhead. The bidding mechanism is based on the standardized resource requirements generated in step 1 (virtual resource units VRU, dynamic priority S(t), delay tolerance D max) as input, a multi-agent game model is constructed, resource allocation decisions are driven by the objective function of maximizing system benefits, and the credit pledge and shard auction mechanism are used to improve the robustness of the system.

[0091] First, determine the winning task and collect the bidding requests B submitted by each agent. i Contains virtual resource requirements (VRUs) i , dynamic priority S(t) i and delay tolerance D maxi .VRU i is the virtual resource requirement of task i, which comes from the VRU output of step 1. S(t) i represents the dynamic priority of task i, which comes from the S(t) output of step 1. The resource allocation problem is transformed into a system revenue optimization model. The objective function is to maximize the sum of the dynamic priorities of the set of winning tasks. The constraint is that the total resource consumption does not exceed the available capacity of the node Θ total Press S(t) i / VRU i Tasks are selected in descending order of their ratios, ensuring a near-optimal solution in polynomial time. In the following formula, T is the set of all bidding tasks. The result A is the set of winning tasks.

[0092]

[0093] Secondly, calculate the actual payment price. After the winning task set A is determined, calculate the actual payment price p of each winning bidder. i , whose value is the difference between the optimal value of the system before and after task i participates in the auction. Specifically, the payment price is equal to the sum of the system benefits of other tasks when task i is not included minus the sum of the system benefits after including task i but excluding its own contribution. This rule forces the agent to quote truthfully, because false reporting will result in a payment price higher than the true value, thereby reducing individual utility. In the following formula, is the set of winning bids when task i is not involved; is the set of winning bids when task i participates (excluding i itself).

[0094]

[0095] Finally, a legal bid is obtained. To prevent Sybil attacks (malicious agents forging multiple identities to split tasks) or collusion to manipulate auction results, a virtual credit pledge mechanism is introduced. The legality of the agent's bid depends on its credit score. iWhether the pledge conditions are met, credit points are dynamically bound to historical task completion rates, resource utilization rates, and timeouts. Malicious behavior (such as frequent timeouts or abnormal resource usage) will result in credit points being deducted, and if the credit points are below the threshold, participation in subsequent auctions will be prohibited. The credit pledge rule controls the credit pledge rate through the global parameter ρ, ensuring that tasks with high resource requirements require higher credit points to suppress attacks and ultimately obtain legal bids. .

[0096]

[0097] At the same time, in order to reduce the communication and computing overhead in high concurrency scenarios, a sharded asynchronous auction protocol is designed. The global resource pool is divided into multiple sub-markets T1, T2, ..., T according to resource type (such as GPU intensive, memory intensive). K , each sub-market runs the auction process independently, and the resource capacity constraint is (satisfy ). Tasks only participate in the auction of sub-markets that match their resource requirements, reducing cross-node communication and global state synchronization overhead. The winning bid set A of each sub-market k After merging, a global allocation result A is formed, and the global constraints of resource allocation are guaranteed to be satisfied through a distributed consistency protocol.

[0098]

[0099] Step 3: Dynamic priority adjustment and resource preemption

[0100] After the non-cooperative auction mechanism completes resource allocation, the system must dynamically adjust priorities and trigger resource preemption based on real-time load and task execution status to address sudden demand and resource competition conflicts for high-value tasks. This step uses the winning task set A, dynamic priority S(t), and virtual resource demand VRU from Step 2 as core inputs to construct a closed-loop feedback control model.

[0101] First, the dynamic priority P(t) is calculated. The calculation of P(t) is based on the dynamic priority S(t) of the winning bid in the auction phase, and is dynamically evolved in combination with the actual resource utilization of the task and the number of historical preemptions. Resource utilization is calculated by mapping the actual occupied physical resources to the VRU, reflecting the actual resource utilization efficiency of the task. When the actual resource occupancy of the task deviates from the VRU, the priority is attenuated in proportion to the deviation; and the priority of frequently preempted tasks is suppressed by the penalty factor to avoid inefficient repeated scheduling. The global control factor λ uniformly controls the resource efficiency deviation and the intensity of the preemption penalty:

[0102] ; Among them, ARO (actual resource occupancy) is obtained by monitoring the usage of physical resources, N preemptis the number of times the task has been preempted in history, and λ is the global control factor (default is 0.1).

[0103] Resource efficiency deviation item Measures the rationality of task resource usage. The greater the deviation, the more significant the priority decay. Preemption penalty Malicious competition can be suppressed by accumulating the number of preemptions.

[0104] Secondly, the preemption decision trigger is implemented. The triggering of the preemption decision depends on the real-time priority difference and the global load status of the system. The global load is calculated by accumulating the actual VRU occupancy of all tasks in the winning task set A, reflecting the total resource utilization of the node. high ∈A cannot be started due to insufficient resources, and there is a low-priority task T low ∈A satisfies When the system triggers the preemption operation. The threshold reuse λ dynamically adjusts the priority difference, and the system load threshold is fixed at 80% (the empirical value avoids overload):

[0105] And the system load ≥80%

[0106] The soft preemption strategy minimizes task interruption by reclaiming resources proportionally. The reclaimed amount is determined by VRU. low Together with λ, it allows the preempted task to continue running in a degraded mode, such as retaining some computing power to maintain basic functions. The preempted task enters the compensation queue, and its priority is dynamically increased through a time-dependent exponential compensation function to accelerate rescheduling. The compensation intensity is determined by λ and VRU. low Automatically adjust to ensure fairness. wait The waiting time of the task in the compensation queue:

[0107]

[0108] Finally, after preemption, a cool-down period begins. The cool-down period mechanism limits the maximum number of preemptions for the same task in a short period of time to prevent system shock. The length is negatively correlated with λ. Frequently preempted tasks automatically have their cooldown time extended, while the λ value can be dynamically adjusted through system policies (for example, increasing λ during high load to accelerate resource recovery):

[0109]

[0110] Step 4: Edge-end collaborative reasoning and result feedback

[0111] After dynamic resource preemption and priority adjustment, the system must split the task into device-side and edge-side subtasks through an edge-to-end collaborative reasoning mechanism to optimize computing efficiency and reduce latency. This step designs an adaptive model splitting strategy and a feedback-driven rescheduling mechanism to achieve efficient collaborative reasoning.

[0112] The task model dynamically selects the split point k based on the computing power of the end-side device (such as the computing power required by the end-side) and the real-time load of the edge node. The end-side performs the first k layers of inference to generate intermediate feature data F k , and then transmitted to the edge node after compression encoding.

[0113] The compression rate η is based on the bandwidth requirement R in step 1 bw Dynamic adjustment to ensure transmission time (D max Maximum tolerable delay for the task):

[0114]

[0115] At the same time, the confidence level of the result C(Y) is calculated by the built-in confidence estimation module of the model. If C(Y) <C thr (the confidence threshold is set by the task reliability requirement in step 1), triggering a cloud re-inference request and simultaneously increasing the task priority P(t) in step 3 to accelerate rescheduling. γ is the priority gain coefficient, which is dynamically configured by the system and is used to control the impact of confidence on task priority:

[0116]

[0117] The re-inference task is submitted to the auction mechanism in step 2, and its virtual resource requirement VRU retry =VRU⋅(1+ω) (ω is the redundancy coefficient), the priority inherits the updated P new (t), ensuring that high-priority tasks quickly obtain resources.

[0118] If C(Y)>C thr , the edge node receives the intermediate feature data F compressed with compression rate η k The obtained compression features The remaining layers are then inferred to generate the final result Y and return it to the client.

[0119] like Figure 2As shown, the high-concurrency task reasoning system corresponding to the high-concurrency task reasoning method described in the present invention includes a high-performance central control module, a strong real-time edge computing module, a business intelligence module, and a communication module. The high-performance central control module includes a high-performance computing server cluster. The strong real-time edge computing module includes an edge server. The business intelligence module is an intelligence cluster used to complete specific order business, including a market intelligence agent, a procurement intelligence agent, a warehouse intelligence agent, a process intelligence agent, a planning intelligence agent, a production intelligence agent, an emergency intelligence agent, etc.

[0120] For example, in the industrial production process, when a new order is generated, it is necessary to complete the order task based on the current actual situation of the company.

[0121] The high-performance central control module includes a high-performance computing server cluster with strong computing power support capabilities. It is responsible for the centralized scheduling of global tasks and complex decision-making calculations, and realizes efficient coordination and dynamic configuration of the entire system resources.

[0122] The strong real-time edge computing module, deployed at the business front end, consists of small, high-density edge servers. It features low-latency, high-concurrency data processing capabilities and supports real-time collection and preprocessing of local data. The module's built-in dedicated AI accelerator enables intelligent reasoning and preliminary analysis at the edge, effectively reducing reliance on the central control module and improving system responsiveness and operational reliability.

[0123] The business agent module is used to complete specific order operations and includes market agents, procurement agents, warehouse agents, process agents, planning agents, production agents, and emergency agents. The market agent dynamically analyzes market demand and the competitive environment to drive accurate forecasting, customer segmentation, and marketing strategy optimization. It also connects with customers to determine order details and issue order tasks. The procurement agent optimizes supplier collaboration, cost control, and procurement plans based on forecast data to ensure efficient material supply. The warehouse agent monitors inventory status in real time, enabling intelligent replenishment, warehouse scheduling, and logistics route optimization. The process agent analyzes product process parameters to optimize production process design and resource allocation. The planning agent coordinates order priorities, develops scientific production schedules, and dynamically adjusts production rhythm. The production agent performs production task monitoring, integrating equipment status and process control to ensure delivery quality and efficiency. The emergency agent responds to supply chain disruptions or abnormal events in real time, providing risk warnings and rapid remediation solutions. Each agent builds an end-to-end closed-loop business system to achieve optimal resource allocation and intelligent management and control of the entire order process.

[0124] The communication module provides high-speed, low-latency, and highly stable network connectivity between devices, building an end-to-end industrial-grade communication assurance system. This module includes core communication equipment such as 5G base stations and industrial switches, supporting massive terminal access and efficient data transmission, ensuring real-time and secure information exchange between terminals, edge devices, and the center.

[0125] The invention also discloses an electronic device.

[0126] Specifically, the electronic device may be a computer device such as a desktop computer, a laptop computer, a PDA, or a cloud server. The computer device may include, but is not limited to, a processor and a memory. The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, graphics processing units (GPU), embedded neural network processors (NPU) or other dedicated deep learning coprocessors, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0127] As a non-transient computer-readable storage medium, the memory can be used to store non-transient software programs, non-transient computer executable programs and modules. The processor executes various functional applications and data processing of the processor by running the non-transient software programs, instructions and modules stored in the memory. The memory may include a program storage area and a data storage area, wherein the program storage area may store a control unit, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.

[0128] The invention also discloses a computer-readable storage medium.

[0129] Specifically, a computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method in the above-mentioned method implementation is implemented. Those skilled in the art will understand that the implementation of all or part of the process in the above-mentioned method implementation of the present application can be completed by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it can include the process of the implementation of each method as described above. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (Flash Memory), a hard disk drive (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memory.

Claims

1. A multi-agent non-cooperative game-driven high-concurrency task reasoning method, characterized by: The following steps are involved: (1) Decompose heterogeneous tasks into standardized task units and quantitatively model the computing resource requirements of tasks based on a dynamic priority adjustment mechanism, a resource demand prediction model, and a virtual resource unit (VRU) normalization method; (2) Based on standardized task units and quantitative modeling results, a task bidding mechanism for a multi-agent non-cooperative game model is constructed. Anti-strategy bidding rules and sharding asynchronous protocols are used for task allocation, and a credit pledge mechanism is used to constrain the resource declaration behavior of agents. (3) Based on the task allocation results, monitor the real-time resource utilization of the system, dynamically adjust the task priority based on the task preemption history status, trigger the resource soft preemption strategy and establish a compensation queue; (4) Based on task priority, an adaptive model splitting strategy is used to allocate inference tasks to end-side devices and edge computing nodes, achieve collaborative execution through data compression transmission, and reschedule tasks based on feedback from execution results; Step 2 includes: Receive bidding requests B submitted by each agent i , each bidding request B i Contains virtual resource requirements (VRUs) i , dynamic priority S(t) i and delay tolerance D maxi ; Press S(t) i / VRU i The winning task set A is selected in descending order of the ratio to ensure that a near-optimal solution is obtained in polynomial time: st ; Where T is the set of all bidding tasks, and the result A is the set of winning tasks; Calculate the actual payment price p for the winning task i i : Where, is the set of winning bids when task i is not involved; is the set of winning bids when task i participates; Introducing a credit pledge mechanism, the legal bid of the intelligent entity must meet , ρ is the credit pledge threshold coefficient, Credit i Dynamically correlated with historical task completion rates, resource utilization rates, and timeout times; Divide the global resource pool into multiple sub-markets T1, T2, ..., T K , each sub-market runs the auction process independently, resource capacity constraints satisfy , the winning bid set A in each sub-market k Merge into global allocation result A through distributed consensus protocol: where s.t. 。 2. The high-concurrency task reasoning method according to claim 1, characterized in that: Decomposing heterogeneous tasks into standardized task units as described in step 1 includes: According to the task type identifier T type and input data features D input , through the predefined operator resource mapping function Generate the computing power requirement R of the subtask compute , where T type Including convolution calculation, matrix operation and stream processing, D input Including data scale, dimensionality and batch size; Decompose hierarchical tasks into serialized subtasks based on the computational graph topology. The computing power requirement of each subtask is determined by the type of the corresponding computation operator and the characteristics of the input data. Decompose parallel tasks into a set of stateless subtasks based on data partitioning rules and hardware parallelism constraints, ensuring that the total resource requirements of the subtasks do not exceed the upper limit of the node computing power capacity.

3. The high-concurrency task reasoning method according to claim 1, characterized in that: The dynamic priority adjustment mechanism described in step 1 is implemented in the following way: The initial priority S0 is defined by the user service level agreement SLA, and the decay rate λ is negatively correlated with the system real-time load; When the task waiting time t delay Reaching delay tolerance D max When the preset ratio threshold is reached, the priority is smoothly enhanced through the Sigmoid function, and the enhancement amplitude is proportional to the computing power requirement R compute Positive correlation; The calculation formula of the priority S(t) is: Where, is the weight coefficient of the delay-sensitive compensation mechanism. The design of the compensation function P is related to the task computing power requirement R compute and delay tolerance D max Strong correlation; The calculation formula of the compensation function P is: 。 4. The high-concurrency task reasoning method according to claim 1, characterized in that: The resource demand forecast described in step 1 includes: Based on computing power requirement R compute As the benchmark value, combined with the dynamic correction value ΔR output by the time series prediction network, the final computing power requirement is generated ; The input of the time series prediction network includes the historical resource utilization sequence {R compute } and node real-time status code E node , where E node Including GPU utilization, memory usage, and network bandwidth fluctuations.

5. The high-concurrency task reasoning method according to claim 1, characterized in that: The normalization process of the virtual resource unit VRU in step 1 includes: The multi-dimensional resource demand is mapped into a single virtual resource unit through the normalized weight coefficient ω, where ω is dynamically calculated by the global resource competition intensity; The global resource competition intensity is defined as the ratio of total resource demand to available supply; The node dynamic capacity Calculated as total capacity minus allocated capacity.

6. The high-concurrency task reasoning method according to claim 1, characterized in that: Step 3 includes: Based on the deviation ratio between the actual resource occupancy ARO and the virtual resource requirement VRU of each task in the winning task set A, the adjusted priority P(t) is dynamically calculated: Where, N preempt is the number of times the task has been preempted in history, λ is the global control factor, and the resource efficiency deviation term Measures the rationality of task resource usage. The greater the deviation, the more significant the priority decay. Preemption penalty Malicious competition can be suppressed by accumulating the number of preemptions; When a high priority task T is detected high ∈A cannot be started due to insufficient resources, and there is a low-priority task T low ∈A satisfies When , the resource preemption operation is triggered, and the threshold reuse λ dynamically adjusts the priority difference: And the system load is ≥80%; where T high It's a high-priority task. is the priority of the high-priority task; T low It's a low-priority task. is the priority of the low-priority task; Reclaim low-priority tasks proportionally The amount of resources recovered is determined by VRU low Decide together with λ whether to allow the preempted task to continue running in degraded mode; Tasks that will be preempted Join the compensation queue, its compensation priority is dynamically improved through the time-dependent exponential compensation function, and the compensation intensity is determined by λ and VRU low Automatic adjustment: Where, t wait The waiting time of the task in the compensation queue; Set a cooldown period mechanism to limit the maximum number of preemptions of the same task within a unit of time. The cooldown period expression is: ; Wherein, the λ value is dynamically adjusted through the system strategy.

7. The high-concurrency task reasoning method according to claim 1, characterized in that: Step 4 includes: The model split point k is dynamically selected based on the computing power of the end-side device and the real-time load of the edge node, where the end-side performs the first k layers of inference to generate intermediate feature data F k ; For the intermediate feature data F k Compression coding is performed, and the compression rate η is based on the bandwidth requirement R bw and the maximum tolerable delay of the task Dynamic adjustment to meet transmission time : ; Calculate the confidence C(Y) of the inference result through the built-in confidence estimation module of the model; When C(Y) is lower than the confidence threshold C thr When , a re-reasoning request is triggered and the task priority P(t) is increased to , where the confidence threshold Cthr is dynamically set according to the task reliability requirement, and γ is the priority gain coefficient, which is dynamically configured by the system to control the impact of confidence on task priority; Submit the re-inference task to the task auction mechanism in step 2, and its virtual resource requirement is set to VRU retry =VRU⋅(1+ω), where ω is the redundancy coefficient, and the priority inherits the updated P new (t), ensure that high-priority tasks quickly obtain resources; When C(Y) is higher than the confidence threshold C thr When the edge node receives the intermediate feature data F compressed with compression rate η k The obtained compression features The remaining layers are then inferred to generate the final result Y and return it to the client.

8. A multi-agent non-cooperative game driven high-concurrency task reasoning system that implements the multi-agent non-cooperative game driven high-concurrency task reasoning method of claim 1, characterized in that: include: High-performance central control module, including a high-performance computing server cluster, for centralized scheduling of global tasks and complex decision-making calculations; A strong real-time edge computing module, including an edge server with a built-in AI accelerator, for real-time collection and pre-processing of local data; The business agent module includes market agent, procurement agent, warehouse agent, process agent, planning agent, production agent and emergency agent, which are used to complete specific order business; The communication module, including 5G base stations and industrial switches, is used to provide high-speed, low-latency network connections between devices.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the multi-agent non-cooperative game-driven high-concurrency task reasoning method described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Large model reasoning scheduling method based on off-network computing power server

    CN119537032A

  • Multi-robot chemical synthesis real-time scheduling control method based on MAS

    CN119811518A