An AI computing power optimization method based on multi-agent collaboration

By using a multi-agent collaborative AI computing power optimization method, the shortcomings of existing computing power scheduling systems in task perception and load prediction are solved, achieving efficient and stable allocation and utilization of resources, and adapting to complex high-concurrency environments.

CN122470382BActive Publication Date: 2026-08-25JIANGSU LUOYAO SMART COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610943587.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-08-25
Estimated Expiration
2046-06-29

AI Technical Summary

Technical Problem

Existing computing power scheduling systems have semantic blind spots in task perception, failing to understand the computation graph topology and memory access patterns of AI tasks. They lack elastic negotiation capabilities, leading to resource prediction errors and fragmentation. Furthermore, inaccurate load prediction can easily cause node overload and scheduling instability.

Method used

A multi-agent collaborative AI computing power optimization method is adopted. Through multi-modal feature fusion encoding, dynamic risk perception matching, and SLA elastic negotiation mechanism, a collaborative representation vector is generated to calculate the preference scores on the task side and the machine side. Combined with a two-sided stable matching algorithm and elastic constraints, the resource allocation is non-blocking and closed-loop correction is achieved.

Benefits of technology

It improves the semantic awareness and matching fairness of resource scheduling, reduces the risk of node overload, alleviates memory fragmentation, improves resource utilization efficiency and system robustness, and adapts to complex high-concurrency environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122470382B_ABST
    Figure CN122470382B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer resource scheduling, in particular to an AI computing power optimization method based on multi-agent cooperation, which comprises the following steps: multi-modal task features are extracted, and a cooperative representation vector is generated in combination with double prediction heads; based on the vector, node resource occupation and a risk penalty coefficient generated by a load prediction distribution entropy value, a bilateral preference score is calculated; a bilateral stable matching algorithm with a capacity constraint is run to output a distribution scheme; for unallocated tasks, a marginal acceptance threshold is used to guide the tasks to be re-matched after one-way relaxation of constraints in an elastic interval, and a bottom distribution is set; and the prediction model parameters are updated online based on actual load deviation. Through multi-modal semantic perception, dynamic risk matching and an elastic negotiation mechanism, the application reduces node overload and display memory fragmentation, alleviates the scheduling congestion problem in a resource shortage state, and further improves resource utilization and system operation robustness in a high-concurrency environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer resource scheduling technology, specifically to an AI computing power optimization method based on multi-agent collaboration. Background Technology

[0002] With the widespread application of large language models and deep neural networks, intelligent computing centers face the challenge of massive heterogeneous tasks and collaborative computing across multiple types of accelerators. Existing computing power scheduling relies heavily on heuristic rules or basic multi-agent reinforcement learning, allocating resources solely based on physical requirements such as memory capacity and computing power. This approach has significant shortcomings in practical applications. First, task awareness suffers from a "semantic blind spot." Existing systems ignore the unique computational graph topology, operator distribution, and drastically different memory access patterns of large models during the pre-filling and decoding phases of AI tasks. Relying solely on static physical resource vectors fails to understand the actual computational behavior of tasks, easily leading to resource estimation errors, memory fragmentation, and frequent secondary scheduling. Second, there is a lack of dynamic negotiation capabilities based on SLA (Service Level Agreement) elasticity. Real-world AI tasks exhibit vastly different sensitivities to quality: some tasks require high precision but tolerate queuing, while some real-time tasks prioritize low latency but accept quantization degradation. Existing systems cannot quantify the elasticity space of "precision-latency," employing rigid rejection strategies when resources are scarce. This fails to guide tasks to relax constraints to facilitate matching, resulting in long-tail task blocking and fragmented, idle computing power. Furthermore, load prediction and scheduling decisions are disconnected. Existing solutions' prediction modules only output deterministic point estimates, causing scheduling decisions to lose awareness of the risks associated with future load uncertainty (entropy), easily leading to short-sighted node overload; and lacking a closed-loop correction mechanism based on actual observed load, it is difficult to maintain scheduling robustness in dynamic environments over the long term.

[0003] In summary, there is an urgent need for a multi-agent collaborative scheduling system that can effectively align computational semantics and resource representation, while taking into account bidirectional stable matching, flexible compromise negotiation, and closed-loop correction mechanisms.

[0004] To address this, a method for optimizing AI computing power based on multi-agent collaboration is proposed. Summary of the Invention

[0005] The purpose of this invention is to provide an AI computing power optimization method based on multi-agent collaboration. By using multimodal feature fusion encoding, dynamic risk perception matching, and SLA elastic negotiation mechanism, this method improves the technical problems existing in computing power scheduling technology, such as the lack of task semantic perception, lack of elastic compromise mechanism, and uncontrollable load prediction risk.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for optimizing AI computing power based on multi-agent collaboration includes: receiving an AI task, obtaining the resource demand vector of the task, and setting an elastic range for the task; extracting the multimodal features of the task, encoding them into a semantic embedding vector by a semantic encoder, and calculating the urgency score and accuracy sensitivity score by the prediction head output, and fusing and concatenating them with the resource demand vector and the semantic embedding vector to form a collaborative representation vector.

[0007] The task agent generates task-side preference scores for each computing power node based on the collaborative representation vector; the computing power agent generates machine-side preference scores based on the node resource occupancy and the load prediction distribution output by the prediction model, and adjusts the risk penalty coefficient during calculation based on the load prediction distribution.

[0008] The arbitration agent constructs a two-sided preference matrix based on the task-side preference score and the machine-side preference score, runs a two-sided stable matching algorithm with capacity constraints, and outputs a stable allocation scheme for non-blocking pairs.

[0009] If no task is assigned in a round, the lowest machine-side preference score among the tasks accepted by the matched computing power node that is closest to the task's elasticity range will be received. After adjusting the constraints within the elasticity range to increase the task-side preference score, the node will be rematched. If no task is assigned after the preset maximum number of rounds, the remaining computing power will be allocated in descending order of computation urgency score.

[0010] After each round of allocation, the computing power agent updates the deviation compensation parameters of the prediction model based on the deviation between the actual load and the predicted distribution.

[0011] Preferably, the acquisition of the collaborative representation vector includes: the semantic encoder employing a pre-set shared network encoding layer; inputting extracted multimodal features, including topology, operator distribution, accuracy, batch size, and inference stage, into the shared network encoding layer to generate the semantic embedding vector; setting a first prediction head and a second prediction head in parallel in the shared network encoding layer, updating parameters using historical queuing time cost and historical model quality degradation data as supervision signals, and independently outputting the computational urgency score and the accuracy sensitivity score; acquiring a task resource requirement vector including memory requirements, computing power requirements, and bandwidth requirements; calculating a weighted fusion vector of the semantic embedding vector and the resource requirement vector through a cross-modal attention network, and concatenating the weighted fusion vector with the computational urgency score and the accuracy sensitivity score in a preset order to generate the collaborative representation vector.

[0012] Preferably, the generation of the task-side preference score and the machine-side preference score includes: the task agent mapping the collaborative representation vector to the task-side preference score sequence for each candidate computing power node using a multilayer perceptron; the computing power agent collecting the resource occupancy percentage of the managed computing power nodes in terms of memory, computing power, and bandwidth, calling a time-series network to predict the distribution matrix of future continuous time steps as the load prediction distribution, and calculating the distribution entropy value and the corresponding statistical expectation value of the distribution matrix; when at least one of the following two conditions is met, namely, the statistical expectation value exceeds a first preset occupancy rate threshold, and the distribution entropy value exceeds a preset warning value and the statistical expectation value exceeds a second preset occupancy rate threshold, a positive linear mapping function is established and a risk penalty coefficient is obtained based on the distribution entropy value; the resource occupancy percentage and the load prediction distribution are input into the machine evaluation logic to generate an initial score; the initial score is subjected to a numerical decay operation using the risk penalty coefficient to obtain the machine-side preference score, so that the machine-side preference score for the task with excessive resource consumption decreases monotonically with the increase of the distribution entropy value.

[0013] Preferably, the output of the stable allocation scheme for non-blocking pairs includes: summarizing the task-side preference scores and machine-side preference scores to construct the cross-mapping bilateral preference matrix; obtaining the current unused GPU memory capacity of each computing power node; each task agent initiating a matching request in descending order of task-side preference scores; after receiving the request, the computing power agent sorts the requests in descending order of machine-side preference scores, sequentially accumulating the GPU memory requirements of the corresponding requesting task according to the sorting order, retaining matching requests whose accumulated GPU memory requirements do not exceed the unused GPU memory capacity, and rejecting matching requests that trigger excessive amounts; cyclically executing the retransmission of requests from rejected tasks to other candidate nodes and the filtering by computing power agents until no new matching requests are generated, outputting the stable allocation scheme for non-blocking pairs, and marking the finally rejected tasks as unallocated tasks.

[0014] Preferably, the step of receiving the lowest machine-side preference score among the tasks accepted by the matched node closest to the task elasticity interval in this round includes: the elasticity interval includes a precision elasticity interval and a latency elasticity interval; extracting the precision elasticity boundary set and latency elasticity boundary set of the unassigned tasks; traversing all computing power nodes that successfully accepted task quotas in this round, extracting the associated precision elasticity boundary set and associated latency elasticity boundary set of the paired tasks of each computing power node in this round; calculating the weighted Euclidean distance in precision and latency dimensions between the associated boundaries of the unassigned tasks and each successfully paired computing power node, marking the node corresponding to the minimum distance value as the matched computing power node closest to the task elasticity interval; extracting the minimum machine-side preference score of all tasks accepted by the closest matched computing power node in this round, and feeding it back to the unassigned tasks as a marginal acceptance threshold.

[0015] Preferably, the allocation of remaining computing power includes: if the difference between the machine-side preference score generated by the unallocated task for the nearest matched computing power node and the received marginal acceptance threshold exceeds a preset tolerance limit, then the limit delay parameter in the delay elastic boundary set and the critical precision parameter in the precision elastic boundary set are extracted; within the elastic range, the constraint is relaxed unidirectionally by adding a first preset constant to the limit delay parameter and subtracting a second preset constant from the critical precision parameter; based on the relaxed critical precision parameter, the collaborative representation vector and task-side preference score are recalculated, and combined with the relaxed limit delay parameter, the matching is re-entered; when the matching cycle reaches the preset limit number of times and there are unallocated remaining tasks, the computation urgency score is read and arranged in descending order, and the remaining available resources of each global computing power node are independently allocated to the remaining tasks until the resources are cleared.

[0016] Preferably, updating the bias compensation parameters of the prediction model includes: at the deadline of the current allocation round time window, collecting actual resource consumption observation values ​​of the computing facilities to constitute the actual load; extracting the statistical expectation corresponding to the predicted distribution output of the previous allocation round, calculating the load residual matrix of the actual load and the statistical expectation; multiplying the load residual matrix into a preset exponential decay constant, updating the bias compensation parameters and converting them into compensation tensors of the same dimension, superimposing them into the bias weight matrix of the hidden layer of the prediction model, and driving the next round of inference results to generate equivalent correction towards the actual resource consumption observation values.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention improves upon traditional computing power scheduling methods that rely solely on explicit physical resource indicators by employing multimodal feature fusion encoding and a dual-prediction-head mechanism. By extracting features such as task topology, operator distribution, and inference stage, and combining them with a prediction model, it fuses computational semantics and physical requirements across modalities to generate a collaborative representation vector. This enables a numerical representation of AI task computational features and SLA (Service Level Agreement) elastic boundaries. This mechanism refines the feature-aware granularity of heterogeneous tasks, helping to reduce queuing and allocation mismatches caused by resource prediction errors.

[0018] 2. This invention employs a capacity-constrained, dual-sided stable matching architecture with integrated dynamic risk awareness, reducing node overload risk and mitigating memory fragmentation. By introducing the distribution entropy value of time-series prediction distribution and the risk penalty coefficient, computing nodes can reflect the future load fluctuation risk when calculating preference scores. Combined with a dual-sided stable matching algorithm that accumulates based on actual memory demand, resource allocation is performed while maintaining the non-blocking characteristics of the allocation scheme, improving the matching fairness of the system and the utilization efficiency of cluster computing resources.

[0019] 3. This invention introduces a task elasticity constraint relaxation and closed-loop prediction correction mechanism to alleviate scheduling blockage under resource scarcity. By calculating and feeding back the marginal acceptance threshold, specific adjustment references are provided for unassigned tasks, guiding them to appropriately relax accuracy or delay constraints within the elastic range to improve matching probability, and combining a global fallback strategy to utilize remaining computing power; at the same time, the bias weights of the prediction model are updated online based on the difference matrix between the actual load and the predicted expectation, providing the system with a lightweight dynamic correction approach, which helps the intelligent computing center operate in complex high-concurrency environments. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the overall solution of the present invention; Figure 2 This is a diagram of the collaborative representation vector acquisition architecture of the present invention; Figure 3 This is the machine-side preference score and risk penalty mapping diagram of the present invention; Figure 4 This is a diagram illustrating the elastic boundary calculation and constraint relaxation mechanism of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figures 1 to 4 This invention provides an AI computing power optimization method based on multi-agent collaboration. The technical solution is as follows: receiving an AI task, obtaining the resource demand vector of the task, and setting an elastic range for the task; extracting the multimodal features of the task, encoding them into a semantic embedding vector by a semantic encoder, and calculating the urgency score and accuracy sensitivity score by the prediction head output, and fusing and concatenating them with the resource demand vector and the semantic embedding vector to form a collaborative representation vector.

[0023] The task agent generates task-side preference scores for each computing power node based on the collaborative representation vector; the computing power agent generates machine-side preference scores based on the node resource occupancy and the load prediction distribution output by the prediction model, and adjusts the risk penalty coefficient during calculation based on the load prediction distribution.

[0024] The arbitration agent constructs a two-sided preference matrix based on the task-side preference score and the machine-side preference score, runs a two-sided stable matching algorithm with capacity constraints, and outputs a stable allocation scheme for non-blocking pairs.

[0025] If no task is assigned in a round, the lowest machine-side preference score among the tasks accepted by the matched computing power node that is closest to the task's elasticity range will be received. After adjusting the constraints within the elasticity range to increase the task-side preference score, the node will be rematched. If no task is assigned after the preset maximum number of rounds, the remaining computing power will be allocated in descending order of computation urgency score.

[0026] After each round of allocation, the computing power agent updates the deviation compensation parameters of the prediction model based on the deviation between the actual load and the predicted distribution.

[0027] Example 1: Further, the acquisition of the collaborative representation vector includes: the semantic encoder employing a pre-set shared network encoding layer; inputting extracted multimodal features, including topology, operator distribution, accuracy, batch size, and inference stage, into the shared network encoding layer to generate the semantic embedding vector; setting a first prediction head and a second prediction head in parallel in the shared network encoding layer, updating parameters using historical queuing time cost and historical model quality degradation data as supervision signals, and independently outputting the computational urgency score and the accuracy sensitivity score; acquiring a task resource requirement vector including memory requirements, computing power requirements, and bandwidth requirements; calculating a weighted fusion vector of the semantic embedding vector and the resource requirement vector through a cross-modal attention network, and concatenating the weighted fusion vector with the computational urgency score and the accuracy sensitivity score in a preset order to generate the collaborative representation vector.

[0028] Specifically, after receiving an AI task, the system extracts its computation graph topology, operator distribution frequency, initial precision declaration, batch size value, and inference stage identifier. Before concatenating these multimodal features, each modality feature needs to be preprocessed: the topology feature is mapped to a fixed-length 64-dimensional vector by graph average pooling; the operator distribution frequency histogram is uniformly normalized to the [0, 1] interval and retains 128 dimensions; the system distinguishes precision features into continuous numerical precision elastic boundaries and discrete categorical data precision types (such as FP32, FP16, or INT8), where the data precision type declaration and inference stage identifier are expanded into 8-dimensional and 4-dimensional vectors respectively through one-hot encoding; the batch size value is processed into a 1-dimensional scalar through logarithmic normalization. The above processed vectors are concatenated sequentially to form an original feature vector with a total dimension of 205. Subsequently, this original feature vector is input into a preset shared network encoding layer. The shared network coding layer consists of three fully connected layers with 256, 128, and 64 neurons in each layer, respectively. The ReLU activation function is used between layers for nonlinear feature extraction and dimensionality reduction. The final output is a dense vector of 64 dimensions, which serves as the semantic embedding vector for this task.

[0029] At the output of the shared network coding layer, a first prediction head and a second prediction head are connected in parallel. These two prediction heads process the semantic embedding vectors through independent fully connected branches and output computational urgency and precision sensitivity scores between 0 and 1 respectively via a sigmoid activation function. During offline training, to ensure the dual prediction heads can accurately quantify the true constraints of the task, the system updates parameters using historical real costs. Specifically, the system constructs a training dataset based on historical task scheduling logs from the intelligent computing center. Key feature fields extracted from the historical scheduling logs after anonymization include, but are not limited to: task ID, model category (e.g., large language model or visual model), input and output sequence lengths, inference stage identifiers (pre-filling stage and decoding stage), allocated GPU specifications, peak memory usage, measured bandwidth usage, computing power utilization, queuing time, SLA target setting, and measured data reflecting model quality (e.g., precision or perplexity). The training dataset is divided according to a time series (e.g., divided into training and validation sets in an 8:2 ratio) to address data distribution drift issues. Each sample contains the processed multimodal features, resource demand vector, and corresponding real waiting time cost and accuracy degradation data. Since the output range of the first and second prediction heads after the Sigmoid activation function is 0 to 1, the real supervision signal needs to be preprocessed: the historical queuing waiting time cost is normalized to the [0, 1] interval using a maximum-minimum normalization method, serving as the real label for the first prediction head. The maximum waiting time is taken as the 95th percentile of historical statistical data to filter out abnormal long-tail extrema; the historical model quality degradation data is calculated based on the measured inference quality indicators in the log (such as the percentage increase in perplexity or the percentage decrease in accuracy after quantization degradation). Specifically, the relative accuracy degradation rate (i.e., the absolute value of the ratio of the current downgraded accuracy to the initial declared accuracy minus 1) is smoothed and compressed using the Sigmoid function and used as the true label of the second prediction head. .

[0030] Specifically, the system constructs a training dataset based on the historical task scheduling logs of the intelligent computing center over the past three months. This dataset contains no fewer than 100,000 task samples, each with complete multimodal features, resource requirement vectors, and corresponding real waiting time cost and accuracy degradation data. Since the output range of the first and second prediction heads after the Sigmoid activation function is 0 to 1, the real supervision signal needs to be preprocessed: the historical queuing waiting time cost is normalized to the [0, 1] interval using a maximum-minimum normalization method, serving as the real label for the first prediction head. The maximum waiting time The 95th percentile of historical statistical data is used to filter out outliers with long tails; historical model quality degradation data, after being smoothed and compressed using the Sigmoid function with the relative accuracy decline rate (i.e., the absolute value of the ratio of current accuracy to initial accuracy minus 1), is used as the true label for the second prediction head. By constructing a joint loss function End-to-end synchronous supervised training is performed on the shared network coding layer and dual prediction heads. The joint loss function is defined as: ; In the formula, and These are the current prediction scores output by the first and second prediction heads, respectively, where MSE represents the mean squared error function. and In this embodiment, to adjust the balance coefficient of the weights of the two prediction tasks, and Both are set to 0.5 to ensure that the gradient contributions of the two prediction tasks are equal during backpropagation.

[0031] The task agent synchronously parses the basic runtime environment requirements declared when the task is submitted, extracts the required video memory capacity, peak computing power, and read / write bandwidth, and after normalization of the maximum and minimum values, constructs a one-dimensional task resource requirement vector containing video memory requirements, computing power requirements, and bandwidth requirements.

[0032] A cross-modal attention network is employed for deep alignment of high-dimensional computational semantics with low-dimensional physical resources. In this step, simply using vector concatenation can easily lead to feature dimensionality overload. Therefore, the previously generated 64-dimensional semantic embedding vector is used as the query matrix for the attention mechanism (i.e., dimension 1). For the obtained one-dimensional task resource requirement vector containing memory, computing power, and bandwidth requirements, the system independently inputs each of its three scalar components and performs feature upscaling operations through the corresponding linear mapping layer, generating corresponding 64-dimensional resource feature vectors for memory, computing power, and bandwidth respectively. Subsequently, these three resource feature vectors are stacked and concatenated along the sequence dimension, thereby transforming the one-dimensional resource requirement vector into a key matrix K and a value matrix V of sequence length 3. The final dimensions of both the key matrix K and the value matrix V are... Furthermore, its parameter weights are set before training using the Xavier initialization method. The weighted fusion vector is calculated using the following cross-modal attention formula: ; In the formula, Let K be the feature dimension of the key matrix K. In this embodiment, we take K as the feature dimension. In the denominator of the formula This is the scaling factor that actually takes effect. The Softmax operation performs numerical normalization along the resource dimension (i.e., the sequence length direction of the key, corresponding to the aforementioned memory, computing power, and bandwidth components with a sequence length of 3), obtaining the dynamic attention weight vector of the query features for the three major physical resources (dimension: Then, a weighted summation is performed with the value matrix V, and the final output weighted fusion vector has a dimension of 64 (i.e., strictly maintaining the dimension of V). Furthermore, the linear mapping weight matrix of the cross-modal attention network... The system fully participates in deep training. Since the weighted fusion vector, concatenated with the calculated urgency score and accuracy sensitivity score, serves as the state input to the task agent, its parameter gradient is not controlled by the joint loss function of the aforementioned dual prediction heads. Instead, it is updated synchronously online or offline via a multilayer perceptron through backpropagation, based on an allocation objective function constructed by the downstream task agent during offline training to improve the historical matching success rate. Through this step, the system can adaptively assign differentiated attention weights to different physical resource dimensions (such as GPU memory or bandwidth) based on the semantic features of the task.

[0033] Finally, the task agent concatenates the calculated weighted fusion vector, the urgency score, and the accuracy sensitivity score according to a preset order of contiguous physical memory addresses, ultimately generating a collaborative representation vector. This vector will serve as the sole state input for the task-initiating agent to subsequently map task-side preference scores.

[0034] This invention improves upon traditional schedulers that rely solely on explicit physical metrics by combining a shared network coding layer with a cross-modal attention mechanism, thus aligning physical resource constraints with the semantic features of high-level model computation. It introduces a dual-parallel prediction head trained under supervised training based on historical real costs, transforming task SLA boundaries into continuous urgency and sensitivity scores, providing a numerical basis for elastic scheduling. This representation generation mechanism refines the granularity of perception for heterogeneous tasks in multi-agent systems, helping to reduce memory allocation mismatches caused by resource prediction errors, thereby improving the scheduling rationality and resource utilization efficiency of intelligent computing centers in concurrent scenarios.

[0035] Further, the generation of the task-side preference score and the machine-side preference score includes: the task agent mapping the collaborative representation vector to the task-side preference score sequence for each candidate computing power node using a multilayer perceptron; the computing power agent collecting the resource occupancy percentage of the managed computing power nodes in terms of memory, computing power, and bandwidth, calling a time-series network to predict the distribution matrix of future continuous time steps as the load prediction distribution, and calculating the distribution entropy value and the corresponding statistical expectation value of the distribution matrix; when at least one of the following two conditions is met, namely, the statistical expectation value exceeds a first preset occupancy rate threshold, and the distribution entropy value exceeds a preset warning value and the statistical expectation value exceeds a second preset occupancy rate threshold, a positive linear mapping function is established and a risk penalty coefficient is obtained based on the distribution entropy value; the resource occupancy percentage and the load prediction distribution are input into the machine evaluation logic to generate an initial score; the initial score is subjected to a numerical decay operation using the risk penalty coefficient to obtain the machine-side preference score, so that the machine-side preference score for the task with excessive resource consumption decreases monotonically as the distribution entropy value increases.

[0036] Specifically, after acquiring the collaborative representation vector as described above, the task agent inputs it as the global state description of the current task into the built-in multilayer perceptron network. The last layer of this multilayer perceptron uses a max-min linear normalization layer (or a sigmoid activation function) to independently map the output values ​​of each dimension to the [0, 1] interval, directly outputting a fixed-length sequence of the full set of task-side preference scores. Specifically, to meet the static engineering constraints of the neural network output layer dimensions, the number of output neurons in the last layer of the multilayer perceptron is rigidly fixed during model construction and training, and its number is equal to the preset total number of nodes in the cluster. To maintain the stability of the model mapping boundary under dynamic node addition and subtraction environments when mapping the collaborative representation vector to the final score sequence, the system strictly adopts a globally unified node enumeration order (e.g., based on fixed physical identifiers of computing power nodes) to generate and align the output score sequence during both the offline end-to-end training phase and the online inference phase. For the candidate node set dynamically composed of candidate computing power nodes that are available in the current round and meet the basic physical resource constraints, the system uses this set to perform masking filtering on the entire sequence: for inactive or unavailable nodes, the output values ​​at their corresponding positions are masked or removed, thereby dynamically extracting the task-side preference score sequence for each candidate computing power node for the current task. Each element in this vector represents, in absolute terms, the matching fitness value of the current task for a specific candidate computing power node; this sequence constitutes the task-side preference score sequence for the current task.

[0037] To achieve end-to-end joint optimization of feature extraction and scheduling decisions, the multilayer perceptron and the preceding cross-modal attention network are jointly trained under supervision using historical scheduling success logs during the offline phase. For each historical task sample, the extraction system extracts features from samples shared in the current cluster. The task-side preference score sequence generated by each candidate computing node It also extracts computing power nodes that were successfully allocated in real-world historical scenarios without overload, and constructs a one-hot encoded real-matching label vector. (If actually allocated to the first If there are 1 node, then (All other corresponding elements are 0). The system establishes a binary cross-entropy allocation objective function. The mathematical formula for optimizing model parameters is as follows: ; In the formula, This represents the total number of candidate computing nodes in the current cluster. For the multilayer perceptron targeting the first The predicted task-side preference score output by each candidate computing node. The corresponding true matching label. The linear mapping weight matrix of the cross-modal attention network. The parameter gradient is determined by the objective function. It dominates and propagates upwards along the network topology across the multilayer perceptron during training, achieving self-closed-loop training with the ability to adaptively assign weights to physical resource features.

[0038] In parallel with the task agent, each computing power agent in the system collects the current resource occupancy percentage of its managed computing power nodes in terms of memory, computing power, and bandwidth at a fixed time frequency (e.g., per second). The computing power agent inputs this real-time occupancy percentage sequence and historical sliding window data into a preset time-series network. In this embodiment, the time-series network is configured with three sub-networks with identical structures and independent parameters for each of the three resource dimensions: memory, computing power, and bandwidth. Each sub-network adopts a two-layer LSTM (Long Short-Term Memory) structure with a hidden layer dimension of 128. The system sets the historical sliding window length to... At each time step, the input to each sub-network is a univariate historical occupancy percentage time series with a resource dimension length of 60. After being input into the network, this time series network does not output a single deterministic prediction value. Instead, it outputs future continuous predictions through fully connected layers and Softmax activation functions connected to the outputs of each sub-network. The load probability distribution matrix over a time step is denoted as the load prediction distribution. Specifically, the predicted distribution for each future time step is as follows: The probability vector is represented by 10 equally spaced discrete intervals (i.e., the resource utilization rate from 0% to 100% is divided into 10 intervals), thus outputting a probability vector of size in each of the three dimensions of memory, computing power, and bandwidth. The probability distribution matrix. Furthermore, before formal deployment, the time-series network is pre-trained offline using historical resource usage data, with negative log-likelihood loss as the objective function for parameter optimization. The pre-training dataset contains measured resource usage time series data of the target intelligent computing center for at least the past 30 days at a 1-second granularity.

[0039] To quantify the uncertainty of future load forecasting, the computing power agent calculates the distribution entropy value H based on the aforementioned distribution matrix. This distribution entropy value is calculated based on the probability mass distribution after discretizing the resource occupancy rate into N equally wide intervals, and is used to quantify the degree of uncertainty in the forecast distribution. Specifically, in the calculation, the forecast distribution for each resource dimension is divided into N discrete intervals according to the aforementioned equally wide method, and the forecast probability (i.e., the probability mass function value) corresponding to each time step in a specific interval is denoted as... Its satisfaction .

[0040] To achieve a stable dimensionality reduction from a multi-dimensional, multi-time-step prediction distribution matrix to a single risk assessment indicator, the system executes explicit aggregation calculation rules: First, it independently calculates the discrete distribution entropy of memory, computing power, and bandwidth dimensions at each future time step. and the corresponding single-step statistical expectation; subsequently, according to future continuous The arithmetic mean is calculated at each time step to obtain the temporal mean distribution entropy and the temporal mean statistical expectation value for each of the three resource dimensions. Finally, to address the most severe system bottleneck risk, the system extracts the maximum value among the temporal mean distribution entropies of the three dimensions of video memory, computing power, and bandwidth, and uses it as the final global distribution entropy value H for computation. Simultaneously, the system forcibly extracts the temporal mean statistical expectation value corresponding to the resource dimension containing the maximum entropy, and uses it as the final statistical expectation value E for computation. After calculating the globally unified distribution entropy value H and statistical expectation value E, the computing power agent compares them with the system-set warning value and occupancy threshold. To unify the parameter setting criteria and ensure the objectivity of the physical meaning, the preset warning value... First preset occupancy threshold With the second preset occupancy threshold (and The threshold values ​​are dynamically determined by the system through statistical analysis of the critical load distribution of computing nodes in normal operation and overload / blocking states during historical allocation rounds. Specifically, the first preset occupancy rate threshold... Set as the 95th percentile of the statistically expected mean among historical overload and congestion samples, with a preset warning value. With the second preset occupancy threshold The values ​​are then set as the maximum value of the distribution entropy in the historical normal operation samples and the mean of the statistical expectation plus one standard deviation. During state comparison, the target risk penalty coefficient is adjusted when none of the following high-risk triggering conditions are met. The default value is zero (i.e., the post-exponential decay factor). ), without any punitive decay effect on the initial score of the machine evaluation; when the conditions are met (Indicating a high degree of certainty regarding the risk of overloading in the future), or meeting the requirements. and When the system triggers a positive linear mapping function to calculate the target risk penalty coefficient (indicating that although there is no absolute overload in the future, the overall load is high and faces a significant risk of upward fluctuation), it indicates that the system is not absolutely overloaded. : ; The formula introduces a maximum function. To ensure that trigger condition one (i.e., only when condition one is met) is met. and When ), the penalty coefficient The minimum value is zero, thus avoiding the erroneous amplification of the initial preference score due to negative values.

[0041] To prevent extreme distributions from causing machine-side preference scores to abnormally drop to zero, the positive linear mapping function has a hard upper limit saturation value. This mechanism ensures that, under the premise of meeting high-risk triggering conditions, the risk penalty coefficient... The value of [the value] increases monotonically with the increase of the distribution entropy value H, and its specific penalty intensity is strictly determined by the risk sensitivity weighting coefficient. Scalar of current task resource consumption A joint decision.

[0042] In the formula, For a fixed risk sensitivity weighting coefficient, This is a scalarized penalty factor applied to the resource consumption (such as GPU memory requirements) of the current requesting task. To comprehensively assess the overall load pressure of heterogeneous tasks, this scalarized penalty factor... The calculation is performed using a weighted normalized summation method for multidimensional resource requirements: ; In the formula, , , These represent the video memory requirements, computing power requirements, and bandwidth requirements allocated to the current task, respectively. , , These represent the upper limit of the total physical resources corresponding to the candidate computing power node; , , The system's preset resource dimension sensitivity weights satisfy... The weight values ​​can be dynamically allocated based on the current global stress level of each resource dimension in the system, or statically assigned based on historical operational experience. For dynamic allocation, the calculation logic for the weight of a specific resource is as follows: Real-time calculation of the ratio of the total global demand for that resource within the current cluster to the total global available resources. And through the normalization formula Dynamically generated. This mapping mechanism ensures that the more resources a task consumes overall, the greater its associated penalty coefficient.

[0043] The machine evaluation logic of the computing power agent specifically adopts a multi-dimensional resource weighted scoring mechanism. This logic first extracts the current resource occupancy percentage of the node in terms of memory, computing power, and bandwidth. Combined with the expected load distribution of the corresponding dimension, it calculates the percentage of future available resources in each dimension. Then, it introduces the task resource requirements of the current task in terms of memory, computing power, and bandwidth, calculating the degree to which the node's future available resources satisfy the actual requirements of the task. Finally, it weights and sums the satisfaction degrees of memory, computing power, and bandwidth according to preset dimension weights, normalizes the result to the [0, 1] interval, and calculates an initial evaluation score for the task. The specific scoring logic is as follows: extract the ratio coefficient of the node's future available capacity to the current task resource requirements in the three resource dimensions. If the ratio of any dimension is less than 1, it is determined that the physical hard constraints are not met and a minimum penalty is directly assigned; if all are met, the score is calculated using the formula... Calculate the weighted sum. Where... , , The dimensionless demand satisfaction score is calculated as follows: the difference between the future available resource balance of a node and the resource requirement of the corresponding dimension of the task, divided by the upper limit of the total physical resources of that node, and the result is mapped to a value using the Sigmoid function. The interval. , , Then it is used as a normalized weighting coefficient in the calculation, satisfying the following conditions: This ensures that the initial score is quantitatively determined entirely by the node's comprehensive multi-dimensional resource reserve level; the more abundant the node's remaining resources to meet task requirements, the higher the initial score. Subsequently, the system uses the calculated risk penalty coefficient. A numerical decay operation is performed on the initial score to obtain the final machine-side preference score. : ; Due to the penalty coefficient The numerical value of distribution entropy is directly introduced in the text. With task resource consumption The product term, after exponential decay calculation, can mathematically guarantee that when the computing node faces increased uncertainty in future load (i.e., an increase in the distribution entropy), the machine-side preference score given to the task that exceeds the resource consumption limit will show a clear monotonically decreasing state, thereby adaptively blocking high-risk, high-consumption allocation behavior.

[0044] This invention predicts the distribution matrix using a temporal network and introduces distribution entropy calculation, transforming the uncertainty of future load on computing nodes into a risk penalty coefficient that can be used in computation. Compared to algorithms that rely solely on currently available resources, this scheme endows the computing agent with a certain degree of risk perception and adjustment capabilities. When potential future load fluctuations are predicted, the penalty mapping function helps form a computing power buffer by lowering the machine-side preference score for large tasks. This attenuation mechanism helps reduce the probability of node overload caused by short-sighted scheduling, improving the system's robustness and task fault tolerance under dynamic and complex loads.

[0045] Furthermore, the output of the stable allocation scheme for non-blocking pairs includes: summarizing the task-side preference scores and machine-side preference scores to construct the cross-mapping bilateral preference matrix; obtaining the current unused GPU memory capacity of each computing power node; each task agent initiating a matching request in descending order of task-side preference scores; after receiving the request, the computing power agent sorts the requests in descending order of machine-side preference scores, sequentially accumulating the GPU memory requirements of the corresponding requesting task according to the sorting order, retaining matching requests whose accumulated GPU memory requirements do not exceed the unused GPU memory capacity, and rejecting matching requests that trigger excessive amounts; cyclically executing the retransmission of requests from rejected tasks to other candidate nodes and the filtering by computing power agents until no new matching requests are generated, outputting the stable allocation scheme for non-blocking pairs, and marking the finally rejected tasks as unallocated tasks.

[0046] Specifically, the arbitration agent, acting as a central coordinating component, collects the task-side preference scores submitted by all task agents in the current allocation round, as well as the machine-side preference scores submitted by all computing power agents, and constructs a cross-mapping bilateral preference matrix in memory. Simultaneously, the arbitration agent obtains the current unused GPU memory capacity scalar value of each computing power node in real time through the system resource monitoring interface. ( (This serves as the identifier for the computing power node) and is used as the absolute physical capacity constraint for this round of matching.

[0047] In the initial phase of the matching loop, each task agent generates a candidate node preference list based on its corresponding task-side preference score sequence, in descending order of score. All task agents in the pending assignment state first send a matching request to the computing power node ranked first in their preference list (i.e., the one with the highest preference score).

[0048] After receiving matching requests from multiple task agents, each computing power agent extracts the machine-side preference scores corresponding to these request tasks and sorts all requests in the current temporary queue in descending order of score. Subsequently, the computing power agents traverse the queue sequentially according to the sorted order to extract the GPU memory requirements of each request task. ( (For the ranking sequence number), and perform an accumulation operation. The computing power agent retains the first few units that satisfy the following accumulation inequality constraint. One matching request: ; When traversing to the th Each task, and adding its memory requirements, would result in a total memory requirement greater than the current node's unused memory capacity. Upon that, the computing power agent immediately triggers the truncation mechanism. The computing power node will retain the previous... The matching status of the first task, and to the first task All ranked tasks from the first to the last will receive a rejection signal (i.e., a rejection that triggers an excess of match requests).

[0049] Upon receiving the rejection signal, the rejected task agent removes the current node from its intention list and resends the matching request to the next-ranked candidate node in the list.

[0050] Upon receiving a new request, the computing power agent merges the new task with the previously retained tasks, re-sorts them according to the machine-side preference scores, and executes the capacity accumulation and filtering logic again (during this process, previously retained low-scoring tasks may be squeezed out of the queue by newly arrived high-scoring tasks).

[0051] The "retransmission of requests - sorting and filtering - dynamic elimination" loop continues to execute until no new matching requests are generated in the system (that is, all tasks are either successfully retained by a node or rejected by all candidate nodes in its list).

[0052] When the iteration stops, the task combinations currently held by each computing node constitute the final allocation result. Because this iterative process of "request retransmission – sorting and filtering – dynamic elimination" monotonically replaces tasks with lower preference scores in each round, the sum of machine-side preference scores corresponding to the task sets held by each computing node exhibits a non-decreasing trend. Under the boundary conditions of limited total system resources and a limited task set, this iterative process will inevitably terminate within a finite number of rounds, thus outputting a stable allocation scheme with non-blocking pairs. It should be noted that, in the context of this invention, a stable allocation scheme with non-blocking pairs specifically refers to a pairing combination (i.e., a single-task replacement blocking pair) within the physical boundary satisfying the upper limit of the prefix memory capacity of each node, where there is no pairing combination that can simultaneously increase the preference scores of both the task and the node through a single task replacement. This stable state with capacity constraints reduces computing power idleness or repeated task migration caused by local forced allocation. The arbitration agent issues and executes this scheme, and tasks that are ultimately rejected after traversing all candidate nodes are explicitly marked as unallocated tasks in the state database for processing in subsequent elastic scheduling stages.

[0053] This invention adjusts the approach of allocating fixed slots based on the "average demand" by sequentially accumulating actual memory requirements in a sorting order, thus promoting dynamic binning of heterogeneous tasks within physical boundaries. This filtering mechanism based on the accumulation of actual memory volume helps alleviate the risks of memory fragmentation and allocation out-of-bounds errors. Simultaneously, the use of a capacity-constrained two-sided matching iteration mechanism takes into account the bidirectional preferences of tasks and nodes, ensuring that the output allocation scheme satisfies the non-blocking pair characteristic. This mechanism reduces idle computing power or repeated task migration caused by local forced allocation, improving the overall fairness of the system and the utilization efficiency of computing resources.

[0054] Further, the step of receiving the lowest machine-side preference score among the tasks accepted by the matched node closest to the task elasticity interval in this round includes: the elasticity interval includes a precision elasticity interval and a latency elasticity interval; extracting the precision elasticity boundary set and latency elasticity boundary set of the unassigned tasks; traversing all computing power nodes that successfully accepted task quotas in this round, extracting the associated precision elasticity boundary set and associated latency elasticity boundary set of the paired tasks of each computing power node in this round; calculating the weighted Euclidean distance in precision and latency dimensions between the associated boundaries of the unassigned tasks and each successfully paired computing power node, marking the node corresponding to the minimum distance value as the matched computing power node closest to the task elasticity interval; extracting the minimum machine-side preference score of all tasks accepted by the closest matched computing power node in this round, and feeding it back to the unassigned tasks as a marginal acceptance threshold.

[0055] Specifically, for target tasks marked as "unassigned tasks," the arbitration agent extracts the SLA (Service Level Agreement) elasticity range declared at the time of submission. This elasticity range is structured and parsed into a precision elasticity range and a latency elasticity range, and further extracted into specific boundary endpoint pairs, thus constructing a precision elasticity boundary set. (Representing the minimum tolerable accuracy and the highest expected accuracy) and the set of delay elasticity boundaries (Represents the ideal delay and the maximum tolerable delay).

[0056] The arbitration agent iterates through all computing power nodes that successfully accepted tasks in this round of matching (i.e., whose remaining available capacity was effectively consumed). For each successfully matched computing power node... The system extracts the precision and latency elasticity boundary sets of all tasks accepted in the current round, and calculates the mean or uses a weighted centroid algorithm to aggregate and generate an "admission profile" of the computing node in the current round, i.e., the associated precision elasticity boundary set. With associated delay elasticity boundary set .

[0057] To find the computing power node that best matches the "threshold" of the unassigned task, the arbitration agent introduces a weighted Euclidean distance model. Before performing distance measurement, the system first performs dimensionless processing on the precision elasticity boundaries and delay elasticity boundaries of the unassigned task and the computing power node (e.g., dividing them by the average elasticity interval width of all tasks in the current cluster in that dimension), to eliminate the influence of dimensional differences. Subsequently, the system concatenates the dimensionless lower and upper limits in a fixed order to form a two-dimensional boundary feature vector. That is, the precision boundary vector of the unassigned task is denoted as... The delay boundary vector is denoted as The aggregated profile of each successfully paired computing power node is also mapped to the corresponding two-dimensional mean boundary vector. and Based on the above vectors, the spatial distance between unassigned tasks and computing nodes is calculated in terms of precision and latency. : ; In the formula, Represents the square of the Euclidean distance between two two-dimensional boundary eigenvectors (e.g. ), and These are the system-preset precision feature weights and delay feature weights, which satisfy... The weight values ​​can be determined by performing a grid search or regression fitting on the system's historical successful scheduling data, in order to adapt to the preferences of different business scenarios for accuracy or latency.

[0058] After completing all traversal calculations, the system retrieves the minimum value in the distance value sequence. The computing power node corresponding to this minimum value is then marked as the matched computing power node closest to the elastic range of the unassigned task. .

[0059] Locate the nearest matched computing power node Next, the arbitrating agent accesses the node's current acceptance list and extracts the machine-side preference scores for all tasks in the list. These scores are then sorted in ascending order, and the first value in the sequence (i.e., the lowest machine-side preference score among the tasks accepted by the node in this round) is extracted. The arbitrating agent defines this minimum value as the marginal acceptance threshold and sends it as feedback signaling to the task agent managing the unassigned task via a message queue.

[0060] This invention extracts the SLA (Service Level Agreement) elastic boundary set and calculates the weighted Euclidean distance to match unassigned tasks with computing power nodes whose parameters are relatively close. By extracting the lowest preference score accepted by this node as a marginal acceptance threshold and feeding it back to the task agent, a quantitative reference standard is provided for unassigned tasks. This mechanism alleviates the situation of blindly retries or excessive compromises after task rejection, provides a specific numerical reference for targeted SLA relaxation in the next iteration, helps promote the convergence of subsequent matching iterations, and improves the allocation success rate of the system under multi-round negotiation.

[0061] Further, the allocation of remaining computing power includes: if the difference between the machine-side preference score generated by the unallocated task for the nearest matched computing power node and the received marginal acceptance threshold exceeds a preset tolerance limit, then the limit delay parameter in the delay elastic boundary set and the critical precision parameter in the precision elastic boundary set are extracted; within the elastic range, the constraint is relaxed unidirectionally by adding a first preset constant to the limit delay parameter and subtracting a second preset constant from the critical precision parameter; based on the relaxed critical precision parameter, the collaborative representation vector and task-side preference score are recalculated, and combined with the relaxed limit delay parameter, the matching is re-entered; when the matching cycle reaches the preset limit number of times and there are unallocated remaining tasks, the computation urgency score is read and arranged in descending order, and the remaining available resources of each global computing power node are independently allocated to the remaining tasks until the resources are cleared.

[0062] Specifically, the task agent receives the marginal acceptance threshold from the arbitration agent. Then, the machine-side preference score generated by the closest matched computing power node for the current task, which was forwarded to the local machine by the arbitration agent in previous rounds, is extracted. The task agent calculates the difference between the two. and compare it with the preset tolerance limit. A comparison is performed. The preset tolerance limit... The value can be dynamically calculated by the system based on the statistical variance of successfully matched scores in historical allocation rounds, or pre-configured by the administrator based on the strictness of the current cluster service level agreement (SLA) (for example, the more lenient the SLA requirements, the larger the tolerance threshold should be). This indicates that the current task's computing power requirements are far from meeting the node's entry threshold, and network fluctuations alone cannot naturally match this. Consequently, the system triggers a one-way relaxation mechanism for constraints.

[0063] After the relaxation mechanism is triggered, the task agent extracts the limiting delay parameter from its set of delay elasticity boundaries. And extract critical accuracy parameters from the set of accuracy elastic boundaries. Provided that the upper and lower limits of the task's global physical elasticity range are not exceeded, the following parameters are updated: ; ; In the formula, The first preset constant is used (e.g., a delay step of 50ms). The second preset constant (e.g., the precision degradation step size, such as the quantization value change from FP16 to INT8) is used. The specific values ​​of the first and second preset constants are set based on the average compromise step size from offline historical matching. When the two are added or subtracted, they are limited by the upper and lower limits of the global physical elasticity range set by the task, ensuring that the relaxed parameters do not exceed the absolute tolerance threshold of the task. This operation realizes the "one-way relaxation" of constraints, that is, giving the task a longer queuing tolerance time and a lower computational precision threshold.

[0064] The task agent is based on the relaxed limit delay parameter. Update the waiting time constraint of the current task in the arbitration queue and adjust the relaxed critical accuracy parameter. Used to trigger feature state evaluation. The system has configured a precision type switching mapping rule here: when the relaxed critical precision parameter... When the data precision falls below the preset safety threshold corresponding to the current data precision type, the data precision type declaration in the input feature is automatically downgraded to the next level of quantization standard (such as automatically switching from FP16 to INT8 one-hot feature), and the physical conversion factor (e.g., 0.5) bound to the downgraded category is called to compress the original memory and bandwidth requirements proportionally.

[0065] During the recalculation of the collaborative representation vector, the system explicitly defines the update boundaries for the feature data: when recombining the aforementioned 205-dimensional original feature vector, only the one-hot features of the data precision type that triggered the degradation are used to update the corresponding input components, while the remaining multimodal features related to the task computation graph topology, operator distribution frequency, batch size, and inference stage identifier remain statically unchanged. The relaxed limit delay parameter does not participate in the feature recombination of the collaborative representation vector but is directly updated in the conditional judgment rules of the arbitral agent. Based on the aforementioned downgraded data precision type features, the boundary features of local updates, and the smaller resource requirement vector obtained from the recalculation, the forward propagation of the multilayer perceptron and dual predictor heads is retried, resulting in a corresponding decrease in the computational urgency score and the precision sensitivity score. Using the updated features, the system reassembles and generates a new collaborative representation vector, and updates the task-side preference scores for each candidate node accordingly. By unidirectionally relaxing constraints and reducing the resource requirement vector, the hard threshold of the task for underlying physical resources is essentially lowered, resulting in an increase in the comprehensive matching fitness of the task's mapping in cross-modal attention networks and multilayer perceptrons. This, in turn, effectively improves the task-side preference score for candidate nodes. Tasks that have completed state updates are re-injected into the scheduling queue to participate in the next round of bidirectional stable matching loops.

[0066] To ensure the convergence of the multi-round flexible negotiation mechanism, the system sets the unidirectional constraint relaxation operation to be triggered only after the current main matching loop has completely stopped (i.e., no new matching requests are generated globally). During the relaxation process, the limit delay parameter only increases in one direction, and the critical precision parameter only decreases in one direction. This unidirectionality ensures that the evolution of the recalculated collaborative representation vector and the task-side preference score has a strict monotonic boundary. This monotonic evolution mechanism, combined with the global counter maintained by the system and the preset upper limit of the number of iterations, avoids infinite loop deadlock caused by disordered oscillations of preference scores, ensuring that the bidirectional stable matching algorithm with capacity constraints converges and outputs the final solution within a finite number of iterations.

[0067] The system maintains a global counter to record the number of matching loops. When the number of loops reaches a preset maximum, a counter is activated. If there are still remaining tasks in the scheduling queue that have failed to match despite multiple compromises, the system terminates the matching algorithm based on the bidirectional preference matrix and triggers the best-effort allocation fallback logic. The arbitration agent reads the updated computational urgency scores of all remaining tasks and establishes a priority forced allocation queue in descending order of values. Simultaneously, it inventories the total available resources of all computing nodes globally (i.e., the fragmented GPU memory that was not fully used in the first round on each node). Following the priority queue order, the arbitration agent independently scans each computing node, strictly prohibiting the merging of physical GPU memory across nodes. If the remaining fragmented resources of a single node can meet the minimum degradation operation threshold of a task (e.g., forcibly performing INT8 quantization compression to adapt to smaller GPU memory for large language model tasks), then the available resources of that single node are forcibly allocated to that remaining task to support its degradation operation, until the available resources of all single nodes globally are completely cleared. In this fallback scenario, "resource clearing" includes two physical states: one is that all available resources of the node have been allocated; the other is that the remaining fragmented resources have physically fallen below the minimum quantitative degradation threshold allowed by any task and are marked by the system as an equivalent clearing state and no longer participate in allocation.

[0068] This invention guides unassigned tasks to perform unidirectional parameter relaxation within a flexible range by establishing conditional triggering logic based on tolerance limits. This adaptive compromise mechanism helps alleviate matching blockage caused by resource contention and improves the matching probability of long-tail tasks in subsequent iterations. Simultaneously, a global fallback mechanism based on computational urgency scores, designed for high-load, resource-scarce scenarios, utilizes available fragmented computing power within the cluster while prioritizing high-urgency tasks. This mechanism supports finite-time convergence of multi-agent scheduling algorithms at the engineering logic level, reducing long-term task waiting times and idle underlying computing resources.

[0069] Furthermore, updating the bias compensation parameters of the prediction model includes: at the deadline of the current allocation round time window, collecting actual resource consumption observation values ​​of the computing facilities to constitute the actual load; extracting the statistical expectation corresponding to the predicted distribution output of the previous allocation round, calculating the load residual matrix of the actual load and the statistical expectation; multiplying the load residual matrix into a preset exponential decay constant, updating the bias compensation parameters and converting them into a compensation tensor of the same dimension, superimposing them into the bias weight matrix of the hidden layer of the prediction model, and driving the next round of inference results to generate an equivalent correction towards the actual resource consumption observation value.

[0070] Specifically, at the end of the current allocation round time window Each computing power agent collects real-time observations of the actual resource consumption of its managed computing facilities in the current cycle through underlying hardware monitoring interfaces (such as graphics card driver APIs or operating system kernel probes). It should be noted that, to avoid false observations (i.e., capturing remnants of the previous cycle) caused by the inherent latency of the underlying physical allocation of the AI ​​computing power cluster (such as the actual allocation of GPU memory and the establishment of large bandwidth), in actual engineering deployments, when the system triggers the collection action at the aforementioned deadline, it preferably includes a brief physical stabilization buffer period (such as a millisecond-level wait) or verifies that the underlying state has reached a steady state through a kernel probe before reading. This engineering process effectively prevents phase lag oscillations in the prediction model during closed-loop correction due to the observed data not reaching a physical steady state, ensuring the effectiveness of the closed-loop update. The collected data, such as actual GPU memory usage, actual computing power utilization, and actual read / write bandwidth, are structured and aligned according to the dimensions of the prediction model to form a true and accurate actual load matrix. .

[0071] The computing power agent retrieves the prediction distribution matrix from memory, which was derived by the prediction model in the previous allocation round (i.e., the deduction phase before the current actual load occurs). For this prediction distribution matrix, its corresponding statistical expectation matrix is ​​calculated through probability-weighted summation. Subsequently, the actual load matrix will be... With statistical expectation matrix By performing element-by-element subtraction, the load residual matrix reflecting the direction and magnitude of the deduction error is calculated. It should be noted that the load residual matrix described in this invention retains the true direction vectors of overestimation or underestimation, accurately quantifying the degree of error of the prediction model in this round of allocation. This is achieved through the calculation formula: ; This load residual matrix not only quantifies the absolute magnitude of the extrapolation error, but also preserves the true direction vectors of overestimation or underestimation, accurately quantifying the degree of overestimation or underestimation of the consumption of various resources in this round of allocation by the prediction model.

[0072] To prevent drastic oscillations in the prediction model caused by single, sporadic load surges, the system introduces a preset exponential decay constant. ( (For example, set it to 0.1). The above load residual matrix... Multiplying by the exponential decay constant yields the smoothed correction for this round. Considering the risk of multiple computing agents simultaneously writing to the same set of shared parameters and causing contention under high concurrency conditions, the system further adopts a decentralized parameter isolation architecture. That is, each computing agent maintains and updates its own deviation compensation parameter matrix in its own independent local memory. During the initialization phase, the time index for setting the deviation compensation parameters is determined. and its initial value Forced to be a zero matrix to prevent the introduction of prior noise. For each subsequent allocation round... Its smooth update formula is: ; Through the pre- The attenuation term allows the system to absorb newly occurring errors while exponentially forgetting old errors, effectively preventing error accumulation and drift under long-term cyclic corrections from a mathematical perspective. Subsequently, the updated deviation compensation parameter matrix is... Based on the data structure requirements of the underlying neural network of the prediction model, it is transformed into a compensation tensor of the same dimension. Specifically, since the aforementioned time-series network is configured with three structurally independent and parameter-decoupled sub-networks for memory, computing power, and bandwidth respectively, in this transformation step, the system strictly splits the above-mentioned deviation compensation parameter matrix into three independent one-dimensional deviation component sequences according to the resource dimension.

[0073] To eliminate the dimensional conflict between physical resource dimensions (such as memory usage percentage, actual bandwidth value, etc.) and the dimensionless parameter space of the hidden layers of the neural network, and to avoid activation function saturation caused by direct superposition, the system performs a standardization operation during the conversion process: extracting the historical load standard deviation vector of each computing node in the corresponding resource dimension. For each dimension after splitting Perform element-wise dimensionless division (i.e.) This maps the physical bias to a compensation tensor of the same order of magnitude as the network bias weights. .

[0074] The specific operation is as follows: The compensation tensor generated after dimensionless processing... The model is replicated and expanded via matrix broadcasting to precisely align the physical tensor shape to the dimension of the corresponding hidden layer neurons. Subsequently, parameters are independently superimposed on the three resource sub-networks (memory, computing power, and bandwidth), and then combined with the original bias weight matrix of the hidden layer of the prediction model (specifically, the post-processing calibration hidden layer specifically set at the end of the network backbone to avoid disrupting the core gating dynamics of the LSTM) in the corresponding sub-network instantiated locally by the computing power agent. Tensor addition is performed to perform superposition. During this process, the... Strictly speaking, it refers to the baseline bias matrix that is fixed after the model has completed offline pre-training, while As a cumulative bias correction that incorporates the exponentially decaying memory of historical time steps, each allocation round is only superimposed once on this fixed baseline matrix to calculate the updated bias weight matrix. : ; The benchmark and incremental decoupling superposition mechanism fundamentally avoids residual drift and repeated accumulation. After superposition, when the prediction model receives new input features for forward inference in the next allocation round, the updated bias weight matrix will mathematically drive its output to automatically generate an equivalent displacement correction towards the actual resource consumption observation value of this round.

[0075] This invention constructs a prediction closed loop that includes inference, feedback, and correction. By extracting the load residual matrix between the predicted expectation and the actual load, and combining it with an exponentially decaying constant for smoothing filtering, noise interference caused by single load mutations is reduced. The compensation tensor is superimposed on the bias weight matrix of the hidden layer of the prediction model, providing a lightweight parameter update path for the model. This mechanism alleviates the systematic bias problem that static prediction models are prone to when introducing new models or when hardware states change in computing clusters, helping to maintain the rationality and environmental adaptability of the scheduling system's prediction of dynamic load over long periods.

[0076] This invention constructs a full-link collaborative scheduling architecture encompassing semantic awareness, risk matching, compromise negotiation, and closed-loop correction. Through multimodal semantic encoding and a dual-predictor mechanism, the granularity of feature awareness for heterogeneous tasks is refined. Combining distributed entropy values ​​with a bidirectional stable matching algorithm with capacity constraints mitigates node overload risks and memory fragmentation. Marginal acceptance threshold feedback and a one-way constraint relaxation mechanism guide tasks to make appropriate compromises, supplemented by a fallback strategy to utilize remaining computing power, alleviating scheduling congestion. A closed-loop compensation mechanism based on actual load deviation provides the system with continuous dynamic correction capabilities. Overall, this invention improves the efficiency of computing power supply and demand matching, overall resource utilization, and long-term dynamic adaptability of intelligent computing centers in complex high-concurrency scenarios.

[0077] Example 2: Further, the step of calculating the weighted fusion vector of the semantic embedding vector and the resource requirement vector via the cross-modal attention network includes: extracting binary identifiers representing the inference stage of the large model from the multimodal features; assigning a stage-aware temperature scaling factor to the current task based on the binary identifiers, wherein a first temperature value is assigned to the identifiers representing the decoding stage, and a second temperature value is assigned to the identifiers representing the pre-filling stage, and the first temperature value is less than the second temperature value; when calculating the attention weight matrix of the cross-modal attention network, based on the dimensionality attention regularization term introduced in the end-to-end training stage, the dot product of the query feature with the inference identifier representing the decoding stage is made to have a dominant dot product in the direction of the key feature corresponding to the memory and bandwidth, the dot product of the query feature and the key feature is divided by the stage-aware temperature scaling factor for numerical scaling, and then the weights are calculated by the activation function and the value features are weighted and summed to generate the weighted fusion vector.

[0078] Specifically, before constructing a cross-modal attention network for feature fusion, the task agent separately parses the binary identifier of the inference stage of the large model from the extracted multimodal features. The system pre-defines a mapping rule: if the identifier is 0 (representing the decoding stage, exhibiting memory-access-bandwidth-intensive features), the system assigns it a smaller initial temperature value. (e.g., 0.5); if the identifier is 1 (representing the Prefill stage, exhibiting computationally intensive characteristics), the system assigns it a larger second temperature value. (e.g., 2.0). The extracted temperature values ​​are collectively referred to as the stage-sensing temperature scaling factor. .

[0079] In the cross-modal attention mechanism computation stage, the system uses semantic embedding vectors as the query matrix. The resource demand vector, after mapping, serves as the key matrix. Sum matrix After performing the dot product operation to obtain the attention score, the system divides the score by the extracted stage-aware temperature scaling factor. The result is then transformed into an attention weight matrix using the Softmax function. (Weighted fusion vector) The forward calculation formula is updated as follows: ; In the formula, Let K be the feature dimension of the key matrix, and This is the base scaling factor. To establish the directionality of the feature dimensions, the cross-modal attention network additionally introduces a dimensional attention regularization term into the joint loss function during the end-to-end training phase. This regularization term constrains the samples identified for decoding during the inference phase, making their query vectors... Key vectors corresponding to video memory and bandwidth The integral of points in the direction is significantly higher than that in terms of computational power. Given that this physical mapping relationship has been established, when... When the values ​​are small, the weight distribution of the Softmax output tends to be sharp, deterministically amplifying the existing score differences and forcing the network to focus its attention on the low-dimensional requirements of memory and bandwidth; when When the weights are large, the weight distribution tends to be smooth, prompting the network to take into account the resource requirements of each component in a balanced way.

[0080] This invention introduces a stage-aware temperature scaling mechanism into the cross-modal feature fusion process, directly influencing the calculation of attention weights with the differentiated characteristics of the underlying physical bottlenecks of large models at different inference stages. Through the combined effect of dimensional attention regularization and stage-aware temperature scaling, a smaller temperature parameter is assigned to the decoding stage to sharpen and lock its existing high attention in the memory access dimension, while a larger parameter is assigned to the pre-filling stage to balance the focus. This mathematical constraint allows the generated collaborative representation vector to more accurately capture the dynamic mapping relationship between computational semantics and physical requirements. This helps reduce resource estimation bias caused by stage transitions in continuous batch inference of large language models, improving the accuracy of multimodal representation mechanisms in depicting the behavioral patterns of complex AI tasks.

[0081] Furthermore, when at least one of the following two conditions is met—that the statistical expected value exceeds a first preset occupancy threshold and that the distribution entropy value exceeds a preset warning value and the statistical expected value exceeds a second preset occupancy threshold—establishing a positive linear mapping function and obtaining a risk penalty coefficient based on the distribution entropy value includes: extracting the temporal expected values ​​corresponding to the load prediction distribution of the managed computing power nodes under the previous consecutive time steps, calculating the difference to generate a temporal statistical gradient; when at least one of the above two conditions is met, determining the positive or negative nature of the temporal statistical gradient; if the temporal statistical gradient is greater than zero, triggering the positive linear mapping function to output the risk penalty coefficient; if the temporal statistical gradient is less than or equal to zero, reducing the distribution entropy value according to a preset decay ratio, and then outputting the reduced risk penalty coefficient through the positive linear mapping function.

[0082] Specifically, in the process of evaluating machine-side preference scores by the computing power agent, the system not only relies on the current distribution entropy value but also introduces load momentum evaluation. The computing power agent maintains a length of [missing value] in memory. The sliding window records the previous The expected time series value corresponding to the load forecast distribution within a time step The system uses a first-order difference approximation to calculate the current time. Time series statistical gradient : ; This time-series statistical gradient is used to characterize the macroscopic trend of future load changes in computing nodes.

[0083] When the calculated expected statistical value of the computing node exceeds the first preset occupancy rate threshold or the distribution entropy value Exceeding the preset warning value Furthermore, when the expected statistical value exceeds the second preset occupancy rate threshold, the system extracts the aforementioned time-series statistical gradient. Perform a sign determination. If... This indicates that the future load is in a period of accelerated increase and has high uncertainty (i.e., a real risk of sudden overload). The system normally triggers the positive linear mapping function to calculate the risk penalty coefficient; if This indicates that although the future load is at a high level or fluctuates greatly, it is generally in a downward trend (i.e., the risk is being released). The system dampens the difference in distributed entropy before calculating the penalty coefficient. ; In the formula, The preset attenuation ratio (e.g., 0.2) is used to output the reduced risk penalty coefficient.

[0084] This invention effectively distinguishes between two distinct physical trends in high-entropy computing nodes: "accelerated load surge" and "load oscillation and decline," by supplementing the risk penalty mechanism with a temporal statistical gradient as a momentum indicator. This mechanism adds a one-dimensional directional constraint to the determination of distributed entropy, enabling the computing agent to impose penalties to avoid overload when the load truly surges, and to appropriately relax restrictions to accommodate new tasks when the load eases. This improves upon the "false positive" phenomenon that easily arises from relying solely on the numerical value of distributed entropy, enhancing the accuracy of risk assessment and the consistency of computing power allocation when dealing with dynamic bursts of traffic.

[0085] Further, the step of multiplying the load residual matrix into a preset exponential decay constant, updating the deviation compensation parameters, converting them into a compensation tensor of the same dimension, and superimposing them onto the bias weight matrix of the hidden layer of the prediction model includes: periodically calculating the Fisher information matrix of the bias weight matrix in the prediction model for historical key steady-state load characteristics, extracting the diagonal elements of the matrix to generate a weight importance diagonal matrix; using the weight importance diagonal matrix to perform element-wise damping adjustment on the preset exponential decay constant, so that the weight dimensions with higher importance in steady-state historical extrapolation are assigned smaller decay constant values; multiplying the load residual matrix into the damped decay constant matrix to generate a compensation update amount, accumulating and updating it into the deviation compensation parameters and converting it into a compensation tensor, and then superimposing it onto the bias weight matrix.

[0086] Specifically, in each passing After a complete allocation round, the system extracts a sample of actual load characteristics marked as "historical key steady state" offline in the background. The computational agent uses the prediction model to calculate the second derivative approximation of the log-likelihood loss on these samples, obtaining the bias weight matrix. The Fisher information matrix is ​​calculated. Specifically, the system uses the empirical Fisher estimation method to calculate the diagonal elements of this matrix: for each sample in the historical key steady-state load feature sample set, the log-likelihood function of the prediction model output with respect to the bias weight matrix is ​​calculated. The first-order gradient of each component is squared element-wise and then statistically averaged over the sample set. To reduce online computational overhead, the system directly extracts this estimation result and constructs a weight importance diagonal matrix with the same dimensions as the bias weight matrix. This matrix quantifies the importance of each specific weight parameter in the model for memorizing existing stable load patterns.

[0087] At the end of the current allocation round time window, the computing power agent obtains the load residual matrix. When calculating update values, the system no longer uses a globally uniform scalar decay constant. Instead, it utilizes a diagonal matrix of weight importance. right Perform element-wise damping adjustment calculations to generate an adaptive attenuation constant matrix. : ; In the formula, This is the preset elastic protection penalty coefficient. Subsequently, the system performs the Hadamard Product operation, converting the load residual matrix... Multiply by the damping attenuation constant matrix Generate the compensation update amount for the current round and add it to the deviation compensation parameters. middle: ; Ultimately The parameters are converted into a compensation tensor and superimposed onto the bias weight matrix to complete one anti-forgetting parameter update.

[0088] This invention introduces elastic weight solidification logic based on the Fisher information matrix into the closed-loop feedback mechanism of the prediction model. By adjusting the decay constant element-wise based on weight importance, the system can spontaneously apply update resistance to weight dimensions that are crucial for remembering historical steady-state patterns when absorbing the latest actual load deviations, while fully absorbing new knowledge in less important dimensions. This effectively alleviates the "catastrophic forgetting" problem that easily occurs during online continuous learning, enabling the prediction model to maintain the stability of its underlying mathematical deductions during long-term adaptation to dynamic heterogeneous loads, extending the lifespan of the scheduling system and the reliability of model predictions.

[0089] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for optimizing AI computing power based on multi-agent collaboration, characterized in that, include: Receive AI tasks, obtain the resource requirement vector of the tasks, and set a flexible range for the tasks; The multimodal features of the task are extracted, encoded into semantic embedding vectors by a semantic encoder, and the urgency score and accuracy sensitivity score are calculated by the output of the prediction head. These are then fused and concatenated with the resource demand vector and the semantic embedding vector to form a collaborative representation vector. The task agent generates task-side preference scores for each computing power node based on the collaborative representation vector; the computing power agent generates machine-side preference scores based on node resource usage and the load prediction distribution output by the prediction model, and adjusts the risk penalty coefficient during computation based on the load prediction distribution; specifically: the task agent maps the collaborative representation vector to the task-side preference score sequence for each candidate computing power node for the current task through a multilayer perceptron; the computing power agent collects the resource usage percentages of the managed computing power nodes in terms of memory, computing power, and bandwidth, and calls a time-series network to predict the distribution matrix of future continuous time steps as the load prediction distribution. The distribution entropy value and corresponding statistical expectation value of the distribution matrix are calculated. When at least one of the following conditions is met, namely, the statistical expectation value exceeds the first preset occupancy rate threshold, and the distribution entropy value exceeds the preset warning value and the statistical expectation value exceeds the second preset occupancy rate threshold, a positive linear mapping function is established and a risk penalty coefficient is obtained based on the distribution entropy value. The resource occupancy percentage and load prediction distribution are input into the machine evaluation logic to generate an initial score. The initial score is then subjected to a numerical decay operation using the risk penalty coefficient to obtain a machine-side preference score, so that the machine-side preference score for tasks with excessive resource consumption decreases monotonically as the distribution entropy value increases. The arbitration agent constructs a two-sided preference matrix based on the task-side preference score and the machine-side preference score, runs a two-sided stable matching algorithm with capacity constraints, and outputs a stable allocation scheme for non-blocking pairs. If no task is assigned in a round, the lowest machine-side preference score among the tasks accepted by the matched computing power node that is closest to the task's elasticity range will be received. After adjusting the constraints within the elasticity range to increase the task-side preference score, the node will be rematched. If no task is assigned after the preset maximum number of rounds, the remaining computing power will be allocated in descending order of computation urgency score. After each round of allocation, the computing power agent updates the deviation compensation parameters of the prediction model based on the deviation between the actual load and the predicted distribution.

2. The AI ​​computing power optimization method based on multi-agent collaboration according to claim 1, characterized in that: The acquisition of the collaborative representation vector includes: the semantic encoder employing a pre-set shared network encoding layer; inputting extracted multimodal features, including topology, operator distribution, accuracy, batch size, and inference stage, into the shared network encoding layer to generate the semantic embedding vector; setting a first prediction head and a second prediction head in parallel in the shared network encoding layer, updating parameters using historical queuing time cost and historical model quality degradation data as supervision signals, and independently outputting the computational urgency score and the accuracy sensitivity score; acquiring a task resource requirement vector including memory requirements, computing power requirements, and bandwidth requirements; calculating a weighted fusion vector of the semantic embedding vector and the resource requirement vector through a cross-modal attention network, and concatenating the weighted fusion vector with the computational urgency score and the accuracy sensitivity score in a preset order to generate the collaborative representation vector.

3. The AI ​​computing power optimization method based on multi-agent collaboration according to claim 1, characterized in that: The output of the stable allocation scheme for non-blocking pairs includes: summarizing the task-side preference scores and machine-side preference scores to construct the cross-mapping bilateral preference matrix; obtaining the current unused GPU memory capacity of each computing power node; each task agent initiating matching requests in descending order of task-side preference scores; after receiving the requests, the computing power agents sort them in descending order of machine-side preference scores, and sequentially accumulate the GPU memory requirements of the corresponding requesting tasks according to the sorting order, retaining matching requests whose accumulated GPU memory requirements do not exceed the unused GPU memory capacity, and rejecting matching requests that trigger excessive amounts; cyclically executing the retransmission of requests from rejected tasks to other candidate nodes and the filtering by computing power agents until no new matching requests are generated, outputting the stable allocation scheme for non-blocking pairs, and marking the finally rejected tasks as unallocated tasks.

4. The AI ​​computing power optimization method based on multi-agent collaboration according to claim 1, characterized in that: The process of receiving the lowest machine-side preference score among the tasks accepted by the matched node closest to the task elasticity interval in this round includes: the elasticity interval includes a precision elasticity interval and a latency elasticity interval; extracting the precision elasticity boundary set and latency elasticity boundary set of the unassigned tasks; traversing all computing power nodes that successfully accepted task quotas in this round, extracting the associated precision elasticity boundary set and associated latency elasticity boundary set of the paired tasks of each computing power node in this round; calculating the weighted Euclidean distance in precision and latency dimensions between the associated boundaries of the unassigned tasks and each successfully paired computing power node, marking the node corresponding to the minimum distance value as the matched computing power node closest to the task elasticity interval; extracting the minimum machine-side preference score of all tasks accepted by the closest matched computing power node in this round, and feeding it back to the unassigned tasks as a marginal acceptance threshold.

5. The AI ​​computing power optimization method based on multi-agent collaboration according to claim 4, characterized in that: The allocation of remaining computing power includes: if the difference between the machine-side preference score generated by the unallocated task for the nearest matched computing power node and the received marginal acceptance threshold exceeds a preset tolerance limit, then the limit delay parameter in the delay elastic boundary set and the critical precision parameter in the precision elastic boundary set are extracted; within the elastic range, the constraint is relaxed unidirectionally by adding a first preset constant to the limit delay parameter and subtracting a second preset constant from the critical precision parameter; based on the relaxed critical precision parameter, the collaborative representation vector and task-side preference score are recalculated, and combined with the relaxed limit delay parameter, the matching is re-entered; when the matching loop reaches the preset limit number of times and there are unallocated remaining tasks, the computation urgency score is read and arranged in descending order, and the remaining available resources of each global computing power node are independently allocated to the remaining tasks until the resources are cleared.

6. The AI ​​computing power optimization method based on multi-agent collaboration according to claim 1, characterized in that: The update of the bias compensation parameters of the prediction model includes: at the deadline of the current allocation round time window, collecting the actual resource consumption observation values ​​of the computing facilities to form the actual load; extracting the statistical expectation corresponding to the predicted distribution output of the previous allocation round, calculating the load residual matrix of the actual load and the statistical expectation; multiplying the load residual matrix into a preset exponential decay constant, updating the bias compensation parameters and converting them into a compensation tensor of the same dimension, superimposing them into the bias weight matrix of the hidden layer of the prediction model, and driving the next round of inference results to generate an equivalent correction towards the actual resource consumption observation values.

Citation Information

Patent Citations

  • Content generation data processing method and system based on multi-agent cooperation

    CN121682718A

  • AI agent dynamic task scheduling method and system based on multi-model collaborative reasoning

    CN122285230A