Cloud-edge collaborative scheduling and agent running method and system based on distributed idle computing power network
By constructing a city-level distributed heterogeneous GPU computing power network, unified quantification and health assessment of heterogeneous computing power were achieved. Combined with an automated operation and maintenance and self-healing system, the problems of low scheduling matching degree and poor stability of heterogeneous computing power networks in existing technologies were solved, thereby improving the utilization rate of idle computing power resources and system stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LINGBO TECH CO LTD
- Filing Date
- 2026-05-19
- Publication Date
- 2026-06-16
AI Technical Summary
Existing computing networks lack a unified standard for heterogeneous computing power and a comprehensive node feature collection system, resulting in insufficient accuracy in node load prediction. This makes it difficult to accurately match the diverse needs of AI agents and large model tasks, leading to low matching degree of computing power resource scheduling, low utilization rate of idle computing power, and a lack of computing power tide perception and dynamic elastic scaling capabilities in traditional scheduling schemes. Furthermore, the scheduling robustness is poor in abnormal scenarios, resulting in insufficient system stability.
Construct a city-level distributed heterogeneous GPU computing power network, build a four-layer collaborative architecture, and realize automatic node registration and performance calibration; collect full-dimensional indicators, realize unified quantification of heterogeneous GPU computing power through HCU, build a node health assessment model and perform time-series load prediction; complete the six-tuple requirement modeling of AI agents and large model tasks, construct a multi-dimensional comprehensive cost function, and realize optimal node scheduling decisions; identify the computing power network tidal pattern, realize elastic scaling of computing power based on MADDPG; deploy an automated operation and maintenance and self-healing system, detect anomalies through the isolated forest algorithm, use the LSTM-Attention model for fault prediction, and save GPU context and migrate containers when a fault is triggered to ensure the canary release and automatic rollback of model versions.
It achieves efficient integration and precise matching of heterogeneous GPU computing power, improves the scheduling matching degree and utilization of idle computing resources, ensures the efficient and stable operation of large model inference tasks in abnormal scenarios, reduces task waiting latency and computing power usage costs, and improves the long-term operational stability and scheduling robustness of the system.
Smart Images

Figure CN122226727A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of cloud-edge collaborative scheduling and intelligent agent operation technology, and particularly relates to a cloud-edge collaborative scheduling and intelligent agent operation method and system based on a distributed idle computing power network. Background Technology
[0002] With the rapid popularization of large-scale model inference and AI agent applications, the demand for heterogeneous GPU computing power continues to surge. However, city-level distributed idle heterogeneous GPU computing power resources have not yet been efficiently integrated and utilized. Existing computing power networks lack unified heterogeneous computing power standardization and a full-dimensional node feature collection system. The accuracy of node load prediction is insufficient, making it difficult to accurately match the diverse needs of AI agents and large-scale model tasks. This results in low matching degree of computing power resource scheduling, low utilization rate of idle computing power, and an inability to form a large-scale, standardized distributed computing power supply capability.
[0003] Traditional computing power scheduling schemes lack the ability to perceive computing power fluctuations and dynamically scale elastically. They also exhibit poor scheduling robustness under abnormal scenarios and have insufficient ability to perceive heterogeneous computation graph segmentation and cross-hardware operator compatibility for large-scale model inference. Furthermore, existing systems lack an integrated automated operation and maintenance and fault self-healing system. Node anomaly detection is lagging, fault prediction capabilities are weak, task migration is easily interrupted when a fault occurs, and handling procedures are cumbersome. Model version updates lack canary release and automatic rollback mechanisms, making it difficult to ensure the continuous and stable operation of the computing power network. This severely restricts the application of distributed idle computing power in large-scale model inference scenarios. Summary of the Invention
[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network, the method comprising: Construct a city-level distributed idle heterogeneous GPU computing power network, build a four-layer collaborative architecture, and complete automatic node registration, hardware reporting, and performance calibration; Construct a multi-dimensional feature system for edge nodes, collect full-dimensional indicators and achieve unified quantification of heterogeneous GPU computing power based on HCU, normalize heterogeneous GPU computing power by relying on architecture compatibility coefficient, construct a node health assessment model and complete time-series load prediction. Complete the six-tuple requirement modeling and parameter analysis for AI intelligent agents and large model tasks, construct a multi-dimensional comprehensive cost function, and make optimal node scheduling decisions based on four-layer scheduling logic; Implement heterogeneous perceptual computation graph segmentation, distributed parallel execution, and operator compatibility adaptation for large-model inference; Identify the tidal patterns of the computing power network, realize elastic scaling of computing power based on MADDPG, and enable waveform interference decentralized scheduling in abnormal scenarios; Deploy an automated operation and maintenance and self-healing system, realize real-time node anomaly detection through the isolated forest algorithm, complete early fault prediction by relying on the LSTM-Attention model, save the GPU context through the GSRP protocol when a fault is triggered, complete container hot migration by combining CRIU technology, and execute task evacuation, node isolation, maintenance work order generation and disposal actions according to hierarchical alarm rules. The central cloud synchronously completes the gray release and automatic rollback of the model version.
[0005] Furthermore, embodiments of the present invention also provide a cloud-edge collaborative scheduling and intelligent agent operation system based on a distributed idle computing power network, comprising: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the aforementioned cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network by executing the machine-executable instructions.
[0006] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, the processor of a computer device reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the computer device to execute the above-described cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network.
[0007] Based on the above, the system can efficiently integrate city-level distributed idle heterogeneous GPU computing resources. Through a unified heterogeneous computing power standard and a full-dimensional node feature collection system, it accurately matches the diverse computing power needs of AI agents and large-scale model inference tasks, significantly improving the scheduling matching degree and utilization rate of idle computing power resources. This enables large-scale, standardized distributed computing power supply, greatly reducing task latency and computing power usage costs for large-scale model inference. The system possesses comprehensive computing power tide awareness and dynamic elastic scaling capabilities, which can adjust computing power allocation in real time according to node load and task requirements. Even in abnormal scenarios, it can maintain efficient and stable inference services, effectively improving the overall operating efficiency and service capabilities of the heterogeneous computing power network.
[0008] The system constructs an integrated automated operation and maintenance and fault self-healing system. Through real-time node anomaly detection and early fault prediction, it enables rapid discovery and proactive handling of computing node faults. Combined with smooth task migration and uninterrupted switching mechanisms, it ensures that inference tasks are not interrupted and services are not slowed down when faults occur. At the same time, it supports gray release and automatic rollback of model versions, ensuring the smooth and safe model update process. It significantly improves the long-term operational stability and scheduling robustness of the computing network, provides continuous and reliable underlying computing power support for large model inference and AI agent applications, and promotes the large-scale application of distributed idle computing power in intelligent inference scenarios. Attached Figure Description
[0009] Figure 1 This is a schematic diagram of the execution flow of the cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network provided in an embodiment of the present invention.
[0010] Figure 2 This is a schematic diagram of exemplary hardware and software components of the cloud-edge collaborative scheduling and intelligent agent operation system based on a distributed idle computing power network provided in this embodiment of the invention. Detailed Implementation
[0011] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating a cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network, according to an embodiment of the present invention. The following is a detailed description of this cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network.
[0012] Step S110: Construct a city-level distributed idle heterogeneous GPU computing power network, build a four-layer collaborative architecture, and complete automatic node registration, hardware reporting, and performance calibration; A four-layer collaborative architecture comprising local nodes, edge nodes, regional scheduling nodes, and a central cloud is constructed, with lightweight operation and maintenance agent modules deployed on each idle heterogeneous GPU node. Upon startup, each node automatically sends a registration request to the regional scheduling node, carrying its unique hardware identifier and basic configuration information. The agent module collects and reports hardware parameters such as GPU architecture and memory capacity in real time. The regional scheduling node distributes a standard set of operators as benchmark testing tasks, calibrating the actual performance of the nodes through the execution results. This completes the registration, reporting, and calibration loop, forming a city-level distributed computing network that enables unified access and management of heterogeneous nodes.
[0013] Step S111: Construct a multi-dimensional feature vector for each edge node. Through three methods—real-time device monitoring, hardware information reporting, and benchmark dynamic calibration—comprehensively collect all dimensions of indicators, including GPU architecture generation, number of GPUs, single-card memory capacity, memory bandwidth, FP32 floating-point computing power, INT8 fixed-point computing power, available storage space, network bandwidth, network latency, overall power consumption, GPU temperature, and GPU utilization, to complete the real-time collection and standardization of feature data.
[0014] The GPU driver interface is used to collect core metrics such as GPU temperature and utilization. Network monitoring tools are used to obtain bandwidth and latency data. The storage management module collects available storage space, and power consumption sensors collect overall system power consumption. Continuous metrics are mapped to the [0,1] interval using minimum-maximum normalization. Discrete features such as process liveness status are converted into binary identifiers. Abnormal spikes in collected data are smoothed and filtered to eliminate format conflicts caused by hardware differences, generating standardized multi-dimensional feature vectors. This provides a unified data foundation for subsequent computational power optimization and health assessment.
[0015] Step S112: Define the HCU heterogeneous computing power unit as a unified computing power quantification standard, clarify the baseline specifications of 1 HCU corresponding to 1 INT8TOPS, 4GB video memory, 500-1000M effective bandwidth and architecture compatibility coefficient of 1.0, pre-build the operator performance matrix of GPUs from different manufacturers and generations and calculate the corresponding architecture compatibility coefficient, combine the weighted calculation rules of INT8 computing power, video memory and effective bandwidth, complete the computing power quantification and normalization calculation of heterogeneous GPUs from different manufacturers and generations, and generate a unified computing power value for nodes; The baseline specifications for 1 HCU are defined as follows: 1 TOPS of INT8 computing power (based on ResNet-50 operator testing), 4GB GDDR6 video memory, 500-1000M effective network bandwidth, and a baseline architecture compatibility coefficient of 1.0. A test set of 100 commonly used AI operators is constructed and executed on GPUs from different manufacturers (NVIDIA, AMD) and generations (RTX 30 series, 40 series). Operator completion success rate and efficiency are scored, and an architecture compatibility coefficient in the range of 0.5-1.0 is calculated. In the weighted calculation, INT8 computing power accounts for 0.4, video memory 0.3, bandwidth 0.2, and the architecture compatibility coefficient 0.1. This rule is used to normalize the computing power of heterogeneous GPUs and output a unified HCU value.
[0016] Step S113: Based on the node's current GPU utilization, GPU temperature, continuous normal operation time, historical task completion reputation score, and node operation stability data, set weight parameters for each dimension, generate a node health score through weighted calculation, and complete the construction and real-time scoring of the node health assessment model. GPU utilization 0.3, GPU temperature 0.2, continuous uptime 0.15, historical task completion reputation score 0.2, and node operational stability 0.15. The reputation score is calculated based on the on-time task completion rate and number of fault-free executions over the past 30 days; stability is assessed based on the frequency of fault occurrences and recovery time over the past 7 days. Data from each dimension is collected in real-time and standardized, then weighted and summed to generate a health score of 0-100. Nodes scoring 80 or above are considered high-quality, while those below 60 are considered low-quality. The score is updated every 30 seconds to dynamically reflect the node's operational status.
[0017] Step S114: Use the TFT time-series fusion Transformer model and perform periodic explicit modeling to integrate load time-series characteristics with daily and weekly periodic patterns to achieve high-precision prediction of the future load state of nodes.
[0018] Time-series data such as GPU utilization and computing load are collected in 5-minute increments, and meta-features such as hourly intervals (0-23), weekdays (1-7), and holidays (1 / 0) are extracted. Missing 5-minute data points are imputed using linear interpolation, and Kalman filtering is used to smooth out anomalies in GPU utilization exceeding 50%. A normalized sample set is then constructed. A lightweight TFT model is built, containing a periodic embedding layer, two attention heads, and a 128-dimensional hidden layer. Periodic features and load features are concatenated and input into the model. A joint loss function of mean squared error and periodic consistency is used, trained with an Adam adaptive optimizer, and the model is incrementally updated hourly, outputting load values and fluctuation trends for the next 12 and 24 hours.
[0019] Step S1141: Continuously collect historical time-series data on edge node GPU utilization, computing load, task queue length, and network bandwidth at a fixed time granularity. Simultaneously extract hourly, day of the week, holiday, and business cycle time meta-features. Use linear interpolation to fill in missing values, and use Kalman filtering to smooth abnormal spikes / drops in data. Complete data normalization and feature alignment to build a high-quality load time-series sample set. With a fixed time granularity of 5 minutes, the edge node operation and maintenance agent module continuously collects time-series data on GPU utilization, computing load, task queue length, and network bandwidth. Simultaneously, it extracts time-series features based on hourly intervals (0-23), days of the week (1-7), holidays (0 / 1 identifier), and business cycles (preset to levels 1-5 according to industry scenarios). Linear interpolation is used to complete data with no more than 3 consecutive missing time points. Kalman filtering is used to smooth abnormal data with GPU utilization spikes exceeding 50% or bandwidth drops exceeding 30%. Z-score normalization is performed, and load features and meta-features are aligned by timestamps. A high-quality load time-series sample set is constructed, with each sample containing 60 time-series steps and 4 meta-features.
[0020] Step S1142: The daily cycle, weekly cycle, and hourly cycle are taken as explicit discrete features and converted into low-dimensional dense vectors through the cycle embedding layer. They are then concatenated and fused with the load time series features and node hardware status features to form a model input tensor containing explicit cycle information. The daily (24-hour), weekly (7-day), and hourly (60-minute) cycles are used as explicit discrete features and input into a 64-dimensional periodic embedding layer. These features are then transformed into low-dimensional dense vectors through linear transformation and activation functions. Simultaneously, node hardware state features such as GPU memory capacity, architecture generation, and network bandwidth limits are extracted and standardized into 32-dimensional vectors. The periodic embedding vector (64-dimensional), load temporal feature vector (60-dimensional), and hardware state feature vector (32-dimensional) are sequentially concatenated and fused to form a 156-dimensional model input tensor. This ensures that the tensor simultaneously contains temporal dependencies, periodic patterns, and hardware adaptation information.
[0021] Step S1143: Initialize a lightweight TFT timing fusion Transformer model adapted to edge node computing power constraints, configure gated loop units, multi-head attention mechanisms, and gated residual connection structures, send the model input tensor into the model, capture local short-term load dependencies through the gated loop units, and focus on long-term correlations and daily / weekly cycle patterns through the multi-head attention mechanism; A lightweight TFT temporal fusion Transformer model is initialized, with the number of parameters limited to no more than 1 million to accommodate edge node computing power constraints. The model is configured with a single hidden layer gated recurrent unit (hidden layer dimension 128), a multi-head attention mechanism with two attention heads, and a gated residual connection structure with batch normalization. After the input tensor is fed into the model, the gated recurrent unit dynamically filters local short-term load features through forget gates and update gates to capture load dependencies between adjacent 5-10 minutes. The multi-head attention mechanism focuses on daily and weekly cycle patterns respectively, and strengthens the influence of key cycle features such as peak periods and holidays through weight allocation, achieving the collaborative capture of temporal dependencies and cycle patterns.
[0022] Step S1144: Construct a joint loss function that includes mean squared error loss and periodic consistency loss, use an adaptive optimizer to complete model training, dynamically adjust the learning rate, number of attention heads, and hidden layer dimension hyperparameters based on the validation set, and reduce prediction error and overfitting risk through iterative optimization; A joint loss function is constructed, where the mean squared error loss accounts for 0.7% to optimize the accuracy of load numerical prediction, and the periodic consistency loss accounts for 0.3% to constrain the consistency of the prediction results with the historical periodic patterns. The AdamW adaptive optimizer is used, with an initial learning rate of 0.001. Hyperparameters are dynamically adjusted based on the validation set (20% of the total samples): the learning rate adaptively decays according to the validation set loss, the number of attention heads is adjusted between 2 and 4, and the hidden layer dimension can be selected as 64 or 128. During training, a Dropout layer with a probability of 0.1 and L2 regularization are introduced. Iterative training stops when the validation set loss shows no decrease for 5 consecutive rounds, effectively reducing prediction error and the risk of overfitting.
[0023] Step S1145: Incrementally collect the latest load time series data and periodic features on an hourly basis, perform lightweight incremental updates on the model, input the current time series and periodic features into the trained TFT model, output the load values and fluctuation trends for multiple future time periods, and complete the high-precision prediction of the future load status of the node.
[0024] Using the hour as the hourly time node, the latest hourly load time-series data and corresponding periodic features are incrementally collected to construct an incremental training sample set. A flexible weight fixation strategy is employed for lightweight incremental updates to the model, freezing the underlying parameters of the gated recurrent units and attention mechanism, and updating only the weights of the top fully connected layer, with each update taking less than 30 seconds. The latest time-series data and periodic features are input into the trained model, which outputs load values and ±5% fluctuation ranges for the next 6, 12, and 24 hours. A sliding window is used to smooth the prediction results, ensuring the output of load values and fluctuation trends for multiple future time periods.
[0025] Step S120: Construct a multi-dimensional feature system for edge nodes, collect full-dimensional indicators and realize unified quantification of heterogeneous GPU computing power based on HCU, normalize heterogeneous GPU computing power based on architecture compatibility coefficient, construct a node health assessment model and complete time-series load prediction. The system collects 12 comprehensive metrics across all dimensions, including GPU architecture, computing power, memory, temperature, and network bandwidth, through multi-source acquisition modules such as GPU driver interfaces, network monitoring tools, and power consumption sensors. Based on the HCU unified computing power standard, it calculates the architecture compatibility coefficient by combining the operator performance matrix of GPUs from different manufacturers, and normalizes the computing power of heterogeneous GPUs through weighted summation. A node health assessment model is constructed and outputs scores in real time by setting weights for dimensions such as GPU utilization (0.3), temperature (0.2), and runtime (0.15). The trained TFT model is deployed to regional scheduling nodes, and combined with the health score and load prediction results, a multi-dimensional feature system for edge nodes, including computing power, health, and load status, is formed.
[0026] Step S130: Complete the six-tuple requirement modeling and parameter analysis of the AI agent and large model task, construct a multi-dimensional comprehensive cost function, and make the optimal node scheduling decision based on the four-layer scheduling logic; AI agents are deployed at regional scheduling nodes to complete the six-tuple requirement modeling and parameter parsing for large-scale model tasks: key task parameters are collected through the model parsing interface, personalized user requirements are received through the interactive interface, and converted into a standardized six-tuple requirement vector by the parsing module. Based on this vector and six cost dimensions such as latency and computing power matching, a lightweight multi-head attention mechanism is constructed as a multi-dimensional comprehensive cost function, dynamically allocating the weights of each dimension. Decisions are executed according to a four-layer scheduling logic: first, the local node is checked (execution if cost ≤ 0.3), then regional edge nodes are traversed (scheduling if cost ≤ 0.5), then the feasibility of task splitting is determined and distributed collaborative reasoning is executed, and finally, the data is uploaded to the central cloud as a fallback, ensuring the efficient implementation of the optimal node scheduling decision.
[0027] Step S131: Define a six-tuple dimension including task type, computing power requirement, latency threshold, accuracy level, security level, and data privacy level, and formulate the corresponding quantitative standards for each dimension; Define the core dimensions of the six-tuple and formulate clear quantitative standards: Task types are divided into real-time inference (label 0) and offline training (label 1); computing power requirements are in units of HCU, divided into gradients of 10, 20, 50, and 100 HCU; latency thresholds are divided into ≤50ms (level 1), 50-200ms (level 2), and >200ms (level 3); accuracy levels are defined as ≥99% (level 1), 95%-99% (level 2), and <95% (level 3); security levels are divided into symmetric encryption (level 1), asymmetric encryption (level 2), and national cryptographic algorithms (SM4, level 3); data privacy levels are divided into public (level 1), internal (level 2), and confidential (level 3). The quantitative standards of each dimension are uniformly adapted to subsequent requirement analysis and scheduling constraint verification.
[0028] Step S132: In the task submission stage, key task parameters and user-specific requirement parameters are collected. The parameters are standardized and converted through the parsing module to generate a structured six-tuple requirement vector, thus completing requirement modeling and parameter parsing. During the task submission phase, the system automatically collects key inherent parameters of the task through the model file parsing interface, including model type (Transformer / CNN, etc.), input data volume (graded by GB), and operator complexity (scored by the number and type of operators). It also receives personalized requirements parameters configured by the user through a visual interface, including latency limits, accuracy requirements, security encryption methods, and data privacy levels. The parameter parsing module employs natural language processing technology combining keyword matching and a rule engine to map descriptions such as "low latency" and "high accuracy" to corresponding quantization intervals. Discrete configuration information is directly converted into standard labels, and after standardization, a 64-dimensional structured six-tuple requirement vector is generated, fully representing the task's constraints, preferences, and core demands.
[0029] Step S133: Construct a task query vector based on six cost dimension features and six-tuple requirements. Adopt a lightweight multi-head attention mechanism. The model autonomously identifies task constraints and dynamically and differentially integrates each cost dimension, decouples the coupling relationship of indicators, and outputs the globally optimal comprehensive evaluation result. It has the ability to adapt to all scenarios and dynamically adapt to heterogeneous GPU nodes. The six cost dimensions—latency, computing power matching, health, model adaptation, security matching, and data transmission—are mapped to the [0,1] interval using Min-Max normalization. Discrete features in the six-tuple requirements are converted into low-dimensional vectors using 32-dimensional embedding encoding, while continuous features are normalized using Z-score. The two preprocessed features are cross-fused and a 128-dimensional task query vector is generated through a two-layer fully connected network. This vector is then input into a lightweight multi-head attention network with two attention heads. The network generates initial weights through similarity calculation, strengthens the weights of core constraint dimensions through a gating adjustment unit, amplifies latency weights for real-time inference tasks, and enhances security matching weights for security tasks. Parallel computation and feature concatenation decouple the indices, outputting a globally optimal comprehensive evaluation result.
[0030] Step S1331: Standardize the six cost dimensions of latency, computing power matching, health, model adaptation, security matching, and data transmission to eliminate dimensional differences and map the values of each dimension to a unified range; at the same time, embed the discrete features in the six-tuple requirements into low-dimensional dense vectors and normalize the continuous features in the six-tuple requirements to complete the structured preprocessing of the two types of input features. The six-tuple requirements include task type, computing power requirements, latency threshold, accuracy level, security level, and data privacy level. For the six cost dimensions—latency, computing power matching, health, model adaptation, security matching, and data transmission—Min-Max normalization is used to eliminate dimensional differences and uniformly map them to the [0,1] interval. Latency is normalized according to the task threshold, computing power matching is calculated based on the proportion of required HCUs, and health directly uses the score normalization result. Discrete features such as task type and security level in the six-tuple requirements are converted into low-dimensional dense vectors by one-hot encoding and input into a 32-dimensional embedding layer. Continuous features such as computing power requirements and latency thresholds are normalized to a standard normal distribution using Z-score normalization. The above process completes the structured preprocessing of the two types of input features, ensuring uniform feature format and comparable values.
[0031] Step S1332: The preprocessed six-tuple demand vector is used as a guide and cross-fused with the six cost dimension feature vectors. Feature mapping is performed through a shallow fully connected network to generate a task query vector that simultaneously contains task constraint preferences and basic cost assessment information. The preprocessed 64-dimensional six-tuple demand vector is used as the guiding feature and is element-wise multiplied and cross-fused with the [0,1] normalized vectors of the six cost dimensions (a total of six dimensions) to generate a 64-dimensional cross-feature vector. This vector is then input into a two-layer shallow fully connected network (128 dimensions in the first layer and 64 dimensions in the second layer), and the ReLU activation function is used for feature mapping. Batch normalization is used to suppress gradient vanishing, and the final output is a 64-dimensional task query vector. The first 32 dimensions of this vector represent task constraint preferences (such as latency requirements and security levels), and the last 32 dimensions represent basic information for cost assessment (such as the current node's computing power matching degree and health status).
[0032] Step S1333: Build a lightweight multi-head attention mechanism network, configure 2-4 attention heads to reduce computational overhead, and each attention head focuses on different types of task constraint dimensions. The initial attention weights are generated by calculating the similarity between the task query vector and the cost dimension feature vector. A lightweight multi-head attention mechanism network is constructed, configuring two attention heads to control computational overhead, with each attention head having no more than 200,000 parameters. The first attention head focuses on three performance-related constraints: latency cost, computational cost matching cost, and data transmission cost. The second attention head focuses on three reliability-related constraints: health, model adaptation, and security matching. The similarity between the task query vector and the feature vectors of each cost dimension is calculated using cosine similarity to generate an initial attention weight matrix. The weight values are mapped to the [0,1] interval, ensuring that each attention head can accurately capture the core requirements of its corresponding category of constraints.
[0033] Step S1334: Introduce a gating adjustment unit to optimize the initial attention weight, suppress redundant dimension interference and strengthen the weight ratio of constraint dimension. Real-time inference tasks automatically amplify the attention weight of the delay cost dimension, and security and confidentiality tasks automatically increase the weight allocation of the security matching cost dimension, so as to realize the dynamic differentiated fusion of each cost dimension. A gating unit with a Sigmoid activation function is introduced to optimize the initial attention weights: the gating unit receives the task type label from the six-tuple requirement. When the label is real-time inference, the weight of the latency cost dimension is automatically amplified (the weight coefficient is increased to 1.5 times the initial value); when the label is security-related, the weight of the security matching cost dimension is increased to 1.5 times the initial value; when the label is offline training, the weight of the computing power matching cost dimension is strengthened. Simultaneously, for redundant dimensions with a correlation of less than 0.1 with the task constraints, the gating unit attenuates their weights to 0.5 times the initial value, achieving a dynamic and differentiated fusion of core constraint strengthening and redundant information suppression.
[0034] Step S1335: By parallel computation of multi-head attention and feature concatenation, the implicit coupling relationship between different cost dimensions is decoupled. Then, the multi-attention head output is integrated through a linear projection layer to generate a globally optimal comprehensive evaluation result. A labeled dataset containing different task types such as real-time inference, offline training, and security-related tasks is constructed. The dataset covers task query vectors, cost dimension features, and corresponding optimal evaluation labels. The mean squared error loss function combined with task type adaptive regularization is used to train and optimize the model. The model size is compressed through quantization training and pruning to adapt to the computing power constraints of edge scheduling nodes. Feature weights are computed in parallel using two attention heads, and the output features are concatenated into a 128-dimensional vector along the channel dimension. Feature rearrangement and dimensional transformation decouple the implicit coupling between different cost dimensions, avoiding excessive influence of single-dimensional anomalies on the overall evaluation. The concatenated features are input into a 64-dimensional linear projection layer, and an activation function outputs the globally optimal comprehensive evaluation result in the 0-1 interval. A labeled dataset containing 100,000 samples is constructed (50,000 for real-time inference, 30,000 for offline training, and 20,000 for security-related data). The model is trained using a mean squared error loss function combined with a task-type adaptive regularization term (regularization coefficient 0.001). INT8 quantization training and structured pruning compress the model size to adapt to the computing power constraints of edge scheduling nodes.
[0035] Step S1336: After the model is deployed, it receives the six-tuple requirements of new tasks and the dynamic cost dimension features of the current node in real time. After structured preprocessing, it is input into the lightweight multi-head attention model to quickly complete the attention weight calculation and feature fusion and output the comprehensive evaluation result. At the same time, it continuously monitors the running status of node load fluctuations and health changes, updates the cost dimension features in real time, and the model automatically adjusts the weight allocation strategy of the attention heads to adapt to the dynamic running status of the nodes and complete the full-scene adaptive matching.
[0036] The quantized and pruned lightweight multi-head attention model is deployed to scheduling nodes in various regions, with the model working in real-time in conjunction with the node status monitoring module and the task requirement parsing module. Upon submission of a new task, the model receives the standardized six-tuple requirement vector and the current node's dynamic cost dimension features in real time, completing preprocessing and inference within 10ms and outputting a comprehensive evaluation result. Simultaneously, node load fluctuations and health changes are collected every 500ms, and the cost dimension feature vector is updated in real-time. The model automatically adjusts the attention head weight allocation strategy: strengthening the computing power matching weight when node load surges and increasing the health weight when health decreases, ensuring adaptive matching across all scenarios under the dynamic operating conditions of heterogeneous GPU nodes.
[0037] Step S134: Select suitable nodes step by step according to the four-layer scheduling logic, check local nodes and edge nodes in the region in turn, determine the feasibility of task splitting and execute distributed collaborative reasoning or cloud scheduling, and finally determine the optimal scheduling path and target node. The first step is to calculate the comprehensive cost of the local node. If the cost is ≤0.3 and all constraints such as computing power ≥ requirement, latency ≤ threshold, and security level meet the requirements are met, local execution is triggered directly. The second step is to traverse all edge nodes in the region, calculate the comprehensive cost of each node, and select the node with the lowest cost ≤0.5 to perform regional scheduling. The third step is to determine the task's decompositionability based on the model layer independence and data dependency characteristics. After decomposition, the cross-node collaboration cost (including communication overhead) is calculated, and the node combination with the lowest collaboration cost is selected to perform distributed collaborative inference. The fourth step is to upload the task to the central cloud if the task cannot be decomposable or the collaboration cost is >0.7. The comprehensive cost including data upload latency is calculated. If the requirements are met, cloud scheduling is performed, and the optimal scheduling path and target node are finally output.
[0038] Step S135: During the scheduling process, the node status and task progress are synchronized in real time, and the calculation results of the comprehensive cost function are dynamically updated to ensure the real-time performance and accuracy of the scheduling decision.
[0039] During scheduling and execution, node heartbeat packets (sent every 1 second) are used to synchronize real-time status information such as GPU utilization, memory usage, health score, and available computing power for each node. Task execution progress is obtained through a task progress feedback interface (updated every 500ms). Based on the synchronized data, the calculation results for latency cost, computing power matching cost, and other dimensions are dynamically updated, and the latest comprehensive cost function value is obtained through reweighting. If a node experiences a sudden increase in load (exceeding 85%) or a decrease in health (below 60 points), the node's suitability is immediately reassessed, and a scheduling path switch is triggered if necessary. This ensures that scheduling decisions always align with the real-time status of nodes and task progress, guaranteeing the real-time performance and accuracy of the decisions.
[0040] Step S140: Implement heterogeneous perceptual computation graph segmentation, distributed parallel execution, and operator compatibility adaptation for large model inference; A three-layer collaborative architecture consisting of edge nodes, regional scheduling, and a central cloud is constructed. For large-scale model inference tasks, the hardware characteristics and status data of heterogeneous GPU nodes are first obtained through a global computing power awareness module. Then, a standardized parsing tool is used to decompose the native computation graph and convert it into a directed acyclic graph (DAG). The computation graph is intelligently partitioned based on a dynamic programming algorithm. After the subgraph partitions are distributed to the adapted heterogeneous nodes, an asynchronous pipeline parallel execution mechanism is initiated. Cross-vendor GPU compatibility is achieved through operator transformation, precision adaptation, and kernel replacement. The node status and communication quality are monitored in real time, and the execution strategy is dynamically adjusted. In case of anomalies, subgraph migration is triggered. Finally, the computation results are aggregated in an orderly manner to ensure the efficient and stable operation of large-scale model inference in a heterogeneous environment.
[0041] Step S141: After receiving the large model inference task, collect the hardware parameters, available video memory, computing power specifications, network bandwidth and inter-node communication latency of the heterogeneous GPU nodes in the entire network in real time to complete the global perception and status archiving of heterogeneous computing power resources; simultaneously parse the original computation graph of the large model, decompose the model level, operator type, data dependency relationship and tensor flow, convert the computation graph into a standardized directed acyclic graph, and complete the structured parsing of the computation graph; After receiving large model inference tasks, lightweight acquisition agents deployed on each node collect hardware parameters (architecture generation, number of cores), available video memory, FP32 / INT8 computing power specifications, network bandwidth, and inter-node communication latency of heterogeneous GPU nodes once per second in real time. This data is then aggregated at the regional scheduling node for status archiving and synchronized to the central cloud. Simultaneously, the ONNX Parser tool is invoked to parse the original computation graph of the large model, deconstructing the model's hierarchical structure, operator types (convolution, fully connected, etc.), tensor data dependencies and flows. The computation graph is then converted into a standardized directed acyclic graph using Protobuf format, clarifying the data interaction interfaces between nodes and completing the structured parsing of the computation graph.
[0042] Step S142: Based on the parsed list of operators and heterogeneous computing resources, perform a pre-computation process, benchmark the execution time of each type of operator on GPUs from different manufacturers and models, calculate the memory usage and computational complexity of the operators, and at the same time measure the communication cost and transmission latency of data transmission between nodes, and construct an operator, hardware performance mapping library and communication cost matrix. Based on the operator list obtained from the analysis and heterogeneous computing resources, a test set was constructed using core operators from typical models such as ResNet-50 and BERT. This set was executed independently 100 times on GPUs from different manufacturers (NVIDIA, AMD) and of different models (RTX30 / 40 series, MI250), and the average execution time was taken as the operator execution time. Simultaneously, the memory usage and computational complexity of a single operator were statistically analyzed. By transferring tensors of different sizes (100MB-10GB) between nodes, data transmission latency and bandwidth loss were measured to generate communication costs. An operator-hardware performance mapping library was constructed with the key "operator type + GPU model" and the value being execution time / memory usage, along with a communication cost matrix with the dimension "number of nodes × number of nodes," providing a quantitative basis for the partitioning decision.
[0043] Step S143: With the optimization objectives of minimizing end-to-end inference latency, minimizing communication costs, and balancing node computing power load, and combining the computing power upper limit and memory capacity constraints of each heterogeneous node, a dynamic programming algorithm is used to traverse all feasible partitioning schemes of the computation graph, select the optimal partitioning point, and split the complete computation graph into multiple independent subgraph partitions to ensure that each partition is adapted to the heterogeneous computing power characteristics of the target node and minimizes data dependency and communication overhead between partitions. The optimization weights for end-to-end inference latency, communication cost, and load balancing are set to 0.4, 0.3, and 0.3, respectively. Combined with the upper limit of computing power (HCU threshold) for heterogeneous nodes and the constraint of GPU memory capacity (not exceeding 80% of available GPU memory), a dynamic programming algorithm is used to traverse all feasible partitioning schemes of the computation graph. The optimal partitioning point is selected with the objective of maximizing the difference between subgraph computation cost and communication cost. The complete computation graph is split into multiple independent subgraph partitions, each containing continuous model levels and operator combinations. This ensures that the computational power requirements of each partition match the HCU value of the target node, and that the number of data dependency edges between partitions is reduced by more than 30%, minimizing cross-node communication overhead.
[0044] Step S144: Distribute the optimally partitioned subgraph to the corresponding heterogeneous GPU nodes, establish a partition execution dependency timing table, start the asynchronous pipeline execution mechanism, so that the computation tasks and data transmission tasks of adjacent partitions overlap and run, and use pipeline parallelism to hide the communication delay between nodes to realize multi-node distributed parallel inference. The optimally partitioned subgraph is distributed to the corresponding heterogeneous GPU nodes via a high-speed RPC protocol. A partition execution dependency timing table is constructed based on the computation graph data dependencies, clearly defining the startup order and data interaction nodes for each partition. An asynchronous pipeline execution mechanism is initiated; when the current subgraph partition computation reaches 70%, the data source transmission for the next dependent subgraph is triggered, allowing computation tasks and data transmission tasks to overlap. The pipeline parallelism hides inter-node communication latency. Each node deploys a task scheduler, executing subgraph computations in an orderly manner according to the timing table and providing real-time progress feedback. This achieves multi-node distributed parallel inference, improving overall inference throughput.
[0045] Step S145: During the partitioning process, the compatibility between the operator and the target GPU hardware is detected in real time. For operators with incompatible architecture or mismatched precision, operator conversion, precision adaptation, and kernel replacement are automatically performed for compatibility processing. For extremely incompatible operators that cannot be converted, the CPU rollback execution mechanism is automatically triggered. Before subgraph partitioning, the kernel check interface provided by the GPU manufacturer is invoked to detect the architectural compatibility and accuracy support of operators with the target GPU hardware in real time. For incompatible operators, the ONNX Runtime conversion tool is used to convert them into equivalent operators supported by the target GPU; for operators with mismatched accuracy, FP32 and FP16 / INT8 accuracy adaptation is automatically completed; for complex operators that fail to convert, they are replaced with vendor-optimized kernels. If the above methods fail to achieve compatibility, the CPU rollback execution mechanism is automatically triggered to allocate the operator to the node's CPU core for execution, ensuring uninterrupted inference flow.
[0046] Step S146: During the inference run, continuously monitor the load status, partition execution progress and communication link quality of each node. If node computing power fluctuations or communication blockages occur, dynamically fine-tune the pipeline timing or recalculate the local partitioning strategy to achieve rapid migration of subgraph partitions for abnormal nodes. During inference execution, GPU utilization, memory usage, and task execution progress data are collected every 500ms via a node monitoring agent, while communication link bandwidth and latency are monitored in real time via a network monitoring module. When a node's computing power fluctuates by more than 20% (load exceeding 85% or falling below 30%) or communication latency exceeds 100ms, the regional scheduling node dynamically fine-tunes the pipeline timing offset or recalculates the local partitioning strategy only for the subgraphs related to the abnormal node. If a node's health score falls below 60, a rapid subgraph partition migration is immediately initiated, migrating unfinished computation tasks to surrounding healthy nodes. Incremental data transmission is used during the migration process to reduce overhead and ensure inference continuity.
[0047] Step S147: After all heterogeneous nodes have completed the inference calculation of their corresponding subgraph partitions, the intermediate calculation results of each partition are aggregated in an orderly manner according to the data dependencies of the original computation graph to complete the final inference output. At the same time, the node computing power resources are recovered and the operators and hardware performance mapping library are updated.
[0048] After all heterogeneous nodes have completed the inference computation for their respective subgraph partitions, the regional scheduling nodes, following the topological order of the original computation graph, systematically aggregate the intermediate computation results from each partition via a high-speed intranet channel, employing the LZ4 compression algorithm to reduce data transmission overhead. The final results are integrated and post-processed at the central cloud node, generating the large model inference output. Simultaneously, the GPU memory, computing resources, and network connections occupied by each node are released, and the subgraph execution process is shut down. Data such as the actual execution time and communication cost of the operators in this inference operation are updated to the operator-hardware performance mapping library, providing more accurate quantitative support for subsequent task partitioning and scheduling.
[0049] Step S150: Identify the tidal pattern of the computing power network, realize elastic scaling of computing power based on MADDPG, and enable waveform interference decentralized scheduling in abnormal scenarios; A comprehensive scheduling system encompassing perception, decision-making, execution, and self-healing is constructed. First, multi-dimensional time-series data is collected and preprocessed. Then, STL decomposition, Fourier transform, and K-means clustering are used to identify tidal patterns in the computing power network, building a pattern library covering peak, off-peak, low-peak, and sudden scenarios. A multi-agent scheduling system is built based on the MADDPG algorithm, where each regional scheduling node acts as an independent agent, dynamically executing computing power expansion, contraction, or maintenance actions based on tidal pattern prediction results. When abnormal scenarios such as node failure or task surges are detected, the system automatically switches to a waveform interferometric decentralized scheduling mode. Global load balancing is achieved through local agent collaboration. After the anomaly is mitigated, normal scheduling resumes and training data is updated.
[0050] Step S151: Collect and preprocess multi-dimensional computing load time series data, divide basic and burst-type tidal patterns through time series decomposition, frequency analysis and clustering, build a full-scenario tidal pattern library, match the current pattern in real time and output load change trend prediction; Continuously collect time-series data on GPU utilization, computing load, task queue length, and network bandwidth across the entire network at the minute-level granularity, and simultaneously extract meta-features such as hourly segments, days of the week, holidays, and business cycles. Linear interpolation is used to complete missing data, and the 3σ criterion is used to filter out abnormal peaks, completing data standardization preprocessing. The STL algorithm is used to decompose the load data into trend, periodic, and random components. Fourier transform is used to identify daily / weekly periodic features, and K-means clustering is used to classify three basic patterns: peak, off-peak, and trough. Based on business scenarios, a sudden tidal sub-pattern is added, constructing a full-scenario tidal pattern library. Real-time collection of current load data is performed, and the cosine similarity algorithm is used to match it with the pattern library, outputting pattern labels and load change trend predictions for the next 1-6 hours.
[0051] Step S152: Construct a multi-agent scheduling system based on MADDPG, define the agent's state, action and reward function, and after offline training, each agent dynamically executes computing power expansion, contraction or maintenance actions in combination with the global tidal mode to realize elastic scaling of computing power. Each regional scheduling node is encapsulated as a MADDPG agent, and an architecture of centralized training platform and decentralized execution terminal is built. Agents share global tidal patterns, resource status, and task information through low-latency communication links. A standardized state space containing multi-dimensional state parameters is constructed, and three types of actions—expansion, contraction, and maintenance of computing power—and quantification standards are standardized. A multi-objective joint reward function is designed. A training dataset is built based on historical tidal data and scheduling records. The model is trained using a centralized commentator and distributed actor model, and deployed to each agent terminal after training. Agents perceive state and tidal patterns in real time, output optimal scheduling actions, realize elastic scaling of computing power, and periodically incrementally update the model to optimize performance.
[0052] Step S1521: Each regional scheduling node in the entire network is independently encapsulated as a MADDPG agent, and a system framework of centralized training and decentralized execution is built. Point-to-point low-latency communication links are established between agents to realize real-time sharing and collaborative perception of global tidal mode labels, computing resource status, and task queuing information. The scheduling nodes in each region of the network are independently encapsulated as MADDPG agents using Docker containers, building an architecture of "central cloud training platform + edge agent execution". Point-to-point low-latency communication links are established between agents based on the gRPC protocol, with communication latency controlled within 10ms. Each agent collects the local region's computing resource status (available HCU, GPU memory utilization) and task queue length in real time, and synchronously receives global tidal mode tags pushed by the central cloud. Through information synchronization every 500ms, global situational awareness is achieved among agents, ensuring that each agent can obtain detailed local data and grasp key scheduling information across the entire network, providing data support for collaborative decision-making.
[0053] Step S1522: Construct a standardized state space that integrates global tidal mode types, available computing power of regional nodes, memory utilization, average task waiting time, node health score, and network transmission bandwidth. Standardize the action space that includes three types of actions: computing power expansion, computing power reduction, and maintaining the current computing power. Clarify the sub-actions and corresponding quantitative standards for computing power expansion and reduction. Design a multi-objective joint reward function with the goals of optimal resource utilization, minimum task latency, and minimum node energy consumption. Give positive rewards to scheduling behaviors that meet resource utilization and task latency requirements, and give negative rewards to scheduling behaviors that result in computing power overload, excessive task latency, and energy waste. A standardized state space is constructed, integrating six core parameters: global tidal pattern type (encoded as discrete values from 0 to 3), available computing power of regional nodes (HCU value), memory utilization, average task waiting time, node health score, and network transmission bandwidth, all normalized to the [0,1] interval. The standardized action space comprises three categories: computing power expansion, contraction, and maintenance. Expansion includes waking up dormant nodes (1-2 nodes each time) and adding heterogeneous nodes (1 node each time). Contraction includes taking offline idle nodes (≤30%) and putting low-load nodes to sleep (utilization <30%). A multi-objective joint reward function is designed: +10 points for resource utilization of 70%-85% and meeting latency targets, -10 points for idle / overloaded status, and -5 points for exceeding energy consumption limits, guiding the agent to learn the globally optimal strategy.
[0054] Step S1523: Collect historical computing power tidal pattern data, computing power scheduling execution records, and node running status data to construct a standardized training dataset and perform normalization preprocessing; initialize the actor network and critic network of the MADDPG model; and configure hyperparameters such as learning rate, batch size, and experience replay capacity. Data on computing power tidal patterns, computing power scheduling execution records, and node running status were collected over the past six months to construct a standardized training dataset containing 100,000 samples. Min-Max normalization was used to map continuous features to the [0,1] interval, and one-hot encoding was performed on discrete features. The actor network (3 fully connected layers, 128 hidden layers) and critic network (3 fully connected layers, 256 hidden layers) of the MADDPG model were initialized, and the core hyperparameters of initial learning rate 0.001, batch size 32, and experience replay capacity of 1 million records were configured. The training dataset was divided into training and validation sets in an 8:2 ratio to lay the data and model foundation for multi-agent joint training.
[0055] Step S1524: Start offline joint training of multiple agents using a training mode of centralized critics and distributed actors. The critic network obtains the state and action information of all agents and outputs a global value assessment. The actor network generates scheduling actions based on local perception information. Optimize network parameters through experience replay mechanism, simulate various tidal patterns and computing power scenarios to enhance the collaborative decision-making ability of agents, and iterate training until the model converges. A training model employing a "centralized critic + distributed actor" approach is adopted. The critic network is deployed on a central cloud training platform, receiving state, action, and reward data uploaded by all agents and outputting a global value assessment. The actor network is deployed on each agent's terminal, generating scheduling actions based solely on local perception information. Historical training data is randomly sampled using an experience replay mechanism to optimize network parameters. During training, peak, off-peak, trough, and sudden tidal patterns are simulated, constructing over 100 computing power scenarios to enhance the agents' collaborative decision-making capabilities and avoid local computing power imbalances. Iterative training stops when the validation set reward value fluctuates by ≤5% for 10 consecutive rounds, yielding converged model parameters.
[0056] Step S1525: Deploy the trained lightweight model to the intelligent agent terminals in each region, configure a real-time state perception module for each intelligent agent, and seamlessly connect with the computing power tidal pattern recognition module to obtain the current computing power network tidal pattern, load prediction results, and global resource status in real time. The trained model is trained using INT8 quantization and structured pruning, compressed to less than 50MB, and deployed to intelligent agent terminals in various regions. Each agent is equipped with a real-time state awareness module, which seamlessly integrates with the computing power tidal pattern recognition module via an API interface. This module collects local node availability data, memory utilization, and task wait times every second, and obtains global tidal pattern labels and load prediction results every minute, synchronously updating global resource status information to ensure the real-time nature and accuracy of agent input data, thus supporting precise scheduling.
[0057] Step S1526: The agent collects local state information and global tidal pattern features in real time, inputs the trained MADDPG model and outputs the optimal scheduling action. When expansion is triggered, it prioritizes waking up dormant nodes in the local area. When there are not enough nodes, it schedules heterogeneous GPU nodes from the global computing power resource pool. When shrinkage is triggered, it takes offline idle nodes or puts low-load nodes into dormancy according to node load priority and health score. When maintenance action is triggered, it maintains the current computing power configuration. Each agent executes the action independently and without conflict. The agent collects local state information and global tidal pattern characteristics in real time, inputs them into the trained MADDPG model, and outputs the optimal scheduling action within 10ms. When a scaling-up action is triggered, dormant heterogeneous GPU nodes within the local area are woken up first. If they fail to wake up within 30 seconds, suitable nodes are scheduled from the global computing power resource pool. When a scaling-down action is triggered, nodes are sorted by load priority (from low to high) and health score (from low to high), and completely idle nodes or nodes with a dormant load of <30% are taken offline. When a maintenance action is triggered, the current computing power configuration remains unchanged. Each agent executes without conflict through a global lock mechanism in the central cloud, ensuring that computing power dynamically adapts to tidal changes.
[0058] Step S1527: During the elastic scaling process, online scheduling data and operational performance data are continuously collected, and the MADDPG model is incrementally fine-tuned and optimized periodically. The reward function weights are dynamically adjusted according to the tidal pattern changes.
[0059] During the elastic scaling of computing power, online scheduling data (action type, execution effect) and operational performance data (resource utilization, task latency, energy consumption) are collected hourly to construct an incremental training dataset. Every week in the early morning, during off-peak hours, the MADDPG model is incrementally fine-tuned, freezing the underlying network parameters and updating only the weights of the top fully connected layers, with each fine-tuning session lasting less than 30 minutes. Simultaneously, the reward function weights are dynamically adjusted based on the tidal pattern changes (exceeding 20%), increasing resource utilization weights when resources are scarce and increasing task latency weights when tasks are intensive, continuously improving the agent's scheduling accuracy.
[0060] Step S153: Define abnormal scenarios and judgment rules, identify abnormalities and broadcast information through distributed detection modules, enable waveform interference decentralized scheduling mechanism, intelligent agents autonomously and collaboratively allocate tasks and adjust computing power, restore normal scheduling and update training dataset after the abnormality is relieved.
[0061] Three abnormal scenarios are predefined: sudden node failure (hardware failure / network interruption), surge in task load (exceeding the prediction limit by 30%), and network link congestion (latency exceeding the threshold by 2 times). Trigger thresholds and judgment rules are set for each scenario. A distributed anomaly detection module is deployed on each node to monitor hardware status and network connectivity in real time. Regional scheduling nodes aggregate data and calculate the rate of change of indicators using a sliding window algorithm. When an anomaly occurs, an alarm is triggered and an anomaly information is broadcast. A waveform interference decentralized scheduling mechanism is enabled. The agent calculates the load interference coefficient based on its local state and neighbor information, autonomously decides on task migration and computing power adjustment, uses incremental transmission to reduce overhead, and resumes normal scheduling and updates the training dataset after the anomaly is mitigated.
[0062] Step S160: Deploy an automated operation and maintenance and self-healing system. Real-time detection of node anomalies is achieved through the isolated forest algorithm. Fault prediction is completed in advance by relying on the LSTM-Attention model. When a fault is triggered, the GPU context is saved through the GSRP protocol. Container hot migration is completed by combining CRIU technology. Task evacuation, node isolation, and maintenance work order generation and handling actions are executed according to the hierarchical alarm rules. The central cloud synchronously completes the gray release and automatic rollback of the model version.
[0063] An automated operation and maintenance (O&M) and self-healing system covering edge, regional, and central cloud deployments are implemented. An isolated forest anomaly detection module is deployed on heterogeneous GPU nodes, and an LSTM-Attention fault prediction model is deployed on regional scheduling nodes. When a fault is triggered, the entire GPU context is saved in real time via the GSRP protocol, and container hot migration is achieved using CRIU technology. A tiered approach is taken based on the anomaly level: low-level anomalies are logged, medium-level anomalies involve evacuating non-core tasks, and high-level anomalies involve isolating nodes and generating repair work orders. The central cloud platform pushes model versions according to a canary release strategy, monitors operational metrics in real time, and automatically rolls back to the previous stable version when anomalies occur, forming a closed-loop process of "detection-prediction-self-healing-release" to ensure the continuous and stable operation of the computing network.
[0064] Step S161: Build a three-layer architecture foundation of edge agent, regional engine and central control, deploy lightweight operation and maintenance agent module on each heterogeneous GPU node, deploy self-healing execution engine on regional scheduling node, build a unified operation and maintenance management platform in the central cloud, open up the full-link data collection channel of node hardware, container operation, GPU status and task execution, build a standardized operation and maintenance data bus, and realize unified aggregation and collaborative management of the operation and maintenance status of the entire network. Lightweight operation and maintenance agents with a size of ≤10MB are deployed on each heterogeneous GPU node to handle data collection and command execution; self-healing execution engines are deployed on regional scheduling nodes to coordinate local fault handling; and a unified operation and maintenance management platform is built in the central cloud to achieve global monitoring. The gRPC protocol is used to establish a full-link data collection channel connecting node hardware, container operation, GPU status, and task execution, constructing a standardized operation and maintenance data bus based on Protobuf. The entire network's operation and maintenance status is synchronized every 500ms, achieving unified aggregation and collaborative management of node operation data across the entire domain.
[0065] Step S162: Construct a real-time node anomaly detection model based on the isolated forest algorithm, integrating multi-dimensional running features such as GPU temperature, memory usage, computing power utilization, network packet loss rate, disk I / O, process survival status, and instruction response latency. After normalizing and preprocessing the feature data, input it into the model, calculate the single-node anomaly score in real time, set a dynamic judgment threshold to filter data noise and instantaneous fluctuations, and mark the node as an abnormal state in milliseconds when the anomaly score continuously exceeds the threshold. A node anomaly detection model is built based on the Isolation Forest algorithm, integrating seven operational features such as GPU temperature and memory usage. The collected feature data undergoes min-max normalization preprocessing, mapping it to the 0-1 range. Anomaly thresholds are dynamically generated using a sliding time window, and single-node anomaly scores are calculated in real-time streaming. A continuous frame stabilization mechanism is implemented, marking a node as an anomaly and reporting it only when the score exceeds the threshold for multiple consecutive frames. The model is trained offline using historical normal and anomaly samples, with lightweight parameters configured to adapt to edge computing power. Detection results are continuously collected to incrementally optimize the model and threshold strategy, ensuring the accuracy and real-time performance of anomaly detection.
[0066] Step S1621: In the lightweight operation and maintenance agent module of each heterogeneous GPU node, a real-time feature acquisition component is built in. Multi-dimensional running feature data is collected at high frequency at the millisecond level. The accurate values of continuous features such as GPU temperature, memory usage, computing power utilization, network packet loss rate, disk I / O rate, and instruction response latency are obtained in real time. Discrete features of process survival status are monitored in real time and converted into binary identification data. All feature data are temporarily stored through a local cache queue to shield the acquisition format conflicts caused by hardware differences. A real-time feature acquisition component is built into the operation and maintenance agent module of each heterogeneous GPU node, collecting multi-dimensional operational features at a frequency of 10 milliseconds: obtaining precise values of continuous data such as temperature and memory usage through the GPU driver interface, and detecting process liveness status through process monitoring tools and converting it into 0 / 1 binary identifiers. All collected data is stored in a local circular cache queue with a capacity of 1000 entries. A unified data format interface is used to shield the acquisition format conflicts of hardware from different manufacturers such as NVIDIA and AMD, ensuring the continuity and consistency of data acquisition.
[0067] Step S1622: Map the continuous numerical features to a unified interval of 0 to 1 using the minimum-maximum normalization method, perform binarization encoding on the discrete features of the process survival status, use linear interpolation to complete the instantaneous missing data, perform preliminary smoothing filtering on the extreme mutation data of a single frame, and generate a standardized feature vector. For the collected continuous numerical features, a minimum-maximum normalization method is used to uniformly map data such as GPU temperature and computing power utilization to the 0-1 interval, eliminating dimensional differences. Discrete features of process survival status are maintained in a binary encoding format. For transiently missing data during acquisition, linear interpolation is used to fill in data at adjacent time points. For extreme abrupt changes in a single frame exceeding the normal fluctuation range by 30%, a moving average method is used for preliminary smoothing and filtering. Finally, a feature vector with uniform dimensions and standardized values is generated to meet the model input requirements.
[0068] Step S1623: Using historically accumulated normal node operation data and labeled abnormal sample data, train the isolated forest anomaly detection model offline, configure lightweight model parameters adapted to edge computing power, set the number of isolated trees and subsampling scale, so that the model learns the distribution pattern of normal node features, generates an anomaly detection model file that can be quickly reasoned, and distributes the trained lightweight model to each node operation and maintenance agent module. Data on normal node operation over the past six months, along with labeled anomaly samples, were collected to construct a training dataset containing 50,000 samples. When training the isolated forest anomaly detection model offline, 50 isolated trees and a 30% subsampling scale were configured, keeping the model parameters ≤500,000 to ensure lightweight characteristics are suitable for edge node computing power. During training, the model learns the distribution patterns of normal node features, generating a model file (≤5MB) capable of rapid inference. This file is then distributed in batches to the operation and maintenance agent modules of each node via the central cloud platform, enabling localized deployment and inference at the edge.
[0069] Step S1624: Dynamically generate anomaly judgment threshold based on sliding time window mechanism, continuously count the distribution of abnormal scores of normal nodes within the window, calculate the mean and standard deviation of scores, and automatically generate adaptive dynamic threshold. Anomaly detection thresholds are dynamically generated based on a 1-minute sliding time window mechanism, continuously tracking the distribution of anomalous scores among normal nodes within the window. With each new frame of data, the mean and standard deviation of the scores within the window are calculated in real time, and an adaptive dynamic threshold is automatically generated according to the rule of "mean + 2 × standard deviation". This threshold adjusts in real time according to changes in node operating status; it automatically increases when nodes are under high load and automatically decreases when under low load, effectively adapting to different operating scenarios and filtering out invalid judgments caused by data noise and instantaneous performance fluctuations at the source.
[0070] Step S1625: The preprocessed real-time feature vector is streamed into the locally deployed isolated forest model. The anomaly score of a single node is output in real time through random feature segmentation and path length calculation. The entire calculation process is completed by relying on the local computing power of the edge node. The preprocessed, standardized feature vectors are streamed into a locally deployed isolated forest model. The model constructs isolated trees by randomly selecting feature dimensions and segmentation thresholds, calculates the path length of a sample in each tree, and then aggregates these to generate a single-node anomaly score. A higher score indicates a greater anomaly risk. The entire inference process relies on the local CPU computing power of the edge nodes, without depending on central cloud resources. A single inference operation takes ≤10 microseconds, achieving microsecond-level rapid detection and ensuring real-time anomaly identification.
[0071] Step S1626: Establish a continuous frame judgment anti-shake mechanism, set a judgment window of fixed duration, and only when the node abnormal score exceeds the dynamic threshold for multiple consecutive frames is it identified as a real abnormality; A continuous frame detection anti-jitter mechanism is established, with a fixed-length detection window of 500 milliseconds containing 50 frames of feature data. After the model outputs an anomaly score for each frame in real time, it continuously counts the number of frames exceeding a dynamic threshold within the window. Only when the number of consecutive frames exceeding the threshold is ≥30 (60%) is it considered a genuine anomaly, avoiding misjudgments caused by non-fault factors such as instantaneous network jitter or peak computing power fluctuations. The anti-jitter mechanism can dynamically adjust the window length according to the node's operating scenario, extending to 1 second when the load is stable and shortening to 300 milliseconds when the load fluctuates.
[0072] Step S1627: After completing the real anomaly determination, the node operation and maintenance agent module marks the current node as an abnormal state within milliseconds, and synchronously reports the anomaly mark, anomaly score and corresponding multi-dimensional feature data to the regional self-healing execution engine and the central cloud operation and maintenance platform in real time. After completing the real anomaly determination, the node operation and maintenance agent module marks the current node as abnormal within 5 milliseconds and reports it in real time to the regional self-healing execution engine and the central cloud operation and maintenance platform via UDP protocol. The reported information includes the node's unique identifier, anomaly score (accurate to two decimal places), anomaly occurrence timestamp, and the original and standardized values of the seven-dimensional feature data at the corresponding time. This provides complete data support for the regional engine to perform fault handling and the central platform to analyze the cause of the anomaly, ensuring the targeted and efficient handling of the fault.
[0073] Step S1628: Continuously collect labeled anomaly detection results and actual node operating status data, periodically perform incremental fine-tuning and optimization of the isolated forest model, and synchronously update the dynamic threshold calculation strategy.
[0074] The central cloud operations and maintenance platform continuously collects anomaly detection results and actual operational status data from all network nodes, summarizing them daily to form a labeled dataset, with the proportion of anomaly samples controlled below 30%. Every week in the early morning, during off-peak hours, incremental fine-tuning and optimization of the Isolation Forest model is performed, using mini-batch gradient descent to update model parameters. Simultaneously, the dynamic threshold calculation rules are updated based on the latest data distribution, and the sliding window duration and standard deviation coefficient are adjusted to continuously improve the model's anomaly detection accuracy and robustness under different hardware models and load scenarios.
[0075] Step S163: Relying on the LSTM-Attention hybrid model to achieve early fault prediction, collect long-term historical runtime sequence data, historical fault records, and alarm information of nodes to construct a prediction sample set, extract long-term dependent features of time series data through the LSTM network, introduce an attention mechanism to focus on fault-related features, and the model outputs the probability of node fault occurrence and fault type in the future period, and classifies the warning level according to the predicted probability. This system leverages an LSTM-Attention hybrid model to achieve early fault prediction. It collects comprehensive historical runtime data, fault records, and alarm information from nodes to construct a prediction sample set. The LSTM network extracts long-term dependency features from the time-series data, and an attention mechanism focuses on strongly correlated fault features, outputting the probability and type of fault occurrence within the next 1-6 hours. Four warning levels—low, medium, high, and extremely high—are assigned based on probability, generating corresponding four-level warning signals which are pushed to the operations and maintenance platform and the self-healing engine. Real-world fault data is collected periodically to construct an incremental dataset, fine-tuning model parameters and attention weight allocation strategies to continuously improve prediction accuracy.
[0076] Step S1631: Continuously collect historical runtime time series data of nodes across all dimensions. The time series data includes GPU temperature, memory usage, computing power utilization, network packet loss rate, disk I / O rate, instruction response latency, and process status indicators. Synchronously collect historical fault records, alarm information, and maintenance record related data of nodes, unify the timestamp granularity of all data, complete the standardization alignment and persistent storage of multi-source heterogeneous data, and build a database of node running status and fault association across the entire domain. The system continuously collects full-dimensional runtime timing data from nodes at a 10-second granularity using a time-series database. This includes seven metrics such as GPU temperature and memory usage, and simultaneously aggregates related data such as historical fault records, alarm logs, and maintenance work orders. All data is unified to a second-level timestamp granularity, and a data alignment algorithm is used to standardize and align multi-source heterogeneous data. The data is persistently stored in a MySQL database, constructing a comprehensive database linking node runtime status and faults. This database supports multi-dimensional retrieval by node identifier, time range, and fault type, providing comprehensive data support for fault prediction models.
[0077] Step S1632: Construct a fault prediction sample set based on the associated database, set a fixed-length input time window and a future prediction time window, use the continuous operation features in the input window as model input, use whether a fault occurs in the prediction window and the specific fault type as sample labels, perform equalization processing on normal operation samples and fault samples, and digitally encode the fault type to form a supervised learning sample set. A fault prediction sample set was constructed based on an associated database, with a 1-hour input time window and a 6-hour future prediction time window. Continuous operational features within the input window were used as model input, while the occurrence and specific type of fault (e.g., hardware failure, network interruption) within the prediction window were used as sample labels. The SMOTE algorithm was employed to balance normal and fault samples, avoiding sample distribution skew. Eight fault types, including GPU overheating and memory overflow, were digitized from 0 to 7, ultimately forming a supervised learning sample set of 100,000 samples, which were divided into training and validation sets in an 8:2 ratio.
[0078] Step S1633: Apply Z-score standardization to continuous operation features, use time-series interpolation to fill in missing values in time-series data, apply wavelet transform filtering to abnormal noise points, and perform one-hot encoding conversion on the discrete features of alarm type and fault category to generate a time-series feature matrix. Continuous operational features in the sample set are standardized using Z-scores to convert them into standard normal distribution data, eliminating dimensional differences. Missing values in the time-series data are filled using time-series interpolation to ensure data continuity; abnormal noise points are filtered using wavelet transform to preserve data trend characteristics. Discrete features such as alarm types and fault categories are converted using one-hot encoding to generate fixed-dimensional encoding vectors. The processed continuous features and discrete encoding vectors are concatenated along the time dimension to generate a time-series feature matrix with uniform dimensions, meeting the model input format requirements.
[0079] Step S1634: Build an LSTM-Attention hybrid prediction model, input the temporal feature matrix into the LSTM network layer, and use the LSTM gating mechanism to mine the long-term dependency relationship and temporal change pattern of the data, and extract high-dimensional abstract temporal features; A hybrid LSTM-Attention prediction model was constructed, with the temporal feature matrix input into the LSTM network layers. The LSTM network was configured with two hidden layers, each containing 128 hidden units. Gating mechanisms such as forget gates, input gates, and output gates were used to uncover long-term dependencies and temporal variation patterns in node runtime data. The input temporal feature matrix was processed frame-by-frame to extract high-dimensional abstract temporal features, focusing on capturing potential feature patterns related to faults, such as sudden changes in GPU utilization and continuous temperature increases. This lays the foundation for the subsequent attention mechanism to focus on core features.
[0080] Step S1635: After the LSTM output layer, an attention mechanism module is connected to calculate the correlation weight between each temporal feature and the occurrence of the fault. High weights are assigned to strongly correlated features and low weights are assigned to weakly correlated features. Through a fully connected layer and an activation function, the probability of the fault occurring in the node and the classification result of the corresponding fault type are output. A Bahdanau-based attention mechanism module is connected after the LSTM output layer to calculate the association weight between each temporal feature and the occurrence of a fault. High weights are assigned to strongly correlated features such as sudden increases in GPU temperature and excessive memory usage, while low weights are assigned to weakly correlated features such as small fluctuations in network bandwidth. The weighted high-dimensional temporal features are then input into a fully connected layer, and the Softmax activation function is used to output the probability of node fault occurrence within a specified future time period, as well as the classification results for eight fault types. The probability output is accurate to three decimal places to ensure the accuracy and interpretability of fault prediction.
[0081] Step S1636: Configure the hyperparameters of adaptive learning rate, batch size, and number of hidden units in LSTM, use the multi-class cross-entropy loss function as the optimization objective, update the model parameters in combination with Adam optimizer, introduce early stopping strategy and regularization method to prevent model overfitting, use the validation set to evaluate the model prediction accuracy and recall, iterate and optimize until the model converges, obtain a lightweight high-precision fault prediction model, and deploy the model to the regional self-healing execution engine; The adaptive learning rate was initially set to 0.001, the batch size to 32, and the number of hidden units in the LSTM layer to 128. A multi-class cross-entropy loss function was used as the optimization objective, combined with the Adam optimizer to iteratively update the model parameters. An early stopping strategy was introduced: training was stopped when the validation set loss showed no decrease for five consecutive rounds. L2 regularization was also added to suppress overfitting. The model's prediction accuracy and recall were evaluated using the validation set, requiring accuracy ≥ 90% and recall ≥ 85%. Iterative optimization continued until model convergence, resulting in a lightweight model with a size ≤ 20MB, which was then deployed to a region-based self-healing execution engine.
[0082] Step S1637: In the online inference stage, the latest runtime sequence features of the node are collected in real time, and after preprocessing, they are streamed into the LSTM-Attention model. The model completes forward inference and outputs the probability of fault occurrence and the predicted fault type. According to the preset rules, the prediction results are divided into four warning levels: low probability, medium probability, high probability, and extremely high probability, and correspondingly generate four levels of signals: prompt warning, attention warning, danger warning, and emergency warning. During the online inference phase, the latest runtime sequence features of nodes are collected in real time at a 10-second granularity. After preprocessing such as Z-score standardization, missing value imputation, and noise filtering, the data is streamed into the LSTM-Attention model. The model completes forward inference within 50 milliseconds and outputs the probability of fault occurrence and the predicted fault type. Warning levels are categorized according to preset rules: probability <30% is low probability (advance warning), 30%-60% is medium probability (attention warning), 60%-90% is high probability (danger warning), and >90% is extremely high probability (emergency warning), generating four corresponding warning signals.
[0083] Step S1638: The early warning signal is synchronously pushed to the central cloud operation and maintenance platform and the regional self-healing engine. The early warning signal includes node identifier, predicted fault time, fault type, and risk probability. Real fault occurrence data and model prediction results are continuously collected to build an incremental training dataset. The LSTM-Attention model is fine-tuned and optimized regularly, and the attention weight allocation strategy is dynamically updated.
[0084] Early warning signals are synchronously pushed to the central cloud operations and maintenance platform and the regional self-healing engine via message queues. The warning information includes core information such as the node's unique identifier, predicted fault occurrence time, fault type, and risk probability. Real-world fault occurrence data and model prediction results are continuously collected, and incremental training datasets are built daily. The LSTM-Attention model is fine-tuned and optimized weekly. The attention weight allocation strategy is dynamically updated to strengthen the weights of features related to newly emerging fault types, continuously improving the model's accuracy and timeliness in predicting faults in complex scenarios.
[0085] Step S164: After detecting the node failure trigger signal, enable the GSRP protocol to perform GPU context security saving, capture the full context information of GPU computing kernel state, video memory tensor data, task execution breakpoints, and hardware mapping relationship, and use incremental snapshot technology for lightweight storage and synchronous backup to the regional redundant storage node. Upon detecting a node failure trigger signal (anomaly score ≥ 0.8 or a warning level of extremely high probability), the GSRP protocol is immediately activated to perform GPU context security saving. Full context information, including GPU computing kernel state, memory tensor data, task execution breakpoints, and hardware mapping relationships, is captured in real time. Incremental snapshot technology is used to store only the data differing from the previous snapshot, resulting in a lightweight storage context file (≤ 500MB). The snapshot is synchronously backed up to a regional redundant storage cluster (at least 3 copies) via a high-speed intranet, ensuring the integrity of the task context and providing data support for subsequent container recovery.
[0086] Step S165: Combine CRIU technology to complete the seamless hot migration of containers. Based on the GPU context saved by GSRP, CRIU is used to perform freeze and checkpoint operations on the task containers of the faulty node, generating a snapshot of the container running status and an image file. These are then transmitted to a pre-selected healthy node via a high-speed intranet channel, where the container state is restored and the GPU context is reloaded. Based on the GPU context saved by GSRP, the CRIU tool is invoked to freeze the task containers on the faulty node, generating a snapshot and image file of the container's running state, including process status, file handles, and network connections. The snapshot file is then transferred to a pre-selected healthy node (health score ≥80, idle computing power ≥50%) via a high-speed RDMA intranet channel (transmission rate ≥10Gbps). On the target node, the CRIU tool restores the container state and reloads the GPU context. The entire migration process takes ≤3 seconds, achieving seamless hot migration of containers between heterogeneous nodes and uninterrupted task execution.
[0087] Step S166: Perform automated handling according to the preset hierarchical alarm rules. Low-level anomalies are recorded in the operation and maintenance log and synchronized to the central platform; medium-level anomalies automatically evacuate non-core tasks of nodes to surrounding healthy nodes; high-level faults are executed by network isolation and computing resource shielding of nodes, and at the same time, standardized maintenance work orders with associated fault node information, detection results and prediction data are automatically generated and pushed to the operation and maintenance terminal. Low-level anomalies (anomaly score 0.6-0.8) only record operation and maintenance logs and synchronize them to the central platform, without affecting task execution; medium-level anomalies (score 0.8-0.9) automatically evacuate non-core tasks (priority ≤3) on the node to surrounding healthy nodes to ensure the stable operation of core tasks (priority ≥4); high-level faults (score ≥0.9 or hardware faults) immediately execute node network isolation (closing external communication ports) and computing resource shielding (removing from the scheduling list), while automatically generating standardized maintenance work orders, associating fault node identifiers, anomaly detection results, prediction data, and other information, and pushing them to the operation and maintenance terminal via SMS and platform messages.
[0088] Step S167: The central cloud operation and maintenance platform performs gray-scale release and automatic rollback management of model versions. It divides all network nodes into multiple gray-scale release units according to region, load, and hardware type, and pushes new version operation and maintenance models and scheduling models in batches. It monitors the stability of node operation, task inference accuracy, and computing power loss indicators in real time. When the indicators are abnormal, it triggers a fully automatic rollback mechanism to switch back to the previous stable version. The entire network is divided into multiple canary release units based on region, load (high / medium / low), and hardware type (NVIDIA / AMD), with each unit containing no more than 10% of the total nodes. The new version of the operation and maintenance model and scheduling model are pushed out incrementally in batches of 10%, 30%, and 60%, with real-time monitoring of node operational stability (failure rate ≤1%), task inference accuracy (error ≤5%), and computing power consumption (additional consumption ≤10%). If any indicator exceeds the threshold, a fully automatic rollback mechanism is immediately triggered, switching back to the previous stable version within one minute and terminating the canary release process.
[0089] Step S168: Synchronize the entire process execution data of anomaly detection, fault prediction, hot migration, hierarchical handling, and canary release to the central cloud, continuously update the training sample library of the isolated forest and LSTM-Attention models, dynamically optimize the detection threshold, prediction parameters and self-healing strategy, and form a closed-loop operation and maintenance system of collection, detection, prediction, self-healing and optimization.
[0090] The entire process execution data, including anomaly detection results, fault prediction data, hot migration logs, tiered handling records, and canary release metrics, is synchronized in real-time to the central cloud database using data synchronization tools, and a daily summary is generated to form an operation and maintenance data report. The training sample libraries for the Isolation Forest and LSTM-Attention models are continuously updated, with no fewer than 10,000 new samples added each month. The anomaly detection threshold calculation strategy, model hyperparameters, and self-healing rules are dynamically optimized, and the model is lightly fine-tuned weekly, forming a closed-loop operation and maintenance system of "collection-detection-prediction-self-healing-optimization," continuously improving the accuracy and efficiency of automated operation and maintenance and fault self-healing.
[0091] Based on the same inventive concept, please refer to Figure 2 This document illustrates a schematic block diagram of a cloud-edge collaborative scheduling and intelligent agent operation system 100 based on a distributed idle computing power network, provided in an embodiment of this application, for executing the aforementioned cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network. The cloud-edge collaborative scheduling and intelligent agent operation system 100 based on a distributed idle computing power network may include a communication unit 110, a machine-readable storage medium 120, and a processor 130.
[0092] In this embodiment, both the machine-readable storage medium 120 and the processor 130 are located in the cloud-edge collaborative scheduling and intelligent agent operation system 100 based on a distributed idle computing power network and are separately configured. However, it should be understood that the machine-readable storage medium 120 may also be independent of the cloud-edge collaborative scheduling and intelligent agent operation system 100 based on a distributed idle computing power network and may be accessed by the processor 130 through a bus interface. Alternatively, the machine-readable storage medium 120 may also be integrated into the processor 130 and may communicate and interact with external systems through the communication unit 110.
[0093] The processor 130 is the control center of the cloud-edge collaborative scheduling and intelligent agent operation system 100 based on a distributed idle computing power network. It connects to various parts of the system via various interfaces and lines. By running or executing software programs and / or modules stored in the machine-readable storage medium 120, and by accessing data stored in the machine-readable storage medium 120, it performs various functions and processes data within the system, thereby providing overall monitoring of the system. Optionally, the processor 130 may include one or more processing cores; for example, it may integrate an application processor and a modem processor, where the application processor primarily handles the operating system, user interface, and applications, while the modem processor primarily handles wireless communication. It is understood that the modem processor may also not be integrated into the processor. The machine-readable storage medium 120 is used to store machine-executable instructions for executing the scheme of this application, and the processor 130 is used to execute the machine-executable instructions stored in the machine-readable storage medium 120 to realize the cloud-edge collaborative scheduling and intelligent agent operation method based on distributed idle computing power network provided in the aforementioned method embodiments.
[0094] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.
[0095] The embodiments of this application have been described above with reference to the accompanying drawings. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network, characterized in that: Includes the following steps: Construct a city-level distributed idle heterogeneous GPU computing power network, build a four-layer collaborative architecture, and complete automatic node registration, hardware reporting, and performance calibration; Construct a multi-dimensional feature system for edge nodes, collect full-dimensional indicators and achieve unified quantification of heterogeneous GPU computing power based on HCU, normalize heterogeneous GPU computing power by relying on architecture compatibility coefficient, construct a node health assessment model and complete time-series load prediction. Complete the six-tuple requirement modeling and parameter analysis for AI intelligent agents and large model tasks, construct a multi-dimensional comprehensive cost function, and make optimal node scheduling decisions based on four-layer scheduling logic; Implement heterogeneous perceptual computation graph segmentation, distributed parallel execution, and operator compatibility adaptation for large-model inference; Identify the tidal patterns of the computing power network, realize elastic scaling of computing power based on MADDPG, and enable waveform interference decentralized scheduling in abnormal scenarios; Deploy an automated operation and maintenance and self-healing system, realize real-time node anomaly detection through the isolated forest algorithm, complete early fault prediction by relying on the LSTM-Attention model, save the GPU context through the GSRP protocol when a fault is triggered, complete container hot migration by combining CRIU technology, and execute task evacuation, node isolation, maintenance work order generation and disposal actions according to hierarchical alarm rules. The central cloud synchronously completes the gray release and automatic rollback of the model version.
2. The cloud-edge collaborative scheduling and agent operation method based on a distributed idle computing power network according to claim 1, characterized in that: The process involves constructing a multi-dimensional feature system for edge nodes, collecting full-dimensional indicators, achieving unified quantification of heterogeneous GPU computing power based on HCU, normalizing heterogeneous GPU computing power by relying on architecture compatibility coefficients, constructing a node health assessment model, and completing time-series load prediction, including: A multi-dimensional feature vector is constructed for each edge node. Through three methods—real-time device monitoring, hardware information reporting, and benchmark dynamic calibration—all dimensions of indicators such as GPU architecture generation, number of GPU cards, single card memory capacity, memory bandwidth, FP32 floating-point computing power, INT8 fixed-point computing power, available storage space, network bandwidth, network latency, overall power consumption, GPU temperature, and GPU utilization are comprehensively collected to complete the real-time collection and standardized processing of feature data. Define the HCU heterogeneous computing power unit as a unified computing power quantification standard, clarify the baseline specifications of 1 HCU corresponding to 1 INT8TOPS, 4GB video memory, 500-1000M effective bandwidth and architecture compatibility coefficient of 1.0, pre-construct the operator performance matrix of GPUs from different manufacturers and generations and calculate the corresponding architecture compatibility coefficient, combine the weighted calculation rules of INT8 computing power, video memory and effective bandwidth, complete the computing power quantification and normalization calculation of heterogeneous GPUs from different manufacturers and generations, and generate a unified computing power value for nodes; Based on the node's current GPU utilization, GPU temperature, continuous normal operation time, historical task completion reputation score, and node operation stability data, weight parameters for each dimension are set, and a node health score is generated through weighted calculation to complete the construction and real-time scoring of the node health assessment model. By employing a TFT timing fusion Transformer model and performing explicit periodic modeling, and integrating load timing characteristics with daily and weekly periodic patterns, high-precision prediction of the future load state of nodes can be achieved.
3. The cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network according to claim 2, characterized in that: The method employs a TFT time-series fusion Transformer model and performs explicit periodic modeling, integrating load time-series characteristics with daily and weekly periodic patterns to achieve high-precision prediction of future node load states, including: Historical time-series data on edge node GPU utilization, computing load, task queue length, and network bandwidth are continuously collected at a fixed time granularity. Time-series features of hourly segments, days of the week, holidays, and business cycles are extracted synchronously. Missing values are filled in using linear interpolation, and abnormal surges / decreases in data are smoothed using Kalman filtering. Data normalization and feature alignment are completed to build a high-quality load time-series sample set. The daily, weekly, and hourly cycles are used as explicit discrete features. They are converted into low-dimensional dense vectors through a cycle embedding layer and then concatenated with load time series features and node hardware status features to form a model input tensor containing explicit cycle information. A lightweight TFT timing fusion Transformer model adapted to edge node computing power constraints is initialized, and a gated loop unit, a multi-head attention mechanism, and a gated residual connection structure are configured. The model input tensor is fed into the model. Local short-term load dependencies are captured by the gated loop unit, and long-term correlations and daily / weekly cycle patterns are focused by the multi-head attention mechanism. A joint loss function including mean squared error loss and periodic consistency loss is constructed. An adaptive optimizer is used to complete the model training. The learning rate, number of attention heads, and hidden layer dimension hyperparameters are dynamically adjusted based on the validation set. The prediction error and overfitting risk are reduced through iterative optimization. The latest load time-series data and periodic features are collected incrementally on an hourly basis. The model is then updated with lightweight incremental updates. The current time-series and periodic features are input into the trained TFT model, which outputs the load values and fluctuation trends for multiple future time periods, thus achieving high-precision prediction of the future load status of nodes.
4. The cloud-edge collaborative scheduling and agent operation method based on a distributed idle computing power network according to claim 1, characterized in that: The process of completing the six-tuple requirement modeling and parameter analysis for AI agents and large model tasks, constructing a multi-dimensional comprehensive cost function, and making optimal node scheduling decisions based on four-layer scheduling logic includes: Define a six-tuple dimension that includes task type, computing power requirement, latency threshold, accuracy level, security level, and data privacy level, and formulate corresponding quantitative standards for each dimension; During the task submission phase, key task parameters and user-specific requirement parameters are collected. The parameters are standardized and converted through the parsing module to generate a structured six-tuple requirement vector, thus completing requirement modeling and parameter parsing. Based on six cost dimensions and six-tuple requirements, a task query vector is constructed. A lightweight multi-head attention mechanism is adopted. The model autonomously identifies task constraints and dynamically and differentially integrates each cost dimension, decouples the coupling relationship of indicators, and outputs the globally optimal comprehensive evaluation result. It has the ability to adapt to all scenarios and dynamically adapt to heterogeneous GPU nodes. The system uses a four-layer scheduling logic to gradually select suitable nodes, check local nodes and edge nodes within the region in turn, determine the feasibility of task splitting and execute distributed collaborative reasoning or cloud scheduling, and finally determine the optimal scheduling path and target node. During the scheduling process, the node status and task progress are synchronized in real time, and the calculation results of the comprehensive cost function are dynamically updated to ensure the real-time performance and accuracy of scheduling decisions.
5. The cloud-edge collaborative scheduling and agent operation method based on a distributed idle computing power network according to claim 4, characterized in that: The task query vector is constructed based on six cost dimensions and six-tuple requirements. A lightweight multi-head attention mechanism is employed, allowing the model to autonomously identify task constraints and dynamically differentiate and fuse various cost dimensions, decoupling indicator coupling relationships and outputting a globally optimal comprehensive evaluation result. It possesses full-scene adaptive capabilities and dynamic adaptability to heterogeneous GPU nodes, including: The six cost dimensions of latency, computing power matching, health, model adaptation, security matching, and data transmission are standardized to eliminate dimensional differences and map the values of each dimension to a unified range. At the same time, the discrete features in the six-tuple requirements are embedded and encoded into low-dimensional dense vectors, and the continuous features in the six-tuple requirements are normalized. This completes the structured preprocessing of the two types of input features. The six-tuple requirements include task type, computing power requirement, latency threshold, accuracy level, security level, and data privacy level. The preprocessed six-tuple demand vector is used as a guide and cross-fused with the six cost dimension feature vectors. The feature is then mapped through a shallow fully connected network to generate a task query vector that simultaneously contains task constraint preferences and basic cost assessment information. A lightweight multi-head attention mechanism network is built, with 2-4 attention heads configured to reduce computational overhead. Each attention head focuses on different types of task constraint dimensions. Initial attention weights are generated by calculating the similarity between the task query vector and the cost dimension feature vector. A gating adjustment unit is introduced to optimize the initial attention weight, suppress the interference of redundant dimensions and strengthen the weight ratio of constraint dimensions. Real-time inference tasks automatically amplify the attention weight of the delay cost dimension, and security and confidentiality tasks automatically increase the weight allocation of the security matching cost dimension, so as to achieve dynamic and differentiated integration of various cost dimensions. By parallel computation and feature concatenation of multi-head attention, the implicit coupling between different cost dimensions is decoupled. Then, the multi-head attention output is integrated through a linear projection layer to generate a globally optimal comprehensive evaluation result. A labeled dataset containing different task types such as real-time inference, offline training, and security and confidentiality is constructed. The dataset covers task query vectors, cost dimension features, and corresponding optimal evaluation labels. The mean squared error loss function combined with task type adaptive regularization is used to train and optimize the model. The model size is compressed through quantization training and pruning to adapt to the computing power constraints of edge scheduling nodes. After deployment, the model receives the six-tuple requirements of new tasks and the dynamic cost dimension features of the current node in real time. After structured preprocessing, it is input into a lightweight multi-head attention model to quickly complete the calculation of attention weights and feature fusion and output a comprehensive evaluation result. At the same time, it continuously monitors the node's load fluctuations and health changes, updates the cost dimension features in real time, and automatically adjusts the weight allocation strategy of the attention heads to adapt to the dynamic operating status of the nodes and complete the full-scenario adaptive matching.
6. The cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network according to claim 1, characterized in that: The implementation of heterogeneous perceptual computation graph segmentation, distributed parallel execution, and operator compatibility adaptation for large model inference includes: After receiving a large model inference task, the system collects hardware parameters, available video memory, computing power specifications, network bandwidth, and inter-node communication latency of heterogeneous GPU nodes across the network in real time, completing global perception and status archiving of heterogeneous computing resources; it also parses the original computation graph of the large model, decomposes the model hierarchy, operator type, data dependency relationship and tensor flow, and converts the computation graph into a standardized directed acyclic graph, completing the structured parsing of the computation graph; Based on the parsed list of operators and heterogeneous computing resources, a pre-computation process is executed. Benchmark tests are conducted on the execution time of each type of operator on GPUs from different manufacturers and models. The memory usage and computational complexity of the operators are calculated. At the same time, the communication cost and transmission latency of data transmission between nodes are measured, and an operator, hardware performance mapping library and communication cost matrix are constructed. With the optimization goals of minimizing end-to-end inference latency, minimizing communication costs, and balancing node computing power load, and combining the computing power limit and memory capacity constraints of each heterogeneous node, a dynamic programming algorithm is used to traverse all feasible partitioning schemes of the computation graph, select the optimal partitioning point, and split the complete computation graph into multiple independent subgraph partitions to ensure that each partition is adapted to the heterogeneous computing power characteristics of the target node and minimizes data dependency and communication overhead between partitions. The optimally partitioned subgraph is distributed to the corresponding heterogeneous GPU nodes, a partition execution dependency time table is established, and an asynchronous pipeline execution mechanism is started, so that the computation tasks and data transmission tasks of adjacent partitions overlap and run. Pipeline parallelism is used to hide the communication latency between nodes, and multi-node distributed parallel inference is realized. During partitioning execution, the compatibility between operators and target GPU hardware is detected in real time. For operators with incompatible architectures or mismatched precision, compatibility processing such as operator conversion, precision adaptation, and kernel replacement is automatically performed. For extremely incompatible operators that cannot be converted, the CPU rollback execution mechanism is automatically triggered. During inference operation, the load status, partition execution progress and communication link quality of each node are continuously monitored. If node computing power fluctuations or communication blockages occur, the pipeline timing is dynamically fine-tuned or the local partitioning strategy is recalculated to enable rapid migration of subgraph partitions for abnormal nodes. After all heterogeneous nodes have completed the inference calculations for their corresponding subgraph partitions, the intermediate calculation results of each partition are aggregated in an orderly manner according to the data dependencies of the original computation graph, and the final inference output is completed. At the same time, the node computing resources are recovered, and the operators and hardware performance mapping library are updated.
7. The cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network according to claim 1, characterized in that: The identification of the computing power network tidal pattern, based on MADDPG to achieve elastic scaling of computing power, and the activation of waveform interference decentralized scheduling in abnormal scenarios, includes: Collect and preprocess multi-dimensional computing load time-series data, and construct a full-scenario tidal pattern library by dividing basic and burst tidal patterns through time series decomposition, frequency analysis and clustering, and match the current pattern in real time and output load change trend prediction. A multi-agent scheduling system is built based on MADDPG. The state, action and reward function of the agent are defined. After offline training, each agent dynamically executes the actions of expanding, shrinking or maintaining computing power in combination with the global tidal mode, so as to realize the elastic scaling of computing power. Define abnormal scenarios and judgment rules, identify abnormalities and broadcast information through a distributed detection module, enable a waveform interference decentralized scheduling mechanism, and enable intelligent agents to autonomously and collaboratively allocate tasks and adjust computing power. After the abnormality is mitigated, normal scheduling is restored and the training dataset is updated.
8. The cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network according to claim 7, characterized in that: The multi-agent scheduling system based on MADDPG defines agent states, actions, and reward functions. After offline training, each agent dynamically executes computing power expansion, contraction, or maintenance actions in conjunction with a global tidal pattern, achieving elastic scaling of computing power, including: Each regional scheduling node in the entire network is independently encapsulated as a MADDPG intelligent agent, and a system framework of centralized training and decentralized execution is built. Point-to-point low-latency communication links are established between the intelligent agents to realize real-time sharing and collaborative perception of global tidal mode labels, computing resource status, and task queuing information. Construct a standardized state space that integrates global tidal mode types, available computing power of regional nodes, memory utilization, average task waiting time, node health score, and network transmission bandwidth. Standardize the action space that includes three types of actions: computing power expansion, computing power reduction, and maintaining the current computing power. Clarify the sub-actions and corresponding quantitative standards for computing power expansion and reduction. Design a multi-objective joint reward function with the goals of optimal resource utilization, minimum task latency, and minimum node energy consumption. Give positive rewards to scheduling behaviors that meet resource utilization and task latency requirements, and give negative rewards to scheduling behaviors that result in computing power overload, excessive task latency, and energy waste. Collect historical computing power tidal pattern data, computing power scheduling execution records, and node running status data to construct a standardized training dataset and perform normalization preprocessing. Initialize the actor network and critic network of the MADDPG model and configure hyperparameters such as learning rate, batch size, and experience replay capacity. The training mode of centralized critics and distributed actors is adopted to start offline joint training of multiple agents. The critic network obtains the state and action information of all agents and outputs global value assessment. The actor network generates scheduling actions based on local perception information. The network parameters are optimized through experience replay mechanism. Various tidal patterns and computing power scenarios are simulated to enhance the collaborative decision-making ability of agents. The training is iteratively carried out until the model converges. The trained lightweight model is deployed to the intelligent agent terminals in various regions. Each intelligent agent is equipped with a real-time state perception module, which is seamlessly connected with the computing power tidal pattern recognition module to obtain the current computing power network tidal pattern, load prediction results, and global resource status in real time. The agent collects local state information and global tidal pattern features in real time, inputs the trained MADDPG model and outputs the optimal scheduling action. When expansion is triggered, it prioritizes waking up dormant nodes in the local area. When there are not enough nodes, it schedules heterogeneous GPU nodes from the global computing power resource pool. When shrinkage is triggered, it takes offline idle nodes or puts low-load nodes into sleep according to node load priority and health score. When maintenance action is triggered, it maintains the current computing power configuration. Each agent executes the action independently and without conflict. During the elastic scaling process, online scheduling data and operational performance data are continuously collected, and the MADDPG model is incrementally fine-tuned and optimized periodically. The reward function weights are dynamically adjusted in sync with changes in the tidal pattern.
9. The cloud-edge collaborative scheduling and agent operation method based on a distributed idle computing power network according to claim 1, characterized in that: The deployed automated operation and maintenance and self-healing system achieves real-time node anomaly detection through the Isolation Forest algorithm, predicts faults in advance using the LSTM-Attention model, saves the GPU context through the GSRP protocol when a fault is triggered, completes container hot migration using CRIU technology, and executes task evacuation, node isolation, and maintenance work order generation and handling actions according to hierarchical alarm rules. The central cloud synchronously completes the canary release and automatic rollback of the model version, including: A three-tier architecture foundation of edge agent, regional engine and central control is built. Lightweight operation and maintenance agent modules are deployed on each heterogeneous GPU node, self-healing execution engine is deployed on regional scheduling nodes, and a unified operation and maintenance management platform is built in the central cloud. The end-to-end data collection channels of node hardware, container operation, GPU status and task execution are opened up, and a standardized operation and maintenance data bus is built to achieve unified aggregation and collaborative management of the operation and maintenance status of the entire network. A real-time node anomaly detection model is built based on the isolated forest algorithm. It integrates multi-dimensional running features such as GPU temperature, memory usage, computing power utilization, network packet loss rate, disk I / O, process survival status, and instruction response latency. The feature data is normalized and preprocessed before being input into the model. The model calculates the single-node anomaly score in real time using streaming. A dynamic judgment threshold is set to filter data noise and instantaneous fluctuations. When the anomaly score continuously exceeds the threshold, the node is marked as an abnormal state in milliseconds. Fault prediction is achieved by relying on the LSTM-Attention hybrid model. Long-term historical runtime data of nodes, historical fault records, and alarm information are collected to construct a prediction sample set. Long-term dependent features of time series data are extracted through the LSTM network, and an attention mechanism is introduced to focus on fault-related features. The model outputs the probability of node fault occurrence and fault type in future time periods, and the warning level is divided according to the predicted probability. After a node failure trigger signal is detected, the GSRP protocol is enabled to perform GPU context security saving, capturing the full context information of GPU computing kernel state, memory tensor data, task execution breakpoints, and hardware mapping relationships. Incremental snapshot technology is used for lightweight storage and synchronous backup to the regional redundant storage node. By combining CRIU technology, container hot migration is completed without any noticeable impact. Based on the GPU context saved by GSRP, CRIU is used to perform freeze and checkpoint operations on the task containers of the faulty node, generating a snapshot of the container's running state and an image file. These are then transmitted to a pre-selected healthy node via a high-speed intranet channel, where the container state is restored and the GPU context is reloaded. Automated handling is performed according to preset hierarchical alarm rules. Low-level anomalies are recorded in the operation and maintenance log and synchronized to the central platform; medium-level anomalies automatically evacuate non-core tasks of nodes to surrounding healthy nodes; high-level faults are implemented by network isolation and computing resource shielding of nodes, while automatically generating standardized maintenance work orders with associated fault node information, detection results, and prediction data and pushing them to the operation and maintenance terminal. The central cloud operation and maintenance platform implements gray-scale release and automatic rollback control of the model version. It divides all network nodes into multiple gray-scale release units according to region, load and hardware type, pushes new version operation and maintenance models and scheduling models in batches, and monitors the node's running stability, task inference accuracy and computing power consumption indicators in real time. When the indicators are abnormal, it triggers a fully automatic rollback mechanism to switch back to the previous stable version. The entire process of anomaly detection, fault prediction, hot migration, hierarchical handling, and canary release is synchronized to the central cloud. The training sample library of the Isolation Forest and LSTM-Attention models is continuously updated, and the detection threshold, prediction parameters, and self-healing strategy are dynamically optimized to form a closed-loop operation and maintenance system of collection, detection, prediction, self-healing, and optimization.
10. A cloud-edge collaborative scheduling and intelligent agent operation system based on a distributed idle computing power network, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the cloud-edge collaborative scheduling and intelligent agent operation method based on a distributed idle computing power network as described in any one of claims 1 to 9 by executing the machine-executable instructions.