Method for evaluating real-time performance of computing power network based on analytic hierarchy process
Through the real-time performance evaluation method of computing power network based on hierarchical analysis, the multi-dimensional performance evaluation data set is generated using quantum heuristic algorithms and digital twin technology, and the weights are dynamically adjusted, which solves the problems of lag in the evaluation result and conflicts in the traditional method, and realizes efficient resource utilization and response of computing power network.
Patent Information
- Application Number
- CN202510856744.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional hierarchical analysis method based on static weights is difficult to adapt to dynamically changing network states in real-time performance evaluation of computing power networks, resulting in lagging evaluation results and lack of flexibility, and the resource allocation conflicts and scheduling efficiency in cross-domain collaborative scheduling are reduced.
The real-time performance evaluation method of computing power network based on hierarchical analysis is adopted, node status data is obtained through distributed sensors and log systems, and the multi-dimensional performance evaluation data set is generated, the AHP weight is dynamically adjusted, and the risk assessment algorithm and reinforcement learning model is combined to achieve real-time optimization and adaptive adjustment of resource scheduling.
The real-time response efficiency and resource utilization of computing power network have been significantly improved, and the scheduling strategy is quickly generated through a dual-track mechanism of real-time monitoring and future prediction, dynamically optimize resource allocation, avoid resource conflicts, and ensure performance-cost-carbon efficiency balance.
Smart Images

Figure CN120378333A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computer networks, and in particular, to a real-time performance evaluation method for a computing power network based on the analytic hierarchy process (AHP). Background Art
[0002] With the rapid expansion of the scale of the computing power network and the complexity of business requirements, the traditional AHP based on static weights faces significant limitations in real-time performance evaluation. Existing technologies mostly rely on expert experience to set fixed weights, making it difficult to adapt to dynamically changing network states (such as sudden task loads, node failures), resulting in lagging evaluation results and lack of flexibility.
[0003] In addition, due to the lack of a unified data sharing and verification mechanism in cross-domain collaborative scheduling, resource allocation conflicts and a decline in scheduling efficiency are likely to occur. As can be seen from the above, how to improve the real-time response efficiency and resource utilization rate of the computing power network remains to be solved. Summary of the Invention
[0004] To improve the real-time response efficiency and resource utilization rate of the computing power network, this application provides a real-time performance evaluation method and system for a computing power network based on the analytic hierarchy process (AHP).
[0005] In a first aspect, this application provides a real-time performance evaluation method for a computing power network based on the analytic hierarchy process (AHP), adopting the following technical solutions: A real-time performance evaluation method for a computing power network based on the analytic hierarchy process (AHP) includes: Real-time obtaining operation state data and task requirement data of computing power network nodes through distributed sensors and a log system, where the operation state data includes CPU utilization rate, memory occupancy rate, network bandwidth, task latency, and energy consumption, and the task requirement data includes task type, priority, and resource requirements; processing the operation state data through a quantum-inspired algorithm deployed on edge nodes to generate local performance evaluation indicators; simulating future network state data based on digital twins, where the future network state data includes bandwidth change trends and task load predictions, and fusing the local performance evaluation indicators with the future network state data to generate a multi-dimensional performance evaluation data set including the current state and future predictions, where the sensors include lightweight quantum computing modules deployed at the edge of nodes; Calculating corresponding resource deviation data based on the operation state data and a preset standard performance reference value, and classifying the current network state in real time through a geological classification model to obtain a network state classification result; if a node failure or high load is detected, reducing the weight adjustment threshold of the AHP judgment matrix and enabling a high-frequency dynamic weight update mode; if a low load or idle node is detected, expanding the weight allocation range and enabling a global resource scheduling strategy; if a sudden task influx is detected, dynamically adjusting the AHP sub-index weights in combination with the task priority data; Based on the resource deviation data, network status classification results, and task priority data, calculate the performance-cost-carbon efficiency conflict level in the computing power scheduling process through a risk assessment algorithm; when the conflict level is a high-risk level, enable the emergency resource recovery strategy and restrict the scheduling of high-energy-consuming nodes; when the conflict level is a medium-risk level, trigger a multi-objective optimization prompt and display the performance-cost-carbon efficiency trade-off plan; when the conflict level is a low-risk level, provide single-index optimization suggestions. Based on the resource deviation data and conflict assessment results, trigger scheduling strategy adjustment signals at different levels, where the scheduling strategy adjustment signals include green resource allocation instructions, yellow warning-dynamic weight correction, and red alert-global scheduling reset; if the pressure sensor detects node anomalies and the conflict assessment is not triggered, directly initiate active scheduling intervention and restrict the resource allocation of high-load nodes. Obtain the real-time performance evaluation data in the multi-dimensional performance evaluation dataset, compare the real-time performance evaluation data with the historical environment parameter database, and dynamically update the AHP weight matrix through a reinforcement learning model based on the environment comparison results; if an abnormal working condition occurs, conduct a root cause analysis of the abnormal working condition, and automatically optimize the parameters of the quantum-classical hybrid algorithm and the blockchain verification weight.
[0006] Optionally, when the multi-dimensional performance evaluation dataset fuses local performance evaluation indicators and future network status data, the method further includes: Synchronize the local performance indicators and future state data in the time dimension through a timestamp alignment mechanism. Obtain the time deviation data of the local performance indicators and the future state data in the time dimension. If the time deviation data exceeds the preset time deviation threshold, perform interpolation compensation on the local performance indicators or downsampling processing on the future state data.
[0007] Optionally, during the process of triggering scheduling strategy adjustment signals at different levels based on the resource deviation data and conflict assessment results, if the pressure sensor detects node anomalies and the conflict assessment is not triggered, supplement the constraints of the active scheduling intervention. The method further includes: When the pressure sensor detects that a node has an abnormal operating state and the abnormal operating state does not reach any risk trigger threshold of the performance-cost-carbon efficiency conflict level, determine whether the resource deviation data of the abnormal node continuously exceeds the preset resource threshold. If at least one indicator in the resource deviation data continuously exceeds the set resource threshold for a preset duration, trigger the active scheduling intervention mechanism. When performing active scheduling intervention, preferentially retain the resource allocation permissions corresponding to nodes with energy consumption lower than the benchmark value, and at the same time implement resource freezing operations on high-load nodes with resource utilization rates exceeding the load limit.
[0008] Optionally, during the process of dynamically adjusting the sub-index weights of the AHP judgment matrix based on resource deviation data, network status classification results, and task priority data, when a sudden influx of tasks is detected, the method further includes: Identifying the urgency of the current task flow according to the task priority data obtained by the logging system; If it is determined that the task priority is high, increase the relative importance weight of the task delay index in the AHP judgment matrix by 15% - 20%, and at the same time reduce the weight of the resource requirement index by 5% - 10%; After completing the weight adjustment, use the consistency ratio test method to verify the consistency of the updated AHP judgment matrix. Only when the consistency ratio CR is less than 0.1, confirm that the secondary weight adjustment is logically reasonable and will be applied to subsequent performance evaluation and scheduling decisions.
[0009] Optionally, during the process of simulating future network state data based on digital twins, the method further includes: Construct a 3D visual topology model corresponding one-to-one to the actual computing power network, and synchronously update the connection status, bandwidth change trend, and task load distribution between nodes through real-time data streams; Deploy a virtual user behavior simulation module in the virtual environment corresponding to the 3D visual topology model, retrieve the corresponding historical task patterns, obtain the corresponding user behavior characteristics, and generate dynamic task demand data based on the historical task patterns and the user behavior characteristics. The dynamic task demand data is used as the input parameter for future network state prediction; Fuse and analyze the dynamically generated task demand data and the operating status data, and output future network state data including extended spatio-temporal dimensions.
[0010] Optionally, when an abnormal working condition occurs and root cause analysis is required, the method further includes: Based on historical operating status data and task scheduling logs, construct a causal inference graph between nodes in the computing power network, where the nodes in the causal inference graph represent computing power resource units, and the edges represent the causal intensity between task delay, energy consumption change, and bandwidth fluctuation. Among them, task delay, energy consumption change, and bandwidth fluctuation are all performance indicators; Identify the corresponding abnormal propagation path based on the causal inference graph, and quantify the causal correlation degree between the abnormal source node and the downstream affected nodes based on the abnormal propagation path; Generate a counterfactual virtual scenario based on the abnormal propagation path, simulate the expected operating status when the key node does not have an abnormality, and compare the expected operating status with the actual operating status; output a root cause analysis report based on the comparison results, and assist in generating optimization measures for the root cause. The optimization measures include priority reallocation, node isolation, or parameter adaptive adjustment.
[0011] In a second aspect, the present application provides a real-time performance evaluation system for a computing power network based on the analytic hierarchy process, adopting the following technical solutions: A real-time performance evaluation system for a computing power network based on the analytic hierarchy process, comprising: A multi-dimensional performance evaluation data set generation module, which obtains the operation status data and task requirement data of the computing power network nodes in real time through a distributed sensor and a log system. The operation status data includes CPU utilization rate, memory occupancy rate, network bandwidth, task latency, and energy consumption. The task requirement data includes task type, priority, and resource requirements. The operation status data is processed by a quantum-inspired algorithm deployed on the edge node to generate local performance evaluation indicators. Based on digital twin to simulate future network state data, the future network state data includes bandwidth change trend and task load prediction. The local performance evaluation indicators are fused with the future network state data and used to generate a multi-dimensional performance evaluation data set including the current state and future prediction. Among them, the sensor includes a lightweight quantum computing module deployed at the node edge; A network state classification result acquisition module, which calculates the corresponding resource deviation data based on the operation status data and a preset standard performance reference value, and classifies the current network state in real time through a geological classification model to obtain the network state classification result. If a node failure or high load is detected, the weight adjustment threshold of the AHP judgment matrix is reduced and a high-frequency dynamic weight update mode is enabled. If a low load or idle node is detected, the weight allocation range is extended and a global resource scheduling strategy is enabled. If a sudden task influx is detected, the AHP sub-index weight is dynamically adjusted in combination with the task priority data; A conflict level calculation module, which calculates the performance-cost-carbon efficiency conflict level in the computing power scheduling process through a risk assessment algorithm based on the resource deviation data, network state classification result, and task priority data. If the conflict level is a high-risk level, an emergency resource recovery strategy is enabled and the scheduling of high-energy-consuming nodes is restricted. If the conflict level is a medium-risk level, a multi-objective optimization prompt is triggered and a performance-cost-carbon efficiency trade-off scheme is displayed. If the conflict level is a low-risk level, a single-index optimization suggestion is provided; A policy adjustment signal scheduling module, which is used to hierarchically trigger scheduling policy adjustment signals based on the resource deviation data and conflict assessment results. Among them, the scheduling policy adjustment signals include green resource allocation instructions, yellow warning-dynamic weight correction, and red alert-global scheduling reset. If the pressure sensor detects node anomalies and no conflict assessment is triggered, active scheduling intervention is directly initiated and the resource allocation of high-load nodes is restricted; The weight matrix update module obtains real-time performance evaluation data from a multi-dimensional performance evaluation dataset, compares the real-time performance evaluation data with a historical environment parameter database, and uses a reinforcement learning model to dynamically update the AHP weight matrix based on the environment comparison result; if an abnormal working condition occurs, it conducts a root cause analysis of the abnormal working condition and automatically optimizes the parameters of the quantum-classical hybrid algorithm and the blockchain verification weight.
[0012] Thirdly, this application provides a real-time performance evaluation system for a computing power network based on the analytic hierarchy process, and adopts the following technical solutions: A real-time performance evaluation system for a computing power network based on the analytic hierarchy process includes a processor, and a program of the real-time performance evaluation method for a computing power network based on the analytic hierarchy process as described in any one of the above is run in the processor.
[0013] Fourthly, this application provides a storage medium, and adopts the following technical solutions: A storage medium stores a program of the real-time performance evaluation method for a computing power network based on the analytic hierarchy process as described in any one of the above.
[0014] In summary, this application includes at least one of the following beneficial technical effects: The real-time response efficiency and resource utilization rate of the computing power network are significantly improved through multi-dimensional technology integration. First of all, based on the quantum-inspired algorithm and digital twin technology, the system realizes real-time dynamic modeling of the node operation state and future load: the lightweight quantum computing module of the edge node accelerates the generation of local performance evaluation indicators (such as task latency, energy consumption), and the 3D topology model constructed by the digital twin combined with virtual user behavior simulation can predict bandwidth fluctuations and task load trends in advance. This "real-time monitoring + future prediction" dual-track mechanism enables the system to quickly generate scheduling strategy adjustment signals (such as green resource allocation, global scheduling reset) based on the multi-dimensional performance evaluation dataset when tasks burst or nodes are abnormal, thereby shortening the decision-making latency and reducing the risk of resource conflicts.
[0015] Secondly, through dynamic weight optimization and causality-driven attribution analysis, the system achieves refined control of resource allocation. The weight adjustment mechanism of the AHP judgment matrix can be adaptively adjusted according to task priorities (such as increasing the weight of high-priority task delays by 15% - 20%), network status (such as enabling high-frequency dynamic updates under high load), and abnormal working conditions (such as locating the root cause through the causal inference graph), ensuring that key resources preferentially meet the needs of high-value tasks. At the same time, the reinforcement learning model dynamically updates the weight matrix by comparing historical environmental parameters, and combines blockchain to verify the credibility of the weights, avoiding the problem of rigid resource allocation caused by static weights. This "prediction - evaluation - optimization - verification" closed-loop control system significantly improves the response efficiency and resource utilization rate of the computing power network for complex scenarios while ensuring the balance of performance - cost - carbon efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flowchart of a real-time performance evaluation method for a computing power network based on the analytic hierarchy process according to an exemplary embodiment.
[0017] Figure 2 is a structural block diagram of a real-time performance evaluation system for a computing power network based on the analytic hierarchy process according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The following details the embodiments of the present application, and the examples of the embodiments are shown in the drawings.
[0019] In the description of this specification, the description of reference terms "certain embodiments", "one embodiment", "some embodiments", "schematic embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0020] The embodiments of the present application disclose a real-time performance evaluation method for a computing power network based on the analytic hierarchy process, referring to Figure 1 , including: S100. Obtain the operation status data and task demand data of the computing power network nodes in real time through distributed sensors and a logging system. The operation status data includes CPU utilization rate, memory occupancy rate, network bandwidth, task latency, and energy consumption. The task demand data includes task type, priority, and resource requirements. Process the operation status data through the quantum-inspired algorithm deployed on the edge nodes and generate local performance evaluation metrics. Based on digital twin, simulate the future network status data, which includes bandwidth change trend and task load prediction. Integrate the local performance evaluation metrics with the future network status data and generate a multi-dimensional performance evaluation data set containing the current status and future predictions. Among them, the sensors include lightweight quantum computing modules deployed at the edge of the nodes.
[0021] Among them, it specifically includes the following steps: Step 1, operation status data collection: Distributed sensors are deployed at the edge of the network nodes to monitor the hardware resource status in real time, including CPU utilization rate (reflecting the computing power load), memory occupancy rate (indicating the tightness of memory resources), network bandwidth (measuring the data transmission capacity), task latency (characterizing the task response efficiency), and energy consumption (relating to energy cost and carbon efficiency). After the above data is collected by the sensors, it is preprocessed by the lightweight quantum computing module of the edge node to ensure low-latency data transmission and preliminary feature extraction.
[0022] Task demand data collection: The logging system records the metadata during the task scheduling process, including task type (such as compute-intensive or communication-intensive), priority (distinguishing urgent tasks from ordinary tasks), and resource requirements (such as the required number of CPU cores and memory size). It should be noted here that the task demand data and the operation status data together constitute the multi-dimensional input for performance evaluation.
[0023] Through the collaborative work of the distributed sensors and the logging system, the system can capture the dynamic behavior of the network nodes in real time, providing an accurate and low-latency data basis for subsequent performance evaluation and scheduling decisions, and avoiding decision-making biases caused by data lag.
[0024] Step 2, generate local performance evaluation metrics: Quantum-inspired algorithm processing: The edge node is built with a lightweight quantum computing module, and uses quantum annealing or variational quantum algorithms (such as VQE) to extract features and optimize the operation status data. For example, encode the trade-off problem between task latency and energy consumption through qubits and quickly solve the optimal local performance metrics (such as the "task latency - energy consumption balance index").
[0025] Local performance evaluation metric generation. The metrics output by the quantum-inspired algorithm include, but are not limited to: resource utilization balance (the collaborative load level of CPU and memory); task response efficiency (the comprehensive score of latency and bandwidth); energy consumption sensitivity (the energy consumption volatility per unit of computational task).
[0026] It should be noted here that the quantum-inspired algorithm can significantly shorten the computational time of local performance evaluation (several times faster than traditional algorithms), and at the same time enhance the robustness of the metrics through quantum feature extraction, providing high-quality input for subsequent multi-dimensional data fusion.
[0027] Step 3: Construct a virtual network model through digital twin technology, predict the future state, and fuse it with real-time data to form a multi-dimensional dataset containing spatio-temporal dimensions, specifically including: Digital twin simulates the future state: Construct a 3D digital twin model consistent with the actual computing power network topology, and synchronize the connection status, bandwidth change trend, and task load distribution between nodes in real time; through the virtual user behavior simulation module, generate dynamic task demand data based on historical task patterns and user behavior characteristics (such as task request frequency, resource preference) as the input for future state prediction; the key parameters for simulating the future network state include: bandwidth fluctuation prediction (based on time series analysis), task load peak prediction (based on machine learning models), node failure probability (based on reliability analysis).
[0028] Multi-dimensional data fusion: Spatially and temporally align and weight-fuse the local performance evaluation metrics (such as resource utilization balance) with the future state data generated by digital twin (such as bandwidth prediction, load peak) to generate a multi-dimensional performance evaluation dataset containing the following dimensions: current state dimension, real-time operating state and task demand data; future prediction dimension, predicted values such as bandwidth trend, load peak, failure probability; spatio-temporal correlation dimension, resource competition relationship between nodes, time window of task migration path.
[0029] It should be noted here that the multi-dimensional dataset not only reflects the current network state, but also provides a forward-looking perspective through future predictions, enabling the scheduling strategy to anticipate risks in advance before task bursts or node anomalies, thus significantly improving the real-time response efficiency and resource utilization of the computing power network.
[0030] Build a real-time evaluation system for the computing power network through the collaborative construction of a lightweight quantum computing module, digital twin, virtual user behavior simulation, and multi-dimensional data fusion mechanism: lightweight quantum computing realizes the generation of low-latency performance indicators on the edge side; digital twin combines historical tasks and user behavior characteristics to improve the accuracy of future predictions; multi-dimensional data fusion unifies real-time and predictive data through timestamp alignment and weighted algorithms. The three work together to form a closed loop from real-time collection to forward-looking modeling, providing high-precision data support for AHP weight adjustment, conflict assessment, and scheduling optimization, and significantly improving the response efficiency and resource utilization rate of the computing power network in complex environments.
[0031] S200, calculate the corresponding resource deviation data based on the operating status data and the preset standard performance reference values, and classify the current network status in real time through the geological classification model to obtain the network status classification result; if a node failure or high load is detected, reduce the weight adjustment threshold of the AHP judgment matrix and enable the high-frequency dynamic weight update mode; if a low load or idle node is detected, expand the weight allocation range and enable the global resource scheduling strategy; if a sudden task influx is detected, dynamically adjust the AHP sub-index weights in combination with the task priority data.
[0032] Among them, it specifically includes the following steps: Step 1, calculate the resource deviation data based on the operating status data and the standard performance reference values: Data input: Take the operating status data obtained in S100 (such as CPU utilization rate, memory occupancy rate, network bandwidth, task latency, energy consumption) as input, and combine the preset standard performance reference values (such as CPU utilization rate threshold 85%, memory occupancy rate threshold 90%).
[0033] Deviation quantification: Generate a multi-dimensional resource deviation data set by calculating the difference or relative deviation rate between the actual value and the standard value (such as CPU utilization rate deviation rate = (actual value - standard value) / standard value).
[0034] Abnormality determination: If the deviation rate of a certain indicator exceeds the preset threshold (such as CPU utilization rate deviation rate > 20%), it is marked as an abnormal state, which is an important basis for subsequent network status classification.
[0035] It should be noted here that the resource deviation data directly reflects the health status and resource allocation efficiency of network nodes. By quantifying the deviation, the system can quickly identify high-load, low-load, or resource waste scenarios, providing an objective basis for dynamically adjusting the AHP weights and avoiding decision-making biases caused by subjective judgments.
[0036] Step 2, classify the current network status in real time through the geological classification model: Model Selection: The LVQ neural network (Learning Vector Quantization) is adopted as the geological classification model. Because of its simple structure, fast training speed and good classification effect, it is suitable for processing high-dimensional resource deviation data.
[0037] Classification Logic: Input the resource deviation data set (such as CPU utilization deviation rate, memory occupancy deviation rate, etc.), and output the network state classification labels (such as "high load", "low load", "idle", "abnormal"); through the competitive learning mechanism, the model maps the input data to the preset category centers to achieve real-time classification.
[0038] Dynamic Update: The model weight vectors are iteratively optimized regularly through historical resource deviation data to ensure the accuracy and adaptability of the classification results.
[0039] By classifying the network state in real time through the geological classification model, the system can quickly identify key scenarios (such as sudden task influx, node failure), thereby triggering the corresponding AHP weight adjustment strategy to improve the pertinence and real-time nature of the scheduling decision-making.
[0040] Step 3: Dynamically adjust the weights of the AHP judgment matrix according to the network state classification results. The weight adjustment strategies specifically include: High Load / Node Failure Scenario: Reduce the weight adjustment threshold of the AHP judgment matrix (such as from the default 10% to 5%), and enable the high-frequency dynamic weight update mode; preferentially increase the weights of indicators such as "task latency" and "resource demand" to quickly respond to resource competition problems under high load.
[0041] Low Load / Idle Node Scenario: Expand the weight allocation range (such as introducing more sub-indicators, such as "energy consumption sensitivity", "task priority"), and enable the global resource scheduling strategy; reduce the "task latency" weight and increase the "carbon efficiency" weight to optimize the energy utilization efficiency.
[0042] Sudden Task Influx Scenario: Combine task priority data (such as the task type and priority obtained in S100) to dynamically adjust the weights of the "task latency" and "resource demand" indicators; if the task priority is high, increase the "task latency" weight by 15% - 20%, and at the same time reduce the "resource demand" weight by 5% - 10%.
[0043] By dynamically adjusting the AHP weights, the system can flexibly balance the priorities of performance, cost and carbon efficiency according to different network states. For example, it gives priority to ensuring the task response speed in high load scenarios and optimizing the energy efficiency in low load scenarios, so as to achieve dynamic resource allocation and conflict avoidance of the computing power network.
[0044] S200 realizes the adaptive optimization of the computing power network resource allocation strategy through resource deviation quantification and real-time network state classification, combined with the AHP weight dynamic adjustment mechanism. Specifically, the system first calculates the resource deviation based on the operation state data and the standard reference value to identify network load anomalies; then, it uses a geological classification model (such as the LVQ neural network) to classify the current state in real time (such as high load, idle, burst task); finally, it dynamically adjusts the sub-index weights of the AHP judgment matrix according to the classification results (such as increasing the weight of task delay and reducing the weight of resource requirements for priority tasks), so as to balance the performance-cost-carbon efficiency conflict in different network scenarios. This step ensures the flexibility and pertinence of the resource scheduling strategy, and significantly improves the response efficiency and resource utilization rate of the computing power network in complex dynamic environments.
[0045] S300, based on the resource deviation data, network state classification results, and task priority data, calculates the performance-cost-carbon efficiency conflict level during the computing power scheduling process through a risk assessment algorithm; if the conflict level is a high-risk level, an emergency resource recovery strategy is enabled and the scheduling of high-energy-consuming nodes is restricted; if the conflict level is a medium-risk level, a multi-objective optimization prompt is triggered and a performance-cost-carbon efficiency trade-off plan is displayed; if the conflict level is a low-risk level, a single-index optimization suggestion is provided.
[0046] Among them, it specifically includes the following steps: Step 1, integrate the resource deviation data, network state classification results, and task priority data: Resource deviation data: The resource deviation data obtained from S200 (such as CPU utilization deviation rate, memory occupancy deviation rate), which reflects the health status and resource allocation efficiency of the current network nodes.
[0047] Network state classification results: Based on the classification labels output by the geological classification model in S200 (such as "high load", "low load", "burst task"), it provides the macroscopic characteristics of the network operation scenario.
[0048] Task priority data: The task requirement data obtained from S100 (such as task type, priority, resource requirement), which clarifies the sensitivity of the task to performance, cost, and carbon efficiency.
[0049] By integrating the resource status, network scenario, and task characteristics, the system can construct a multi-dimensional conflict assessment input framework, providing a basis for subsequent risk quantification and avoiding the one-sidedness of single-index assessment.
[0050] Step 2, calculate the performance-cost-carbon efficiency conflict level through a risk assessment algorithm: Risk assessment algorithm logic: Index weighted calculation: The improved Analytic Hierarchy Process (AHP) is used to perform weighted scoring on three aspects: performance (such as task latency), cost (such as energy consumption), and carbon efficiency (such as carbon emissions per unit task). The weights are provided by the AHP judgment matrix dynamically adjusted by S200.
[0051] Conflict quantification formula: Conflict level = (performance requirement - actual performance) + (cost budget - actual cost) + (carbon efficiency target - actual carbon efficiency). If the deviation in any dimension exceeds the threshold, conflict escalation is triggered.
[0052] Risk classification rules: High risk, the sum of the three conflicts ≥ upper threshold (such as sum > 15%), indicating that the system is on the verge of performance collapse or resource waste; Medium risk, lower threshold < sum < upper threshold (such as 5% < sum < 15%), requiring trade-off and optimization; Low risk, sum ≤ lower threshold (such as ≤ 5%), allowing local optimization.
[0053] By quantifying the conflict level, the system can accurately identify high-risk scenarios (such as resource contention under high load), and provide a decision-making basis for subsequent hierarchical scheduling strategies, avoiding rigid resource allocation.
[0054] Step 3, trigger the corresponding scheduling strategy according to the conflict level: Handling of high-risk levels: Enable the emergency resource recovery strategy to forcibly recover resources from high-energy-consuming nodes (such as shutting down non-critical tasks), and prioritize ensuring the performance requirements of high-priority tasks; Limit the scheduling of high-energy-consuming nodes. Through the blockchain verification weight mechanism, prohibit high-energy-consuming nodes from participating in new task allocation to reduce carbon efficiency conflicts.
[0055] Handling of medium-risk levels: Trigger multi-objective optimization prompts to display the trade-off plan of performance-cost-carbon efficiency to the scheduling system (such as "increase the tolerance of task latency to reduce energy consumption"); Dynamically adjust the weight allocation. Combining task priority data, temporarily increase the carbon efficiency weight (such as increasing by 5%) to guide resource allocation towards greening.
[0056] Handling of low-risk levels: Provide single-index optimization suggestions, and propose local optimization measures (such as adjusting the task scheduling order) for the weakest dimension (such as insufficient performance); Maintain global stable scheduling, keep the current AHP weight unchanged, and only fine-tune abnormal nodes (such as resource freezing).
[0057] The hierarchical strategy ensures that the system takes different responses in different risk scenarios: quickly intervene in high-risk scenarios to avoid collapse, balance multiple demands in medium-risk scenarios, and maintain efficient operation in low-risk scenarios, thus realizing the dynamic adaptive scheduling of the computing power network.
[0058] S300 precisely evaluates the performance-cost-carbon efficiency conflict level through the combination of conflict quantification and AHP dynamic weights, uses blockchain verification linkage to ensure the trustworthy execution of resource recovery strategies in high-risk scenarios, and realizes the real-time nature of multi-objective optimization prompts based on task priorities, effectively solving the problem of rigid resource allocation caused by static weight deviation, decision-making lag, and trust deficiency in traditional scheduling systems.
[0059] S400, based on resource deviation data and conflict assessment results, hierarchically triggers scheduling strategy adjustment signals. Among them, the scheduling strategy adjustment signals include green resource allocation instructions, yellow warning - dynamic weight correction, and red alert - global scheduling reset; if the pressure sensor detects node anomalies and conflict assessment is not triggered, directly initiate active scheduling intervention and restrict resource allocation for high-load nodes.
[0060] Specifically, it includes the following steps: Step 1, hierarchically trigger scheduling strategy adjustment signals based on resource deviation and conflict assessment results: Scheduling strategy hierarchical trigger logic: High-risk level processing: Enable emergency resource recovery strategies, forcibly recover resources of high-energy-consuming nodes (such as shutting down non-critical tasks), and prioritize ensuring the performance requirements of high-priority tasks; restrict the scheduling of high-energy-consuming nodes, and through the blockchain verification weight mechanism, prohibit high-energy-consuming nodes from participating in new task allocation to reduce carbon efficiency conflicts.
[0061] Medium-risk level processing: Trigger multi-objective optimization prompts, display the trade-off plan of performance-cost-carbon efficiency to the scheduling system (such as "increase task latency tolerance to reduce energy consumption"); dynamically adjust weight allocation, combine task priority data, and temporarily increase the carbon efficiency weight (such as increasing by 5%) to guide resource allocation towards greening.
[0062] Low-risk level processing: Provide single-index optimization suggestions, and propose local optimization measures (such as adjusting the task scheduling order) for the weakest dimension (such as insufficient performance); maintain global stable scheduling, keep the current AHP weights unchanged, and only make fine-tuning for abnormal nodes (such as resource freezing).
[0063] Through hierarchical strategies, ensure that the system adopts different responses in different risk scenarios: quickly intervene in high-risk scenarios to avoid crashes, balance multiple demands in medium-risk scenarios, and maintain efficient operation in low-risk scenarios, thereby realizing the dynamic adaptive scheduling of the computing power network.
[0064] Step 2, the pressure sensor detects node anomalies and triggers active scheduling intervention: Abnormal operation status determination: The pressure sensor continuously monitors indicators such as the CPU temperature, power consumption, and load fluctuation of the node; if a certain indicator (such as CPU temperature) continuously exceeds the set threshold (such as 90°C) and lasts for more than the preset duration (such as 5 minutes), it is determined as abnormal.
[0065] Active scheduling intervention mechanism: Nodes with energy consumption lower than the benchmark value (such as nodes in the energy-saving mode) retain their resource allocation permissions to ensure that critical tasks are not affected; for nodes with resource utilization exceeding the load limit (such as CPU utilization > 95%), resource freezing operations (such as suspending task scheduling) are implemented to prevent overload and collapse.
[0066] Through the real-time monitoring and active intervention of the pressure sensor, the system can avoid potential risks in advance before reaching the conflict level trigger condition, and prevent the global resource allocation imbalance caused by local node abnormalities.
[0067] Step 3, Dynamically update the AHP weight matrix based on the comparison of historical environmental parameters: Comparison of real-time performance evaluation data and historical environmental parameters: Compare and analyze the multi-dimensional performance evaluation data set (including current status and future prediction data) generated by S100 with the historical environmental parameter database (such as the load peak and task type distribution in the past 24 hours); identify the differences between the current environment and the historical environment (such as the influx of sudden tasks and the sharp drop in bandwidth).
[0068] Weight update driven by the reinforcement learning model: The reinforcement learning model (such as the deep Q network) dynamically adjusts the weights of each sub-index in the AHP judgment matrix based on the comparison results (such as increasing the weight of "task delay" to cope with sudden tasks); the weight update needs to meet the consistency test (such as CR < 0.1) to ensure that the adjusted weights are logically reasonable.
[0069] Quantum-classical hybrid algorithm and blockchain verification weight optimization: If abnormal working conditions occur (such as node failures), automatically optimize the parameters of the quantum-classical hybrid algorithm (such as adjusting the number of quantum annealing iterations) to improve the accuracy of local performance evaluation; ensure the credibility and immutability of weight adjustment through the blockchain verification weight mechanism (such as smart contracts).
[0070] By comparing historical data and dynamically updating the weight matrix through reinforcement learning, the system can adapt to complex and changeable operating environments, avoid the rigidity of resource allocation caused by static weights, and at the same time, the blockchain verification mechanism ensures the credibility of policy adjustments.
[0071] Step 4, Attribution analysis and parameter optimization of abnormal working conditions: Causal Inference Graph Construction: Based on historical operation status data and task scheduling logs, construct a causal graph among computing power network nodes (nodes represent resource units, and edges represent the causal intensity of performance metrics); identify abnormal propagation paths (such as "high latency of node A → bandwidth blockage of node B"), and quantify the causal correlation degree between the abnormal source node and the affected nodes.
[0072] Counterfactual Virtual Scenario Generation: Simulate the expected operation status when "key nodes do not have abnormalities" (such as the bandwidth utilization rate when node A is normal), and compare it with the actual operation status for analysis; output an attribution report based on the comparison results (such as "the failure of node A caused a 30% increase in task latency").
[0073] Parameter and Policy Optimization: Adjust system parameters according to the attribution report (such as increasing the redundant resource allocation of node A); generate optimization measures (such as "priority reallocation", "node isolation", "parameter adaptive adjustment"), and automatically deploy them to the scheduling policy.
[0074] Through causal inference and counterfactual analysis, the system can accurately locate the root cause of abnormalities and generate targeted optimization solutions, significantly improving the fault recovery ability and long-term stability of the computing power network.
[0075] S500, obtain real-time performance evaluation data from the multi-dimensional performance evaluation dataset, compare the real-time performance evaluation data with the historical environment parameter database, and dynamically update the AHP weight matrix through the reinforcement learning model based on the environment comparison results; if abnormal working conditions occur, conduct attribution analysis on the abnormal working conditions, and automatically optimize the parameters of the quantum-classical hybrid algorithm and the blockchain verification weight.
[0076] Among them, it specifically includes the following steps: Step 1, Monitor the execution effect of the scheduling policy in real time and generate feedback data: Performance Metric Monitoring: Real-time collect key performance indicators (KPIs) such as task completion time, resource utilization rate, and task latency, and compare them with the target values set by the scheduling policy.
[0077] Cost and Carbon Efficiency Tracking: Record the energy consumption and carbon emissions per unit task after the execution of the scheduling policy, and verify whether the cost and carbon efficiency constraint conditions are met.
[0078] Abnormal Behavior Detection: Identify abnormal behaviors (such as frequent task migrations, node overloads) during the execution of the scheduling policy through log analysis and traffic monitoring.
[0079] Through comprehensive data feedback, the system can quantify the actual effect of the scheduling policy, locate execution deviations, and avoid resource waste or performance degradation caused by policy design defects or environmental changes.
[0080] Step 2, Dynamically adjust the scheduling policy parameters based on feedback data: Fine-tuning of the weight matrix: If the feedback shows that the performance metrics (such as task latency) do not meet the standards, dynamically increase the weight of "task priority" in the AHP weight matrix to prioritize resource allocation for high-priority tasks.
[0081] Correction of resource allocation thresholds: If the feedback shows that the node load frequently exceeds the limit, adjust the resource allocation threshold (such as reducing the maximum number of tasks per node) to avoid local overload.
[0082] Update of conflict handling rules: For abnormal behaviors (such as frequent task migrations), optimize the conflict handling rules (such as introducing a migration cost penalty factor) to reduce ineffective scheduling operations.
[0083] Through dynamic parameter adjustment, the system can quickly respond to environmental changes (such as sudden task surges or node failures), ensuring the robustness and efficiency of the scheduling policy in complex scenarios.
[0084] Step 3, Introduce a reinforcement learning model to optimize the global scheduling policy: Definition of the state space: Encode the current network state (such as node load, task queue length, resource deviation rate) as the state vector of reinforcement learning.
[0085] Design of the reward function: Define a multi-objective reward function that comprehensively considers performance (+10 / task completed), cost (-5 / unit energy consumption), and carbon efficiency (-8 / unit carbon emission) to guide the model to learn the optimal policy.
[0086] Policy training and deployment: Train the model through a Deep Q-Network (DQN) to output the optimal scheduling actions (such as "migrate tasks to low-energy nodes") and deploy them to the actual scheduling system.
[0087] Based on the above steps, it is possible to autonomously learn the optimal scheduling rules in complex scenarios from historical feedback, solve the problem that traditional static policies are difficult to adapt to dynamic environments, and significantly improve resource utilization and scheduling efficiency.
[0088] Step 4, Blockchain verification and policy credibility guarantee: Policy change on the chain: Record each adjustment of the scheduling policy (such as weight matrix update, resource allocation threshold modification) as a transaction and write it into the blockchain to generate an immutable audit log.
[0089] Multi-node consensus verification: Conduct consensus verification on policy changes through multiple nodes in the consortium chain (such as network management nodes, security audit nodes) to prevent malicious tampering.
[0090] Automated execution of smart contracts: Deploy smart contract rules (such as "if the performance metrics do not improve after the policy adjustment, roll back to the previous version") to ensure that policy adjustments comply with the preset security boundaries.
[0091] Blockchain technology provides a transparent and trustworthy auditing mechanism for the adjustment of scheduling policies, avoiding policy failures caused by human intervention or malicious attacks, and ensuring the secure operation of the computing power network.
[0092] Through multi-dimensional data fusion and dynamic decision-making mechanisms, the resource scheduling efficiency and stability of the computing power network have been significantly improved. S100 combines distributed sensors and quantum-inspired algorithms to achieve low-latency real-time performance evaluation; S200 dynamically adjusts the AHP weights based on resource deviation and geological classification models to precisely balance the priorities of performance, cost, and carbon efficiency; S300 triggers through conflict quantification and hierarchical strategies to ensure the timeliness and credibility of resource recovery in high-risk scenarios; S400 and S500 introduce reinforcement learning and blockchain verification to dynamically optimize the global scheduling policy and ensure the transparency of policy adjustments. The overall process realizes closed-loop control from real-time monitoring to forward-looking prediction, solves the problem of rigid resource allocation in traditional scheduling systems due to static weight deviation, decision-making lag, and lack of trust, and enables the computing power network to have higher adaptability and stability in complex dynamic environments.
[0093] Through the collaboration of quantum computing acceleration, digital twin prediction, and blockchain trustworthy verification, an intelligent scheduling system for the computing power network is constructed. The lightweight quantum computing module quickly generates robust performance metrics at the edge side, reducing data processing latency; the digital twin combines historical tasks and user behavior simulations to improve the accuracy of future state predictions, providing a forward-looking basis for scheduling policies; the reinforcement learning model dynamically optimizes the global policy to adapt to sudden tasks and environmental changes; the blockchain verification mechanism ensures the immutability of weight adjustments and policy changes, enhancing system security. This technology combination not only breaks through the static and localized limitations of traditional scheduling systems but also significantly improves resource utilization, fault recovery capabilities, and green levels through multi-objective trade-offs and real-time feedback mechanisms, providing core support for the efficient operation of complex computing power networks.
[0094] In the embodiment of this application, when the multi-dimensional performance evaluation data set fuses local performance evaluation indicators and future network state data, the method further includes: Step 1: The system selects a high-precision time source from all sensors and digital twin modules as the master time reference, such as the sensor deployed at the core node. The master time reference should have a high sampling frequency (e.g., 100 times per second) and a low clock deviation (error less than 1 millisecond), and be calibrated through the Network Time Protocol or the Precision Time Protocol to ensure time synchronization with other devices. Subsequently, the system maps the local performance metrics (such as resource utilization balance) collected by other sensors or the future state data predicted by the digital twin (such as bandwidth trend) to the master time reference through linear interpolation or Lagrange interpolation algorithm. For example, if the sampling frequency of a sensor at an edge node is low (50 times per second), its data will be interpolated and filled into the high-frequency time points of the master time reference, thus ensuring the alignment of all data in the time dimension. This process is similar to uniformly adjusting the video frames with different frame rates to the same playback rate to ensure the synchronization of the picture and audio.
[0095] Step 2: After completing the timestamp alignment, the system detects the time deviation between the local performance metrics and the future state data; the time deviation refers to the difference between the aligned data and the original data on the time axis, for example, a sensor has a data acquisition lag due to network latency. If the detected time deviation exceeds the preset threshold (e.g., 50 milliseconds), the system will activate the compensation mechanism. For local performance metrics, the system will adopt the interpolation compensation method to infer the performance value of the missing time point through the data at known time points. For example, if the sensor of a node does not return data at a certain time point, the system will calculate the intermediate value based on the performance metrics (such as CPU utilization) at the previous and next time points to fill in the missing data. For future state data (such as bandwidth prediction), the system will reduce the data density through downsampling to match the time resolution of the local performance metrics. For example, if the digital twin model predicts a high-frequency bandwidth change trend (100 times per second), but the local performance metrics only support an update frequency of 50 times per second, the system will retain one prediction value every two time points to ensure their consistency in the time dimension. This process is similar to compressing a high-definition video into a low-resolution version to adapt to the performance limitations of the playback device while retaining the key information.
[0096] In the embodiment of the present application, during the process of hierarchically triggering the scheduling policy adjustment signal based on the resource deviation data and the conflict assessment result, if the pressure sensor detects node anomalies and does not trigger a conflict assessment, the constraint conditions for active scheduling intervention are supplemented. The method further includes: Step 1: When the pressure sensor detects that the node operating state is abnormal but does not trigger the performance-cost-carbon efficiency conflict level, the system further analyzes the resource deviation data of the node. The resource deviation data includes the differences between the actual values and the standard values of indicators such as CPU utilization, memory occupancy, network bandwidth, and task latency.
[0097] The system checks whether these metrics continuously exceed the preset resource thresholds (e.g., CPU utilization > 95%, memory occupancy > 90%) within a preset time window (e.g., 5 minutes). If at least one metric continuously exceeds the standard within the preset duration, it is determined that there is a potential risk for the node, and the active scheduling intervention mechanism needs to be activated; for example, if the CPU utilization of a certain node exceeds 95% for 5 consecutive minutes, the system will trigger an intervention to avoid the imbalance of global resource allocation caused by local overload. This process is similar to the "threshold warning" mechanism in industrial equipment, which continuously monitors the change trend of key metrics, identifies abnormalities in advance, and takes actions.
[0098] Step 2, when the resource deviation data meets the condition of continuous exceeding the standard, the system will activate the active scheduling intervention mechanism. At this time, the scheduling policy will prioritize ensuring the resource allocation of critical tasks while restricting the resource usage of high-load nodes. Specifically: Give priority to retaining the resource allocation rights of low-power consumption nodes: The system will screen out nodes with power consumption lower than the benchmark value (for example, nodes with unit task power consumption lower than the industry average), ensuring that their resource allocation rights are not affected. This measure aims to maintain the stable operation of critical tasks while optimizing energy efficiency. For example, when the power supply is tight, the system will prioritize ensuring the resource allocation of low-power consumption nodes to avoid carbon efficiency conflicts caused by energy waste.
[0099] Freeze the resource allocation of high-load nodes: For nodes with resource utilization exceeding the load limit (such as CPU utilization > 95%), the system will perform a resource freezing operation, that is, suspend the new task scheduling of this node and gradually release its allocated resources. For example, if a node's CPU is overloaded due to a sudden influx of tasks, the system will freeze its resource allocation to prevent task backlog from causing node crashes. This strategy is similar to the node resource isolation mechanism in Kubernetes, which dynamically adjusts resource allocation to avoid the spread of local failures to the entire network.
[0100] Step 3, the execution of active scheduling intervention needs to follow the following constraints: Gradualness of resource freezing: When freezing the resources of high-load nodes, the system will first recycle the resources of non-critical tasks (such as low-priority tasks), rather than directly terminating all tasks. For example, if a node is running high-priority tasks and low-priority tasks at the same time, the system will retain the resources of high-priority tasks and only freeze the allocation of low-priority tasks.
[0101] Dynamic adjustment of low-power consumption nodes: The system will regularly update the benchmark value of low-power consumption nodes (for example, adjust the benchmark according to real-time electricity prices or energy supply conditions) to ensure that the resource allocation strategy matches the current environmental conditions. For example, during the low electricity consumption period at night, the system may relax the benchmark value of low-power consumption nodes to make full use of cheap electricity.
[0102] Recovery mechanism after intervention: When the resource deviation data of an abnormal node returns to the normal range (for example, the CPU utilization rate drops below 85%), the system will gradually unfreeze its resource allocation permissions and re-incorporate it into the scheduling pool. For example, if a node's resources are frozen due to short-term overload, after its load returns to normal, the system will re-evaluate its performance metrics and restore the scheduling permissions according to the AHP weight adjustment strategy.
[0103] This mechanism realizes the proactive intervention of abnormal nodes without triggering conflict evaluation through dynamic threshold judgment and resource freezing / restoring strategies, effectively balancing system stability and resource utilization efficiency.
[0104] In the embodiment of this application, during the process of dynamically adjusting the sub-index weights of the AHP judgment matrix based on resource deviation data, network status classification results, and task priority data, when detecting the influx of burst tasks, the method further includes: Step 1, when detecting the influx of burst tasks, the system will obtain the priority data of the current task flow through the log system to judge the urgency of the tasks. The log system will record the label of each task (such as "high priority", "medium priority", "low priority") or a dynamic score (such as a value generated according to the task deadline, resource demand intensity, etc.). If it is determined that the task priority is high (for example, exceeding a preset threshold, such as the priority score > 8), the system will trigger the weight adjustment process. This process is similar to detecting the product type (such as standard parts or customized parts) through sensors in a factory production line and adjusting the processing priority according to the type to ensure that high-value products are processed first.
[0105] Step 2, after confirming that the task priority is high, the system will dynamically adjust the sub-index weights in the AHP judgment matrix: Increase the weight of the task delay index: Increase the relative importance weight of the task delay index (such as task response time, waiting time) by 15% - 20%. For example, if the original weight of the task delay index in the judgment matrix is 0.3, the adjusted weight becomes 0.345 - 0.36 (the specific value depends on the adjustment range). This adjustment reflects the strong demand of high-priority tasks for quick response, similar to opening a green channel for emergency vehicles in traffic scheduling to give priority to ensuring their passing efficiency.
[0106] Decrease the weight of the resource demand index: Decrease the weight of the resource demand index (such as CPU occupancy rate, memory consumption) by 5% - 10%. For example, if the original weight is 0.25, it becomes 0.225 - 0.2375 after adjustment. This adjustment indicates that in the scenario of high-priority tasks, the system is more inclined to sacrifice some resource efficiency to meet the task timeliness requirements, similar to reserving capacity for key equipment in power dispatching, even if the overall resource utilization rate drops slightly.
[0107] Step 3: After completing the weight adjustment, the system will use the consistency ratio test method to verify whether the updated AHP judgment matrix is logically reasonable. The specific process is as follows: Calculate the consistency ratio: The system will first determine the maximum eigenvalue of the adjusted judgment matrix, and then calculate the consistency index in combination with the order of the matrix (i.e., the dimension of the matrix). For example, if the order of the matrix is 4, the system will derive the specific value of the consistency index based on the relationship between the maximum eigenvalue and the order. Subsequently, the system will obtain the random consistency index corresponding to the matrix order through table lookup (this index is a pre-calculated and stored reference value), and further calculate the consistency ratio. If the consistency ratio is less than 0.1, it indicates that the logical consistency of the adjusted judgment matrix is acceptable; if it is greater than or equal to 0.1, the weight allocation needs to be corrected.
[0108] Logical rationality confirmation and application: Only when the consistency ratio is less than 0.1 will the system confirm that this weight adjustment is logically reasonable and apply the adjusted weights to subsequent performance evaluation and scheduling decisions. For example, if the adjusted consistency ratio is 0.08, it indicates that the weight allocation conforms to the expert judgment logic and can be used to guide resource scheduling; if the consistency ratio is 0.12, it is necessary to roll back to the original weights or readjust to avoid decision-making biases caused by logical contradictions. This process is similar to an engineer verifying the rationality of the structural stress distribution through finite element analysis when designing a bridge, and only approving the construction plan when the simulation results meet the safety factor.
[0109] In the embodiment of the present application, during the process of simulating future network state data based on digital twins, the method further includes: Step 1: During the process of simulating future network state data based on digital twins, the system will first construct a 3D visual topology model that corresponds one-to-one with the actual computing power network. This model dynamically models by collecting the topology structure of the physical network (such as node positions, connection relationships), device configuration information (such as bandwidth limits, computing capabilities), and real-time operating status (such as current task loads, link utilization rates). For example, in the scenarios of campus networks or enterprise data centers, the system will map physical devices such as servers, switches, and routers to nodes in the virtual model, and intuitively display their geographical locations and connection relationships through a three-dimensional spatial layout.
[0110] To achieve the synchronization of the model with the real network, the system will continuously receive real-time data streams from sensors, network management platforms, or the Telemetry protocol, and update the connection status between nodes (such as link failures or recoveries), bandwidth change trends (such as congestion caused by burst traffic), and task load distributions (such as a node's CPU utilization soars due to high-concurrency requests) according to the data. This process is similar to the real-time vibration data synchronization in a bridge health monitoring system, ensuring the dynamic consistency between the model and the physical entity by continuously updating the model parameters.
[0111] Step 2, based on the 3D visualization topology model, the system deploys a virtual user behavior simulation module to generate dynamic task demand data required for future network state prediction. The core logic of this module includes: Retrieving historical task patterns: The system extracts historical task data (such as user access frequency, task type distribution, peak period patterns) from the log system or database and identifies typical task patterns based on statistical analysis. For example, in an e-commerce scenario, the system may find that the shopping cart settlement requests during "Double Eleven" show periodic peaks.
[0112] Generating user behavior characteristics: Combining collaborative filtering algorithms or machine learning models (such as LSTM), the system extracts user behavior characteristics (such as preferred product categories, operation path durations, payment method selections) from historical task data. For example, for a certain user group, the system may find that they tend to complete a purchase after browsing 3 product pages.
[0113] Simulating dynamic task demands: Based on historical task patterns and user behavior characteristics, the virtual user behavior simulation module generates dynamic task demand data (such as simulating the access volume, task type proportion, resource request intensity within the next 1 hour). For example, when predicting the future network state, the system may simulate a scenario where "500 users will simultaneously initiate video conference requests at 10 am".
[0114] This process is similar to the simulation test of a factory production line. By presetting historical production data and equipment behavior patterns, it predicts production capacity bottlenecks or failure risks in advance.
[0115] Step 3, the system performs fusion analysis on the simulated dynamic task demand data and real-time operation status data to output future network state data with extended spatio-temporal dimensions. Specifically: Spatio-temporal dimension modeling: The system constructs a multi-dimensional spatio-temporal model by combining the time characteristics (such as the periodicity of task initiation) and spatial characteristics (such as the geographical distribution of task requests) of the dynamic task demand data. For example, for a sudden influx of tasks in a certain area, the system predicts its load diffusion effect on surrounding nodes.
[0116] Data fusion and prediction: Through the simulation engine of the digital twin platform, the system inputs the dynamic task demand data (such as simulated user behavior) and real-time operation status data (such as the current node load) into the prediction model to generate future network state data (such as bandwidth occupancy rate, task delay, node overload risk). For example, when predicting the network state in the next 5 minutes, the system may output a conclusion that "the CPU utilization rate of node A will increase from 70% to 95%, and resource scheduling intervention needs to be triggered".
[0117] Finally, the system will present the prediction results in the form of spatio-temporal dimension expansion. For example, in a 3D visualization topology model, it will mark the network bottleneck area at a future time point (such as highlighting the link with insufficient bandwidth in red), or generate a timeline view to show the change trend of key performance indicators (such as task latency). This process is similar to the typhoon path prediction in weather forecasting, which warns of potential risks in advance through spatio-temporal data modeling.
[0118] In the embodiments of this application, when an abnormal working condition occurs and attribution analysis is required, the method further includes: Step 1, the system constructs a causal inference graph between computing power network nodes based on historical operating state data and task scheduling logs. Among them, the nodes in the causal inference graph represent computing power resource units (such as servers, storage devices), and the edges represent the causal relationships between performance indicators (such as task latency, bandwidth fluctuation). For example, if the CPU overload of node A causes an increase in the latency of node B, a causal edge is formed between the two, and the weight value reflects the causal intensity (quantified through a statistical model). The causal inference graph is automatically generated through time series analysis and task dependencies, and is used to describe the dynamic associations between nodes.
[0119] Step 2, the system locates the abnormal propagation path based on the causal inference graph and quantifies the causal association degree between the abnormal source node and the downstream nodes. For example, if an abnormal bandwidth fluctuation is detected in node C, the system will trace its upstream nodes (such as resource contention in node B) and downstream nodes (such as the latency of node D) to form a complete propagation chain. Combining a statistical model (such as linear regression) or physical simulation, the system quantifies the influence degree of the abnormal source on the downstream nodes in each path (such as the CPU overload of node A causes a 20% increase in the latency of node B, and the causal association degree is 0.8).
[0120] Step 3, the system generates a counterfactual virtual scenario based on the abnormal propagation path to simulate the expected operating state when the key nodes do not have abnormalities. For example, assuming that the CPU overload of node A does not occur, adjust its resource allocation strategy, and predict the task latency and energy consumption distribution. Subsequently, compare the simulation results with the actual operating state in multiple dimensions (such as latency difference, resource utilization change) to identify the root cause of the abnormality (such as resource contention in node A causes a 30% increase in latency), and generate an attribution report.
[0121] According to the attribution report, the system proposes root cause-oriented optimization measures: Priority reallocation, adjust the task scheduling priority to ensure the execution of key tasks first (such as reducing the resource allocation ratio of low-priority tasks in node A); node isolation, isolate the abnormal source nodes with high causal association degree (such as migrating the high-fluctuation tasks of node B to redundant nodes); parameter adaptive adjustment, dynamically adjust network parameters (such as optimizing the bandwidth allocation strategy or increasing the cache capacity) to alleviate the abnormal impact.
[0122] Through causal inference graph-driven dynamic analysis, combined with counterfactual simulation to accurately locate the root cause, and linking multi-dimensional optimization measures (task scheduling, resource management, network layer), significantly improve the fault response efficiency and stability of the computing power network.
[0123] An embodiment of this application discloses a real-time performance evaluation system for a computing power network based on the analytic hierarchy process, referring to Figure 2 , including: A multi-dimensional performance evaluation data set generation module 001, which obtains the operation status data and task demand data of the computing power network nodes in real time through a distributed sensor and a log system. The operation status data includes CPU utilization rate, memory occupancy rate, network bandwidth, task latency, and energy consumption. The task demand data includes task type, priority, and resource requirements; processes the operation status data through a quantum-inspired algorithm deployed on the edge node and generates local performance evaluation indicators; based on digital twin to simulate future network status data, the future network status data includes bandwidth change trend and task load prediction, and fuses the local performance evaluation indicators with the future network status data and uses them to generate a multi-dimensional performance evaluation data set including the current state and future prediction, where the sensor includes a lightweight quantum computing module deployed at the edge of the node; A network status classification result acquisition module 002, which calculates the corresponding resource deviation data based on the operation status data and a preset standard performance reference value, and classifies the current network status in real time through a geological classification model to obtain the network status classification result; if a node failure or high load is detected, reduce the weight adjustment threshold of the AHP judgment matrix and enable the high-frequency dynamic weight update mode; if a low load or idle node is detected, expand the weight allocation range and enable the global resource scheduling strategy; if a sudden task influx is detected, dynamically adjust the AHP sub-index weight in combination with the task priority data; A conflict level calculation module 003, which calculates the performance-cost-carbon efficiency conflict level in the computing power scheduling process through a risk assessment algorithm based on the resource deviation data, network status classification result, and task priority data; if the conflict level is a high-risk level, enable the emergency resource recovery strategy and restrict the scheduling of high-energy-consuming nodes; if the conflict level is a medium-risk level, trigger a multi-objective optimization prompt and display the performance-cost-carbon efficiency trade-off plan; if the conflict level is a low-risk level, provide a single-index optimization suggestion; A policy adjustment signal scheduling module 004, which is used to hierarchically trigger scheduling policy adjustment signals based on the resource deviation data and conflict assessment results, where the scheduling policy adjustment signals include green resource allocation instructions, yellow warning-dynamic weight correction, and red alert-global scheduling reset; if the pressure sensor detects node anomalies and no conflict assessment is triggered, directly start active scheduling intervention and restrict the resource allocation of high-load nodes; The weight matrix update module 005 obtains the real-time performance evaluation data in the multi-dimensional performance evaluation dataset, compares the real-time performance evaluation data with the historical environment parameter database, and uses a reinforcement learning model to dynamically update the AHP weight matrix based on the environment comparison result; if an abnormal working condition occurs, an attribution analysis is performed on the abnormal working condition, and the parameters of the quantum-classical hybrid algorithm and the blockchain verification weight are automatically optimized.
[0124] An embodiment of the present application also discloses a real-time performance evaluation system for a computing power network based on the analytic hierarchy process, including a processor, and a program of the real-time performance evaluation method for a computing power network based on the analytic hierarchy process described in any one of the above is run in the processor.
[0125] An embodiment of the present application also discloses a storage medium storing a program of the real-time performance evaluation method for a computing power network based on the analytic hierarchy process described in any one of the above.
[0126] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A real-time performance evaluation method for a computing power network based on the analytic hierarchy process, characterized in that Including: Obtain the operation status data and task requirement data of the computing power network nodes in real time through a distributed sensor and a logging system. The operation status data includes CPU utilization rate, memory occupancy rate, network bandwidth, task latency, and energy consumption. The task requirement data includes task type, priority, and resource requirements. Process the operation status data through a quantum-inspired algorithm deployed on an edge node and generate local performance evaluation metrics. Based on digital twin to simulate future network status data, where the future network status data includes bandwidth change trend and task load prediction, fuse the local performance evaluation metrics with the future network status data and generate a multi-dimensional performance evaluation data set including the current state and future prediction. Among them, the sensor includes a lightweight quantum computing module deployed at the edge of the node; Calculate the corresponding resource deviation data based on the operation status data and a preset standard performance reference value, and classify the current network status in real time through a geological classification model to obtain the network status classification result. If a node failure or high load is detected, reduce the weight adjustment threshold of the AHP judgment matrix and enable the high-frequency dynamic weight update mode. If a low load or idle node is detected, expand the weight allocation range and enable the global resource scheduling strategy. If a sudden task influx is detected, dynamically adjust the AHP sub-index weights in combination with the task priority data; Based on the resource deviation data, network status classification result, and task priority data, calculate the performance-cost-carbon efficiency conflict level in the computing power scheduling process through a risk assessment algorithm. If the conflict level is a high-risk level, enable the emergency resource recovery strategy and restrict the scheduling of high-energy-consuming nodes. If the conflict level is a medium-risk level, trigger a multi-objective optimization prompt and display the performance-cost-carbon efficiency trade-off plan. If the conflict level is a low-risk level, provide a single-index optimization suggestion; Based on the resource deviation data and conflict assessment result, trigger a scheduling strategy adjustment signal at different levels. Among them, the scheduling strategy adjustment signal includes a green resource allocation instruction, a yellow warning-dynamic weight correction, and a red alert-global scheduling reset. If the pressure sensor detects a node anomaly and the conflict assessment is not triggered, directly start the active scheduling intervention and restrict the resource allocation of high-load nodes; Obtain the real-time performance evaluation data in the multi-dimensional performance evaluation data set, compare the real-time performance evaluation data with the historical environment parameter database, and dynamically update the AHP weight matrix through a reinforcement learning model based on the environment comparison result. If an abnormal working condition occurs, conduct an attribution analysis of the abnormal working condition and automatically optimize the parameters of the quantum-classical hybrid algorithm and the blockchain verification weight.
2. The real-time performance evaluation method of the computing power network based on the analytic hierarchy process according to claim 1, characterized in that, When the multi-dimensional performance evaluation data set fuses the local performance evaluation metrics and the future network status data, the method further includes: Synchronize the local performance metrics and the future state data in the time dimension through a timestamp alignment mechanism; Obtain the time deviation data between the local performance metrics and the future state data in the time dimension. If the time deviation data exceeds the preset time deviation threshold, perform interpolation compensation on the local performance metrics or down-sampling processing on the future state data.
3. The real-time performance evaluation method of the computing power network based on the analytic hierarchy process according to claim 2, characterized in that, In the process of hierarchically triggering the scheduling policy adjustment signal based on resource deviation data and conflict assessment results, if the pressure sensor detects node anomalies and no conflict assessment is triggered, the constraints for active scheduling intervention are supplemented. The method further includes: When the pressure sensor detects that a node has an abnormal operating state and the abnormal operating state does not reach any risk trigger threshold of the performance-cost-carbon efficiency conflict level, it is determined whether the resource deviation data of the abnormal node continuously exceeds the preset resource threshold; If the time for at least one indicator in the resource deviation data to continuously exceed the set resource threshold reaches the preset duration, an active scheduling intervention mechanism is triggered; When performing active scheduling intervention, the resource allocation permissions corresponding to nodes with energy consumption lower than the benchmark value are preferentially retained, and at the same time, resource freezing operations are performed on high-load nodes with resource utilization rates exceeding the load limit.
4. The real-time performance evaluation method of the computing power network based on the analytic hierarchy process according to claim 1, characterized in that In the process of dynamically adjusting the sub-index weights of the AHP judgment matrix based on resource deviation data, network status classification results, and task priority data, when a sudden influx of tasks is detected, the method further includes: Identifying the urgency of the current task flow based on the task priority data obtained from the log system; If the task priority is determined to be high, the relative importance weight of the task delay indicator is increased in the AHP judgment matrix, with an increase range of 15% - 20%, and at the same time, the weight of the resource requirement indicator is decreased, with a decrease range of 5% - 10%; After the weight adjustment is completed, the consistency ratio test method is used to verify the consistency of the updated AHP judgment matrix. Only when the consistency ratio CR is less than 0.1, it is confirmed that the secondary weight adjustment is logically reasonable and will be applied to subsequent performance evaluation and scheduling decisions.
5. The real-time performance evaluation method of the computing power network based on the analytic hierarchy process according to claim 1, wherein, In the process of simulating future network status data based on digital twins, the method further includes: Constructing a 3D visualization topology model corresponding one-to-one to the actual computing power network, and synchronously updating the connection status, bandwidth change trend, and task load distribution between nodes through real-time data streams; Deploying a virtual user behavior simulation module in the virtual environment corresponding to the 3D visualization topology model, retrieving the corresponding historical task patterns, obtaining the corresponding user behavior characteristics, and generating dynamic task demand data based on the historical task patterns and the user behavior characteristics. The dynamic task demand data is used as the input parameter for future network status prediction; Performing fusion analysis on the dynamically generated dynamic task demand data and the operating status data, and outputting future network status data including extended spatio-temporal dimensions.
6. The real-time performance evaluation method of the computing power network based on the analytic hierarchy process according to claim 1, characterized in that, When an abnormal working condition occurs and attribution analysis is required, the method further includes: Based on historical operating status data and task scheduling logs, constructing a causal inference graph between nodes in the computing power network. Among them, the nodes in the causal inference graph represent computing power resource units, and the edges represent the causal intensity between task delay, energy consumption change, and bandwidth fluctuation. Among them, task delay, energy consumption change, and bandwidth fluctuation are all performance indicators; Identifying the corresponding abnormal propagation path based on the causal inference graph, and quantifying the causal correlation degree between the abnormal source node and the downstream affected nodes based on the abnormal propagation path; Generate a counterfactual virtual scenario based on the abnormal propagation path, simulate the expected operating state corresponding to the case where the key node does not have an abnormality, and compare and analyze the expected operating state with the actual operating state; output an attribution report based on the comparison result to assist in generating root cause-oriented optimization measures, and the optimization measures include priority reallocation, node isolation, or parameter adaptive adjustment.
7. A real-time performance evaluation system for a computing power network based on the analytic hierarchy process, characterized in that, Including: A multi-dimensional performance evaluation dataset generation module that obtains the operating state data and task requirement data of the computing power network nodes in real time through a distributed sensor and a logging system. The operating state data includes CPU utilization rate, memory occupancy rate, network bandwidth, task latency, and energy consumption. The task requirement data includes task type, priority, and resource requirements; processes the operating state data through a quantum-inspired algorithm deployed on the edge node and generates local performance evaluation indicators; based on digital twin to simulate future network state data, the future network state data includes bandwidth change trend and task load prediction, and fuses the local performance evaluation indicators with the future network state data and uses them to generate a multi-dimensional performance evaluation dataset including the current state and future prediction, where the sensor includes a lightweight quantum computing module deployed at the node edge; A network state classification result acquisition module that calculates the corresponding resource deviation data based on the operating state data and a preset standard performance reference value, and classifies the current network state in real time through a geological classification model to obtain the network state classification result; if a node failure or high load is detected, reduce the weight adjustment threshold of the AHP judgment matrix and enable the high-frequency dynamic weight update mode; if a low load or idle node is detected, expand the weight allocation range and enable the global resource scheduling strategy; if a sudden task influx is detected, dynamically adjust the AHP sub-index weights in combination with the task priority data; A conflict level calculation module that calculates the performance-cost-carbon efficiency conflict level in the computing power scheduling process through a risk assessment algorithm based on the resource deviation data, network state classification result, and task priority data; if the conflict level is a high-risk level, enable the emergency resource recovery strategy and restrict the scheduling of high-energy-consuming nodes; if the conflict level is a medium-risk level, trigger a multi-objective optimization prompt and display the performance-cost-carbon efficiency trade-off plan; if the conflict level is a low-risk level, provide a single-index optimization suggestion; A policy adjustment signal scheduling module that hierarchically triggers scheduling policy adjustment signals based on the resource deviation data and conflict assessment results, where the scheduling policy adjustment signals include green resource allocation instructions, yellow warning-dynamic weight correction, and red alert-global scheduling reset; if the pressure sensor detects a node abnormality and the conflict assessment is not triggered, directly start the active scheduling intervention and restrict the resource allocation of high-load nodes; The weight matrix update module obtains real-time performance evaluation data from a multi-dimensional performance evaluation dataset, compares the real-time performance evaluation data with a historical environmental parameter database, and uses a reinforcement learning model to dynamically update the AHP weight matrix based on the environmental comparison result; if an abnormal working condition occurs, an attribution analysis is performed on the abnormal working condition, and the parameters of the quantum-classical hybrid algorithm and the blockchain verification weight are automatically optimized.
8. A real-time performance evaluation system for a computing power network based on the analytic hierarchy process, characterized in that, It includes a processor, and a program of the real-time performance evaluation method of the computing power network based on the analytic hierarchy process as described in any one of claims 1-6 runs in the processor.
9. A storage medium, characterized in that, It stores a program of the real-time performance evaluation method of the computing power network based on the analytic hierarchy process as described in any one of claims 1-6.
Citation Information
Patent Citations
Service quality benchmarking test evaluation method and device of 4G / 5G mobile communication network, computer equipment and storage medium
CN113114500A
Network anomaly detection and automatic processing method and system, storage medium and equipment
CN119892476A
Node stability assessment method and system for terminal computing power network
CN120075094A
Cited By
Public transportation travel service big data processing method and system
CN120579797A
Internet of Things card sampling priority scheduling system driven by behavior analysis
CN120676466A
Tunnel anomaly detection and response method and equipment based on rail robot
CN120688702A
Cloud computing task scheduling method and system
CN120723414A
Network security protection method and system based on block chain
CN120856462A