Multi-Agent Based Network Resource Hierarchical Coordination Method and System

Through the multi-agent network resource hierarchical coordination method, the problems of insufficient monitoring accuracy, inflexible resource pool division, slow scheduling response, inaccurate load classification and insufficient energy consumption optimization in the existing technology are solved, and efficient, intelligent and energy-saving network resource management is achieved, improving network resource utilization and system performance.

CN120091245BActive Publication Date: 2025-07-11HENAN YINHENG SOFTWARE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510541152.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-11
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing network resource scheduling technology has insufficient monitoring accuracy, inflexible resource pool division, slow scheduling decision-making response, inaccurate load classification, insufficient multi-level scheduling coordination, and insufficient energy consumption optimization in high-speed network environments, resulting in low network resource utilization and limited system performance improvement.

Method used

The network resource hierarchical coordination method based on multiple agents is adopted, and the node load and link utilization are monitored through the optical switching monitoring module, the resource control pool is built using virtualization technology, the resource control agent and the central coordination control unit are deployed, and feature extraction and multi-level classification adjustment are performed in combination with artificial intelligence algorithms, and the tactical layer and operation layer dual-layer adjustment strategies are implemented, and multi-dimensional energy consumption control is carried out to realize distributed decision-making and efficient resource management.

Benefits of technology

It realizes high-precision network status monitoring, flexible resource configuration, fast scheduling response, accurate load identification and energy consumption optimization, which significantly improves network resource utilization and overall system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091245B_ABST
    Figure CN120091245B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of resource scheduling control, and discloses a multi-agent-based hierarchical coordination method and system for network resources. The method includes: monitoring the load of network nodes to obtain resource status parameters; virtually constructing a resource pool and dividing adjustment domains; deploying agents to construct an adaptive control system; extracting load characteristics to generate an analysis report; executing a two-layer adjustment strategy and recording the execution process; analyzing energy consumption data to optimize control strategy parameters. This application can achieve efficient utilization of network resources and energy consumption optimization while ensuring the performance of the computer network system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of resource scheduling control, and particularly to a multi-agent based hierarchical coordination method and system for network resources. Background Art

[0002] With the rapid development of computer network technology, the scale and complexity of the network are continuously increasing, and network resource scheduling has become a key factor affecting system performance, reliability, and energy consumption. Existing computer network resource scheduling technologies mainly include static resource allocation methods, threshold-based dynamic scheduling methods, and prediction-based resource scheduling methods. Static resource allocation methods allocate network resources to different tasks and users according to preset rules, which are simple to implement and convenient to manage, but lack flexibility; threshold-based dynamic scheduling methods trigger resource reallocation when the system load exceeds a preset threshold by monitoring the system load, and can adapt to a certain degree of load changes; prediction-based resource scheduling methods use historical data to predict future loads and adjust resource allocation in advance, with a certain degree of foresight. These methods have achieved certain effects in traditional network environments and are widely used in resource management of data centers, cloud computing, and distributed systems.

[0003] However, existing network resource scheduling technologies face various limitations and challenges. First, traditional monitoring systems have problems of insufficient accuracy and high monitoring overhead in high-speed network environments, and it is difficult to accurately capture instantaneous changes in network states; second, existing virtualized resource pools often adopt static partitioning methods and lack the ability to dynamically adjust according to network topologies and traffic characteristics, resulting in low resource utilization; third, existing scheduling systems mostly adopt centralized decision-making methods and face response delays and single-point failure risks in large-scale network environments; fourth, the detection and classification accuracy of dynamic workloads is insufficient, and it is difficult to accurately identify load change trends and characteristics, affecting the accuracy of scheduling decisions; fifth, the execution of scheduling strategies lacks a multi-level coordination mechanism and cannot simultaneously take into account short-term and long-term resource scheduling requirements; finally, existing technologies generally ignore energy consumption optimization and fail to effectively reduce energy consumption while ensuring performance, resulting in resource waste and increased operating costs. These deficiencies seriously restrict the efficient utilization of network resources and the improvement of the overall system performance. Summary of the Invention

[0004] This application provides a multi-agent based hierarchical coordination method and system for network resources, which is used to achieve efficient utilization of network resources and energy consumption optimization while ensuring the performance of computer network systems.

[0005] In a first aspect, the present application provides a hierarchical coordination method for network resources based on multi - agents. The hierarchical coordination method for network resources based on multi - agents includes: monitoring the node processor load and link utilization rate in a computer network through an optical switching monitoring module to obtain a set of network resource status control parameters; constructing a resource control pool using virtualization technology based on the set of network resource status control parameters and dividing resource adjustment domains to obtain a resource control status report and an adjustment domain division scheme; deploying resource control agents and a central coordination control unit in each adjustment domain according to the resource control status report and the adjustment domain division scheme to obtain a multi - agent adaptive control system; using the multi - agent adaptive control system to perform feature extraction and multi - level classification adjustment on network workloads to obtain a workload control analysis report; executing a two - layer adjustment strategy at the tactical layer and the operation layer through an optical switching controller based on the workload control analysis report to obtain an adjustment execution control record; performing multi - dimensional control parameter analysis on processor energy consumption and network transmission energy consumption based on the adjustment execution control record to obtain an energy efficiency adjustment report and control strategy parameters.

[0006] In a second aspect, the present application provides a hierarchical coordination system for network resources based on multi - agents. The hierarchical coordination system for network resources based on multi - agents includes:

[0007] A monitoring module, configured to monitor the node processor load and link utilization rate in a computer network through an optical switching monitoring module to obtain a set of network resource status control parameters;

[0008] A division module, configured to construct a resource control pool using virtualization technology based on the set of network resource status control parameters and divide resource adjustment domains to obtain a resource control status report and an adjustment domain division scheme;

[0009] A control module, configured to deploy resource control agents and a central coordination control unit in each adjustment domain according to the resource control status report and the adjustment domain division scheme to obtain a multi - agent adaptive control system;

[0010] An extraction module, configured to perform feature extraction and multi - level classification adjustment on network workloads using the multi - agent adaptive control system to obtain a workload control analysis report;

[0011] An adjustment module, configured to execute a two - layer adjustment strategy at the tactical layer and the operation layer through an optical switching controller based on the workload control analysis report to obtain an adjustment execution control record;

[0012] An analysis module, configured to perform multi - dimensional control parameter analysis on processor energy consumption and network transmission energy consumption based on the adjustment execution control record to obtain an energy efficiency adjustment report and control strategy parameters.

[0013] In a third aspect, a multi-agent-based network resource hierarchical coordination device is provided, including: a memory and at least one processor, wherein instructions are stored in the memory; the at least one processor invokes the instructions in the memory to enable the multi-agent-based network resource hierarchical coordination device to execute the above-mentioned multi-agent-based network resource hierarchical coordination method.

[0014] In a fourth aspect, a computer-readable storage medium is provided, in which instructions are stored. When it runs on a computer, it enables the computer to execute the above-mentioned multi-agent-based network resource hierarchical coordination method.

[0015] In the technical solution provided by this application, by monitoring the node processor load and link utilization rate in a computer network through an optical switching monitoring module, high-precision and low-interference network status monitoring is achieved, solving the problems of insufficient accuracy and high monitoring overhead in traditional monitoring systems, providing an accurate data basis for resource scheduling, and at the same time, the optical switching monitoring technology ensures that the monitoring process does not affect normal network transmission; based on the network resource status control parameter set, a resource control pool is constructed using virtualization technology and resource adjustment domains are divided, breaking the limitations of traditional physical resource boundaries, achieving unified management and flexible configuration of resources, improving resource utilization rate, and the division of resource adjustment domains takes into account resource affinity and service relevance, creating conditions for the efficient scheduling of resources; according to the resource control status report and the adjustment domain division scheme, resource control agents and a central coordination control unit are deployed in each adjustment domain, constructing a multi-agent adaptive control system, which adopts a distributed decision-making mechanism, solving the problems of slow response and high single-point failure risk in a centralized scheduling system, and high-speed collaborative decision-making is achieved among multiple agents through an optical switching network, ensuring the fast response ability of the scheduling system to network state changes; the multi-agent adaptive control system is used to extract features and perform multi-level classification and adjustment on network workloads, and accurate identification and classification of workloads are achieved through artificial intelligence algorithms. The algorithm features play a key role in the process of workload classification. The hybrid classifier combines the interpretability of decision trees, the classification efficiency of support vector machines, and the feature extraction ability of deep neural networks, significantly improving classification accuracy; according to the workload control analysis report, a two-layer adjustment strategy of the tactical layer and the operation layer is executed through an optical switching controller, realizing the coordination of long-term resource planning and immediate resource allocation. The two-layer scheduling mechanism makes full use of the low-latency characteristics of the optical switching network to achieve millisecond-level reconfiguration of resources; based on the adjustment execution control record, multi-dimensional control parameter analysis is performed on the processor energy consumption and network transmission energy consumption. Through a multi-variable energy consumption prediction model and a multi-objective optimization framework, effective reduction of energy consumption is achieved while ensuring performance. Artificial intelligence algorithms play an important role in the energy consumption prediction and optimization process. Through comprehensive analysis of various variables such as device power consumption characteristics, resource utilization rate, and environmental temperature, a high-precision energy consumption prediction model is constructed. The genetic algorithm effectively balances the performance and energy efficiency goals in the multi-objective optimization process, achieving global optimization of the scheduling strategy. The entire technical solution combines optical switching technology, virtualization technology, and artificial intelligence algorithms organically to construct a set of efficient, intelligent, and energy-saving computer network resource scheduling methods, solving problems such as insufficient monitoring accuracy, inflexible resource pool division, slow scheduling decision response, inaccurate load classification, insufficient multi-level scheduling coordination, and insufficient energy consumption optimization in the prior art, and significantly improving network resource utilization rate and overall system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings used in the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a schematic diagram of an embodiment of the multi-agent-based network resource hierarchical coordination method in the embodiments of the present application;

[0018] Figure 2 It is a schematic diagram of an embodiment of the multi-agent-based network resource hierarchical coordination system in the embodiments of the present application;

[0019] Figure 3 It is a structural schematic block diagram of the multi-agent-based network resource hierarchical coordination device in the embodiments of the present invention. Detailed implementation manners

[0020] The embodiments of the present application provide a multi-agent-based network resource hierarchical coordination method and system. The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "comprising" or "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0021] For ease of understanding, the following describes the specific process of the embodiments of the present application. Please refer to Figure 1 , an embodiment of the multi-agent-based network resource hierarchical coordination method in the embodiments of the present application includes:

[0022] Step S101, monitor the node processor load and link utilization rate in the computer network through an optical switching monitoring module to obtain a set of network resource status control parameters;

[0023] Step S102, based on the set of network resource status control parameters, use virtualization technology to construct a resource control pool and perform resource adjustment domain division to obtain a resource control status report and an adjustment domain division plan;

[0024] Step S103: Deploy resource control agents and a central coordination control unit in each adjustment domain according to the resource control status report and the adjustment domain division scheme to obtain a multi-agent adaptive control system;

[0025] Step S104: Use the multi-agent adaptive control system to perform feature extraction and multi-level classification adjustment on the network workload to obtain a workload control analysis report;

[0026] Step S105: Execute the double-layer adjustment strategy of the tactical layer and the operation layer through the optical switch controller based on the workload control analysis report to obtain an adjustment execution control record;

[0027] Step S106: Perform multi-dimensional control parameter analysis on the processor energy consumption and network transmission energy consumption based on the adjustment execution control record to obtain an energy efficiency adjustment report and control strategy parameters.

[0028] It can be understood that the execution entity of this application can be a multi-agent-based network resource hierarchical coordination system, or a terminal or a server. Specifically, it is not limited here. This application embodiment is described by taking the server as the execution entity as an example.

[0029] Specifically, the optical switching monitoring module monitors the node processor load and link utilization in the computer network. The optical switching monitoring module is a monitoring device based on fiber optic technology, deployed at key nodes of the network, and realizes lossless monitoring of data streams through wavelength division multiplexing technology. These monitoring modules collect key indicators such as link utilization, packet transmission delay, and queue length, and set the collection frequency to be dynamically adjustable. When the network traffic fluctuates greatly, the sampling frequency will automatically increase to the millisecond level to capture instantaneous traffic changes. In practical applications, when burst traffic appears in the network, the sampling frequency is adjusted from the standard once per second to once per millisecond, thus accurately capturing the network traffic peak data. After adding timestamp, optical path identifier, and topological location information to these collected raw data, they are transmitted to the central data processing unit through a dedicated optical transmission channel, and form a set of network resource status control parameters after filtering, denoising, and standardization processing. Abstract the physical network resources, use resource description language to standardize the description of computing resources, storage resources, and network bandwidth resources, and form a resource metadata model. Subsequently, establish a mapping relationship table from physical resources to virtual resources, and assign a unique identifier to each virtual resource unit through the optical switching controller to obtain a resource mapping data structure. According to the topological information and traffic characteristics in the network resource status control parameter set, calculate the resource correlation matrix to obtain the resource affinity index. Based on these affinity indexes, perform hierarchical clustering analysis on the resource control pool, and classify resources with a correlation higher than the preset threshold of 0.75 into the same resource adjustment domain to form an initial adjustment domain division result. Optimize the boundary of this result, calculate the internal connectivity and external isolation degree of each adjustment domain, and adjust the attribution of boundary nodes to obtain an optimized resource adjustment domain division. For different service types, create resource configuration templates, and finally generate a resource control status report and an adjustment domain division plan. Assign a unique identifier to each resource adjustment domain and determine its physical boundary to obtain an adjustment domain control mapping table. Analyze the resource data in the resource control status report to determine the resource control complexity level of each adjustment domain, and form a hierarchical control agent model. According to this model, deploy resource control agents with a three-layer structure of a sensing layer, a decision-making layer, and an execution layer in each resource adjustment domain to form a hierarchical agent control network. Load the reinforcement learning algorithm into the decision-making layer of the hierarchical agent control network, set the resource utilization rate, task completion time, and system response delay as reward signals, and the weight ratio is 3:2:1 to form an agent learning decision rule. Deploy a central coordination control unit at the network center node, import the agent learning decision rule into the central coordination control unit to form a global coordination control topology. Based on this topology, establish an inter-agent communication mechanism, construct an agent cooperation control mechanism and a status synchronization protocol, and finally form a multi-agent adaptive control system.

[0030] Extract the time characteristics, spatial characteristics, and type characteristics of network workloads from a multi-agent adaptive control system to form a multi-dimensional workload feature set. Conduct time series analysis and pattern recognition on these features, calculate the statistical distribution parameters of the features using the sliding window method, and construct a workload pattern feature library. Based on this feature library, build a multi-level classification framework to hierarchically classify workloads according to resource demand types, business characteristics, and priorities. Input the classification structure into a hybrid classifier for accurate classification, and trigger the manual confirmation process when the classification confidence is lower than 0.8. Set up a change detector based on the classification results, monitor workload changes using the CUSUM algorithm, and generate a load change trigger signal when the cumulative change amount exceeds the threshold. Combine the load change trigger signal and the workload pattern feature library to establish a workload prediction model, predict the future workload change trend, and finally generate a workload control analysis report. Develop a medium- and long-term tactical layer resource plan based on the load prediction data in the workload control analysis report, set the resource adjustment threshold parameters, and form a tactical layer scheduling plan table. Dynamically adjust and calculate the size of the virtual resource pool and the boundary of the adjustment domain based on this table to generate a resource reservation instruction set and a resource domain adjustment plan. Calculate the operation layer resource allocation strategy based on this plan and real-time load data, including calculating the computing resource allocation matrix, network path selection matrix, and priority queue configuration table, to form a real-time scheduling instruction sequence. Decompose the instruction sequence into atomic operations and construct a dependency graph, convert it into device-level control commands through an optical switching controller and execute the resource reconfiguration operation, record the execution status, and finally generate a regulation execution control record. Finally, conduct multi-dimensional control parameter analysis on the processor energy consumption and network transmission energy consumption based on the regulation execution control record. Extract various energy consumption data from the regulation execution control record, group and standardize it according to the resource adjustment domain and time dimension, calculate the energy consumption baseline value, and construct a multi-variable energy consumption prediction model. Use this model to evaluate the energy efficiency of different workload types and resource scheduling strategies, calculate energy efficiency ratio indicators, energy consumption distribution balance degrees, etc., identify energy consumption hotspots and energy efficiency troughs, and pay special attention to the energy consumption during the optoelectronic conversion process. Input the optimization target into a multi-objective optimization framework, set the weight ratio of the performance target and the energy efficiency target to 6:4, optimize the scheduling strategy parameters, and finally generate an energy efficiency regulation report and control strategy parameters.

[0031] In an actual application scenario, when a large data center network faces a sudden increase in traffic, this method first captures the traffic change through the optical switching monitoring module, and then quickly adjusts the boundary of the resource adjustment domain to allocate more resources to the high-load area. The multi-agent system works collaboratively to predict the traffic trend based on historical load patterns and adjusts the resource configuration in advance. Through the double-layer scheduling of the tactical layer and the operation layer, it not only ensures the rationality of long-term resource planning but also ensures the timely response of short-term resource allocation. Finally, the energy consumption analysis shows that under the condition of processing the same load, the optimized resource scheduling strategy reduces resource conflicts and idling and effectively balances the energy consumption distribution of each area.

[0032] In the embodiments of the present application, by monitoring the node processor load and link utilization rate in a computer network through an optical switching monitoring module, high-precision and low-interference network status monitoring is achieved, solving the problems of insufficient accuracy and high monitoring overhead in traditional monitoring systems, providing an accurate data basis for resource scheduling, and at the same time, the optical switching monitoring technology ensures that the monitoring process does not affect normal network transmission; based on the network resource status control parameter set, a resource control pool is constructed using virtualization technology and resource adjustment domains are divided, breaking the limitations of traditional physical resource boundaries, achieving unified management and flexible configuration of resources, improving resource utilization rate, and the division of resource adjustment domains takes into account resource affinity and service relevance, creating conditions for the efficient scheduling of resources; according to the resource control status report and the adjustment domain division scheme, resource control agents and a central coordination control unit are deployed in each adjustment domain to construct a multi-agent adaptive control system. This system adopts a distributed decision-making mechanism, solving the problems of slow response and high single-point failure risk in a centralized scheduling system. The multi-agents achieve high-speed collaborative decision-making through an optical switching network, ensuring the fast response ability of the scheduling system to network state changes; using the multi-agent adaptive control system to extract features and perform multi-level classification and adjustment on network workloads, accurate identification and classification of workloads are achieved through artificial intelligence algorithms. The algorithm features play a key role in the process of workload classification. The hybrid classifier combines the interpretability of decision trees, the classification efficiency of support vector machines, and the feature extraction ability of deep neural networks, significantly improving classification accuracy; according to the workload control analysis report, a two-layer adjustment strategy at the tactical layer and the operation layer is executed through an optical switching controller, achieving the coordination of long-term resource planning and immediate resource allocation. The two-layer scheduling mechanism makes full use of the low-latency characteristics of the optical switching network to achieve millisecond-level reconfiguration of resources; based on the adjustment execution control record, multi-dimensional control parameter analysis of processor energy consumption and network transmission energy consumption is carried out. Through a multi-variable energy consumption prediction model and a multi-objective optimization framework, effective reduction of energy consumption is achieved while ensuring performance. Artificial intelligence algorithms play an important role in the process of energy consumption prediction and optimization. By comprehensively analyzing various variables such as device power consumption characteristics, resource utilization rate, and environmental temperature, a high-precision energy consumption prediction model is constructed. The genetic algorithm effectively balances the performance and energy efficiency objectives in the multi-objective optimization process, achieving the global optimization of the scheduling strategy. The entire technical solution constructs an efficient, intelligent, and energy-saving computer network resource scheduling method through the organic combination of optical switching technology, virtualization technology, and artificial intelligence algorithms, solving the problems of insufficient monitoring accuracy, inflexible resource pool division, slow scheduling decision response, inaccurate load classification, insufficient multi-level scheduling coordination, and insufficient energy consumption optimization in the prior art, and significantly improving the network resource utilization rate and the overall performance of the system.

[0033] In a specific embodiment, the process of executing step S101 may specifically include the following steps:

[0034] Deploy distributed optical switching monitoring modules at key network nodes, configure the connections of key network nodes, and obtain the distribution topology of monitoring nodes;

[0035] Dynamically adjust and set the acquisition frequency of monitoring nodes, automatically calibrate the sampling rate according to network traffic volatility, and obtain an adaptive sampling configuration;

[0036] Collect link utilization rate, packet transmission delay, queue length, packet loss rate, and optical signal intensity through the optical switching monitoring module to obtain original monitoring data;

[0037] Add timestamp, optical path identifier, and topological location information to the original monitoring data to obtain tokenized original data;

[0038] Transmit the tokenized original data to the central data processing unit through a dedicated optical transmission channel to obtain a centralized data stream;

[0039] Filter, denoise, and standardize the centralized data stream to obtain standardized network resource status data;

[0040] Perform optical signal quality anomaly detection based on the standardized network resource status data, trigger an alarm when a decrease in optical signal quality or optical path congestion is detected, and obtain a set of network resource status control parameters.

[0041] Specifically, the optical switching monitoring module monitors the node processor load and link utilization in the computer network, and obtains the network resource status control parameter set. The specific implementation process includes seven sub-steps, and each sub-step is a key link in data collection and processing. First, a distributed optical switching monitoring module is deployed at key network nodes, and the key network nodes are connected and configured to obtain the monitoring node distribution topology. The optical switching monitoring module is a device that can directly connect to the fiber optic network and monitor optical signals, featuring high-speed and lossless detection. During the deployment process, according to the complexity and scale of the network topology, monitoring modules are installed at key positions such as core switches, routers, server clusters, and access layer devices. The monitoring module is directly connected to the network backbone link through optical fibers and adopts a bypass listening method without interfering with normal network transmission. Each monitoring module is assigned a unique identification ID, and its physical location and the network segment information it is responsible for monitoring are recorded to form a monitoring node distribution topology map, including node locations, coverage ranges, and connection relationships between nodes. Next, the sampling frequency of the monitoring nodes is dynamically adjusted and set, and the sampling rate is automatically calibrated according to the volatility of network traffic to obtain an adaptive sampling configuration. This step analyzes historical traffic data, calculates the change rate and fluctuation amplitude of the traffic, and determines the basic sampling frequency. When the network traffic is stable, the sampling frequency is set to a lower value, such as once every 10 seconds; when it is detected that the traffic volatility exceeds the preset threshold, the sampling frequency is increased step by step, up to once every millisecond. The adaptive sampling algorithm calculates the ratio of the standard deviation to the mean (coefficient of variation) of the traffic within the current time window. When this value exceeds 0.5, the sampling frequency is increased, and when it is lower than 0.2, the sampling frequency is decreased. This dynamic adjustment mechanism effectively controls the data volume and processing pressure while ensuring data accuracy.

[0042] Then, the optical switching monitoring module is used to collect indicators such as link utilization, packet transmission delay, queue length, packet loss rate, and optical signal intensity to obtain the original monitoring data. The link utilization is calculated by measuring the ratio of the data transmission volume per unit time to the maximum capacity of the link; the packet transmission delay is measured by sending test packets and recording the round-trip time; the queue length records the number of packets to be processed in the buffer of the network device; the packet loss rate calculates the difference ratio between the number of sent and received packets; the optical signal intensity directly measures the power level of the optical signal in the optical fiber. These original data are stored in binary format, and each data point contains the measurement value, measurement time, and source device ID. Timestamp, optical path identifier, and topology location information are added to the original monitoring data to obtain the tokenized original data. The timestamp is accurate to the millisecond level and uses the UTC format to ensure the timing of the data; the optical path identifier includes the optical fiber path ID and wavelength channel number, which are used to track the physical path through which the data flows; the topology location information includes the level, area, and functional role of the monitoring point in the network topology. The tokenization process converts the original binary data into a structured data format for subsequent analysis and processing.

[0043] The tokenized original data is transmitted to the central data processing unit through a dedicated optical transmission channel to obtain a centralized data stream. The dedicated optical transmission channel is a management channel independent of the service network, using an independent wavelength to ensure that the transmission of monitoring data is not affected by network load. The data transmission adopts an encryption method to protect the security of monitoring data. The central data processing unit receives the data streams from each monitoring point and performs preliminary integration according to the timestamp and node ID to form a centralized data stream arranged in time series.

[0044] The centralized data stream is filtered, denoised, and normalized to obtain normalized network resource status data. The filtering process uses a median filtering algorithm to remove burst outliers; the denoising process uses a wavelet transform method to separate signals and noise and retain valid information; the normalization process converts data with different metrics into a unified numerical range, usually using the maximum-minimum normalization method to map the data to the [0, 1] interval. The processed data is smoother and more consistent, facilitating comparative analysis. Finally, optical signal quality anomaly detection is performed based on the normalized network resource status data. When a decrease in optical signal quality or optical path congestion is detected, an early warning is triggered to obtain a set of network resource status control parameters. The anomaly detection uses a method based on statistics and thresholds to calculate the historical average and standard deviation of metrics such as optical signal intensity, bit error rate, and link utilization, and sets multi-level early warning thresholds. When the metric value deviates from the mean by more than 2 standard deviations, a mild early warning is issued, and when it exceeds 3 standard deviations, a severe early warning is issued. At the same time, the change trend of the metric is detected. When the metric continues to deteriorate for a certain time window, an early warning is triggered even if the severe early warning threshold is not reached. The early warning information, together with the relevant resource status data, is incorporated into the set of network resource status control parameters as an important basis for subsequent resource scheduling.

[0045] Taking a data center network as an example, by deploying optical switching monitoring modules on four core switches, a monitoring node distribution topology covering the entire backbone network was formed. During normal operation, the monitoring module collects data at a frequency of once every 5 seconds. When a large-scale data migration task was detected in the North Area server group one morning, the traffic variation coefficient rapidly rose from 0.15 to 0.68, and the monitoring system automatically increased the sampling frequency of the North Area monitoring points to twice per second. The data collected showed that the utilization rate of the North Area backbone link jumped from an average of 25% to 78%, the packet delay increased from 1.2 milliseconds to 3.8 milliseconds, and the queue length increased from an average of 12 packets to 47 packets. After adding the time stamp in the UTC+8 time zone, the optical path identifier of "NC-TRUNK-03", and the topology location identifier of "Core-North-Primary" to these raw data, they were transmitted to the central processing unit through the managed wavelength channel. After the data was processed, it was found that the optical signal strength in the North Area dropped from the standard value of -3 dBm to -4.8 dBm, exceeding the range of two standard deviations from the historical average. The system generated a warning for the degradation of the optical signal quality and incorporated this information together with the processed network load data into the resource status control parameter set. These accurate monitoring data provided a clear basis for subsequent resource scheduling, ensuring that network resources could be reasonably allocated according to actual needs.

[0046] In a specific embodiment, the process of executing step S102 may specifically include the following steps:

[0047] Abstract the physical network resources, use resource description language to standardize the description of computing resources, storage resources, and network bandwidth resources, and obtain a resource metadata model;

[0048] Based on the resource metadata model, establish a mapping relationship from physical resources to virtual resources, and assign identifiers to each virtual resource unit through an optical switching controller to obtain a resource mapping data structure;

[0049] Calculate the resource correlation matrix according to the topology information and the resource mapping data structure in the network resource status control parameter set to obtain the resource affinity index;

[0050] Based on the resource affinity index and the resource metadata model, perform hierarchical clustering analysis on the resource control pool, and classify resources with a correlation higher than the preset threshold of 0.75 into the same resource adjustment domain to obtain the initial adjustment domain division result;

[0051] Perform boundary optimization analysis on the initial adjustment domain division result and the resource mapping data structure, calculate the internal connectivity and external isolation degree of each adjustment domain, and adjust the attribution of boundary nodes to obtain the optimized resource adjustment domain division;

[0052] Create computing-intensive, bandwidth-intensive, and storage-intensive resource configuration templates for the optimized resource adjustment domain division and resource affinity metrics, define the proportional parameters for various types of resources, and obtain a resource configuration template library;

[0053] Based on the resource configuration template library and the optimized resource adjustment domain division, conduct resource status integration analysis to generate a resource control status report and an adjustment domain division plan. The resource control status report includes the total virtual resource volume, current allocation status, and resource utilization statistical data for each resource adjustment domain. The adjustment domain division plan includes the physical boundaries of each resource adjustment domain, the inter-domain connection relationships, and the resource scheduling permission settings.

[0054] Specifically, use virtualization technology to construct a resource control pool and conduct resource adjustment domain division based on the network resource status control parameter set. Abstract the physical network resources, and use a resource description language to standardize the description of computing resources, storage resources, and network bandwidth resources to obtain a resource metadata model. The resource description language is a structured language specifically used to describe the characteristics of computer network resources, which converts the key attributes of physical resources into a standardized data structure. For computing resources, describe its CPU core count, main frequency, cache size, instruction set type, etc.; for storage resources, describe its capacity, read / write speed, access latency, persistence characteristics, etc.; for network bandwidth resources, describe its maximum transmission rate, current available bandwidth, latency characteristics, connection topology, etc. These descriptions are organized in XML or JSON format to form a unified resource metadata model, which contains the identification information, performance parameters, status information, and location information of the resources. Based on the resource metadata model, establish the mapping relationship from physical resources to virtual resources, and assign identifiers to each virtual resource unit through an optical switch controller to obtain a resource mapping data structure. This step abstracts physical resources into virtual resource units, and each virtual resource unit represents a certain amount of physical resources, such as CPU time slices, memory blocks, or bandwidth segments. The optical switch controller is responsible for assigning a globally unique identifier to each virtual resource unit, using a multi-level identification structure, including a resource type code, a physical location number, and a serial number. The resource mapping data structure is stored in a graph form, with nodes representing virtual resource units and edges representing the dependencies or connection relationships between resources, and at the same time records the corresponding relationship between each virtual resource unit and physical resources, including the mapping ratio, resource status, and occupancy situation.

[0055] Calculate the resource correlation matrix based on the topology information and resource mapping data structure in the network resource status control parameter set, and obtain the resource affinity index. The resource correlation matrix is an n×n matrix, where n is the number of virtual resource units, and each element in the matrix represents the degree of correlation between two resources. The dimensions for calculating correlation include physical location distance, communication frequency, degree of shared dependent resources, and similarity of access patterns. The physical location distance is calculated based on the number of hops or physical distance between nodes in the topology information; the communication frequency analyzes the frequency of data exchange between resources according to historical traffic data; the degree of shared dependent resources calculates the number of third-party resources jointly dependent on two resources; the similarity of access patterns analyzes the similarity of the time distribution characteristics of resource access. By weighted combination of the scores of these four dimensions, the final resource affinity index is obtained, and the value range of this index is from 0 to 1. The larger the value, the higher the affinity between resources.

[0056] Based on the resource affinity index and the resource metadata model, perform hierarchical clustering analysis on the resource control pool, and classify resources with a correlation higher than the preset threshold of 0.75 into the same resource adjustment domain to obtain the initial adjustment domain division result. The hierarchical clustering analysis adopts a bottom-up aggregation strategy. Initially, each virtual resource unit is regarded as an independent adjustment domain, and then gradually merge the adjustment domain pairs with the highest affinity until the affinity between all adjustment domain pairs is lower than the preset threshold of 0.75. During the merging process, consider the resource type and performance parameters in the resource metadata model to ensure that the heterogeneity of resources within each adjustment domain does not exceed the set value, and avoid assigning resources with too large performance differences to the same adjustment domain. The initial adjustment domain division result is represented in a tree structure, where each leaf node is a virtual resource unit, and the non-leaf node is an adjustment domain, recording the resource list within the domain, the affinity between domains, and the boundary information of the domain.

[0057] Perform boundary optimization analysis on the initial adjustment domain division result and the resource mapping data structure, calculate the internal connectivity and external isolation degree of each adjustment domain, and adjust the attribution of boundary nodes to obtain the optimized resource adjustment domain division. The internal connectivity measures the total affinity between resources within an adjustment domain, and the external isolation degree measures the negative sum of the affinity between an adjustment domain and other adjustment domains. For each node on the boundary of an adjustment domain, calculate its contribution to the internal connectivity of the current domain it belongs to and the change in internal connectivity after moving to an adjacent domain. If the move can simultaneously improve the internal connectivity and external isolation degree of both domains, then adjust the attribution of this node. This process is repeated until no node move can improve the optimization goal or reach the maximum number of iterations. The optimized resource adjustment domain division has higher cohesion and lower coupling.

[0058] For the optimized resource adjustment domain division and resource affinity metrics, create compute-intensive, bandwidth-intensive, and storage-intensive resource configuration templates, define the proportional parameters for various types of resources, and obtain a resource configuration template library. A configuration template is a pre-defined resource allocation scheme for different types of workloads, used to quickly respond to resource requests. The compute-intensive template allocates a higher proportion of CPU resources and an appropriate amount of memory, suitable for a large number of computing tasks; the bandwidth-intensive template allocates a higher network bandwidth resource and buffer, suitable for data transmission tasks; the storage-intensive template allocates a larger storage space and I / O bandwidth, suitable for data storage and retrieval tasks. Each template defines the proportional parameters of resources, such as the ratio of compute:storage:bandwidth, and also considers the resource affinity metrics to ensure that resources with high affinity tend to be allocated together. The resource configuration template library contains multiple preset templates and custom templates, and each template records the resource type, proportional parameters, applicable workload types, and performance metrics. Based on the resource configuration template library and the optimized resource adjustment domain division, conduct a resource status integration analysis to generate a resource control status report and an adjustment domain division plan. The resource control status report is a detailed resource inventory and status record, including the total virtual resource volume, current allocation status, and resource utilization statistics data for each resource adjustment domain. The total virtual resource volume counts the total quantity and performance parameters of various resources within the domain; the current allocation status records the occupancy of resources, including the allocated and free resource ratios; the resource utilization statistics data analyzes the usage efficiency of resources, including average utilization, peak utilization, and utilization fluctuation. The adjustment domain division plan is an execution blueprint for guiding resource scheduling, including the physical boundaries of each resource adjustment domain, the inter-domain connection relationships, and the resource scheduling permission settings. The physical boundaries define the scope of the adjustment domain in the network topology; the inter-domain connection relationships describe the communication paths and bandwidths between different adjustment domains; the resource scheduling permission settings specify the allocation and recovery permissions of each level of scheduling entity for resources.

[0059] Taking a multi-region data center network as an example, first, the physical network resources are abstracted. Forty physical servers, twelve storage arrays, and eight core switches are described as a resource metadata model. Attributes such as the computing power, memory size, and network interface bandwidth of each server are standardized and recorded; parameters such as the capacity, read / write speed, and redundancy level of the storage array are structurally described; metrics such as the port number, switching capacity, and forwarding delay of the switch are quantitatively represented. Based on these metadata, the physical resources are mapped into 320 computing resource units, 96 storage resource units, and 64 network bandwidth resource units through an optical switching controller. Each resource unit is assigned a unique identifier. For example, "COMP-R3-S15-C01" represents the computing resource of the first server in the 15th rack of the third computer room. According to network traffic analysis and resource access patterns, a 480×480 resource correlation matrix is calculated, and it is found that the affinity between some resource groups is particularly high. For example, the affinity between the Web server and the application server is 0.82, and the affinity between the application server and the database server is 0.78. Through hierarchical clustering analysis, the resources are initially divided into eight adjustment domains, but it is found that the attribution of some boundary resources is not reasonable enough. After boundary optimization, the edge network resources are adjusted from the computing domain to the network domain, improving the overall connectivity. According to business requirements, various resource configuration templates are created, such as a compute-intensive template suitable for real-time transaction processing (compute:storage:bandwidth ratio of 5:1:3) and a storage-intensive template suitable for big data analysis (compute:storage:bandwidth ratio of 2:5:3). The finally generated resource control status report details the resource status of each adjustment domain. For example, the first adjustment domain has 80 computing resource units, 56 of which are currently allocated, and the average utilization rate is 65%; the adjustment domain division scheme clearly stipulates the physical scope, interconnection bandwidth, and scheduling authority of each domain.

[0060] In a specific embodiment, the process of executing step S103 may specifically include the following steps:

[0061] Assign a unique identifier to each resource adjustment domain based on the adjustment domain division scheme and determine its physical boundary to obtain an adjustment domain control mapping table;

[0062] Perform correlation analysis on the resource data in the resource control status report and the adjustment domain control mapping table to determine the resource control complexity level of each adjustment domain and obtain a hierarchical control agent model;

[0063] Deploy resource control agents with a three-layer structure of a sensing layer, a decision-making layer, and an execution layer in each resource adjustment domain according to the hierarchical control agent model and the adjustment domain control mapping table to obtain a hierarchical agent control network;

[0064] Load the reinforcement learning algorithm into the decision-making layer of the hierarchical agent control network, and use the hierarchical control of the parameters in the agent model as the initial value to obtain the agent learning decision rule;

[0065] Deploy a central coordination and control unit at the network central node, and import the agent learning decision rule and the adjustment domain control mapping table into the central coordination and control unit to obtain the global coordination and control topology;

[0066] Based on the global coordination and control topology and the hierarchical agent control network, establish an inter-agent communication mechanism, calculate the trust matrix between agents, and obtain the agent cooperation control mechanism;

[0067] Apply the agent cooperation control mechanism to the hierarchical agent control network, construct a state synchronization protocol based on the adjustment domain control mapping table, and use the optical switching network to transmit the agent state data to obtain a multi-agent adaptive control system.

[0068] Specifically, based on the adjustment domain division scheme, assign a unique identifier to each resource adjustment domain and determine its physical boundary to obtain the adjustment domain control mapping table. In this process, each adjustment domain is assigned a structured identifier, usually using a hierarchical naming method, such as the format of "RD-region number-type code-sequence number", where the region number represents the physical region, the type code represents the main resource type, and the sequence number is the sequential number. At the same time, according to the physical boundary information in the adjustment domain division scheme, clearly define the physical device range, network segment, and geographical location covered by each adjustment domain. These information are organized into the adjustment domain control mapping table, which is a relational data structure, including fields such as domain identifier, physical boundary description, list of physical devices contained, inter-domain connection, and resource scheduling permission. Conduct an association analysis on the resource data in the resource control status report and the adjustment domain control mapping table to determine the resource control complexity level of each adjustment domain, and obtain the hierarchical control agent model. The resource control complexity is a comprehensive index to measure the difficulty of adjustment domain resource management, which is calculated by analyzing multiple factors, including the number of resources within the domain, resource type diversity, resource status change frequency, load fluctuation degree, and external dependence. Specifically in the calculation, first extract the total amount of resources, type distribution, and historical utilization rate data of each adjustment domain from the resource control status report, and combine the physical boundary and permission information in the adjustment domain control mapping table to calculate the resource control complexity score. Then divide the adjustment domain into three complexity levels of high, medium, and low according to the complexity score, and customize control agent models with different capabilities for different levels. The high-complexity domain is equipped with full-functional agents with complex decision-making capabilities; the medium-complexity domain is equipped with standard agents with basic autonomous decision-making capabilities; the low-complexity domain is equipped with lightweight agents that mainly execute superior instructions. These information form the hierarchical control agent model, which describes the type, function range, and decision-making authority of the agents required for each adjustment domain.

[0069] Resource control agents with a three - layer structure of a perception layer, a decision - making layer, and an execution layer are deployed in each resource regulation domain according to the hierarchical control agent model and the regulation domain control mapping table, obtaining a hierarchical agent control network. The perception layer is responsible for collecting resource status data, including resource utilization rate, performance metrics, and abnormal events. By deploying monitoring agents at key nodes within the regulation domain, data is collected regularly and preliminarily processed. The decision - making layer is the core of the agent, responsible for analyzing the data from the perception layer and formulating resource scheduling decisions according to the strategy. Its complexity matches the control complexity level of the regulation domain. The execution layer is responsible for converting the decisions into specific resource scheduling operations, including resource allocation, recycling, migration, and load balancing, etc. These three layers within each agent are connected through standardized interfaces, forming a complete information flow and control flow. All agents within the regulation domain constitute a hierarchical agent control network, and the agents in the network form a tree - shaped or mesh - shaped structure according to the hierarchical relationship and physical location of the regulation domain. The reinforcement learning algorithm is loaded into the decision - making layer of the hierarchical agent control network, using the parameters in the hierarchical control agent model as the initial values to obtain the learning decision rules for the agents. Reinforcement learning is a machine - learning method that continuously optimizes decisions through a trial - and - error and reward mechanism, especially suitable for dynamic decision - making problems such as resource scheduling. In practical applications, algorithms such as Q - learning or DeepQ Network (DQN) are used. The scheduling state, action, and reward function are defined as follows: the state includes the current resource utilization rate, task queue length, etc.; the action includes operations such as resource allocation, task scheduling, and load migration; the reward function comprehensively considers the resource utilization rate, task completion time, and system response delay, and calculates the total reward value according to the weight ratio of 3:2:1. When the algorithm is initialized, the parameters in the hierarchical control agent model are used as the starting point. For example, agents in high - complexity domains are configured with larger neural networks and more historical memories, and simplified algorithms are configured for low - complexity domains. Through continuous learning and updating, each agent gradually forms decision rules suitable for the characteristics of its management domain.

[0070] A central coordination control unit is deployed at the network center node. The learning decision rules of the agents and the regulation domain control mapping table are imported into the central coordination control unit to obtain a global coordination control topology. The central coordination control unit is the coordination center of the entire multi - agent system, responsible for global policy formulation and cross - domain resource coordination. It is deployed at the central position of the physical network, usually a core node with high - performance computing capabilities and extensive connectivity. The coordination unit imports the learning decision rules of the agents in each regulation domain to form a rule library, and at the same time imports the regulation domain control mapping table to understand the resource distribution and management structure of the entire network. Based on this information, the coordination unit constructs a global coordination control topology, which describes the connection relationship, communication method, and coordination authority between the coordination unit and the agents in each regulation domain, forming a star - shaped or hierarchical control structure.

[0071] Based on the global coordination control topology and the hierarchical agent control network, establish an inter-agent communication mechanism, calculate the trust matrix between agents, and obtain the agent cooperation control mechanism. The inter-agent communication mechanism defines the protocols and methods for agents to exchange information, including regular status synchronization, event-triggered notifications, and proactive query responses, etc. The communication content includes resource status updates, decision intention announcements, and cooperation requests, etc. On this basis, calculate the trust matrix between agents. Each element of this matrix represents the degree of trust of one agent in another agent. The initial value is preset based on the relationship between adjustment domains, and then dynamically adjusted according to historical interaction records. The calculation method considers factors such as historical cooperation success rate, information accuracy, and response timeliness. When the historical cooperation success rate is lower than the preset threshold, the trust weight of the corresponding agent is automatically reduced. The trust matrix is the core of the agent cooperation control mechanism, which determines the adoption weights of information and suggestions from all parties during the cooperation process. Apply the agent cooperation control mechanism to the hierarchical agent control network, construct a status synchronization protocol based on the adjustment domain control mapping table, use the optical switching network to transmit agent status data, and obtain a multi-agent adaptive control system. The status synchronization protocol stipulates the timing, content, and methods for agents to synchronize status information. According to the inter-domain relationship in the adjustment domain control mapping table, determine the synchronization priority and frequency. When the amplitude of the status change exceeds the preset threshold, trigger immediate synchronization; otherwise, perform regular synchronization according to the set period. The synchronization content includes resource status changes, policy adjustments, and decision conflicts, etc. Use the optical switching network for data transmission to ensure high-speed and low-latency communication effects, especially for status changes that require immediate response. The optical switching network uses wavelength division multiplexing technology to allocate dedicated optical wavelength channels for inter-agent communication, avoiding conflicts with traffic flows. The finally formed multi-agent adaptive control system is a distributed resource scheduling decision network that can automatically adjust the scheduling strategy according to network load changes.

[0072] Taking an enterprise-level data center network as an example, according to the adjustment domain division scheme, the network is divided into 6 resource adjustment domains, namely the computing core area, the storage core area, the network front-end area, the application service area, the database area, and the edge access area. Each area is assigned a unique identifier. For example, "RD-C-01" represents the first computing core area, and "RD-S-01" represents the first storage core area. At the same time, the physical boundaries of each area are clearly defined. For example, the computing core area includes all computing servers on racks 1-3. According to the data in the resource control status report, it is analyzed that the computing core area contains 120 computing units, with large average load fluctuations, diverse resource types, and its complexity score is calculated as 85 points, which is classified as a high complexity level; while the edge access area only contains 32 network units, with stable load, single resources, and a complexity score of 35 points, which is classified as a low complexity level. For the high-complexity computing core area, a full-functional agent is deployed, with a perception layer configured with an eight-core processor, a decision-making layer of a 12-layer neural network, and an execution layer of a full instruction set; for the low-complexity edge access area, a lightweight agent is deployed, with only a perception layer configured for basic monitoring, a decision-making layer of a simplified Q-learning, and an execution layer of a limited instruction set. A central coordination and control unit is deployed on the management server in the central computer room, and the decision-making rules and adjustment domain mapping tables of each agent are imported to form a star-shaped coordination topology. Through the initial cooperation test, the trust matrix between agents is calculated, and it is found that the historical cooperation success rate between the agent in the computing core area and the agent in the application service area is 92%, and the trust value is set to 0.9; while the historical cooperation between the computing core area and the edge access area only has a success rate of 65%, and the trust value is set to 0.6. Based on these trust values, when cross-domain resource scheduling occurs, the agent in the computing core area is more inclined to adopt the suggestions of the agent in the application service area. Finally, a dedicated agent communication channel is established through the optical switching network, and synchronization is immediately triggered when the resource state changes by more than 10% to ensure that all agents can make coordinated decisions based on consistent information, forming a multi-agent adaptive control system.

[0073] In a specific embodiment, the process of executing step S104 may specifically include the following steps:

[0074] Extract the time characteristics, spatial characteristics, and type characteristics of the network workload from the multi-agent adaptive control system to obtain a multi-dimensional workload feature set;

[0075] Perform time series analysis and pattern recognition on the multi-dimensional workload feature set, and use the sliding window method to calculate the statistical distribution parameters of the features to obtain a workload pattern feature library;

[0076] Construct a multi-level classification framework based on the workload pattern feature library, hierarchically classify the workload according to resource demand type, business characteristics, and priority to obtain a workload classification structure;

[0077] Input the workload classification structure into the hybrid classifier, and combine decision trees, support vector machines, and deep neural networks to perform accurate workload classification. When the classification confidence is lower than 0.8, trigger the manual confirmation process to obtain the classified result dataset;

[0078] Based on the classified result dataset, set up a change detector, and use the CUSUM algorithm to monitor changes in the mean value, variance, and distribution characteristics of the workload. When the cumulative change amount exceeds the set threshold, obtain the workload change trigger signal;

[0079] According to the workload change trigger signal and the workload pattern feature library, establish a workload prediction model, and use time series analysis and machine learning methods to predict the workload change trend within the next 60 minutes to obtain the workload prediction data;

[0080] Integrate and analyze the classified result dataset, the workload change trigger signal, and the workload prediction data, calculate the resource requirements and scheduling priorities of various workloads, and obtain the workload control analysis report.

[0081] Specifically, extract the time features, space features, and type features of the network workload from the multi-agent adaptive control system to obtain a multi-dimensional workload feature set. The time features are the features that describe the change law of the workload over time, including the arrival rate, duration, and periodicity. Among them, the arrival rate represents the number of new workloads per unit time, the duration represents the duration of the workload from start to end, and the periodicity represents the periodic pattern of the workload appearance; the space features are the features that describe the distribution of the workload in the network space, including the resource demand distribution, access pattern, and spatial aggregation degree. Among them, the resource demand distribution represents the demand ratio of the workload for different types of resources, the access pattern represents the path and method of the workload accessing network resources, and the spatial aggregation degree represents the aggregation degree of the workload in the network topology; the type features are the features that describe the attributes of the workload itself, including the computing intensity, data transmission intensity, and storage intensity, reflecting the workload's emphasis on different resources. During the extraction process, each agent collects the original workload data from its responsible area, analyzes the resource usage pattern, calculates the feature values of each dimension, and then summarizes them to form a multi-dimensional workload feature set containing three major types of features: time, space, and type.

[0082] Perform time series analysis and pattern recognition on the multi-dimensional workload feature set, calculate the statistical distribution parameters of the features using the sliding window method, and obtain the workload pattern feature library. Time series analysis is a method for analyzing the changing patterns of data time series. By analyzing the changing trends of workload features over time, periodic, trend, and random components are identified. Pattern recognition is to discover recurring patterns and regularities from data. By comparing the workload features in different time periods, similar workload patterns are found. The sliding window method is a commonly used time series data processing technique. By setting a time window of a fixed size and sliding this window on the time axis, the statistical features of the data within the window are calculated. In this method, the set window sizes vary from 1 minute to 24 hours, and appropriate window sizes are selected for different features. For the data within each window, statistical distribution parameters such as mean, variance, kurtosis, and skewness are calculated to describe the central tendency, dispersion degree, and distribution shape of the features. These statistical parameters, together with the original features, constitute the workload pattern feature library, which contains the feature patterns of various workloads at different time scales and provides a basis for subsequent classification. Based on the workload pattern feature library, a multi-level classification framework is constructed to hierarchically classify the workloads according to resource demand types, business characteristics, and priorities, obtaining the workload classification structure. The multi-level classification framework is a hierarchical classification method from coarse to fine. In this method, a three-level classification structure is adopted: the first level is classified based on resource demand types, and the workloads are divided into compute-intensive, data transfer-intensive, storage-intensive, and hybrid types; the second level is classified based on business characteristics, and the workloads are further divided into real-time interactive, batch processing, stream processing, and transaction processing types, etc.; the third level is classified based on priorities, considering business importance, time urgency, and service quality requirements, and the workloads are divided into critical level, high priority, normal priority, and low priority. The classification process adopts a top-down decision-making method, first determining the first-level category according to the resource demand ratio, then determining the second-level category according to the business characteristics, and finally determining the third-level category according to the priority index. Finally, a tree-like workload classification structure is formed, clearly describing the hierarchical relationships of various workloads.

[0083] Input the workload classification structure into the hybrid classifier, which combines decision trees, support vector machines, and deep neural networks for accurate workload classification. When the classification confidence is below 0.8, trigger the manual confirmation process to obtain the classification result dataset. The hybrid classifier is a composite classifier that combines the advantages of multiple classification algorithms. In this method, three different types of classification algorithms are integrated: decision trees are suitable for handling cases with clear classification rules, support vector machines are good at handling high-dimensional data with clear boundary distinctions, and deep neural networks are suitable for handling complex non-linear relationships. During the classification process, first input the feature vectors of the workload into the three classifiers for independent classification respectively; then integrate the three classification results through weighted voting, and the voting weights are dynamically adjusted according to the performance of each classifier on historical data; finally, calculate the confidence of the comprehensive classification result. When the confidence is higher than 0.8, directly adopt the classification result, and when the confidence is lower than 0.8, trigger the manual confirmation process, and let professionals judge the accuracy of the classification result and make necessary adjustments. The accurately classified workloads, together with their class labels, feature values, and classification confidences, constitute the classification result dataset.

[0084] Set a change detector based on the classification result dataset, and use the CUSUM algorithm to monitor changes in the mean, variance, and distribution characteristics of the workload. When the cumulative change amount exceeds the set threshold, obtain the load change trigger signal. The change detector is a component used to monitor data changes in real time and respond quickly. In this method, it mainly monitors three key indicators of the workload: the mean reflects the overall intensity of the workload, the variance reflects the degree of fluctuation of the workload, and the distribution characteristics reflect the structural changes of the workload. The CUSUM (Cumulative Sum) algorithm is a classic change point detection algorithm, and its core idea is to calculate the cumulative deviation of the data from the reference value. When the cumulative deviation exceeds the threshold, it is determined that a change has occurred. In the implementation of this method, CUSUM detectors are set for the mean, variance, and distribution characteristics respectively. Each detector records the difference between the current value and the reference value, and accumulatively adds these differences. When the cumulative sum exceeds the preset threshold, a change detection signal is triggered. The advantage of the CUSUM algorithm is that it can detect slow but continuous changes and is suitable for monitoring the gradual change trend of network loads. When a significant change is detected, the change detector generates a load change trigger signal, which includes the type, amplitude, duration, and scope of the change.

[0085] A workload prediction model is established based on the load change trigger signal and the workload pattern feature library. Time series analysis and machine learning methods are used to predict the workload change trend in the next 60 minutes, and load prediction data is obtained. The workload prediction model is a mathematical model for predicting future load changes based on historical data. In this method, two types of methods, time series analysis and machine learning, are combined: time series analysis focuses on mining the periodicity and trend of the load and is suitable for processing loads with obvious time patterns; machine learning methods focus on learning complex non-linear relationships and are suitable for processing loads affected by multiple factors. The prediction process is divided into three stages: First, according to the type of the load change trigger signal, relevant historical patterns are selected from the workload pattern feature library as references; then, the selected historical data is preprocessed to remove noise and extract trend and periodic components; finally, the processed data is input into the prediction model to generate the load prediction values in the next 60 minutes, including the total load volume, the demand for various resources, and the distribution change trend. The prediction results include prediction values, prediction intervals, and confidence levels, forming the load prediction data.

[0086] The classification result data set, the load change trigger signal, and the load prediction data are integrated and analyzed to calculate the resource demand and scheduling priority of various workloads, and a workload control analysis report is obtained. Integration analysis is a process of comprehensively processing multi-source data to obtain a more comprehensive understanding. In this method, three data sources are mainly processed: the classification result data set provides the current classification of the workload; the load change trigger signal provides important information about load changes; the load prediction data provides a trend prediction of future loads. The core of the integration analysis is to calculate the resource demand and scheduling priority of various workloads: the resource demand is calculated according to the load characteristics and historical resource consumption rules, including CPU demand, memory demand, storage demand, and bandwidth demand; the scheduling priority is calculated according to the business importance, urgency, and resource competition situation of the load, which determines the order and strategy of resource allocation. The analysis results form a workload control analysis report, which includes the current load classification result, the load change detection result, the future load prediction, and the resource requirements and scheduling suggestions for various loads, providing a decision-making basis for the next resource scheduling.

[0087] Taking the network environment of an e-commerce platform as an example, workload characteristics are extracted from a multi-agent adaptive control system. It is found that the network traffic of this platform has obvious temporal characteristics (two peaks appear at 10-12 o'clock and 19-21 o'clock every day, and the traffic on weekends is higher than that on weekdays), spatial characteristics (user access is mainly concentrated on the front end of the website and the payment system), and type characteristics (browsing behavior is mainly data transmission, and transaction behavior is mainly computing and storage operations). Through time series analysis of these characteristics, the traffic mean and standard deviation of each period are calculated using a 12-hour sliding window, and typical load patterns such as "daytime browsing mode", "evening shopping mode", and "weekend promotion mode" are identified, thus constructing a workload pattern feature library. Based on this feature library, a three-level classification framework is constructed, dividing the website traffic into first-level categories such as "browsing category", "transaction category", and "background processing category", second-level categories such as "ordinary browsing", "product search", and "order processing", and third-level sub-categories such as "VIP user transactions" and "ordinary user transactions". The real-time traffic is input into a hybrid classifier, which combines a decision tree (to handle clear user behavior patterns), an SVM (to handle the boundary division of user groups), and a deep neural network (to handle complex transaction behavior patterns) to accurately classify each access request. When a promotion activity is detected to start on a certain day, the CUSUM algorithm monitors that the mean value of the transaction category load continues to rise within 30 minutes, and the cumulative change amount exceeds the set threshold, triggering a load change signal. The system immediately extracts historical patterns similar to the promotion activity from the feature library and predicts that the transaction load will continue to increase by 30% within the next 60 minutes, and the load of the payment system will start to climb after 15 minutes. By comprehensively analyzing this information, it is calculated that the VIP user transaction processing should be preferentially allocated 150 units of computing resources, 90 units of storage resources, and 120 units of bandwidth. The demand and priority of ordinary order processing are adjusted accordingly, generating a detailed workload control analysis report, which provides an accurate decision-making basis for subsequent dynamic resource scheduling.

[0088] In a specific embodiment, the process of executing step S105 may specifically include the following steps:

[0089] Formulate a medium- and long-term tactical layer resource plan according to the load prediction data in the workload control analysis report, set resource adjustment threshold parameters, and obtain a tactical layer scheduling plan table;

[0090] Based on the tactical layer scheduling plan table, perform dynamic adjustment calculations on the size of the virtual resource pool and the boundary of the adjustment domain, generate a resource reservation instruction set, and obtain a resource domain adjustment plan;

[0091] According to the resource domain adjustment plan and the real-time load data in the workload control analysis report, calculate the operation layer resource allocation strategy, including the computing resource allocation matrix, network path selection matrix, and priority queue configuration table, and obtain a real-time scheduling instruction sequence;

[0092] Decompose the real-time scheduling instruction sequence into atomic operation instructions, construct an instruction dependency graph and calculate the execution order, automatically insert synchronization points when dependency conflicts occur, and obtain the optimized instruction execution flow;

[0093] Convert the optimized instruction execution flow into device-level control commands through an optical switching controller, start a parallel execution mechanism for operations with an execution time exceeding 50 milliseconds, and obtain a device control command sequence;

[0094] Transmit the device control command sequence to the resource control agent in the corresponding resource adjustment domain, perform resource reconfiguration operations, and record the execution status of each operation. When the number of execution failures exceeds three times, trigger a fallback mechanism to obtain execution status tracking data;

[0095] Analyze and integrate the execution status tracking data, record the execution results, influence scope, and resource status changes of all scheduling operations. When the resource status change exceeds 20% of the expectation, mark it as an abnormal event to obtain an adjustment execution control record.

[0096] Specifically, formulate a medium- and long-term tactical layer resource plan according to the load prediction data in the workload control analysis report, set resource adjustment threshold parameters, and obtain a tactical layer scheduling plan table. The tactical layer resource plan is a resource allocation plan for the next few hours to days. Compared with the real-time scheduling of the operation layer, it pays more attention to global resource utilization and long-term balance. The load prediction data provides the intensity prediction and resource demand prediction of various workloads within the future time window. Based on these prediction data, the tactical layer plan calculates the total amount and distribution of resources required in different time periods. The resource adjustment threshold parameters are the condition settings for triggering resource reconfiguration, including the load change rate threshold, resource utilization threshold, and service quality threshold. When the monitored actual load change exceeds these thresholds, trigger the corresponding resource adjustment operation. By comprehensively analyzing the prediction data and historical resource utilization efficiency, generate a tactical layer scheduling plan table. This plan table is a time-resource matrix, where the horizontal axis is the time period division, the vertical axis is the resource type and adjustment domain, and the values in the matrix represent the amount of a specific type of resource allocated to a specific adjustment domain in a specific time period.

[0097] Based on the tactical layer scheduling plan table, dynamically adjust and calculate the size of the virtual resource pool and the boundary of the adjustment domain, generate a resource reservation instruction set, and obtain a resource domain adjustment plan. The adjustment of the virtual resource pool size is to dynamically expand or shrink the capacity of the virtual resource pool according to the resource demand planned at the tactical layer; the adjustment of the boundary of the adjustment domain is to re-divide the scope of the adjustment domain according to the change of the load distribution, and optimize the resource management structure. The adjustment calculation process first analyzes the resource demand in each time period in the tactical layer scheduling plan table, calculates the ideal size of the virtual resource pool, and then compares it with the current size of the virtual resource pool to determine the amount of resources to be expanded or contracted. For the boundary of the adjustment domain, calculate the load distribution and resource utilization of each adjustment domain. When the load of a certain adjustment domain is too high or too low, adjust its boundary and reallocate resources. According to the calculation results, generate a resource reservation instruction set, including resource addition instructions, resource release instructions, and resource transfer instructions. Each instruction specifies the operation type, target resource, time window, and quantity. These instructions are combined to form a resource domain adjustment plan, which clearly stipulates the adjustment operations of the virtual resource pool and the adjustment domain.

[0098] According to the resource domain adjustment plan and the real-time load data in the workload control analysis report, calculate the operation layer resource allocation strategy, including calculating the resource allocation matrix, network path selection matrix, and priority queue configuration table, and obtain a real-time scheduling instruction sequence. The operation layer resource allocation is for the current and near-term real-time resource scheduling, directly responding to the changes in real-time workload. The resource allocation matrix describes the mapping relationship between computing tasks and processor nodes. By considering the computing requirements of tasks, the processing capabilities of nodes, and the current load, the most suitable processing node is assigned to each computing task. The network path selection matrix describes the selection of data flow transmission paths in the network. By analyzing the network topology, link utilization, and transmission delay, the optimal transmission path is selected for data flows with different priorities. The priority queue configuration table stipulates the processing priorities and resource preemption permissions of different workloads to ensure that critical services are given priority. Convert the calculation results of these three policy components into a series of specific scheduling operation instructions to form a real-time scheduling instruction sequence. Each instruction contains information such as operation type, target resource, source task, target task, and parameter settings.

[0099] Decompose the real-time scheduling instruction sequence into atomic operation instructions, construct an instruction dependency graph and calculate the execution order. When there are conflicts in the dependency relationships, automatically insert synchronization points to obtain an optimized instruction execution flow. Atomic operation instructions are basic operation units that cannot be further divided. Each atomic operation completes an independent and basic resource scheduling action, such as allocating a single resource, starting a single task, or adjusting a single parameter. The instruction dependency graph is a directed graph structure, where nodes represent atomic operations and edges represent the dependency relationships between operations. The dependency relationships include data dependency (one operation uses the result of another operation), resource dependency (multiple operations on the same resource), and timing dependency (operations must be executed in a specific order). Based on the dependency graph, use the topological sorting algorithm to calculate the execution order of operations to ensure that the dependency relationships are satisfied. When conflicts are found in the dependency relationships, such as circular dependencies or resource contention, automatically insert synchronization points at critical positions to force operations to be executed in a specific order or wait for resources to be released. Through this process, the original instruction sequence is converted into a more optimized instruction execution flow, which not only ensures the correct execution order of operations but also improves the parallelism and execution efficiency. Convert the optimized instruction execution flow into device-level control commands through an optical switching controller. For operations with an execution time exceeding 50 milliseconds, start a parallel execution mechanism to obtain a device control command sequence. The optical switching controller is the core control device in the network, capable of dynamically configuring optical paths to achieve high-speed and low-latency network resource reconfiguration. During the conversion process, the optical switching controller parses the optimized instruction execution flow, looks up the device command mapping table according to the instruction type and target resources, and converts the high-level scheduling instructions into control commands executable by specific devices, such as switch port configuration commands, server resource allocation commands, or storage system access permission setting commands. For large operations with an expected execution time exceeding 50 milliseconds, the system decomposes them into multiple sub-operations that can be executed in parallel and generates a parallel execution plan for these sub-operations to improve the execution efficiency. The converted device control command sequence contains detailed execution time arrangements, target device identifiers, operation parameters, and verification conditions, providing direct guidance for actual execution.

[0100] Transmit the device control command sequence to the resource control agent in the corresponding resource adjustment domain, execute the resource reconfiguration operation, and record the execution status of each operation. When the number of execution failures exceeds three times, trigger the fallback mechanism to obtain the execution status tracking data. During the command transmission process, send the command sequence to the resource control agent in the corresponding adjustment domain through the network management channel. After receiving the command, the agent first verifies the validity of the command and whether the current environmental conditions meet the execution requirements, and then executes the resource reconfiguration operation in the planned order. During the execution process, record the start time, completion time, execution result, and resource status changes of each operation to form a detailed execution log. When an operation fails to execute, the agent will try to execute it again. If the consecutive failures exceed three times, trigger the fallback mechanism, revoke the executed operations, restore to the state before execution, and report the execution failure to the superior. All the execution status information is aggregated to form the execution status tracking data, which details the execution process and results of each operation and provides a basis for subsequent analysis and optimization.

[0101] Analyze and integrate the execution status tracking data, record the execution results, influence scope, and resource status changes of all scheduling operations. When the resource status change exceeds 20% of the expectation, mark it as an abnormal event to obtain the adjustment execution control record. In the analysis and integration process, first clean and standardize the execution status tracking data to ensure the consistency and integrity of the data format, and then classify and sort it according to the operation type, adjustment domain, and execution time period, and calculate key indicators such as operation success rate, average execution time, and resource change rate. For each scheduling operation, record its execution result (success, failure, or partial success), influence scope (the resource scope involved and the number of affected tasks), and the resulting resource status changes (such as resource utilization rate, queue length change, etc.). During the analysis process, compare with the expected effect. When the change in the status of a certain resource exceeds 20% of the expected value (whether it exceeds or falls short of the expectation), mark it as an abnormal event and record the detailed information for subsequent investigation. Finally, generate the adjustment execution control record, which is a comprehensive execution report containing the success and failure analysis of the execution, resource status change statistics, list of abnormal events, and performance evaluation data.

[0102] Taking the network resource scheduling of a large cloud service provider as an example, according to the load prediction data in the workload control analysis report, the load of the northern computing cluster will continue to grow within the next 24 hours, and the peak is expected to occur 12 hours later. Based on this prediction, resource adjustment parameters with an upper threshold of 85% and a lower threshold of 25% for resource utilization are set, and a tactical layer scheduling plan is formulated. It is planned to expand the computing resources of the virtual resource pool in the north by 20% after 6 hours, and at the same time transfer 15% of the computing resources from the eastern region with lower load to the north. Based on this plan, the size of the virtual resource pool is accurately calculated, and it is determined that 60 computing resource units need to be added in the north, generating a resource reservation instruction set containing instructions such as "reserve 60 computing resource units in the north", "release 45 computing resource units in the east", and "adjust the north boundary to include rack No. 3". As the real-time load data changes, a detailed operation layer resource allocation strategy is calculated, such as preferentially allocating video processing tasks to nodes with GPUs, allocating database query tasks to nodes with larger memory, and preferentially routing data transmission through optical fibers No. 5 and No. 8, forming a real-time scheduling instruction sequence containing 32 instructions. After dependency analysis of these instructions, they are decomposed into 72 atomic operations, an instruction dependency graph is constructed, and it is found that there are two resource competition points (multiple tasks simultaneously applying for the same GPU resource). Therefore, synchronization points are inserted to require tasks to access these resources serially, generating an optimized instruction execution flow. Through the optical switch controller, these instructions are converted into specific device commands, such as device-level instructions like "configure switch S12 port 1 - 8 to enable VLAN 24" and "allocate 8-core CPU of server N05 to task T17". During the execution process, the status of each operation is recorded, such as "switch S12 configuration successful" and "server N05 resource allocation failed, reason: insufficient memory". For failed operations, the system retries, and after three consecutive failures, a rollback is triggered to revoke the resource allocation of the corresponding server. The final generated execution control record shows that the overall success rate of this scheduling operation is 94%, and the resource utilization rate has increased from the original 65% to 78%. However, the actual allocated amount of computing resources differs from the expected value by 25%, which is marked as an abnormal event, and the reason is further investigated and the prediction model is adjusted.

[0103] In a specific embodiment, the process of executing step S106 may specifically include the following steps:

[0104] Extract the processor energy consumption data, network transmission energy consumption data, storage system energy consumption data, and refrigeration system energy consumption data from the adjustment execution control record to obtain the original energy consumption data set;

[0105] Group and standardize the original energy consumption data set according to the resource adjustment domain and time dimension, calculate the energy consumption baseline value of each adjustment domain, and obtain the standardized energy consumption data matrix;

[0106] Construct a multivariate energy consumption prediction model based on the standardized energy consumption data matrix, taking the device power consumption characteristics, resource utilization rate, and environmental temperature as input variables. When the model fitting degree is lower than 0.85, re-collect the training data to obtain the energy consumption prediction function;

[0107] Use the energy consumption prediction function to evaluate the energy efficiency of different workload types and resource scheduling strategies, calculate the energy efficiency ratio index, energy consumption distribution balance degree, and peak energy consumption control effect of each combination, and obtain the energy efficiency evaluation result table;

[0108] Identify the energy consumption hotspots and energy efficiency valleys according to the energy efficiency evaluation result table, conduct a special analysis of the energy consumption in the optoelectronic conversion process, and perform an optimization mark when the proportion of optoelectronic conversion in the total energy consumption exceeds 30% to obtain the energy efficiency optimization target list;

[0109] Input the energy efficiency optimization target list into the multi-objective optimization framework, set the weight ratio of the performance target and the energy efficiency target to 6:4, and perform genetic algorithm optimization iteration calculations on the scheduling strategy parameters to obtain the balanced strategy parameter set;

[0110] Generate an energy efficiency adjustment report based on the balanced strategy parameter set, including the energy consumption analysis results, optimization suggestions, and predicted energy-saving effects of each adjustment domain, and at the same time update the control strategy parameters to obtain the energy efficiency adjustment report and control strategy parameters.

[0111] Specifically, the process of extracting the processor energy consumption data, network transmission energy consumption data, storage system energy consumption data, and refrigeration system energy consumption data from the adjustment execution control record is achieved through a data extraction algorithm. First, the adjustment execution control record contains detailed records generated by the system during the execution of resource scheduling, and these records contain energy consumption-related data of various devices. The data extraction algorithm locates and extracts various energy consumption data by identifying specific markers in the adjustment execution control record. Processor energy consumption data refers to the electrical energy consumed by computing resources during the execution of tasks, and these data are usually recorded in watts per unit time; network transmission energy consumption data refers to the electrical energy consumed during the transmission of data in the network, including the energy consumption of network devices such as switching devices and routers; storage system energy consumption data records the electrical energy consumed by storage devices during read and write operations; refrigeration system energy consumption data records the electrical energy consumed by the refrigeration system used for device cooling. After extraction, these data form an original energy consumption data set, which contains energy consumption records of different types of devices at different time points and under different load conditions.

[0112] Grouping and normalizing the original energy consumption dataset according to the resource adjustment domain and time dimension means grouping the energy consumption data according to different resource adjustment domains and aligning them on the time axis. The resource adjustment domain refers to the resource set formed during the division of the resource control pool, and resources with similar resource characteristics and management requirements are divided into the same adjustment domain. After grouping, the data in each group is normalized to eliminate the order-of-magnitude differences between different devices and different-scale resource domains, making the data comparable. The normalization process uses the Z-score normalization method to convert the original data into a standard normal distribution with a mean of 0 and a standard deviation of 1. On this basis, the energy consumption baseline value of each adjustment domain is calculated. The energy consumption baseline value refers to the energy consumption level under standard load conditions and serves as a baseline for evaluating energy efficiency. After these processes, a normalized energy consumption data matrix is formed, and each element in this matrix represents the normalized energy consumption value of a specific adjustment domain at a specific time point.

[0113] The process of constructing a multivariate energy consumption prediction model based on the normalized energy consumption data matrix is to use the device power consumption characteristics, resource utilization rate, and environmental temperature as input variables to construct the prediction model. The device power consumption characteristics refer to the power consumption curve of the device under different working conditions, the resource utilization rate refers to the degree to which resources are used, and the environmental temperature refers to the temperature value of the environment where the device is located. The model construction uses the multiple regression analysis method to establish the mathematical relationship between the input variables and energy consumption. During the model training process, the historical data in the normalized energy consumption data matrix is used for parameter fitting. When the model fitting degree (usually represented by the R² value) is lower than 0.85, it indicates that the model prediction accuracy is insufficient, and the training data is re-collected for model optimization. After the model training is completed, an energy consumption prediction function is obtained, and this function can predict the energy consumption level under specific conditions based on the input variables.

[0114] In the process of using the energy consumption prediction function to evaluate the energy efficiency of different workload types and resource scheduling strategies, three key indicators are calculated: the energy efficiency ratio indicator, the energy consumption distribution balance degree, and the peak energy consumption control effect. The energy efficiency ratio indicator refers to the amount of work completed per unit of energy consumption, and the calculation method is the amount of work divided by the energy consumption value; the energy consumption distribution balance degree measures the distribution of energy consumption among different devices and different time periods, usually represented by the standard deviation of the energy consumption distribution; the peak energy consumption control effect measures the ability of the system to control the peak energy consumption under high load conditions, usually represented by the ratio of the peak energy consumption to the average energy consumption. By evaluating the combinations of different workload types (such as compute-intensive, data transmission-intensive, hybrid) and different resource scheduling strategies, an energy efficiency evaluation result table is generated, which records the energy efficiency evaluation index values for various combinations.

[0115] The process of identifying energy consumption hot spots and energy efficiency valleys based on the energy efficiency evaluation result table is to find areas with abnormally high energy consumption (hot spots) and areas with particularly low energy efficiency (valleys) through data analysis algorithms. In particular, the energy consumption of the photoelectric conversion process is specifically analyzed. Photoelectric conversion refers to the conversion process between optical signals and electrical signals, which is a high energy consumption link in the optical switching network. When the proportion of photoelectric conversion energy consumption in total energy consumption exceeds 30%, it is marked as an optimization target and added to the energy efficiency optimization target list. The energy efficiency optimization target list contains all optimized energy consumption hot spots and efficiency valleys, as well as specific optimization directions.

[0116] The energy efficiency optimization target list is input into the multi-objective optimization framework, and the weight ratio of performance target and energy efficiency target is set to 6:4, which means that in the optimization process, the system performance target accounts for 60% of the weight and the energy efficiency target accounts for 40% of the weight. The multi-objective optimization framework uses genetic algorithms for optimization calculations. Genetic algorithms are an optimization algorithm that simulates the biological evolution process. It continuously iterates through operations such as selection, crossover, and mutation to find the optimal solution. In this process, the scheduling strategy parameters (such as resource allocation ratio, task priority, scheduling cycle, etc.) are used as optimization variables. Through multiple iterative calculations, a set of strategy parameters that balance performance and energy efficiency is obtained.

[0117] An energy efficiency regulation report is generated based on the balance strategy parameter set, which includes the energy consumption analysis results, specific optimization suggestions and predicted energy saving effects of each regulation domain. At the same time, the optimized strategy parameters are updated to the control system as new control strategy parameters to guide subsequent resource scheduling decisions.

[0118] For example: In a network resource scheduling system of a large data center, the adjustment execution control record stores the energy consumption data of each device in the past 24 hours. Through a data extraction algorithm, the processor energy consumption data of 50 servers (one sampling point every 15 minutes), the transmission energy consumption data of 30 network devices, the energy consumption data of 20 storage devices, and the energy consumption data of 5 refrigeration units are extracted. These data are grouped according to the 8 resource adjustment domains divided in the early stage, and each adjustment domain contains several computing, network, and storage devices. The grouped data is processed by Z-score standardization to eliminate the order-of-magnitude differences between different devices. The energy consumption baseline values of each adjustment domain are calculated. For example, the processor energy consumption baseline of adjustment domain 1 is 450 watts, and the network device energy consumption baseline is 120 watts. When constructing a multivariable energy consumption prediction model, the server utilization rate (0 - 100%), network traffic (Mbps), storage read / write rate (IOPS), and ambient temperature (in degrees Celsius) are used as input variables, and a prediction model is established through multiple regression analysis. The initial model fitting degree is 0.78, which is lower than the threshold of 0.85. Therefore, the sampling points are increased and retrained to obtain a model with a fitting degree of 0.89. This model is used to evaluate different workloads (such as Web applications, data analysis, video transcoding) and scheduling strategy combinations, and the energy efficiency ratio, energy consumption balance degree, and peak control effect of each combination are calculated. The analysis results show that in the current scheduling strategy, the optoelectronic conversion energy consumption of the data analysis workload accounts for 34% of the total energy consumption, so it is marked as an optimization target. Through genetic algorithm optimization, the resource allocation strategy is adjusted, and the compute-intensive tasks are concentrated on specific servers for execution, reducing unnecessary data transmission and optoelectronic conversion. Finally, a detailed energy efficiency adjustment report and optimized control strategy parameters are generated.

[0119] The above describes the multi-agent-based network resource hierarchical coordination method in the embodiments of the present application. Next, the multi-agent-based network resource hierarchical coordination system in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the multi-agent-based network resource hierarchical coordination system in the embodiments of the present application includes:

[0120] A monitoring module, configured to perform optical switching monitoring on the node processor load and link utilization rate in the computer network to obtain a set of network resource status control parameters;

[0121] A partitioning module, configured to construct a resource control pool using virtualization technology based on the set of network resource status control parameters and perform resource adjustment domain partitioning to obtain a resource control status report and an adjustment domain partitioning scheme;

[0122] A control module, configured to deploy a resource control agent and a central coordination control unit in each adjustment domain according to the resource control status report and the adjustment domain partitioning scheme to obtain a multi-agent adaptive control system;

[0123] An extraction module, configured to use a multi-agent adaptive control system to perform feature extraction and multi-level classification adjustment on network workloads, and obtain a workload control analysis report;

[0124] An adjustment module, configured to execute a two-layer adjustment strategy at the tactical layer and the operation layer through an optical switch controller according to the workload control analysis report, and obtain an adjustment execution control record;

[0125] An analysis module, configured to perform multi-dimensional control parameter analysis on processor energy consumption and network transmission energy consumption based on the adjustment execution control record, and obtain an energy efficiency adjustment report and control strategy parameters.

[0126] Through the collaborative cooperation of the above-mentioned various components, by monitoring the node processor load and link utilization rate in the computer network through the optical switching monitoring module, high-precision and low-interference network status monitoring is achieved, solving the problems of insufficient accuracy and large monitoring overhead in traditional monitoring systems, providing an accurate data basis for resource scheduling, and at the same time, the optical switching monitoring technology ensures that the monitoring process does not affect normal network transmission; based on the network resource status control parameter set, a resource control pool is constructed using virtualization technology and resource adjustment domains are divided, breaking the limitations of traditional physical resource boundaries, realizing the unified management and flexible configuration of resources, improving resource utilization rate, and the division of resource adjustment domains takes into account resource affinity and service relevance, creating conditions for the efficient scheduling of resources; according to the resource control status report and the adjustment domain division plan, resource control agents and a central coordination control unit are deployed in each adjustment domain, constructing a multi-agent adaptive control system. This system adopts a distributed decision-making mechanism, solving the problems of slow response and high single-point failure risk in the centralized scheduling system. The multi-agents achieve high-speed collaborative decision-making through the optical switching network, ensuring the fast response ability of the scheduling system to network state changes; using the multi-agent adaptive control system to extract features and perform multi-level classification and adjustment on network workloads, accurate identification and classification of workloads are achieved through artificial intelligence algorithms. The algorithm features play a key role in the workload classification process. The hybrid classifier combines the interpretability of decision trees, the classification efficiency of support vector machines, and the feature extraction ability of deep neural networks, significantly improving classification accuracy; based on the workload control analysis report, double-layer adjustment strategies at the tactical layer and the operation layer are executed through the optical switching controller, realizing the coordination of long-term resource planning and instant resource allocation. The double-layer scheduling mechanism makes full use of the low-latency characteristics of the optical switching network, realizing millisecond-level reconfiguration of resources; based on the adjustment execution control record, multi-dimensional control parameter analysis of processor energy consumption and network transmission energy consumption is carried out. Through the multi-variable energy consumption prediction model and the multi-objective optimization framework, effective reduction of energy consumption is achieved while ensuring performance. Artificial intelligence algorithms play an important role in the energy consumption prediction and optimization process. Through the comprehensive analysis of various variables such as device power consumption characteristics, resource utilization rate, and environmental temperature, a high-precision energy consumption prediction model is constructed. The genetic algorithm effectively balances the performance and energy efficiency objectives in the multi-objective optimization process, realizing the global optimization of the scheduling strategy. The entire technical solution combines optical switching technology, virtualization technology, and artificial intelligence algorithms organically to construct a set of efficient, intelligent, and energy-saving computer network resource scheduling methods, solving the problems of insufficient monitoring accuracy, inflexible resource pool division, slow scheduling decision response, inaccurate load classification, insufficient multi-level scheduling coordination, and insufficient energy consumption optimization in the existing technology, and significantly improving the network resource utilization rate and the overall performance of the system.

[0127] Above Figure 2From the perspective of modular functional entities, the multi-agent-based network resource hierarchical coordination system in the embodiments of the present invention is described in detail. Next, the multi-agent-based network resource hierarchical coordination device in the embodiments of the present invention is described in detail from the perspective of hardware processing.

[0128] Figure 3 FIG. 4 is a schematic structural diagram of a multi-agent-based network resource hierarchical coordination device provided by an embodiment of the present invention. The multi-agent-based network resource hierarchical coordination device 300 may vary greatly due to configuration or performance, and may include one or more processors (central processing units, CPUs) 310 (for example, one or more processors) and a memory 320, and one or more storage media 330 for storing application programs 333 or data 332 (for example, one or more mass storage device terminals). Among them, the memory 320 and the storage media 330 may be transient storage or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the multi-agent-based network resource hierarchical coordination device 300. Further, the processor 310 may be configured to communicate with the storage media 330 and execute a series of instruction operations in the storage media 330 on the multi-agent-based network resource hierarchical coordination device 300 to implement the steps of the above multi-agent-based network resource hierarchical coordination method.

[0129] The multi-agent-based network resource hierarchical coordination device 300 may further include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 3 the shown structure of the multi-agent-based network resource hierarchical coordination device does not limit the multi-agent-based network resource hierarchical coordination device provided by the present invention, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0130] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions are run on a computer, the computer is caused to execute the steps of the multi-agent-based network resource hierarchical coordination method.

[0131] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, systems and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0132] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a multi-agent-based network resource hierarchical coordination device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0133] The above is the case. The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A multi-agent based hierarchical coordination method for network resources, characterized in that, The method includes: Monitoring the node processor load and link utilization rate in the computer network by an optical switching monitoring module to obtain a set of network resource status control parameters; Based on the set of network resource status control parameters, using virtualization technology to construct a resource control pool and perform resource adjustment domain division to obtain a resource control status report and an adjustment domain division scheme; Deploying a resource control intelligent agent and a central coordination control unit in each adjustment domain according to the resource control status report and the adjustment domain division scheme to obtain a multi-agent adaptive control system; Using the multi-agent adaptive control system to perform feature extraction and multi-level classification adjustment on the network workload to obtain a workload control analysis report; According to the workload control analysis report, executing a two-layer adjustment strategy at the tactical layer and the operation layer through an optical switching controller to obtain an adjustment execution control record, including: formulating a medium- and long-term tactical layer resource plan according to the load prediction data in the workload control analysis report, setting resource adjustment threshold parameters to obtain a tactical layer scheduling plan table; dynamically adjusting and calculating the size of the virtual resource pool and the adjustment domain boundary based on the tactical layer scheduling plan table to generate a resource reservation instruction set to obtain a resource domain adjustment scheme; calculating an operation layer resource allocation strategy according to the resource domain adjustment scheme and the real-time load data in the workload control analysis report, including calculating a resource allocation matrix, a network path selection matrix, and a priority queue configuration table to obtain a real-time scheduling instruction sequence; decomposing the real-time scheduling instruction sequence into atomic operation instructions, constructing an instruction dependency graph and calculating the execution order, and automatically inserting synchronization points when dependency conflicts occur to obtain an optimized instruction execution flow; converting the optimized instruction execution flow into a device-level control command through the optical switching controller, starting a parallel execution mechanism for operations with an execution time exceeding 50 milliseconds to obtain a device control command sequence; transmitting the device control command sequence to the resource control intelligent agent in the corresponding resource adjustment domain to execute a resource reconfiguration operation and record the execution status of each operation, triggering a fallback mechanism when the number of execution failures exceeds three times to obtain execution status tracking data; analyzing and integrating the execution status tracking data, recording the execution results, influence scope, and resource status changes of all scheduling operations, and marking as an abnormal event when the resource status change exceeds 20% of the expectation to obtain an adjustment execution control record; Based on the adjustment execution control record, performing multi-dimensional control parameter analysis on the processor energy consumption and network transmission energy consumption to obtain an energy efficiency adjustment report and control strategy parameters.

2. The multi-agent based network resource hierarchical coordination method according to claim 1, wherein The monitoring of the node processor load and link utilization rate in the computer network by the optical switching monitoring module to obtain a set of network resource status control parameters includes: Deploying a distributed optical switching monitoring module at network key nodes, configuring connections for the network key nodes to obtain the distribution topology of the monitoring nodes; Dynamically adjusting and setting the acquisition frequency of the monitoring nodes, automatically calibrating the sampling rate according to the network traffic volatility to obtain an adaptive sampling configuration; Collect link utilization rate, packet transmission delay, queue length, packet loss rate, and optical signal strength through the optical switching monitoring module to obtain original monitoring data; Add timestamp, optical path identifier, and topological location information to the original monitoring data to obtain tokenized original data; Transmit the tokenized original data through a dedicated optical transmission channel to the central data processing unit to obtain a centralized data stream; Perform filtering, denoising, and normalization processing on the centralized data stream to obtain normalized network resource status data; Perform optical signal quality anomaly detection based on the normalized network resource status data, and trigger an alarm when the optical signal quality deteriorates or the optical path is congested to obtain a network resource status control parameter set; 3. The multi-agent based network resource hierarchical coordination method according to claim 1, characterized in that Based on the network resource status control parameter set, use virtualization technology to construct a resource control pool and perform resource adjustment domain division to obtain a resource control status report and an adjustment domain division plan, including: Perform abstraction processing on physical network resources, and use a resource description language to standardize the description of computing resources, storage resources, and network bandwidth resources to obtain a resource metadata model; Establish a mapping relationship from physical resources to virtual resources based on the resource metadata model, and assign identifiers to each virtual resource unit through an optical switching controller to obtain a resource mapping data structure; Calculate a resource correlation matrix based on the topological information in the network resource status control parameter set and the resource mapping data structure to obtain a resource affinity index; Based on the resource affinity index and the resource metadata model, perform hierarchical clustering analysis on the resource control pool, and classify resources with a correlation higher than the preset threshold of 0.75 into the same resource adjustment domain to obtain an initial adjustment domain division result; Perform boundary optimization analysis on the initial adjustment domain division result and the resource mapping data structure, calculate the internal connectivity and external isolation degree of each adjustment domain, and adjust the attribution of boundary nodes to obtain an optimized resource adjustment domain division; For the optimized resource adjustment domain division and the resource affinity index, create resource configuration templates for computing-intensive, bandwidth-intensive, and storage-intensive resources, and define the proportional parameters of various resources to obtain a resource configuration template library; Perform resource status integration analysis based on the resource configuration template library and the optimized resource adjustment domain division to generate a resource control status report and an adjustment domain division plan, where the resource control status report includes the total amount of virtual resources, the current allocation status, and resource utilization statistical data of each resource adjustment domain, and the adjustment domain division plan includes the physical boundaries, inter-domain connection relationships, and resource scheduling permission settings of each resource adjustment domain.

4. The multi-agent based network resource hierarchical coordination method according to claim 1, wherein Deploy resource control agents and a central coordination control unit in each adjustment domain according to the resource control status report and the adjustment domain division plan to obtain a multi-agent adaptive control system, including: Assign a unique identifier to each resource adjustment domain based on the adjustment domain division plan and determine its physical boundary to obtain an adjustment domain control mapping table; Perform an association analysis on the resource data in the resource control status report and the adjustment domain control mapping table to determine the resource control complexity level of each adjustment domain, and obtain a hierarchical control agent model; Deploy resource control agents with a three-layer structure of a perception layer, a decision-making layer, and an execution layer in each resource adjustment domain according to the hierarchical control agent model and the adjustment domain control mapping table to obtain a hierarchical agent control network; Load a reinforcement learning algorithm into the decision-making layer of the hierarchical agent control network, and use the parameters in the hierarchical control agent model as initial values to obtain an agent learning decision rule; Deploy a central coordination control unit at the network center node, and import the agent learning decision rule and the adjustment domain control mapping table into the central coordination control unit to obtain a global coordination control topology; Establish an inter-agent communication mechanism based on the global coordination control topology and the hierarchical agent control network, calculate the trust matrix between agents, and obtain an agent cooperation control mechanism; Apply the agent cooperation control mechanism to the hierarchical agent control network, construct a state synchronization protocol based on the adjustment domain control mapping table, and use an optical switching network to transmit agent state data to obtain a multi-agent adaptive control system.

5. The multi-agent-based network resource hierarchical coordination method according to claim 1, characterized in that The multi-agent adaptive control system is used to extract features and perform multi-level classification adjustment on network workloads to obtain a workload control analysis report, including: Extract the time features, space features, and type features of the network workload from the multi-agent adaptive control system to obtain a multi-dimensional workload feature set; Perform time series analysis and pattern recognition on the multi-dimensional workload feature set, and use the sliding window method to calculate the statistical distribution parameters of the features to obtain a workload pattern feature library; Construct a multi-level classification framework based on the workload pattern feature library, hierarchically classify the workloads according to resource demand types, service characteristics, and priorities to obtain a workload classification structure; Input the workload classification structure into a hybrid classifier, and combine decision trees, support vector machines, and deep neural networks to perform accurate classification of the workload. When the classification confidence is lower than 0.8, trigger an artificial confirmation process to obtain a classification result dataset; Set a change detector based on the classification result dataset, and use the CUSUM algorithm to monitor changes in the mean value, variance, and distribution characteristics of the workload. When the cumulative change amount exceeds the set threshold, obtain a load change trigger signal; Establish a workload prediction model based on the load change trigger signal and the workload pattern feature library, and use time series analysis and machine learning methods to predict the workload change trend within the next 60 minutes to obtain load prediction data; Integrate and analyze the classification result dataset, the load change trigger signal, and the load prediction data, calculate the resource requirements and scheduling priorities of various types of workloads, and obtain a workload control analysis report.

6. The multi-agent-based network resource hierarchical coordination method according to claim 1, wherein Perform multi-dimensional control parameter analysis on the processor energy consumption and network transmission energy consumption based on the adjustment execution control record to obtain an energy efficiency adjustment report and control strategy parameters, including: Extract the processor energy consumption data, network transmission energy consumption data, storage system energy consumption data, and refrigeration system energy consumption data from the adjustment execution control record to obtain an original energy consumption data set; Group and standardize the original energy consumption data set according to the resource adjustment domain and time dimension, calculate the energy consumption baseline value of each adjustment domain, and obtain a standardized energy consumption data matrix; Construct a multi-variable energy consumption prediction model based on the standardized energy consumption data matrix, use the device power consumption characteristics, resource utilization rate, and environmental temperature as input variables, and re-collect training data when the model fitting degree is lower than 0.85 to obtain an energy consumption prediction function; Use the energy consumption prediction function to evaluate the energy efficiency of different workload types and resource scheduling strategies, calculate the energy efficiency ratio index, energy consumption distribution balance degree, and peak energy consumption control effect of each combination, and obtain an energy efficiency evaluation result table; Identify energy consumption hotspots and energy efficiency troughs based on the energy efficiency evaluation result table, conduct a special analysis of the energy consumption of the optoelectronic conversion process, and perform an optimization mark when the proportion of optoelectronic conversion in the total energy consumption exceeds 30% to obtain an energy efficiency optimization target list; Input the energy efficiency optimization target list into a multi-objective optimization framework, set the weight ratio of the performance target and the energy efficiency target to 6:4, and perform genetic algorithm optimization iteration calculation on the scheduling strategy parameters to obtain a balanced strategy parameter set; Generate an energy efficiency adjustment report based on the balanced strategy parameter set, including the energy consumption analysis results, optimization suggestions, and predicted energy-saving effects of each adjustment domain, and at the same time update the control strategy parameters to obtain an energy efficiency adjustment report and control strategy parameters.

7. A multi-agent-based hierarchical coordination system for network resources, characterized in that, For implementing the multi-agent-based network resource hierarchical coordination method according to any one of claims 1-6, the multi-agent-based network resource hierarchical coordination system includes: A monitoring module for monitoring the node processor load and link utilization rate in a computer network through an optical switching monitoring module to obtain a network resource status control parameter set; A partitioning module for constructing a resource control pool using virtualization technology based on the network resource status control parameter set and performing resource adjustment domain partitioning to obtain a resource control status report and an adjustment domain partitioning scheme; A control module for deploying resource control agents and a central coordination control unit in each adjustment domain according to the resource control status report and the adjustment domain partitioning scheme to obtain a multi-agent adaptive control system; An extraction module for using the multi-agent adaptive control system to perform feature extraction and multi-level classification adjustment on network workloads to obtain a workload control analysis report; An adjustment module for executing a two-layer adjustment strategy at the tactical layer and the operation layer through an optical switching controller according to the workload control analysis report to obtain an adjustment execution control record; An analysis module for performing multi-dimensional control parameter analysis on the processor energy consumption and network transmission energy consumption based on the adjustment execution control record to obtain an energy efficiency adjustment report and control strategy parameters.

8. A network resource hierarchical coordination device based on multi-agent, characterized in that, It includes a memory and a processor, the memory stores a computer program that can run on the processor, and when the processor executes the computer program, it implements the multi-agent-based network resource hierarchical coordination method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by a processor, the processor is caused to execute the multi-agent-based hierarchical coordination method for network resources according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • SDN-based all-optical exchange data center network control system and implementation method thereof

    CN106941633A

  • O-RAN-oriented multilevel heterogeneous resource scheduling method

    CN118535304A