Multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power cluster

By constructing a global physical performance topology graph and a dynamic virtual computing topology graph, and combining a topology matching optimization decision module and a continuous learning evolution module, the problems of low scheduling efficiency and unscientific reconstruction decisions in existing technologies are solved, achieving efficient computing power utilization and energy efficiency optimization.

CN121979679APending Publication Date: 2026-05-05BEIJING YUANJING TECHNOLOGY INFORMATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YUANJING TECHNOLOGY INFORMATION CO LTD
Filing Date
2026-01-23
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing AI computing power cluster scheduling systems do not fully exploit the inherent graph computing characteristics of AI training tasks, and do not uniformly encode and collaboratively optimize the real-time performance of physical devices and network status. This makes it difficult to achieve optimal computing power utilization efficiency and energy efficiency ratio, and lacks a scientific reconfiguration decision-making mechanism, resulting in task execution interruption and resource waste.

Method used

A multi-dimensional scheduling and energy efficiency optimization system is adopted. A global physical efficiency topology graph and a dynamic virtual computing topology graph are constructed through a physical efficiency topology management module and a virtual computing topology extraction module. The topology matching optimization decision module is combined to solve the dual graph matching problem, realize the topology reconstruction of computing power flow, and optimize the scheduling decision through a continuous learning evolution module.

Benefits of technology

It realizes the transformation of computing power scheduling from resource slot allocation to topology collaborative optimization, improves the overall computing power utilization of the cluster and the effective computing power output per unit power consumption, and ensures service continuity and dynamic scheduling adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121979679A_ABST
    Figure CN121979679A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses an artificial intelligence computing power cluster-oriented multi-dimensional scheduling and energy efficiency optimization system, which comprises a physical efficiency topology management module, a virtual computing topology extraction module, a topology matching optimization decision module, an online collaborative reconstruction execution module and a continuous learning evolution module, based on static attribute data and dynamic monitoring data of all physical devices in a computing power cluster, a global physical efficiency topological graph is constructed, a dynamic virtual computing topological graph is extracted, a topological matching degree optimization model is constructed and solved, an optimal mapping target for mapping virtual operators to the physical devices and an expected matching degree gain are obtained, and the optimal mapping target and the expected matching degree gain are obtained. When the expected matching degree gain exceeds a gain threshold value, computing power flow topology reconstruction is completed, and continuous learning and evolution are carried out on the topology matching degree optimization model; according to the invention, the multi-dimensional scheduling and energy efficiency optimization efficiency for the artificial intelligence computing power cluster can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters. Background Technology

[0002] In the field of AI computing cluster scheduling, existing technologies generally adopt traditional resource allocation approaches, treating hardware resources such as CPUs, memory, and GPUs of physical devices as independent "slots." Scheduling decisions focus solely on whether the remaining resources meet task requirements, lacking a deep consideration of the coupling between the physical topology characteristics of the cluster and the computational characteristics of the tasks. This type of scheduling method does not fully explore the inherent graph computation characteristics of AI training tasks, nor does it uniformly encode and collaboratively optimize the real-time performance of physical devices and network status. As a result, the matching of tasks and resources remains at a basic adaptation level, making it difficult to achieve the optimal balance between computing power utilization efficiency and energy efficiency ratio. At the same time, existing solutions lack the ability to adapt to dynamic changes and cannot perceive the dynamic evolution of device load fluctuations, temperature changes, and task computing communication modes in real time. This can easily lead to problems such as wasted computing power, overheating and frequency reduction of hot devices, or cross-device communication bottlenecks.

[0003] Furthermore, existing computing power scheduling systems lack scientific reconfiguration decision-making mechanisms and seamless migration solutions, making it difficult to address the core pain points of "when to reconfigure" and "how to reconfigure seamlessly." On the one hand, traditional systems have not established an effective expected gain assessment system, either blindly reallocating resources, resulting in migration costs exceeding performance improvements, or missing optimization opportunities due to failure to adjust mapping relationships in a timely manner. On the other hand, when it is necessary to adjust the mapping relationship between tasks and devices, overall migration or interrupted migration methods are often adopted, which not only generate a large amount of state data transmission overhead but also cause task execution interruptions, seriously affecting service continuity and user experience. These shortcomings make it difficult for existing scheduling systems to adapt to the dynamic scheduling needs of large-scale AI clusters, restricting the improvement of the overall performance and energy efficiency optimization level of computing power clusters. Therefore, how to improve the overall performance and energy efficiency optimization level of computing power clusters has become an urgent problem to be solved. Summary of the Invention

[0004] This invention provides a multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing clusters to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, this invention provides a multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing clusters. The system comprises a physical performance topology management module, a virtual computing topology extraction module, a topology matching optimization decision module, an online collaborative reconfiguration execution module, and a continuous learning and evolution module, wherein: The physical performance topology management module is used to construct and update the global physical performance topology map in real time based on the static attribute data and dynamic monitoring data of all physical devices in the computing power cluster. The virtual computing topology extraction module is used to extract and update the dynamic virtual computing topology of artificial intelligence computing tasks in real time through lightweight runtime instrumentation technology. The topology matching optimization decision module is used to construct and solve a topology matching degree optimization model based on the global physical performance topology graph and the dynamic virtual computing topology graph, so as to obtain the optimal mapping target and expected matching degree gain for mapping virtual operators to physical devices. The online collaborative reconstruction execution module is used to determine the set of operator subgraphs to be migrated based on the difference between the optimal mapping target and the current actual mapping when the expected matching degree gain exceeds the preset gain threshold, and migrate the set of operator subgraphs to be migrated to the target physical device in a streaming state migration mode to complete the computing power flow topology reconstruction. The continuous learning and evolution module is used to continuously learn and evolve the topology matching degree optimization model based on the actual performance and energy efficiency feedback data after topology reconstruction.

[0006] In a preferred embodiment, the step of constructing and updating the global physical performance topology map in real time based on the static attribute data and dynamic monitoring data of all physical devices in the computing power cluster includes: Based on the dynamic monitoring data and the static attribute data, calculate the real-time comprehensive cost-effectiveness ratio score vector for each physical device; Based on the interconnection topology and real-time network monitoring data, the real-time effective communication cost of each physical link is calculated. Using physical devices as nodes, physical connections as edges, real-time comprehensive cost-effectiveness ratio score vectors as node weights, and real-time effective communication costs as edge weights, a weighted global physical performance topology graph is constructed and continuously updated.

[0007] In a preferred embodiment, calculating the real-time comprehensive cost-effectiveness score vector for each physical device includes: The standardized instantaneous computing power score, unit power consumption computing power score, and heat dissipation efficiency score are calculated based on the comprehensive cost-effectiveness scoring equation set. The mathematical expression of the comprehensive cost-effectiveness scoring equation set is as follows: ; In the formula, P score To standardize the instantaneous computing power score, f current f is the current operating frequency of the device. base γ is the reference frequency of the device. throttle E is the frequency reduction factor. score P is used to score computing power per unit of power consumption.current For real-time power consumption of the device, C score To score heat dissipation efficiency, T max T is the upper limit of the chip junction temperature. junction For real-time core temperature, T coolant_in Intake air temperature; The standardized instantaneous computing power score, the unit power consumption computing power score, and the heat dissipation efficiency score are combined in an orderly manner to form a three-dimensional vector, resulting in a real-time comprehensive cost-effectiveness ratio score vector.

[0008] In a preferred embodiment, calculating the real-time effective communication cost of each physical link includes: Based on the interconnection topology and the real-time network monitoring data, real-time network performance measurement data for each physical link is obtained; Based on the formula for calculating the effective bandwidth of a link, the bandwidth utilization and transmission error rate in the real-time network performance measurement data are processed to obtain the effective bandwidth of each physical link. Based on the communication delay comprehensive estimation algorithm, the one-way delay measurement value and packet loss rate in the real-time network performance measurement data are processed to obtain the communication delay of each physical link; Based on the shortest path principle of graph theory, with the effective bandwidth as the capacity constraint and the communication delay as the path cost, the maximum available bandwidth and minimum communication delay between any two physical device nodes in the global physical performance topology graph are calculated, and these are respectively used as the effective bandwidth and communication delay between the pair of nodes.

[0009] In a preferred embodiment, the method for extracting and updating the dynamic virtual computing topology graph of the artificial intelligence computing task in real time using lightweight runtime instrumentation technology includes: A topology perceptron is mounted within the runtime framework of an artificial intelligence computing task to capture runtime information of the computation graph at a preset sampling frequency; Identify the key operator nodes in the computation graph and generate feature vectors representing computational density based on the shape and data type of their input and output tensors; Monitor the tensor transfer on the data dependency edges between operators, and record the amount of transmitted data and the frequency of communication triggering as the data flow characteristics of the edges; Based on the evolution of the task execution phase, the node set, edge set, and their associated characteristics are dynamically updated to form a dynamic virtual computing topology graph.

[0010] In a preferred embodiment, the step of constructing and solving a topology matching degree optimization model based on the global physical performance topology graph and the dynamic virtual computing topology graph to obtain the optimal mapping target and expected matching degree gain for mapping virtual operators to physical devices includes: Define the mapping function from virtual operator nodes to physical device nodes; Construct an overall topology matching score function and an optimization model that aims to maximize the overall topology matching score function; Under resource capacity constraints and data dependency reachability constraints, the simulated annealing algorithm is used to solve the optimization model, obtain the optimal mapping target, and calculate the expected matching degree gain.

[0011] In a preferred embodiment, the mathematical expression of the overall topology matching degree scoring function is as follows: ; In the formula, S(M) is the overall topology matching score function, S node (M) represents the total node matching score, S edge (M) represents the total edge matching cost, λ is the communication cost weight coefficient, and V v E is a set of virtual operators. v C is the set of dependency edges between virtual operators. v W is the computational density eigenvector of operator v. v (M(v)) is the real-time comprehensive cost-effectiveness ratio score vector of physical device node M(v), D e .v and D e .f represents the data transmission volume and communication frequency of the dependent edge e, respectively, and BW eff (M(v i ),M(v j )) and Lat(M(v i ),M(v j )) are respectively device node M(v i ) and M(v j The effective bandwidth and communication delay between ) are given by sim, where sim is the similarity calculation function, and M(v i ) and M(v j ) represent the virtual operator v i and v j The physical device node to which it is mapped, v i and v j It is a virtual operator; In a preferred embodiment, the step of determining the set of operator subgraphs to be migrated based on the difference between the optimal mapping target and the current actual mapping when the expected matching gain exceeds a preset gain threshold includes: Based on the optimal mapping target and the current actual mapping, the mapping positions of all virtual operators in the dynamic virtual computing topology are compared to obtain a set of operators to be moved. Based on the data dependency edges of the dynamic virtual computing topology graph, connectivity analysis is performed on the operators in the set of operators to be moved to obtain at least one connected component, wherein each connected component constitutes a candidate migration subgraph. Based on the amount of state data to be migrated by the operators contained in each candidate migration subgraph, and the predicted communication overhead incurred in migrating the subgraph to the corresponding target device in the optimal mapping target under the current mapping, a migration cost evaluation is performed to obtain the migration cost value of each candidate migration subgraph. Based on the migration cost and the expected matching gain, a cost-benefit trade-off is performed to determine the final set of operator subgraphs to be migrated.

[0012] In a preferred embodiment, the step of migrating the set of operator subgraphs to be migrated to the target physical device using a streaming state migration method to complete the computing power flow topology reconstruction includes: Based on the optimal mapping target and the current actual mapping, operators whose mappings have changed are identified in the dynamic virtual computing topology graph, and connected subgraphs formed by these operators under the data dependency relationship in the dynamic virtual computing topology graph are extracted to obtain a set of operator subgraphs to be migrated. Coordinate the task processes corresponding to the operators involved in the set of operator subgraphs to be migrated, perform cooperative checkpoint operations, and obtain a consistent checkpoint state; On each target physical device specified by the optimal mapping target, a new computing process is started, the checkpoint state is loaded, and the communication connection between the process and the physical devices where other unmigrated operators in the cluster are located is rebuilt according to the optimal mapping target. Resume computation execution of all operators in the set of operator subgraphs to be migrated from the interruption point, so that the migrated operators can be seamlessly connected to the data stream, and update the global task mapping state to the current actual mapping equal to the optimal mapping target, thus completing the topology reconstruction of the computing power flow.

[0013] In a preferred embodiment, the step of continuously learning and evolving the topology matching degree optimization model based on the actual performance and energy efficiency feedback data after topology reconstruction includes: The computational task of completing the topology reconstruction is monitored, and the actual performance and energy efficiency indicators during the stable operation period after reconstruction are collected to obtain the feedback dataset. Based on the feedback dataset, evaluate the actual matching gain obtained after executing the optimal mapping objective; Based on the difference between the actual matching degree gain and the expected matching degree gain, the parameters of the topology matching degree optimization model are adjusted to update the model.

[0014] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention achieves a fundamental shift in computing power scheduling from "resource slot allocation" to "topology collaborative optimization" through an innovative dual-graph real-time perception and dynamic matching topology scheduling scheme, resulting in significant performance and energy efficiency improvements. This scheme creatively combines the graph computing characteristics of AI training tasks with the topological characteristics of physical clusters. Through a physical performance topology management module, it uniformly encodes multi-dimensional information such as real-time computing power, power consumption, heat dissipation status, and network link bandwidth and latency of devices into node and edge weights of a global physical performance topology graph. Simultaneously, it uses a virtual computing topology extraction module to accurately capture the operator computing density and data dependency characteristics of tasks, constructing a dynamic virtual computing topology graph. By using a topology matching optimization model to solve the dual-graph matching problem, it can find the optimal physical device for each virtual operator, ensuring precise alignment between task computing requirements and cluster resource efficiency. This effectively avoids problems such as wasted computing power and frequency reduction of hotspot devices caused by neglecting dynamic performance differences and task communication modes in traditional scheduling, significantly improving the overall computing power utilization of the cluster and the effective computing power output per unit power consumption.

[0015] 2. This invention, based on an online computing power flow reconstruction mechanism using expected gain assessment and streaming migration, successfully addresses the core pain points of traditional scheduling: "difficulty in grasping the timing of reconstruction" and "the impact of migration on service continuity." The system calculates the expected matching degree gain through a topology matching optimization decision module, triggering reconstruction only when the gain exceeds a preset threshold. This ensures that every scheduling adjustment brings significant benefit improvement and avoids resource consumption caused by ineffective migration. At the migration execution level, connectivity analysis determines the set of subgraphs to be migrated, and cost-benefit trade-offs are made in conjunction with migration cost assessment. Seamless migration of subgraphs is achieved through streaming state migration—maintaining a consistent state through collaborative checkpoints, rebuilding communication connections on the target device, and resuming execution from the interruption point. The entire process does not interrupt task operation, ensuring service continuity. Simultaneously, the continuous learning and evolution module continuously optimizes the topology matching degree model parameters through feedback on the difference between actual and expected gains. This allows the system's scheduling decision-making capabilities to continuously evolve with operational experience, forming a self-improving intelligent scheduling closed loop. Under long-term operation, it can continuously optimize cluster scheduling effects and adapt to various complex and dynamic AI computing task scenarios. Attached Figure Description

[0016] Figure 1 This is a system architecture diagram of a multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters, provided in an embodiment of the present invention. The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments belong to some, but not all, embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “said” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms, and “multiple” generally includes at least two unless the context clearly indicates otherwise.

[0019] Depending on the context, the word "if" or "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0020] Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.

[0021] In practice, the server-side equipment deployed in a multi-dimensional scheduling and energy efficiency optimization system for AI computing power clusters may consist of one or more devices. This system can be implemented as: a business instance, a virtual machine, or hardware devices. For example, it can be implemented as a business instance deployed on one or more devices in a cloud node. Simply put, it can be understood as software deployed on a cloud node to provide a multi-dimensional scheduling and energy efficiency optimization system for AI computing power clusters to various user terminals. Alternatively, it can be implemented as a virtual machine deployed on one or more devices in a cloud node, with application software installed to manage various user terminals. Or, it can also be implemented as a server-side system composed of numerous identical or different types of hardware devices, with one or more hardware devices configured to provide a multi-dimensional scheduling and energy efficiency optimization system for AI computing power clusters to various user terminals.

[0022] In terms of implementation, a multi-dimensional scheduling and energy efficiency optimization system for AI computing power clusters and the user terminal are mutually compatible. Specifically, if the system is implemented as an application installed on a cloud service platform, the user terminal acts as a client establishing a communication connection with that application; or if the system is implemented as a website, the user terminal acts as a webpage; or if the system is implemented as a cloud service platform, the user terminal acts as a mini-program within an instant messaging application.

[0023] like Figure 1 The figure shown is a system architecture diagram of a multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters provided by an embodiment of the present invention.

[0024] The multi-dimensional scheduling and energy efficiency optimization system 100 for artificial intelligence computing power clusters described in this invention can be set up in a cloud server. In terms of implementation, it can be used as one or more service devices, or as an application installed in the cloud (e.g., a mobile service operator's server, server cluster, etc.), or it can be developed into a website. Depending on the implemented functions, the multi-dimensional scheduling and energy efficiency optimization system 100 for artificial intelligence computing power clusters may include a physical performance topology management module 101, a virtual computing topology extraction module 102, a topology matching optimization decision module 103, an online collaborative reconfiguration execution module 104, and a continuous learning and evolution module 105. The modules described in this invention can also be called units, referring to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, stored in the memory of the electronic device.

[0025] In this embodiment of the invention, in a multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters, each of the above modules can be implemented independently and can call other modules. Here, "calling" can be understood as a module connecting to multiple modules of another type and providing corresponding services to those connected modules. In this embodiment of the invention, the multi-dimensional scheduling and energy efficiency optimization system architecture for artificial intelligence computing power clusters can be adjusted in scope without modifying the program code by adding modules and directly calling them, achieving cluster-based horizontal expansion. This allows for quick and flexible expansion of the system. In practical applications, the above modules can be set in the same device or different devices, or in virtual devices, such as service instances in a cloud server.

[0026] The following describes, with reference to specific embodiments, the various components and specific workflows of a multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters: The physical performance topology management module is used to construct and update the global physical performance topology map in real time based on the static attribute data and dynamic monitoring data of all physical devices in the computing power cluster. In this embodiment of the invention, the step of constructing and updating a global physical performance topology map in real time based on static attribute data and dynamic monitoring data of all physical devices in the computing power cluster includes: Based on the dynamic monitoring data and the static attribute data, calculate the real-time comprehensive cost-effectiveness ratio score vector for each physical device; Based on the interconnection topology and real-time network monitoring data, the real-time effective communication cost of each physical link is calculated. Using physical devices as nodes, physical connections as edges, real-time comprehensive cost-effectiveness ratio score vectors as node weights, and real-time effective communication costs as edge weights, a weighted global physical performance topology graph is constructed and continuously updated.

[0027] It should be noted that static attribute data includes: device model and specifications, nominal computing power, memory capacity, and interconnection topology. Dynamic monitoring data includes: real-time calculation of load, power consumption, core temperature, network traffic, packet loss rate, and heat dissipation parameters of the device's microenvironment.

[0028] In this embodiment of the invention, calculating the real-time comprehensive cost-effectiveness ratio score vector for each physical device includes: The standardized instantaneous computing power score, unit power consumption computing power score, and heat dissipation efficiency score are calculated based on the comprehensive cost-effectiveness scoring equation set. The mathematical expression of the comprehensive cost-effectiveness scoring equation set is as follows: ; In the formula, P score To standardize the instantaneous computing power score, f current f is the current operating frequency of the device. base γ is the reference frequency of the device. throttle E is the frequency reduction factor. score P is used to score computing power per unit of power consumption. current For real-time power consumption of the device, C score To score heat dissipation efficiency, T max T is the upper limit of the chip junction temperature. junction For real-time core temperature, T coolant_in Intake air temperature; The standardized instantaneous computing power score, the unit power consumption computing power score, and the heat dissipation efficiency score are combined in an orderly manner to form a three-dimensional vector, resulting in a real-time comprehensive cost-effectiveness ratio score vector.

[0029] It should be noted that the standardized instantaneous computing power score represents the ratio of the actual computing power of the device under the current operating state to its nominal baseline capacity, reflecting whether the device is operating at the rated frequency. The frequency reduction factor is used to capture dynamic frequency reductions caused by overheat protection and other reasons. The closer the standardized instantaneous computing power score is to 1, the closer the computing power currently provided by the device is to its nominal optimal value; less than 1 indicates that its computing power output has been reduced due to load or temperature and other reasons. The unit power consumption computing power score represents the energy efficiency of the device, that is, the relative computing power level that can be supported by consuming one unit of power. The higher the value, the more effective computing power the device can provide under the same power consumption, and the better the energy efficiency ratio. The heat dissipation efficiency score represents the thermal safety margin or heat dissipation performance of the device under the current heat dissipation conditions. The closer the value is to 1, the more efficient the heat dissipation system is, the lower the core temperature is, and the higher the thermal reliability of the device. A value that is too low indicates an overheating risk, which may lead to performance degradation or shorten the lifespan of the device.

[0030] Furthermore, the upper limit of the chip junction temperature and the device reference frequency are derived from the device model and specifications in the static attribute data.

[0031] It should be noted that the frequency reduction factor is a real-time dynamic parameter between 0 and 1. When the real-time core temperature reaches or exceeds a certain threshold that triggers frequency reduction, the value of the frequency reduction factor will be adjusted according to the preset frequency reduction strategy: when the real-time core temperature is lower than the frequency reduction start temperature, the frequency reduction factor is 1.0; when the real-time core temperature is between the frequency reduction start temperature and the upper temperature limit, the frequency reduction factor is 0.6; when the real-time core temperature is higher than the upper temperature limit, the device may force frequency reduction or shutdown, at which time the frequency reduction factor is 0. The frequency reduction start temperature and the upper temperature limit are derived from static attribute data, which are hardware attributes of the device and are related to the performance of the device itself.

[0032] In this embodiment of the invention, calculating the real-time effective communication cost of each physical link includes: Based on the interconnection topology and the real-time network monitoring data, real-time network performance measurement data for each physical link is obtained; Based on the formula for calculating the effective bandwidth of a link, the bandwidth utilization and transmission error rate in the real-time network performance measurement data are processed to obtain the effective bandwidth of each physical link. Based on the communication delay comprehensive estimation algorithm, the one-way delay measurement value and packet loss rate in the real-time network performance measurement data are processed to obtain the communication delay of each physical link; Based on the shortest path principle of graph theory, with the effective bandwidth as the capacity constraint and the communication delay as the path cost, the maximum available bandwidth and minimum communication delay between any two physical device nodes in the global physical performance topology graph are calculated, and these are respectively used as the effective bandwidth and communication delay between the pair of nodes.

[0033] It should be noted that obtaining real-time network performance measurement data is based on the interconnection topology between physical devices. A network monitoring agent deployed in the cluster periodically sends probe packets to each physical link and collects responses, thereby synchronously acquiring raw performance metrics including bandwidth utilization, transmission error rate, one-way latency measurements, and packet loss rate. This operation aims to provide accurate and real-time link status input for subsequent calculations.

[0034] It should be noted that the mathematical expression for calculating the effective bandwidth of a link is as follows: In the formula, BW eff For the effective bandwidth of a single physical link, BW max U is the nominal maximum bandwidth of the link determined based on the interconnection topology. t To monitor bandwidth utilization in real time, E r This refers to the real-time monitored packet loss rate.

[0035] Furthermore, the essence of the formula for calculating the effective bandwidth of a link is to subtract the portion of invalid bandwidth caused by current traffic occupancy and transmission errors from the maximum theoretical bandwidth of the link, thereby calculating the actual bandwidth capacity of the link that can be used for effective data transmission at the current moment.

[0036] Furthermore, the nominal maximum bandwidth of the link determined based on the interconnection topology is derived from static attribute data.

[0037] It should be noted that the mathematical expression for the comprehensive communication delay estimation algorithm is as follows: In the formula, Lat is the communication delay of a single physical link, and RT is... T To detect the round-trip time of the packet, P LP The delay penalty term is calculated based on the packet loss rate. α and β are weighting coefficients set according to the characteristics of the network protocol stack, both of which are half. The delay penalty term is obtained by multiplying the real-time packet loss rate by the round-trip time of the probe packet.

[0038] It should be noted that the calculation process of the graph theory-based shortest path algorithm is as follows: The global physical performance topology graph is considered as a weighted undirected graph, where the weight of each edge represents the communication delay, and the effective bandwidth is used as the capacity attribute of that edge. When calculating the communication performance between any two physical device nodes A and B, firstly, the set of paths from A to B where the minimum effective bandwidth of each edge is greater than the required bandwidth threshold is found; then, the path with the minimum total communication delay is found in this set. Finally, the bottleneck bandwidth of this path is the effective bandwidth between nodes A and B, denoted as BW. eff The total delay of the path (A,B) is the communication delay between them, denoted as Lat(A,B). This step aggregates the performance of a single link into the end-to-end performance between any pair of nodes, providing a direct input for subsequent topology matching optimization.

[0039] The virtual computing topology extraction module is used to extract and update the dynamic virtual computing topology of artificial intelligence computing tasks in real time through lightweight runtime instrumentation technology. In this embodiment of the invention, the method for extracting and updating the dynamic virtual computing topology graph of the artificial intelligence computing task in real time through lightweight runtime instrumentation technology includes: A topology perceptron is mounted within the runtime framework of an artificial intelligence computing task to capture runtime information of the computation graph at a preset sampling frequency; Identify the key operator nodes in the computation graph and generate feature vectors representing computational density based on the shape and data type of their input and output tensors; Monitor the tensor transfer on the data dependency edges between operators, and record the amount of transmitted data and the frequency of communication triggering as the data flow characteristics of the edges; Based on the evolution of the task execution phase, the node set, edge set, and their associated characteristics are dynamically updated to form a dynamic virtual computing topology graph.

[0040] It should be noted that the preset sampling frequency is 1 second. By probing the operators that are currently being executed or have just finished executing, as well as the active data dependency edges between these operators, and by accessing the runtime context, the actual execution information of the operators in the current batch can be obtained, including the operator type, metadata of input and output tensors, and data transfer relationships between operators. This runtime-aware approach can capture information that is difficult to obtain from static analysis, such as dynamic graphs and conditional branch execution.

[0041] It should be noted that the identification of key operator nodes involves filtering operators that are computationally intensive, memory-intensive, or act as data flow bottlenecks. For each identified key operator, a three-dimensional feature vector representing the computational intensity is generated. Specifically, this includes: estimating the number of floating-point operations based on the operator type and the shape of the input tensor to obtain the estimated number of floating-point operations; calculating the number of bytes of memory occupied based on the volume and data type of the main output tensor to obtain the memory bandwidth requirement index, i.e., memory usage estimation; and assigning a normalized computational intensity weight based on the data type to obtain the normalized intensity coefficient, i.e., the data type weight. This weight is a weight that is pre-set according to the data type. For example, the weight of FP32 is set to 1.0, FP16 is set to 0.5, and INT8 is set to 0.25.

[0042] It should be noted that the data flow feature is used to quantify the communication behavior on data dependency edges between operators. It consists of a tuple of two core indicators: data transmission volume and communication frequency. These two indicators together determine the communication bandwidth requirement generated by the data flow edge, providing a direct input for subsequent evaluation of cross-device communication costs.

[0043] It should be noted that dynamic updates are achieved through continuous incremental exploration and phase awareness. In each sampling period, the system compares the newly captured operators and data dependencies with the existing topology graph, incrementally adds newly emerging nodes and edges, and updates the feature values ​​of existing nodes and edges. When some nodes or edges are no longer observed for several consecutive periods, they are marked and removed from the current active topology graph.

[0044] Furthermore, in each sampling period, the method for updating the feature values ​​of existing nodes and edges is as follows: After capturing the computation graph running information through the topology perceptron, for existing nodes and edges in the dynamic virtual computation topology graph, their associated feature values, the computation density feature vector of the node, and the data transmission volume and communication frequency of the edge will be directly updated to the latest values ​​obtained from the current sampling calculation.

[0045] It should be noted that the nodes of the dynamic virtual computing topology graph are computing operators, the edges are data dependencies between operators, the nodes have computing density feature vectors, and the edges have data flow features.

[0046] The topology matching optimization decision module is used to construct and solve a topology matching degree optimization model based on the global physical performance topology graph and the dynamic virtual computing topology graph, so as to obtain the optimal mapping target and expected matching degree gain for mapping virtual operators to physical devices. In this embodiment of the invention, the step of constructing and solving a topology matching degree optimization model based on the global physical performance topology graph and the dynamic virtual computing topology graph to obtain the optimal mapping target and expected matching degree gain for mapping virtual operators to physical devices includes: Define the mapping function from virtual operator nodes to physical device nodes; Construct an overall topology matching score function and an optimization model that aims to maximize the overall topology matching score function; Under resource capacity constraints and data dependency reachability constraints, the simulated annealing algorithm is used to solve the optimization model, obtain the optimal mapping target, and calculate the expected matching degree gain.

[0047] It should be noted that the mapping function is mathematically defined as a function derived from the set of virtual operator nodes V. v To the physical device node set V p The mapping relationship, i.e., M:V v →V p For each virtual operator v in the dynamic virtual computing topology graph, M(v) explicitly specifies the physical device node to which the operator is assigned and runs.

[0048] It should be noted that the method for solving the optimization model using the simulated annealing algorithm to obtain the optimal mapping target and calculate the expected matching degree gain is as follows: First, the current actual mapping is randomly generated and the temperature parameter is set. During the iteration process, a new neighborhood solution M is generated by randomly changing the mapping position of an operator. new Calculate the score difference ΔS between the old and new solutions. If ΔS > 0, accept the new solution; if ΔS ≤ 0, reject the new solution with probability. Accept, where T is the temperature, this mechanism helps to escape local optima; then reduce the temperature according to the cooling coefficient, repeat the above process until the temperature no longer changes; finally, output the highest score mapping encountered in the search process as the optimal mapping target, and the expected matching degree gain is obtained by calculating the difference between the overall topology matching degree score function of the optimal mapping target and the initial mapping score before optimization.

[0049] In this embodiment of the invention, the mathematical expression of the overall topology matching degree scoring function is as follows: ; In the formula, S(M) is the overall topology matching score function, S node (M) represents the total node matching score, S edge (M) represents the total edge matching cost, λ is the communication cost weight coefficient, and V v E is a set of virtual operators. v C is the set of dependency edges between virtual operators. v W is the computational density eigenvector of operator v. v (M(v)) is the real-time comprehensive cost-effectiveness ratio score vector of physical device node M(v), D e .v and D e .f represents the data transmission volume and communication frequency of the dependent edge e, respectively, and BW eff(M(v i ),M(v j )) and Lat(M(v i ),M(v j )) are respectively device node M(v i ) and M(v j The effective bandwidth and communication delay between ) are given by sim, where sim is the similarity calculation function, and M(v i ) and M(v j ) represent the virtual operator v i and v j The physical device node to which it is mapped, v i and v j It is a virtual operator; It should be noted that the mathematical expression for the similarity calculation function is as follows: In the formula, sim(C v W v (M(v))) represents the similarity, C v W is the computational density eigenvector of operator v. v (M(v)) is the real-time comprehensive cost-effectiveness ratio score vector of physical device node M(v).

[0050] It should be noted that the communication cost weighting coefficient is set to 0.5 by default.

[0051] The online collaborative reconstruction execution module is used to determine the set of operator subgraphs to be migrated based on the difference between the optimal mapping target and the current actual mapping when the expected matching degree gain exceeds the preset gain threshold, and migrate the set of operator subgraphs to be migrated to the target physical device in a streaming state migration mode to complete the computing power flow topology reconstruction. In this embodiment of the invention, the step of determining the set of operator subgraphs to be migrated based on the difference between the optimal mapping target and the current actual mapping when the expected matching degree gain exceeds a preset gain threshold includes: Based on the optimal mapping target and the current actual mapping, the mapping positions of all virtual operators in the dynamic virtual computing topology are compared to obtain a set of operators to be moved. Based on the data dependency edges of the dynamic virtual computing topology graph, connectivity analysis is performed on the operators in the set of operators to be moved to obtain at least one connected component, wherein each connected component constitutes a candidate migration subgraph. Based on the amount of state data to be migrated by the operators contained in each candidate migration subgraph, and the predicted communication overhead incurred in migrating the subgraph to the corresponding target device in the optimal mapping target under the current mapping, a migration cost evaluation is performed to obtain the migration cost value of each candidate migration subgraph. Based on the migration cost and the expected matching gain, a cost-benefit trade-off is performed to determine the final set of operator subgraphs to be migrated.

[0052] It should be noted that the preset gain threshold is 5% of the theoretical maximum score of the overall topology matching score function under ideal configuration.

[0053] Furthermore, subsequent processes are only triggered when the expected matching degree gain exceeds a preset gain threshold. The core purpose is to implement cost filtering. The online reconstruction operation itself consumes computing, I / O, and network resources and may temporarily affect task performance. If the expected benefits of optimization are very small, or even lower than the instantaneous cost and disturbance brought by the reconstruction operation itself, then this scheduling adjustment is inefficient or unnecessary. By setting a gain threshold, a benefit threshold is established to ensure that only those optimization schemes that are expected to bring significant overall performance improvement will be actually executed.

[0054] It should be noted that the mapping position comparison refers to comparing the optimal mapping objective function obtained by the topology matching optimization decision module with the actual mapping function that is currently in effect during system operation, one operator at a time. For each virtual operator in the dynamic virtual computing topology graph, it is checked whether its optimal mapping objective is the same as the physical device node pointed to by the current actual mapping. If they are different, the operator is marked as an operator that needs to change its position. All operators whose positions have changed are gathered together to form the set of operators to be moved.

[0055] It should be noted that connectivity analysis refers to using the operators in the set of operators to be moved as nodes and the actual data dependency edges between these operators in the dynamic virtual computing topology graph as connections, to search for connected components in graph theory, and to find those operator groups that are tightly coupled in the data flow and have direct or indirect data transmission relationships with each other. Each found connected component constitutes a candidate migration subgraph.

[0056] It should be noted that migration cost evaluation is a quantitative calculation process. For each candidate migration subgraph, its migration cost value is calculated using the following cost function: In the formula, J mig (G k G represents the migration cost. k Let SG represent the k-th candidate migration subgraph. k This represents the set of runtime state data required to migrate this subgraph, such as model parameters, intermediate activation values, optimizer states, etc., |SG k | Indicates the size of the state data, Comm(G) k M currTo maintain the current mapping of other operators M curr Without changing the subgraph G k Migrate from its current physical location to the optimal mapping target M opt The communication overhead generated by the cross-device data transfer required on the specified target device, where α and β are preset weighting coefficients, both of which are one-half.

[0057] In this embodiment of the invention, the step of migrating the set of operator subgraphs to be migrated to the target physical device using a streaming state migration method to complete the computing power flow topology reconstruction includes: Based on the optimal mapping target and the current actual mapping, operators whose mappings have changed are identified in the dynamic virtual computing topology graph, and connected subgraphs formed by these operators under the data dependency relationship in the dynamic virtual computing topology graph are extracted to obtain a set of operator subgraphs to be migrated. Coordinate the task processes corresponding to the operators involved in the set of operator subgraphs to be migrated, perform cooperative checkpoint operations, and obtain a consistent checkpoint state; On each target physical device specified by the optimal mapping target, a new computing process is started, the checkpoint state is loaded, and the communication connection between the process and the physical devices where other unmigrated operators in the cluster are located is rebuilt according to the optimal mapping target. Resume computation execution of all operators in the set of operator subgraphs to be migrated from the interruption point, so that the migrated operators can be seamlessly connected to the data stream, and update the global task mapping state to the current actual mapping equal to the optimal mapping target, thus completing the topology reconstruction of the computing power flow.

[0058] It should be noted that identifying operators whose mappings have changed means comparing the optimal mapping target obtained by the solution with the actual mapping that is currently in effect during system operation, one operator at a time. For each virtual operator in the dynamic virtual computing topology graph, if the comparison results are not equal, the operator is marked as having a mapping change. Extracting a connected subgraph refers to taking all operators whose mappings change as the initial node set, expanding along the data dependency edges between operators in the dynamic virtual computation topology graph, and finding the maximal connected components formed by all operators connected to each other through data dependency edges. Each such connected component is a connected subgraph.

[0059] It should be noted that the core of the collaborative checkpointing operation is to save a globally consistent runtime state snapshot for all computational processes involved in the operator subgraph to be migrated, under the coordination of the system. The specific operation is as follows: First, a preparation instruction is sent to the relevant processes. Each process then pauses further computation after completing the current smallest computational unit, and serializes and packages the key state data such as model parameters and optimizer intermediate variables in its memory, along with the current computation progress identifier. This process is carried out almost synchronously under the control of the system to ensure that all saved states logically correspond to the same computation completion point. The final packaged complete data packet is the "consistent checkpoint state".

[0060] It should be noted that starting a new computing process and rebuilding the communication connection is the execution phase of the migration. On the target physical device, the system initializes the corresponding computing kernel and memory space according to the operator type and configuration recorded in the checkpoint state, forming a brand new execution instance with the same state as the source. Subsequently, according to the globally optimal connection relationship described by the optimal mapping target, network communication links are established between the operators in this new process and all other unmigrated operators.

[0061] It should be noted that resuming execution from the interruption point is the final and switching phase of the migration. After successful loading on the target device and the communication connection is established, these new processes are instructed to continue execution from the "interruption point position" recorded in the checkpoint. For example, this means forward propagation starting from the first data sample of the next batch. In the inference pipeline, this means processing starting from the next request in the queue. Since the state is completely consistent and the data flow endpoint has been switched to the new position, this recovery process is "seamless" for the overall data input and output of the task. External observers cannot perceive the interruption. Finally, the global task mapping state is updated to ensure that the current actual mapping equals the optimal mapping target.

[0062] The continuous learning and evolution module is used to continuously learn and evolve the topology matching degree optimization model based on the actual performance and energy efficiency feedback data after topology reconstruction.

[0063] In this embodiment of the invention, the step of continuously learning and evolving the topology matching degree optimization model based on the actual performance and energy efficiency feedback data after topology reconstruction includes: The computational task of completing the topology reconstruction is monitored, and the actual performance and energy efficiency indicators during the stable operation period after reconstruction are collected to obtain the feedback dataset. Based on the feedback dataset, evaluate the actual matching gain obtained after executing the optimal mapping objective; Based on the difference between the actual matching degree gain and the expected matching degree gain, the parameters of the topology matching degree optimization model are adjusted to update the model.

[0064] It should be noted that collecting actual performance and energy efficiency indicators is a fundamental step in obtaining the basis for model evolution. The feedback dataset includes, but is not limited to, the task's iteration completion time after reconstruction, the total power consumption and energy efficiency ratio of the physical device cluster involved in the corresponding time period, and the actual latency and bandwidth utilization of cross-device communication. These data come from the system's continuous monitoring of the runtime status after the reconstruction operation is completed. They objectively record the comprehensive effect of this scheduling decision in the real physical environment.

[0065] It should be noted that evaluating the actual matching degree gain is a key process in transforming objective monitoring data into quantifiable evaluation. Specifically, the collected actual performance and energy efficiency data are substituted into the overall topology matching degree scoring function for calculation to obtain the actual score after adopting the optimal mapping target. Then, it is compared with the score calculated based on historical data before reconstruction, and the difference is the actual matching degree gain.

[0066] It should be noted that feedback adjustment and model updating based on discrepancies are the core of achieving a closed-loop technology and model self-evolution. The discrepancy refers to the deviation between the expected matching gain and the actual matching gain. This deviation reflects the degree of agreement between the model's previous predictions and the actual situation. Based on this deviation, the system adjusts the parameters within the model, such as adjusting the communication cost weight coefficient in the overall topology matching score function. If the actual gain is lower than expected, the communication cost weight coefficient is appropriately increased, making future optimizations more inclined towards mapping schemes that reduce communication overhead; conversely, the opposite adjustment is performed. Through this feedback mechanism, the model's parameters are continuously calibrated, making its predictions of future dynamic environments more accurate. As a result, the scheduling and optimization decision-making capabilities of the entire system continuously evolve with the accumulation of operational experience, forming a self-improving intelligent closed loop.

[0067] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0068] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters, characterized in that, The system includes a physical performance topology management module, a virtual computing topology extraction module, a topology matching optimization decision module, an online collaborative reconstruction execution module, and a continuous learning and evolution module, wherein: The physical performance topology management module is used to construct and update the global physical performance topology map in real time based on the static attribute data and dynamic monitoring data of all physical devices in the computing power cluster. The virtual computing topology extraction module is used to extract and update the dynamic virtual computing topology of artificial intelligence computing tasks in real time through lightweight runtime instrumentation technology. The topology matching optimization decision module is used to construct and solve a topology matching degree optimization model based on the global physical performance topology graph and the dynamic virtual computing topology graph, so as to obtain the optimal mapping target and expected matching degree gain for mapping virtual operators to physical devices. The online collaborative reconstruction execution module is used to determine the set of operator subgraphs to be migrated based on the difference between the optimal mapping target and the current actual mapping when the expected matching degree gain exceeds the preset gain threshold, and migrate the set of operator subgraphs to be migrated to the target physical device in a streaming state migration mode to complete the computing power flow topology reconstruction. The continuous learning and evolution module is used to continuously learn and evolve the topology matching degree optimization model based on the actual performance and energy efficiency feedback data after topology reconstruction.

2. The multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters as described in claim 1, characterized in that, The method for constructing and updating a global physical performance topology map in real time based on static attribute data and dynamic monitoring data of all physical devices in the computing power cluster includes: Based on the dynamic monitoring data and the static attribute data, calculate the real-time comprehensive cost-effectiveness ratio score vector for each physical device; Based on the interconnection topology and real-time network monitoring data, the real-time effective communication cost of each physical link is calculated. Using physical devices as nodes, physical connections as edges, real-time comprehensive cost-effectiveness ratio score vectors as node weights, and real-time effective communication costs as edge weights, a weighted global physical performance topology graph is constructed and continuously updated.

3. The multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters as described in claim 2, characterized in that, The calculation of the real-time comprehensive cost-effectiveness score vector for each physical device includes: The standardized instantaneous computing power score, unit power consumption computing power score, and heat dissipation efficiency score are calculated based on the comprehensive cost-effectiveness scoring equation set. The mathematical expression of the comprehensive cost-effectiveness scoring equation set is as follows: ; In the formula, P score To standardize the instantaneous computing power score, f current f is the current operating frequency of the device. base γ is the reference frequency of the device. throttle E is the frequency reduction factor. score P is used to score computing power per unit of power consumption. current For real-time power consumption of the device, C score To score heat dissipation efficiency, T max T is the upper limit of the chip junction temperature. junction For real-time core temperature, T coolant_in Intake air temperature; The standardized instantaneous computing power score, the unit power consumption computing power score, and the heat dissipation efficiency score are combined in an orderly manner to form a three-dimensional vector, resulting in a real-time comprehensive cost-effectiveness ratio score vector.

4. A multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters as described in claim 2, characterized in that, The calculation of the real-time effective communication cost for each physical link includes: Based on the interconnection topology and the real-time network monitoring data, real-time network performance measurement data for each physical link is obtained; Based on the formula for calculating the effective bandwidth of a link, the bandwidth utilization and transmission error rate in the real-time network performance measurement data are processed to obtain the effective bandwidth of each physical link. Based on the communication delay comprehensive estimation algorithm, the one-way delay measurement value and packet loss rate in the real-time network performance measurement data are processed to obtain the communication delay of each physical link; Based on the shortest path principle of graph theory, with the effective bandwidth as the capacity constraint and the communication delay as the path cost, the maximum available bandwidth and minimum communication delay between any two physical device nodes in the global physical performance topology graph are calculated, and these are respectively used as the effective bandwidth and communication delay between the pair of nodes.

5. A multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters as described in claim 1, characterized in that, The method for extracting and updating the dynamic virtual computing topology graph of artificial intelligence computing tasks in real time through lightweight runtime instrumentation technology includes: A topology perceptron is mounted within the runtime framework of an artificial intelligence computing task to capture runtime information of the computation graph at a preset sampling frequency; Identify the key operator nodes in the computation graph and generate feature vectors representing computational density based on the shape and data type of their input and output tensors; Monitor the tensor transfer on the data dependency edges between operators, and record the amount of transmitted data and the frequency of communication triggering as the data flow characteristics of the edges; Based on the evolution of the task execution phase, the node set, edge set, and their associated characteristics are dynamically updated to form a dynamic virtual computing topology graph.

6. The multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters as described in claim 1, characterized in that, The step of constructing and solving a topology matching degree optimization model based on the global physical performance topology graph and the dynamic virtual computing topology graph to obtain the optimal mapping target and expected matching degree gain for mapping virtual operators to physical devices includes: Define the mapping function from virtual operator nodes to physical device nodes; Construct an overall topology matching score function and an optimization model that aims to maximize the overall topology matching score function; Under resource capacity constraints and data dependency reachability constraints, the simulated annealing algorithm is used to solve the optimization model, obtain the optimal mapping target, and calculate the expected matching degree gain.

7. A multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters as described in claim 6, characterized in that, The mathematical expression for the overall topology matching score function is as follows: ; In the formula, S(M) is the overall topology matching score function, S node (M) represents the total node matching score, S edge (M) represents the total edge matching cost, λ is the communication cost weight coefficient, and V v E is a set of virtual operators. v C is the set of dependency edges between virtual operators. v W is the computational density eigenvector of operator v. v (M(v)) is the real-time comprehensive cost-effectiveness ratio score vector of physical device node M(v), D e .v and D e .f represents the data transmission volume and communication frequency of the dependent edge e, respectively, and BW eff (M(v i ),M(v j )) and Lat(M(v i ),M(v j )) are respectively device node M(v i ) and M(v j The effective bandwidth and communication delay between ) are given by sim, where sim is the similarity calculation function, and M(v i ) and M(v j ) represent the virtual operator v i and v j The physical device node to which it is mapped, v i and v j It is a virtual operator.

8. A multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters as described in claim 1, characterized in that, The method for determining the set of operator subgraphs to be migrated based on the difference between the optimal mapping target and the current actual mapping when the expected matching gain exceeds a preset gain threshold includes: Based on the optimal mapping target and the current actual mapping, the mapping positions of all virtual operators in the dynamic virtual computing topology are compared to obtain a set of operators to be moved. Based on the data dependency edges of the dynamic virtual computing topology graph, connectivity analysis is performed on the operators in the set of operators to be moved to obtain at least one connected component, wherein each connected component constitutes a candidate migration subgraph. Based on the amount of state data to be migrated by the operators contained in each candidate migration subgraph, and the predicted communication overhead incurred in migrating the subgraph to the corresponding target device in the optimal mapping target under the current mapping, a migration cost evaluation is performed to obtain the migration cost value of each candidate migration subgraph. Based on the migration cost and the expected matching gain, a cost-benefit trade-off is performed to determine the final set of operator subgraphs to be migrated.

9. A multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters as described in claim 1, characterized in that, The process of migrating the set of operator subgraphs to be migrated to the target physical device using a streaming state migration method to complete the computing power flow topology reconstruction includes: Based on the optimal mapping target and the current actual mapping, operators whose mappings have changed are identified in the dynamic virtual computing topology graph, and connected subgraphs formed by these operators under the data dependency relationship in the dynamic virtual computing topology graph are extracted to obtain a set of operator subgraphs to be migrated. Coordinate the task processes corresponding to the operators involved in the set of operator subgraphs to be migrated, perform cooperative checkpoint operations, and obtain a consistent checkpoint state; On each target physical device specified by the optimal mapping target, a new computing process is started, the checkpoint state is loaded, and the communication connection between the process and the physical devices where other unmigrated operators in the cluster are located is rebuilt according to the optimal mapping target. Resume computation execution of all operators in the set of operator subgraphs to be migrated from the interruption point, so that the migrated operators can be seamlessly connected to the data stream, and update the global task mapping state to the current actual mapping equal to the optimal mapping target, thus completing the topology reconstruction of the computing power flow.

10. A multi-dimensional scheduling and energy efficiency optimization system for artificial intelligence computing power clusters as described in claim 1, characterized in that, The method for continuously learning and evolving the topology matching degree optimization model based on actual performance and energy efficiency feedback data after topology reconstruction includes: The computational task of completing the topology reconstruction is monitored, and the actual performance and energy efficiency indicators during the stable operation period after reconstruction are collected to obtain the feedback dataset. Based on the feedback dataset, evaluate the actual matching gain obtained after executing the optimal mapping objective; Based on the difference between the actual matching degree gain and the expected matching degree gain, the parameters of the topology matching degree optimization model are adjusted to update the model.