Multivariate heterogeneous computing power scheduling method and scheduling platform for artificial intelligence public computing power center

By using the initial cluster mapping table to perform in-cluster fission and inter-cluster merging and adjustments in clusters in the multivariate heterogeneous environment of the artificial intelligence public computing power center, the problems of low task execution efficiency and delay fluctuations caused by inflexible computing power scheduling are solved, load balancing and safe isolation are achieved, and task processing efficiency and operation and maintenance efficiency are improved.

CN120540804AActive Publication Date: 2025-08-26GUANGZHOU ZIYUN JIXING INFORMATION TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510624797.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-26
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

In a multivariate heterogeneous environment, the existing technology has low task execution efficiency and large delay fluctuations in a multivariate heterogeneous environment, and cannot dynamically perceive and adjust computing power resources in real time.

Method used

By obtaining the initial cluster mapping table, fission adjustment within the cluster and merge adjustment between clusters, the calculation power stability parameters and exchange parameters are based on load balancing and safe isolation are achieved, and the adjustment trajectory map is generated for dynamic adjustment.

Benefits of technology

It significantly improves task processing efficiency, reduces delay fluctuations, solves the problem of uneven resource allocation, and improves operation and maintenance efficiency and fault recovery speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120540804A_ABST
    Figure CN120540804A_ABST
Patent Text Reader

Abstract

The invention discloses a multivariate heterogeneous computing power scheduling method and scheduling platform for an artificial intelligence public computing power center, and belongs to the technical field of electric digital data processing, and the method comprises the following steps: S1, obtaining a preset initial clustering mapping table, thereby obtaining an initial intra-cluster heterogeneous computing unit; s2, a cluster stability judgment result is obtained, if the cluster stability judgment result is qualified, intra-cluster fission adjustment is not executed, and S3 is directly executed, otherwise, intra-cluster fission adjustment is executed, and S3 is executed after adjustment; s3, obtaining a cluster safety judgment result, and performing corresponding inter-cluster adjustment based on the cluster safety judgment result; and S4, after inter-cluster adjustment, counting the computing power utilization rate of each heterogeneous computing unit, performing judgment to obtain a thermal migration execution judgment result, and performing corresponding thermal migration scheduling adjustment based on the thermal migration execution judgment result. The problems of low task execution efficiency and large time delay fluctuation in a multivariate heterogeneous environment caused by inflexible computing power scheduling in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic digital data processing technology, and in particular to a multi-heterogeneous computing power scheduling method and scheduling platform for an artificial intelligence public computing power center. Background Art

[0002] With the rapid development of artificial intelligence technology, the demand for computing power is growing exponentially. As a key infrastructure supporting AI research and development and application, artificial intelligence public computing power centers face huge challenges. The existing multi-heterogeneous computing power scheduling method constructs operator-level computing power tables and adopts a fixed two-layer clustering strategy to achieve the initial mapping between data processing models and computing power devices.

[0003] For example, the invention patent with announcement number CN116700934B discloses a method, apparatus, device, and storage medium for scheduling multi-heterogeneous computing power equipment, including: obtaining an operator-level computing power table corresponding to the multi-heterogeneous computing power equipment; the operator-level computing power table is used to characterize the performance of operators on different types of computing power equipment; according to the operator-level computing power table, operators related to the data processing model are deployed to the corresponding computing power equipment through a two-level clustering method to obtain a mapping relationship between the data processing model and the multi-heterogeneous equipment; and the multi-heterogeneous computing power equipment is scheduled based on the mapping relationship between the data processing model and the multi-heterogeneous equipment. This method can maximize the utilization of various underlying hardware resources, effectively improve the processing efficiency of high-throughput data processing models, and enhance data processing performance.

[0004] For example, the patent application with publication number CN119917267A discloses a multi-computing power heterogeneous computing service management system and method, including: a user terminal, a task scheduling module, a resource management module, and a hybrid computing power calculation module; it also discloses a multi-computing power heterogeneous computing service management method, which is implemented based on a multi-computing power heterogeneous computing service management system. A multi-computing power heterogeneous computing service management system and method are disclosed, which effectively integrates different computing resources and uses machine learning to obtain the optimal allocation plan for computing resources based on resource allocation requests, thereby reasonably allocating computing tasks and computing resources.

[0005] However, in the process of implementing the technical solutions of the invention in the embodiments of the present application, the present application found that the above technology has at least the following technical problems:

[0006] In the existing technology, by relying on static modeling and preset matching, it is impossible to dynamically perceive and adjust computing resources in real time when facing device load fluctuations, network instability or sudden changes in task types, and it is difficult to support stability control of task execution. Therefore, there are problems such as low task execution efficiency and large latency fluctuations in a multi-heterogeneous environment due to inflexible computing power scheduling. Summary of the Invention

[0007] The embodiments of the present application solve the problems of low task execution efficiency and large latency fluctuations in a multi-heterogeneous environment caused by inflexible computing power scheduling in the prior art by providing a multi-heterogeneous computing power scheduling method and scheduling platform for an artificial intelligence public computing power center, thereby achieving improved task processing efficiency.

[0008] This embodiment of the present application provides a multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center, comprising the following steps: S1. Upon receiving a computing power scheduling signal, the artificial intelligence public computing power center obtains a preset initial clustering mapping table, thereby obtaining heterogeneous computing units within the initial cluster; S2. Obtaining computing power stability parameters for the heterogeneous computing units within the initial cluster and obtaining a cluster stability determination result. If the cluster stability determination result is qualified, intra-cluster fission adjustment is not performed and the process directly proceeds to S3; otherwise, intra-cluster fission adjustment is performed and the process proceeds to S3 after the adjustment; S3. Obtaining computing power exchange parameters between adjacent initial clusters and obtaining a cluster safety determination result, and performing corresponding inter-cluster adjustments based on the cluster safety determination result; S4. After the inter-cluster adjustments, the computing power utilization of each heterogeneous computing unit is calculated and determined to obtain a hot migration execution determination result, and corresponding hot migration scheduling adjustments are performed based on the hot migration execution determination result.

[0009] Furthermore, it also includes: when a task processing exception signal is received, an adjustment trajectory map is generated and displayed. The specific method is: based on the final cluster mapping table, a final clustering topology baseline is established as the graph structure framework of the adjustment trajectory map; based on the computing power stability parameters of the heterogeneous computing units in the initial cluster and the computing power exchange parameters between adjacent initial clusters, a multidimensional feature vector is constructed to obtain an adjustment vector sequence with a time order as the event sequence of the adjustment trajectory map; based on the graph structure framework and event sequence of the adjustment trajectory map, an adjustment trajectory map is generated and displayed.

[0010] An embodiment of the present application provides a multi-heterogeneous computing power scheduling platform for an artificial intelligence public computing power center, including: an initial allocation module, a cluster fission analysis module, a cluster merging analysis module, and a hot migration analysis module; wherein the initial allocation module is used to obtain a preset initial clustering mapping table when the artificial intelligence public computing power center receives a computing power scheduling signal, thereby obtaining heterogeneous computing units within the initial cluster; the cluster fission analysis module is used to obtain computing power stability parameters of the heterogeneous computing units within the initial cluster and obtain a cluster stability determination result. If the cluster stability determination result is qualified, the intra-cluster fission adjustment is not performed and the cluster merging analysis module is directly entered; otherwise, the intra-cluster fission adjustment is performed and the cluster merging analysis module is entered after the adjustment; the cluster merging analysis module is used to obtain computing power exchange parameters between adjacent initial clusters, obtain a cluster safety determination result, and perform corresponding inter-cluster adjustments based on the cluster safety determination result; the hot migration analysis module is used to count the computing power utilization rate of each heterogeneous computing unit after the inter-cluster adjustment, and determine the hot migration execution determination result, and perform corresponding hot migration scheduling adjustments based on the hot migration execution determination result.

[0011] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0012] 1. The multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center provided by the present invention obtains an initial cluster mapping table, performs intra-cluster fission adjustment and inter-cluster merging adjustment, and achieves load balancing and security isolation of multi-heterogeneous computing units based on computing power stability parameters and computing power exchange parameters, thereby significantly improving task processing efficiency and reducing latency fluctuations, thereby achieving improved task processing efficiency and effectively solving the problems of low task execution efficiency and large latency fluctuations in a multi-heterogeneous environment caused by inflexible computing power scheduling in the prior art.

[0013] 2. The present invention calculates the operator resource consumption based on the state parameters of each operator and maps the task scheduling ratio according to the interval of the computing power utilization ratio of the low-load sub-cluster, thereby accurately migrating high-resource-consuming operators to the low-load sub-cluster, thereby solving the problem of uneven resource allocation between heterogeneous computing units and reducing the task processing failure rate.

[0014] 3. The present invention establishes a final cluster topology baseline based on the final cluster mapping table, and constructs a multi-dimensional feature vector based on the initial cluster computing power stability parameters and the adjacent initial cluster computing power exchange parameters to generate an adjustment trajectory map, thereby realizing the full-link tracing of abnormal events and visualization of the dynamic adjustment process, thereby improving operation and maintenance efficiency and fault recovery speed.

[0015] 4. The present invention obtains the cluster merging processing value of each cluster to be merged by analyzing the network delay and security level matching degree of each cluster to be merged, thereby arranging the cluster merging processing values ​​of each cluster to be merged in descending order, and cluster merging two adjacent clusters to be merged in sequence according to the arrangement order to obtain an adjusted clustering mapping table, which is convenient for generating an adjustment trajectory map. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A flowchart of a multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center provided in an embodiment of the present application;

[0017] Figure 2 A macroscopic framework diagram of a multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center provided in an embodiment of the present application;

[0018] Figure 3 A detailed flowchart of the multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center provided in an embodiment of the present application;

[0019] Figure 4 A cluster topology diagram of the multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center provided in an embodiment of the present application;

[0020] Figure 5 A schematic diagram of a cloud platform for a multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center provided in an embodiment of the present application;

[0021] Figure 6 A schematic diagram of the structure of a multi-heterogeneous computing power scheduling platform for an artificial intelligence public computing power center provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] The embodiments of the present application solve the problems of low task execution efficiency and large latency fluctuations in a multi-heterogeneous environment caused by inflexible computing power scheduling in the prior art by providing a multi-heterogeneous computing power scheduling method and scheduling platform for an artificial intelligence public computing power center, thereby achieving an improvement in task processing efficiency.

[0023] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0024] like Figure 1As shown, it is a flow chart of a multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center provided by an embodiment of the present application, and the method includes the following steps: S1. When the artificial intelligence public computing power center receives a computing power scheduling signal, it obtains a preset initial clustering mapping table to obtain heterogeneous computing units in the initial cluster; S2. Obtain computing power stability parameters of the heterogeneous computing units in the initial cluster to obtain a cluster stability judgment result. If the cluster stability judgment result is qualified, the intra-cluster fission adjustment is not performed and directly enters S3; otherwise, the intra-cluster fission adjustment is performed and enters S3 after the adjustment; S3. Obtain computing power exchange parameters between adjacent initial clusters to obtain a cluster security judgment result, and perform corresponding inter-cluster adjustments based on the cluster security judgment result; S4. After the inter-cluster adjustment, the computing power utilization rate of each heterogeneous computing unit is counted, and a judgment is performed to obtain a hot migration execution judgment result, and corresponding hot migration scheduling adjustments are performed based on the hot migration execution judgment result.

[0025] In this embodiment, it should be noted that the initial clustering mapping table is obtained from the database. The initial clustering mapping table records the initial grouping of various heterogeneous computing units (such as GPUs, CPUs, NPUs, etc.) in the AI ​​public computing power center. It lists in detail which specific heterogeneous computing units (such as GPUs, CPUs, NPUs, etc.) each cluster contains, as well as information such as the basic attributes and locations of these units.

[0026] The initial clustering mapping table provides the initial distribution of all heterogeneous computing units within the computing center, quickly locating the cluster to which each unit belongs, and providing a clear understanding of the currently available computing resources and their distribution. As the starting point for task scheduling, the initial clustering mapping table provides an initial reference framework for subsequent task allocation, allowing for a rational start in task allocation.

[0027] The initial cluster mapping table is updated and the final cluster mapping table is generated because the initial cluster mapping table only reflects the cluster structure state at the start of task scheduling. However, after tasks undergo multiple steps such as fission, merging, and hot migration, the initial table is no longer true. Therefore, by constructing the final cluster mapping table, we can ensure the consistency of the cluster mapping table and the task scheduling system. The final cluster mapping table records the actual devices, resource indicators, topological connections, and task distribution in each cluster. The real-time update process provides a basis for intra-cluster fission adjustments, corresponding inter-cluster adjustments, and corresponding hot migration scheduling adjustments. If the initial cluster table is still used, it will cause abnormal alarms such as false alarms or missed alarms.

[0028] Furthermore, computing power stability parameters of the heterogeneous computing units in the initial cluster are obtained to obtain a cluster stability judgment result. The specific method is: obtaining computing power stability parameters of the heterogeneous computing units in the initial cluster within a preset time period, the computing power stability parameters including computing power utilization standard deviation and delay jitter duration standard deviation; obtaining the computing power utilization standard deviation benchmark value and delay jitter duration standard deviation benchmark value preset in the database, and comparing them with the computing power stability parameters to obtain a cluster stability judgment result; if the computing power utilization standard deviation is less than the computing power utilization standard deviation benchmark value and the delay jitter duration standard deviation is less than the delay jitter duration standard deviation benchmark value, the cluster stability judgment result is qualified, otherwise the cluster stability judgment result is unqualified.

[0029] In this embodiment, it should be noted that the standard deviation of computing power utilization refers to the standard deviation of the computing power utilization of each heterogeneous computing unit in the initial cluster, which can be obtained by obtaining the computing power utilization of each heterogeneous computing unit and then taking the standard deviation. The standard deviation of delay jitter duration refers to the standard deviation of the delay jitter duration of each heterogeneous computing unit in the initial cluster, which can be obtained by obtaining the delay jitter duration of each heterogeneous computing unit (the delay jitter duration can be obtained by network traffic analysis tools such as Wireshark) and then taking the standard deviation.

[0030] By analyzing the standard deviation of computing power utilization to measure the load balancing of heterogeneous computing units in the initial cluster, and using the standard deviation of delay jitter to reflect the volatility of the communication links of heterogeneous computing units in the initial cluster, the dual criteria are combined to avoid misjudgment and effectively improve the accuracy of cluster status analysis. When the standard deviation of computing power utilization is less than the benchmark value of computing power utilization standard deviation and the standard deviation of delay jitter duration is less than the benchmark value of delay jitter duration standard deviation, the cluster stability judgment result is qualified and no cluster fission adjustment is performed, avoiding system fluctuations caused by frequent clustering and improving scheduling efficiency and resource utilization.

[0031] The standard deviation of computing power utilization indicates the load balance of heterogeneous computing units within the initial cluster. If the standard deviation of computing power utilization is too large, it means that some nodes are overloaded, some nodes are idle, and resource allocation is unbalanced. If the standard deviation of computing power utilization is small, it indicates balanced resource allocation and stable operation.

[0032] The standard deviation of delay jitter duration represents the stability of the initial intra-cluster communication link. The delay jitter duration reflects the delay fluctuations that occur during the communication transmission process of the task. The larger the standard deviation of delay jitter duration, the more unstable the data flow is and the more likely it is to cause congestion or delay.

[0033] By introducing the standard deviation of computing power utilization and latency jitter duration as criteria for cluster stability, we can quantitatively monitor and perform threshold-driven analysis of resource fluctuations within the cluster, effectively identifying the discrete nature of the operating states of heterogeneous computing units within the cluster and network stability. Compared to traditional methods that rely solely on average values ​​or single parameters, this improves the accuracy of cluster operating status assessments, thereby avoiding unnecessary resource reconfiguration operations.

[0034] Furthermore, intra-cluster fission adjustment is performed, and the specific method is as follows: obtaining the computing power utilization of each heterogeneous computing unit in the initial cluster; obtaining the computing power utilization splitting threshold preset in the database, and comparing it with the computing power utilization of each heterogeneous computing unit in the initial cluster, thereby adjusting each heterogeneous computing unit with a computing power utilization above the computing power utilization splitting threshold from the initial cluster to a high-load sub-cluster, and adjusting each heterogeneous computing unit with a computing power utilization less than the computing power utilization splitting threshold from the initial cluster to a low-load sub-cluster; obtaining each operator state parameter under each heterogeneous computing unit within a preset time period after the intra-cluster fission adjustment, analyzing to obtain each operator state value, thereby matching to obtain each operator resource consumption; obtaining the computing power utilization of each heterogeneous computing unit in the low-load sub-cluster and the computing power utilization of each heterogeneous computing unit in the high-load sub-cluster, analyzing to obtain the computing power utilization ratio of the low-load sub-cluster, thereby performing task migration and operator migration, and marking the high-load sub-cluster and the low-load sub-cluster after the intra-cluster fission adjustment as fission adjustment clusters.

[0035] In this embodiment, if Figure 2 As shown, this is a macro-framework diagram of the multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center provided in an embodiment of the present application. It can be analyzed that when the artificial intelligence public computing power platform receives the computing power scheduling signal, it reads the heterogeneous computing unit composition of each cluster from the initial cluster mapping table, and judges whether the current cluster meets the stability requirements, that is, whether the cluster is qualified. If so (qualified), intra-cluster fission adjustment will not be performed. If not (unqualified), intra-cluster fission adjustment will be performed to obtain a cluster safety judgment result, thereby performing corresponding cluster component adjustments, and then performing corresponding hot migration scheduling adjustments to obtain the final cluster mapping table.

[0036] By obtaining the computing power utilization of each heterogeneous computing unit and splitting the initial cluster into high-load sub-clusters and low-load sub-clusters based on preset thresholds, the computing power scheduling platform has the ability to quickly respond to sudden computing pressure and avoid computing bottlenecks. By introducing operator state parameter analysis and resource consumption matching, operator-level migration judgment is realized, and key operators with high resource usage can be prioritized for migration, significantly improving the refinement and rationality of computing power utilization and reducing overall load imbalance. After dividing the low-load sub-clusters, the reasonable proportion of migration tasks that can be carried out is determined based on their computing power utilization ratio to ensure that the scale of migration tasks is controlled and avoid instability caused by sudden task migration.

[0037] The computing power utilization ratio of the low-load sub-cluster is obtained by the following specific steps: adding the total computing power utilization value of each heterogeneous computing unit in the low-load sub-cluster and the total computing power utilization value of each heterogeneous computing unit in the high-load sub-cluster to obtain the total computing power utilization rate; then dividing the total computing power utilization value of each heterogeneous computing unit in the low-load sub-cluster by the total computing power utilization rate to obtain the computing power utilization ratio of the low-load sub-cluster.

[0038] Get the computing power utilization ratio of the low-load sub-cluster by:

[0039]

[0040] Where ZB represents the utilization ratio of the low-load sub-cluster computing power, k represents the number of heterogeneous computing units in the low-load sub-cluster, k = 1, 2, ..., k max , k max represents the total number of heterogeneous computing units in the low-load sub-cluster, L k represents the computing power utilization of the kth heterogeneous computing unit in the low-load sub-cluster, and m represents the number of the heterogeneous computing unit in the high-load sub-cluster, m = 1, 2, ..., m max , m max represents the total number of heterogeneous computing units in the high-load sub-cluster, Y m Indicates the computing power utilization of the mth heterogeneous computing unit in the high-load sub-cluster.

[0041] Furthermore, the resource consumption of each operator is obtained. The specific method is as follows: the state parameters of each operator under each heterogeneous computing unit include the maximum computing power utilization, maximum video memory occupancy, historical failure rate, maximum latency, and maximum packet loss rate of each operator under each heterogeneous computing unit; a state reference set preset in the database is obtained, and the difference degree of each operator state parameter under each heterogeneous computing unit is compared based on the state reference set preset in the database. After obtaining the difference comparison result, a corresponding weighting factor is introduced to quantify the difference comparison result to obtain the state value of each operator under each heterogeneous computing unit; each operator state interval preset in the database and the reference resource consumption corresponding to each operator state interval are obtained, and compared with the state value of each operator under each heterogeneous computing unit. If the operator state value of an operator is within a certain operator state interval, the reference resource consumption corresponding to the interval is obtained as the resource consumption of the operator, thereby obtaining the resource consumption of each operator under each heterogeneous computing unit; the state reference set includes the ideal computing power utilization value, the ideal video memory occupancy value, the allowable historical failure rate value, the allowable latency value, and the allowable packet loss rate value.

[0042] like Figure 3 As shown, Figure 3A detailed flow chart of a multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center provided in an embodiment of the present application, wherein, based on the cluster safety judgment result, it is determined whether it is qualified. If it is (qualified), the inter-cluster merging adjustment is not performed, and the initial clustering mapping table is marked as the cluster adjustment mapping table. If not (unqualified), the inter-cluster merging adjustment is performed to obtain the cluster to be merged, and the processing value of the cluster to be merged is analyzed to perform cluster merging processing, and then the computing power utilization rate of each heterogeneous computing unit is counted to obtain the hot migration judgment result, and it is determined whether to migrate. If it is (hot migration adjustment is performed), adjustment is performed to obtain the final cluster mapping table. Otherwise, hot migration adjustment is not performed, and the cluster adjustment mapping table is marked as the final cluster mapping table.

[0043] In this embodiment, computing power utilization can be measured through NVIDIA-SMI (for GPU), Ascend Toolkit (for Ascend NPU), Cambricon Metrics (for Cambricon), etc., video memory occupancy can be obtained through NVIDIA-SMI or a dedicated driver interface API, historical failure rate can be obtained through a log analysis system (such as ELK), latency can be obtained by capturing the running time of each operator call, and packet loss rate can be detected by network detection tools (such as Ping).

[0044] Get the state value of each operator under each heterogeneous computing unit. The specific method is:

[0045]

[0046] Where ZT i,j Indicates the jth operator state value under the i-th heterogeneous computing unit, i represents the number of the heterogeneous computing unit, i = 1, 2, ..., i max ,i max Indicates the total number of heterogeneous computing units, j indicates the operator number under a heterogeneous computing unit, j = 1, 2, ..., j max ,j max Indicates the total number of operators under a heterogeneous computing unit, ZY i,j represents the maximum computing power utilization of the jth operator under the i-th heterogeneous computing unit, σZY represents the ideal computing power utilization, and ZC i,j represents the maximum memory occupancy of the jth operator under the i-th heterogeneous computing unit, σZC represents the ideal value of the memory occupancy, and ZG i,j represents the historical failure rate of the jth operator under the i-th heterogeneous computing unit, σZG represents the allowable value of the historical failure rate, and ZS i,j represents the maximum delay of the jth operator under the i-th heterogeneous computing unit, σZS represents the allowed delay value, and ZD i,jrepresents the maximum packet loss rate of the jth operator under the i-th heterogeneous computing unit, σZD represents the allowed value of the packet loss rate, α1 represents the weighting factor of the maximum computing power utilization, α2 represents the weighting factor of the maximum memory occupancy, α3 represents the weighting factor of the historical failure rate, α4 represents the weighting factor of the maximum latency, and α5 represents the weighting factor of the maximum packet loss rate.

[0047] The maximum computing power utilization weighting factor, the maximum video memory occupancy weighting factor, the historical failure rate weighting factor, the maximum latency weighting factor, and the maximum packet loss rate weighting factor can be obtained from a database. For example, the maximum computing power utilization weighting factor can be obtained by obtaining the historical maximum computing power utilization values ​​stored in the database, as well as the maximum computing power utilization weighting factors corresponding to the historical maximum computing power utilization values, thereby constructing a maximum computing power utilization mapping set, wherein the mapping set has a one-to-one or many-to-one correspondence. The maximum computing power utilization weighting factor can be obtained by inputting the required maximum computing power utilization data into the maximum computing power utilization mapping set. The other weighting factors are obtained in the same manner as the maximum computing power utilization weighting factor and can all be matched in the corresponding mapping sets. The maximum video memory occupancy weighting factor corresponds to the maximum video memory occupancy mapping set, the historical failure rate weighting factor corresponds to the historical failure rate mapping set, the maximum latency weighting factor corresponds to the maximum latency mapping set, and the maximum packet loss rate weighting factor corresponds to the maximum packet loss rate mapping set.

[0048] By analyzing the state parameters of each operator within each heterogeneous computing unit, including the maximum computing power utilization, maximum memory occupancy, historical failure rate, maximum latency, and maximum packet loss rate, the state values ​​of each operator within each heterogeneous computing unit are derived. This takes into account the interplay between these parameters. For example, computing power utilization reflects the demand for computing resources during operator execution and is a direct indicator of performance load. Memory occupancy indicates the pressure an operator places on storage resources. When memory resources are limited, the task parallelization capability of high-computing power devices is inhibited, indirectly affecting the improvement of computing power utilization. The historical failure rate reflects whether an operator has been constrained by device stability or data anomalies during past operations. An excessively high historical failure rate often indicates resource scheduling mismatches. Furthermore, maximum latency and packet loss rates are often caused by network quality, node congestion, or inter-device communication load. These phenomena can in turn cause operators to repeatedly retransmit or wait during execution, further increasing their computing power utilization and memory occupancy.

[0049] By quantitatively evaluating multiple state parameters of operators running under each heterogeneous computing unit, and combining the preset state reference set in the database to calculate the resource consumption of each operator under each heterogeneous computing unit, we can achieve refined control of operator scheduling at the task granularity, accurately identify high-consumption operators and optimize their migration paths, thereby effectively improving the utilization efficiency of computing resources and the stability of operation, and reducing computing bottlenecks and system jitter caused by resource mismatch.

[0050] Furthermore, task migration and operator migration are performed. The specific method is as follows: obtain the low computing power utilization ratio intervals preset in the database and the reference task scheduling ratio corresponding to each low computing power utilization ratio interval, and compare them with the computing power utilization ratio of the low-load sub-cluster. If the computing power utilization ratio of the low-load sub-cluster is within a preset low computing power utilization ratio interval, obtain the reference task scheduling ratio corresponding to the interval as the task scheduling ratio; according to the task scheduling ratio, call the corresponding task from the high-load sub-cluster at the task scheduling ratio, and mark the task as a task to be migrated; obtain the preset low computing power utilization ratio intervals in the database and compare them with the reference task scheduling ratio corresponding to the low-load sub-cluster. The operator resource consumption threshold is set and compared with the resource consumption of each operator in the high-load sub-cluster. The operators in the high-load sub-cluster whose operator resource consumption is above the operator resource consumption threshold are marked as operators to be migrated; the tasks to be migrated and the operators to be migrated are matched with heterogeneous computing units with the same computing architecture in the low-load sub-cluster respectively, among which the tasks to be migrated are preferentially assigned to the heterogeneous computing units with the lowest computing power utilization in the low-load sub-cluster for execution, and the operators to be migrated are sequentially assigned to heterogeneous computing units with corresponding load capacity according to their resource consumption, thereby realizing multi-granularity scheduling and completing task migration and operator migration.

[0051] In this embodiment, it should be noted that the tasks to be migrated and the operators to be migrated should be based on the compatibility principle of giving priority to computing units of the same type, that is, heterogeneous computing units with the same computing architecture (for example, if the heterogeneous computing unit of the task to be migrated is a CPU, the task should be migrated to the computing unit whose heterogeneous computing unit is a CPU), ensuring that the operator does not need to be recompiled or converted at the container level; at the same time, sorting and selection are performed based on real-time computing power utilization, so as to maximize the pressure on high-load devices and improve the overall system throughput efficiency.

[0052] By setting the correspondence between the computing power utilization ratio of the low-load sub-cluster and the reference task scheduling ratio, and combining the operator resource consumption threshold to screen the migration objects, and matching them to heterogeneous computing units with consistent architecture and lowest utilization, accurate offloading and migration scheduling of tasks and operators can be achieved, ensuring the rationality of the overall computing power distribution and the stability of scheduling after migration.

[0053] Furthermore, computing power exchange parameters between adjacent initial clusters are obtained to obtain a cluster security determination result. The specific method is as follows: obtaining computing power exchange parameters between adjacent initial clusters within a preset time period, the computing power exchange parameters between adjacent initial clusters including the average computing power exchange delay and security level matching degree between adjacent initial clusters; obtaining the computing power exchange average delay threshold between adjacent initial clusters and the security level matching degree threshold between adjacent initial clusters preset in the database; comparing the computing power exchange average delay between adjacent initial clusters with the computing power exchange average delay threshold between adjacent initial clusters, and comparing the security level matching degree between adjacent initial clusters with the security level matching degree threshold between adjacent initial clusters to obtain a cluster security determination result; if the computing power exchange average delay between adjacent initial clusters is less than the computing power exchange average delay threshold between adjacent initial clusters and the security level matching degree between adjacent initial clusters is greater than the security level matching degree threshold between adjacent initial clusters, then the cluster security determination result is qualified; otherwise, the cluster security determination result is unqualified.

[0054] In this embodiment, it should be noted that adjacent initial clusters refer to two initial clustering units that have a direct communication relationship in the physical network connection or logical topology path in the artificial intelligence public computing power center, that is, two initial clusters with a direct data exchange path are called adjacent initial clusters.

[0055] It should be noted that security level matching refers to the degree of consistency in security levels between two or more computing units or clusters. The security level matching between adjacent initial clusters can be determined based on the "Guidelines for the Classification of Security Protection Levels for Computer Information Systems." The average latency of computing power exchange between adjacent initial clusters can be measured using RFC2544 testing equipment such as the Xinertai network tester. The matching value range is [0% to 100%].

[0056] Furthermore, corresponding inter-cluster adjustments are performed based on the cluster safety determination results. The specific method is as follows: based on the cluster safety determination results, analysis is performed. If the cluster safety determination result is qualified, the inter-cluster merging adjustment is not performed, and the initial clustering mapping table is marked as the cluster adjustment mapping table; if the cluster safety determination result is unqualified, the inter-cluster merging adjustment is performed, and the cluster performing the inter-cluster merging adjustment is marked as the cluster to be merged; the network delay and security level matching degree of each cluster to be merged are obtained; the network delay allowable value and the security level matching lower limit value preset in the database are obtained, and the difference degree analysis is performed with the network delay and security level matching degree of each cluster to be merged respectively. After obtaining the difference analysis results, the corresponding empowerment factor is introduced to quantify the difference analysis results to obtain the cluster merging processing value of each cluster to be merged; the cluster merging processing values ​​of each cluster to be merged are arranged in descending order, and the two adjacent clusters to be merged are cluster-merged in sequence according to the arrangement order, and after the cluster merging process, the initial clustering mapping table is updated to obtain the adjusted clustering mapping table, and it is marked as the cluster adjustment mapping table.

[0057] In this embodiment, it should be noted that the cluster merging processing values ​​of each cluster to be merged are arranged in descending order, and the two adjacent clusters to be merged are cluster merged in sequence according to the arrangement order. The optimal merging pair that best suits the current overall state (that is, two adjacent clusters to be merged in sequence) can be found from these clusters to be merged. For example, if cluster A is adjacent to cluster B, and cluster B is adjacent to cluster C, if cluster A, cluster B and cluster C are all clusters to be merged, if cluster A is merged with cluster B, and cluster B is merged with cluster C, cluster merging overlap and duplication will occur. Therefore, it is necessary to evaluate the cluster merging processing value of each cluster to be merged and arrange the cluster merging processing value of each cluster to be merged in descending order, and cluster merge the two adjacent clusters to be merged in sequence according to the arrangement order. The merging process is performed pair by pair from the sorted adjacent clusters to be merged, which can avoid overlap and duplication and achieve optimal optimization.

[0058] like Figure 4 As shown, Figure 4The cluster topology diagram of the multi-heterogeneous computing power scheduling method for the artificial intelligence public computing power center provided in the embodiment of the present application shows that there are various types of heterogeneous computing resources. The computing nodes marked above include: GPU nodes (such as: A800 series, H800 series), heterogeneous AI nodes (such as: Ascend 910), general computing nodes (such as: X86 servers, JFMHG-01, etc.), reflecting the multi-heterogeneous characteristics. The switching network structure is clearly layered: the network is divided into three layers: the core switch layer (#01 and #02), the middle layer is 25G switches (#01 to #06) and 1G switches (#01 to #03), the upper layer connects the computing nodes, and there is a high-speed interconnection network (IB switch) that can provide low-latency, high-bandwidth interconnection channels, which is suitable for scenarios such as deep learning training that are sensitive to communication performance. The cluster topology diagram can be used as the physical basis for constructing the initial cluster mapping table and the final cluster mapping table, reflecting the actual existence of multi-channel computing power exchange paths. The existence of multi-path forwarding in the network can support the collection of computing power exchange parameters (such as latency, matching degree, etc.) between computing nodes for subsequent cluster security judgment.

[0059] like Figure 5 As shown, it is a schematic diagram of the cloud platform for the multi-heterogeneous computing power scheduling method for the artificial intelligence public computing power center provided in the embodiment of the present application. From the image, it can be seen that the service entrance, platform service content, and OS service pool of the platform, etc. of the platform realize the security protection of the platform.

[0060] It should be noted that the cluster merging processing values ​​of each cluster to be merged are arranged in descending order, and the two adjacent clusters to be merged are cluster merged in the order of arrangement. It is not the original physical "topological adjacency", but the two adjacent cluster pairs after sorting are merged. The purpose is to give priority to merging the cluster pairs with the highest value and avoid wasting resources by merging low-value clusters.

[0061] Get the cluster merging processing value of each cluster to be merged. The specific method is:

[0062]

[0063] Where, HL t Indicates the cluster merging processing value of the tth cluster to be merged, t represents the number of the cluster to be merged, t=1,2,...,t max , t max Indicates the total number of clusters to be merged, HY t represents the network delay of the tth cluster to be merged, HA t represents the security level matching degree of the tth cluster to be merged, σHY represents the allowed value of network delay, σHA represents the lower limit of security level matching degree, β1 represents the network delay weighting factor, and β2 represents the security level matching weighting factor.

[0064] The network delay weighting factor and the security level matching weighting factor can be obtained from the database. For example, the network delay weighting factor can be obtained by obtaining the historical network delay stored in the database, and the network delay weighting factor corresponding to the historical network delay, thereby constructing a network delay mapping set, wherein there is a one-to-one or many-to-one correspondence in the mapping set. The network delay weighting factor can be obtained by inputting the required network delay data into the network delay mapping set. The security level matching weighting factor is obtained in the same way as the network delay weighting factor, and can also be obtained by matching in the corresponding mapping set, wherein the security level matching weighting factor corresponds to the security level matching mapping set.

[0065] Furthermore, after the inter-cluster adjustment, the computing power utilization of each heterogeneous computing unit is counted, and a judgment is made to obtain the hot migration execution judgment result. The specific method is: obtain the computing power utilization of each heterogeneous computing unit at the first sampling moment, the computing power utilization at the final sampling moment, and the computing power utilization change value of each heterogeneous computing unit within the preset time period after the adjustment; obtain the computing power utilization change threshold preset in the database, and compare it with the computing power utilization change value of each heterogeneous computing unit. If there is a computing power utilization change value of a certain heterogeneous computing unit that is above the computing power utilization change threshold, and the computing power utilization of the heterogeneous computing unit at the final sampling moment is greater than the computing power utilization of the heterogeneous computing unit at the first sampling moment, the computing power utilization of the heterogeneous computing unit is compared with the computing power utilization change value of the heterogeneous computing unit. If the computing power utilization rate at a sampling moment is less than 1%, the hot migration execution judgment result of the heterogeneous computing unit is migration, and hot migration scheduling adjustment is performed. After the hot migration scheduling adjustment, the cluster adjustment mapping table is updated to obtain a final cluster mapping table. Otherwise, the hot migration execution judgment result is not migration, and hot migration scheduling adjustment is not performed. At the same time, the cluster adjustment mapping table is marked as the final cluster mapping table; the hot migration scheduling adjustment is performed, and the specific method is as follows: obtain the heterogeneous computing unit that needs to be hot migration scheduling adjustment, and mark it as the heterogeneous computing unit to be migrated; obtain the spare heterogeneous computing unit preset in the database, and migrate the task content of the heterogeneous computing unit to be migrated to the spare heterogeneous computing unit.

[0066] In this embodiment, the cluster adjustment mapping table is updated. The specific method is: after the hot migration scheduling adjustment is completed, a new computing power mapping relationship is obtained (including the heterogeneous computing unit information of task migration in and out), and the latest status data such as the computing power utilization rate and network topology structure of each computing unit are obtained. According to the task distribution after hot migration, the mapping relationship between heterogeneous computing units and tasks is updated. According to the new mapping relationship, the network connection configuration information between the computing units is updated to ensure that the data transmission path and communication relationship are consistent with the new mapping relationship. The metadata information such as the timestamp of this hot migration scheduling adjustment, the computing units involved, and the migrated task list are recorded. All the above information is sorted and replaced with the information in the original mapping table to obtain the final cluster mapping table, and it is saved in the database.

[0067] Migrate the task content of the heterogeneous computing unit to be migrated to the backup heterogeneous computing unit. Specifically, for each task running in the heterogeneous computing unit to be migrated, identify the heterogeneous computing type it currently depends on (such as GPU, NPU or domestic processor), and search for heterogeneous computing units of the same type in the backup heterogeneous computing units as the target migration node. Sort the computing power utilization of the heterogeneous computing units to be migrated from high to low, and give priority to migrating tasks with high computing power utilization to reduce the pressure on high-load nodes. In addition, each task to be migrated establishes a one-to-one correspondence with the target migration node during the migration process to ensure that the task completely restores its computing environment and execution status on the target node.

[0068] By continuously sampling and dynamically comparing the computing power utilization of each heterogeneous computing unit after inter-cluster adjustments, it is possible to identify in real time computing nodes with a significant upward trend in computing power load and, based on this, determine whether to trigger hot migration, effectively avoiding task delays, interruptions, or resource waste caused by node overload. Compared to static resource scheduling, this method offers dynamic response capabilities, enhancing the high throughput and low latency capabilities of task execution.

[0069] Furthermore, it also includes: when a task processing exception signal is received, an adjustment trajectory map is generated and displayed. The specific method is: based on the final cluster mapping table, a final clustering topology baseline is established as the graph structure framework of the adjustment trajectory map; based on the computing power stability parameters of the heterogeneous computing units in the initial cluster and the computing power exchange parameters between adjacent initial clusters, a multidimensional feature vector is constructed to obtain an adjustment vector sequence with a time order as the event sequence of the adjustment trajectory map; based on the graph structure framework and event sequence of the adjustment trajectory map, an adjustment trajectory map is generated and displayed.

[0070] In this embodiment, a final cluster topology baseline is established based on the final cluster mapping table. The specific method is as follows: based on the final cluster mapping table, the current clustering information of each heterogeneous computing unit (including the cluster to which each heterogeneous computing unit belongs, the type of each heterogeneous computing unit (e.g., GPU, CPU, etc.), and its network location) is obtained, and the network topology diagram of the artificial intelligence public computing power center is constructed based on the current clustering information of each heterogeneous computing unit. The devices (all network entities involved in constructing the artificial intelligence public computing power center topology, including: heterogeneous computing units (e.g., GPU / CPU nodes), network switches at all levels (e.g., edge switches, core switches, IB switches), and intermediate network components that implement node interconnection) are laid out according to their network locations and connection relationships, and the physical or logical connection paths between each device (all network entities involved in constructing the artificial intelligence public computing power center topology, including: heterogeneous computing units (e.g., GPU / CPU nodes), network switches at all levels (e.g., edge switches, core switches, IB switches), and intermediate network components that implement node interconnection) are identified, such as connections through network devices such as switches and routers. The computing power parameters of each heterogeneous computing unit (such as the computing power of the GPU, the number of CPU cores, etc.) are integrated into the topology diagram to form a network topology model containing computing power information. The constructed network topology diagram and computing power distribution are defined as the final clustering topology baseline.

[0071] A multidimensional feature vector is constructed based on the computing power stability parameters of the heterogeneous computing units in the initial cluster and the computing power exchange parameters between adjacent initial clusters. The specific method is as follows: the standard deviation of the computing power utilization and the standard deviation of the delay jitter duration of each heterogeneous computing unit in the initial cluster within a preset period are obtained from the database as the intra-cluster dimension components of the feature vector, and the average communication delay between adjacent initial clusters and the security level matching degree between adjacent initial clusters are obtained as the inter-cluster dimension components of the feature vector. Both are normalized to eliminate dimensional differences to obtain the processed parameters, and are combined into a set of multidimensional feature vectors in a preset order. The multidimensional feature vectors are numbered and marked in chronological order to form a set of feature sequences with timestamps, which are used to generate the event change basis of the trajectory map.

[0072] A time-ordered adjustment vector sequence refers to a vector set formed by arranging the multidimensional feature vectors constructed in each time slice (or sampling period) in the order of their sampling time. This set can continuously reflect the dynamic changes in the state of the computing power cluster.

[0073] By constructing an adjustment trajectory graph, we can quickly trace back every cluster adjustment operation during the multi-heterogeneous computing power scheduling process after a task processing anomaly occurs, including key steps such as intra-cluster fission, inter-cluster merging, and hot migration. Leveraging the graph structure framework established by the final cluster mapping table, combined with a time-ordered sequence of adjustment vectors constructed from computing power stability parameters and computing power exchange parameters, we can visually reconstruct and analyze the computing power evolution path before and after the anomaly, providing greater stability, transparency, and controllability for AI public computing power centers.

[0074] like Figure 6 As shown, it is a structural schematic diagram of a multi-heterogeneous computing power scheduling platform for an artificial intelligence public computing power center provided by an embodiment of the present application, including an initial allocation module, a cluster fission analysis module, a cluster merging analysis module and a hot migration analysis module; wherein, the initial allocation module is used to obtain a preset initial clustering mapping table when the artificial intelligence public computing power center receives a computing power scheduling signal, thereby obtaining heterogeneous computing units within the initial cluster; the cluster fission analysis module is used to obtain computing power stability parameters of heterogeneous computing units within the initial cluster and obtain a cluster stability judgment result. If the cluster stability judgment result is qualified, the intra-cluster fission adjustment is not performed and the cluster merging analysis module is directly entered; otherwise, the intra-cluster fission adjustment is performed and the cluster merging analysis module is entered after the adjustment; the cluster merging analysis module is used to obtain computing power exchange parameters between adjacent initial clusters, obtain a cluster safety judgment result, and perform corresponding inter-cluster adjustments based on the cluster safety judgment result; the hot migration analysis module is used to count the computing power utilization of each heterogeneous computing unit after the inter-cluster adjustment, and make a judgment to obtain a hot migration execution judgment result, and perform corresponding hot migration scheduling adjustments based on the hot migration execution judgment result.

[0075] To summarize, this embodiment achieves load balancing and security isolation of multiple heterogeneous computing units by obtaining an initial cluster mapping table, performing intra-cluster fission adjustment and inter-cluster merging adjustment, and based on computing power stability parameters and computing power exchange parameters, thereby significantly improving task processing efficiency and reducing latency fluctuations, thereby achieving improved task processing efficiency and effectively solving the problems of low task execution efficiency and large latency fluctuations in a multiple heterogeneous environment caused by inflexible computing power scheduling in the prior art.

[0076] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0077] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0078] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0079] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0080] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0081] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center, characterized in that: The following steps are involved: S1. When the AI ​​public computing power center receives the computing power scheduling signal, it obtains a preset initial cluster mapping table, thereby obtaining heterogeneous computing units within the initial cluster; S2. Obtain computing power stability parameters of the heterogeneous computing units within the initial cluster and obtain a cluster stability determination result. If the cluster stability determination result is qualified, no intra-cluster fission adjustment is performed and the process directly proceeds to S3. Otherwise, intra-cluster fission adjustment is performed and the process proceeds to S3 after the adjustment. S3. Obtain computing power exchange parameters between adjacent initial clusters, obtain cluster security determination results, and perform corresponding inter-cluster adjustments based on the cluster security determination results; S4. After the inter-cluster adjustment, the computing power utilization of each heterogeneous computing unit is counted, and a judgment is made to obtain a hot migration execution determination result, and corresponding hot migration scheduling adjustments are made based on the hot migration execution determination result.

2. The multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center according to claim 1, characterized in that: The method for obtaining computing power stability parameters of heterogeneous computing units in the initial cluster and obtaining cluster stability determination results is as follows: Obtain computing power stability parameters of heterogeneous computing units in the initial cluster within a preset time period, wherein the computing power stability parameters include a computing power utilization standard deviation and a delay jitter duration standard deviation; Obtain the preset standard deviation benchmark values ​​of computing power utilization and latency jitter duration in the database, and compare them with the computing power stability parameters to obtain the cluster stability judgment result; If the computing power utilization standard deviation is less than the computing power utilization standard deviation reference value and the latency jitter standard deviation is less than the latency jitter standard deviation reference value, the cluster stability determination result is qualified; otherwise, the cluster stability determination result is unqualified.

3. The multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center according to claim 1, characterized in that: The specific method for performing intra-cluster fission adjustment is as follows: Obtain the computing power utilization of each heterogeneous computing unit in the initial cluster; Obtain a computing power utilization split threshold preset in the database and compare it with the computing power utilization of each heterogeneous computing unit in the initial cluster. Thus, each heterogeneous computing unit with a computing power utilization above the computing power utilization split threshold is split from the initial cluster to a high-load sub-cluster, and each heterogeneous computing unit with a computing power utilization below the computing power utilization split threshold is split from the initial cluster to a low-load sub-cluster. Obtain the state parameters of each operator under each heterogeneous computing unit within the preset time period after the intra-cluster fission adjustment, analyze and obtain the state value of each operator, and thus match and obtain the resource consumption of each operator; Obtain the computing power utilization of each heterogeneous computing unit in the low-load sub-cluster and the computing power utilization of each heterogeneous computing unit in the high-load sub-cluster, analyze and obtain the computing power utilization ratio of the low-load sub-cluster, and perform task migration and operator migration accordingly. After performing intra-cluster fission adjustment, both the high-load sub-cluster and the low-load sub-cluster are marked as fission adjustment clusters.

4. The multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center according to claim 3, characterized in that: The specific method for obtaining the resource consumption of each operator is as follows: The status parameters of each operator under each heterogeneous computing unit include the maximum computing power utilization, maximum memory occupancy, historical failure rate, maximum latency, and maximum packet loss rate of each operator under each heterogeneous computing unit; Obtain a preset state reference set in the database, compare the difference between the preset state reference set and the state parameters of each operator under each heterogeneous computing unit, and introduce corresponding weighting factors to quantify the difference comparison results after obtaining the difference comparison results to obtain the state value of each operator under each heterogeneous computing unit; Obtain each operator state interval preset in the database and the reference resource consumption corresponding to each operator state interval, and compare them with the operator state value of each heterogeneous computing unit. If the operator state value of an operator is within a certain operator state interval, obtain the reference resource consumption corresponding to the interval as the resource consumption of the operator, thereby obtaining the resource consumption of each operator under each heterogeneous computing unit; The state reference set includes an ideal value of computing power utilization, an ideal value of video memory occupancy, an allowable value of historical failure rate, an allowable value of latency, and an allowable value of packet loss rate.

5. The multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center according to claim 3, characterized in that: The specific method for performing task migration and operator migration is as follows: Obtain each preset low computing power utilization ratio interval in the database and the reference task scheduling ratio corresponding to each low computing power utilization ratio interval, and compare them with the low-load sub-cluster computing power utilization ratio. If the low-load sub-cluster computing power utilization ratio is within a preset low computing power utilization ratio interval, obtain the reference task scheduling ratio corresponding to the interval as the task scheduling ratio; Based on the task scheduling ratio, the corresponding task is transferred from the high-load sub-cluster at the task scheduling ratio and marked as a task to be migrated; Obtain the operator resource consumption threshold preset in the database and compare it with the resource consumption of each operator in the high-load sub-cluster. Mark operators in the high-load sub-cluster whose operator resource consumption exceeds the operator resource consumption threshold as operators to be migrated. The tasks and operators to be migrated are matched with heterogeneous computing units with the same computing architecture in the low-load sub-cluster, and are preferentially assigned to the heterogeneous computing units with the lowest computing power utilization in the low-load sub-cluster for execution, thereby completing task migration and operator migration.

6. The multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center according to claim 1, characterized in that: The method for obtaining the computing power exchange parameters between adjacent initial clusters and obtaining the cluster security determination result is as follows: Obtaining computing power exchange parameters between adjacent initial clusters within a preset time period, wherein the computing power exchange parameters between adjacent initial clusters include an average computing power exchange delay and a security level matching degree between adjacent initial clusters; Obtaining the average delay threshold for computing power exchange between adjacent initial clusters and the security level matching threshold between adjacent initial clusters preset in the database; Compare the average delay of computing power exchange between adjacent initial clusters with the average delay threshold of computing power exchange between adjacent initial clusters, and compare the security level matching degree between adjacent initial clusters with the security level matching degree threshold between adjacent initial clusters to obtain the cluster security judgment result; If the average delay of computing power exchange between adjacent initial clusters is less than the average delay threshold of computing power exchange between adjacent initial clusters and the security level matching degree between adjacent initial clusters is greater than the security level matching degree threshold between adjacent initial clusters, the cluster security determination result is qualified; otherwise, the cluster security determination result is unqualified.

7. The multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center according to claim 1, characterized in that: The specific method for performing corresponding inter-cluster adjustments based on the cluster security determination result is as follows: Analyze based on the cluster safety determination result. If the cluster safety determination result is qualified, the inter-cluster merging adjustment is not performed, and the initial cluster mapping table is marked as a cluster adjustment mapping table; If the cluster safety determination result is unqualified, then perform inter-cluster merging and adjustment, and mark the cluster that performs inter-cluster merging and adjustment as a cluster to be merged; Obtain the network delay and security level matching degree of each cluster to be merged; Obtain the preset network delay allowable value and security level matching lower limit in the database, and perform difference analysis with the network delay and security level matching of each cluster to be merged. After obtaining the difference analysis results, introduce the corresponding weighting factor to quantify the difference analysis results and obtain the cluster merging processing value of each cluster to be merged; The cluster merging processing values ​​of each cluster to be merged are arranged in descending order, and two adjacent clusters to be merged are cluster merged in sequence according to the arrangement order. After the cluster merging process, the initial clustering mapping table is updated to obtain an adjusted clustering mapping table, which is marked as a cluster adjustment mapping table.

8. The multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center according to claim 1, characterized in that: After the inter-cluster adjustment, the computing power utilization of each heterogeneous computing unit is counted and judged to obtain the hot migration execution judgment result. The specific method is: Obtaining the computing power utilization rate of each heterogeneous computing unit at the first sampling moment, the computing power utilization rate at the final sampling moment, and the computing power utilization change value of each heterogeneous computing unit within the adjusted preset time period; Obtain a computing power utilization change threshold preset in the database and compare it with the computing power utilization change value of each heterogeneous computing unit. If the computing power utilization change value of a certain heterogeneous computing unit is above the computing power utilization change threshold, and the computing power utilization of the heterogeneous computing unit at the final sampling time is greater than the computing power utilization of the heterogeneous computing unit at the first sampling time, then the hot migration execution determination result of the heterogeneous computing unit is migration, and hot migration scheduling adjustment is performed. After the hot migration scheduling adjustment, the cluster adjustment mapping table is updated to obtain a final cluster mapping table. Otherwise, the hot migration execution determination result is not migration, and the hot migration scheduling adjustment is not performed. At the same time, the cluster adjustment mapping table is marked as the final cluster mapping table. The method for performing the hot migration scheduling adjustment is as follows: obtaining a heterogeneous computing unit that needs to be hot migrated and marking it as a heterogeneous computing unit to be migrated; Obtain the backup heterogeneous computing unit preset in the database, and migrate the task content of the heterogeneous computing unit to be migrated to the backup heterogeneous computing unit.

9. The multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center according to claim 1, characterized in that: Also includes: When a task processing exception signal is received, an adjustment trajectory map is generated and displayed. The specific method is as follows: Based on the final cluster mapping table, the final cluster topology baseline is established as the graph structure framework for adjusting the trajectory map; A multi-dimensional feature vector is constructed based on the computing power stability parameters of the heterogeneous computing units within the initial cluster and the computing power exchange parameters between adjacent initial clusters. A time-ordered adjustment vector sequence is obtained as the event sequence of the adjustment trajectory map. Based on the graph structure framework and event sequence of the adjustment trajectory map, an adjustment trajectory map is generated and displayed.

10. A platform using the multi-heterogeneous computing power scheduling method for an artificial intelligence public computing power center as described in any one of claims 1 to 9, characterized in that: include: Initial allocation module, cluster fission analysis module, cluster merging analysis module and thermal migration analysis module; The initial allocation module is configured to obtain a preset initial clustering mapping table after the artificial intelligence public computing power center receives the computing power scheduling signal, thereby obtaining heterogeneous computing units within the initial cluster; The cluster fission analysis module is used to obtain the computing power stability parameters of the heterogeneous computing units in the initial cluster and obtain the cluster stability determination result. If the cluster stability determination result is qualified, the intra-cluster fission adjustment is not performed and the cluster merging analysis module is directly entered. Otherwise, the intra-cluster fission adjustment is performed and the cluster merging analysis module is entered after the adjustment. The cluster merging analysis module is used to obtain computing power exchange parameters between adjacent initial clusters, obtain cluster security determination results, and perform corresponding inter-cluster adjustments based on the cluster security determination results; The heat migration analysis module is used to count the computing power utilization of each heterogeneous computing unit after inter-cluster adjustment, and to make judgments to obtain the heat migration execution judgment results, and to make corresponding heat migration scheduling adjustments based on the heat migration execution judgment results.

Citation Information

Patent Citations

  • Multi-element heterogeneous computing power equipment scheduling method, device and equipment and storage medium

    CN116700934A

  • Control method for avoiding aggressive migration of large-core processor and small-core processor

    CN116755864A

  • Service cluster scheduling method and device, equipment and medium

    CN118138589A

  • Heterogeneous multi-domain processor-oriented reaction flow simulation application load balance scheduling optimization method and system

    CN119106622A

  • Multi-element computing power heterogeneous computing power service management system and method

    CN119917267A