Scale analysis method, computing power scheduling method and device for intelligent computing power cluster
By acquiring business needs and cluster characteristics, the scale of the intelligent computing power cluster is assessed, solving the problem that existing technologies cannot measure the scalability of intelligent computing networks. This enables precise resource allocation and efficient resource utilization, ensuring that the computing power cluster meets business needs.
Patent Information
- Application Number
- CN202511237747.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-12-09
AI Technical Summary
Existing technologies cannot comprehensively and accurately measure the scalability of intelligent computing network architectures, making it difficult to determine whether intelligent computing systems have sufficient scalability to meet future computing needs when planning them.
By acquiring information on the computing power requirements of the target business scenario, and combining the cluster hardware resources, network switching equipment, and server configuration characteristics, the current computing power level and storage level of each cluster are determined, and the number of intelligent computing cards and servers that the network can support is evaluated to obtain information on the scale matching capability of the computing power cluster, and finally the scale analysis results of each cluster are determined.
It enables accurate assessment of the scale of computing power clusters, avoids under- or over-allocation of resources, improves resource utilization efficiency, reduces operating costs, ensures that computing power clusters meet business needs, and enhances enterprise competitiveness.
Smart Images

Figure CN121093607A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method for scale analysis, a method for scheduling computing power, and a device for intelligent computing power clusters. Background Technology
[0002] In today's era of rapid digital and intelligent development, the application scope of artificial intelligence (AI) technology is constantly expanding. From intelligent voice interaction and image recognition to autonomous driving and financial risk prediction, many fields rely on powerful computing capabilities. As a key indicator of computing power, the scale of AI computing power is crucial for its application scenarios. Sufficient AI computing power ensures the efficient operation of various complex AI algorithms and models, thereby achieving more accurate predictions, smarter decision-making, and a better user experience. For example, in the field of medical image diagnosis, large-scale AI computing power can accelerate the analysis of massive amounts of medical image data, helping doctors to discover diseases more quickly and accurately; in weather forecasting, powerful computing power can support more complex meteorological model calculations, improving the accuracy and timeliness of forecasts.
[0003] However, intelligent computing network architecture is highly complex, involving numerous hardware devices and intricate software systems. Due to the complexity of the architecture and the technical and engineering challenges of scaling up, quantifying the scalability of intelligent computing is difficult, lacking unified and scientific quantitative indicators to accurately measure its performance. Traditional evaluation methods often only assess aspects such as hardware performance indicators and software system functionalities, failing to comprehensively and accurately reflect the scalability of intelligent computing systems in practical applications. This makes it difficult to accurately determine whether a system possesses sufficient scalability to meet ever-changing future computing demands when planning its construction and development, thus hindering the development and application of intelligent computing technology. Summary of the Invention
[0004] In view of this, the present disclosure provides a method for scale analysis, a method for scheduling computing power, and an apparatus for intelligent computing power clusters, which can solve the problems existing in the prior art, such as the inability to achieve unified measurement of heterogeneous computing power resources.
[0005] In a first aspect, embodiments of this disclosure provide a method for analyzing the scale of intelligent computing power clusters, including:
[0006] Obtain information on the computing power requirements of the target business scenario;
[0007] Obtain the cluster hardware resource dimension characteristics, network switching device dimension characteristics, and server configuration dimension characteristics of each computing power cluster;
[0008] Based on the demand information and the characteristics of the cluster hardware resources, determine the current computing power level and current storage level of each cluster;
[0009] Based on the network switching device dimension characteristics and the server configuration dimension characteristics, the maximum number of smart computing cards that the parameter plane network can support and the maximum number of servers that the sample plane network can support are determined.
[0010] Based on the cluster hardware resource dimension characteristics, the maximum number of smart computing cards that the parameter plane network can support, and the maximum number of servers that the sample plane network can support, information on the computing power cluster scale matching capability is obtained.
[0011] Based on the current computing power level, current storage level, and computing power cluster size matching capability information of each cluster, the scale analysis results of each computing power cluster are determined.
[0012] Secondly, this disclosure also provides a method for scheduling computing power in an intelligent computing cluster, comprising:
[0013] The analysis results of the corresponding computing power cluster are obtained based on the scale analysis method of the intelligent computing power cluster described above;
[0014] All computing power clusters whose analysis results are greater than a preset threshold are considered as currently available computing power resources.
[0015] Thirdly, this disclosure also provides a computer device, which adopts the following technical solution:
[0016] The computer device includes:
[0017] At least one processor; and,
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to execute either the scale analysis method for the intelligent computing power cluster or the computing power scheduling method for the intelligent computing power cluster described above.
[0020] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to execute any of the above-described methods for analyzing the scale of intelligent computing power clusters or for scheduling the computing power of intelligent computing power clusters.
[0021] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.
[0022] The scale analysis method for intelligent computing power clusters disclosed in this application first obtains the demand information for computing power scale in the target business scenario. Then, it obtains the cluster hardware resource dimension characteristics, network switching device dimension characteristics, and server configuration dimension characteristics of each computing power cluster. Next, based on the demand information and cluster hardware resource dimension characteristics, it determines the current computing power level and current storage level of each cluster. Then, based on the network switching device dimension characteristics and server configuration dimension characteristics, it determines the maximum number of intelligent computing cards that the parameter plane network can support and the maximum number of servers that the sample plane network can support. Then, based on the cluster hardware resource dimension characteristics, the maximum number of intelligent computing cards that the parameter plane network can support, and the maximum number of servers that the sample plane network can support, it obtains the computing power cluster scale matching capability information. Finally, based on the current computing power level, current storage level, and computing power cluster scale matching capability information of each cluster, it determines the scale analysis result for each computing power cluster. This method comprehensively considers multiple dimensions such as cluster hardware resource dimension characteristics, network switching device dimension characteristics, and server configuration dimension characteristics, and quantifies the characteristics of each dimension, ultimately obtaining a scale analysis result that comprehensively and accurately reflects the actual situation of the computing power cluster.
[0023] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 A flowchart illustrating the scale analysis method for intelligent computing power clusters provided in this embodiment of the disclosure.
[0026] Figure 2 This is a flowchart illustrating the method for determining the current computing power level and current storage level of each cluster provided in this embodiment of the disclosure.
[0027] Figure 3 A flowchart illustrating the method for determining the scale analysis results of each computing cluster provided in this embodiment of the disclosure.
[0028] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation
[0029] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0030] Reference Figure 1 This application discloses a method for analyzing the scale of intelligent computing power clusters, including:
[0031] S100: Obtain information on the computing power requirements of the target business scenario.
[0032] The demand information includes the current computing power demand value F. u Current computing power per smart computing card (P) u The current number of smart computing cards installed on a single smart computing server (N) u The demand for computing power expansion in the next three years, value F f The computing power of a single smart computing card to be expanded (P) f Number of intelligent computing cards (N) to be installed on a single intelligent computing server to be expanded f Current storage requirement value T u The current available storage capacity of a single storage server is B. u Storage expansion demand in the next three years (T) f And the available storage capacity B of a single storage server to be expanded f .
[0033] Clearly defining the computing power requirements of the target business scenario is the foundation of the entire scale analysis. Only by accurately understanding the business's requirements for computing power can we conduct targeted subsequent cluster analysis and planning, avoid over-configuration or under-configuration of computing resources, thereby improving resource utilization efficiency and reducing costs.
[0034] S200 obtains the cluster hardware resource dimension characteristics, network switching device dimension characteristics, and server configuration dimension characteristics of each computing power cluster.
[0035] Among them, the cluster hardware resource dimension features include the number N of GPU cards currently connected to the parameter plane. G The number of servers currently connected to the sample surface is N. S The current number of intelligent computing servers connected to the sample surface is N. S1 The number of current sample surfaces connected to storage servers, N S2 .
[0036] The network switching device dimension features include device quantity, port quantity, and port speed. Specifically, the device quantity dimension features include the number of TOR / Leaf switches (L), Spine switches (S), and Super Spine switches (C); the port quantity dimension features include the number of ports per TOR / Leaf switch (P). L Number of ports per Spine switch (P)S Number of ports per Super Spine switch (P) C Port rate dimension features include the single-port rate α1 on the access side of the TOR / Leaf switch and the single-port rate α2 on the server side.
[0037] Server configuration dimension features include port multi-dimensional features and intelligent computing card configuration dimension features; among them, port multi-dimensional features include the number of ports β1 used by each server to connect to the parameter plane and the number of ports β2 used by each server to connect to the sample plane, and intelligent computing card configuration dimension features include the number of intelligent computing cards Γ per server.
[0038] A comprehensive understanding of the characteristics of a computing cluster helps in the in-depth analysis of its performance and potential. This information is crucial for determining the cluster's current computing power and storage level, assessing network support capabilities, and providing detailed data support for accurately judging whether the cluster meets business needs.
[0039] S300 determines the current computing power level and current storage level of each cluster based on demand information and cluster hardware resource characteristics.
[0040] Clearly defining the current computing power and storage level of each cluster allows for a direct understanding of the degree of match between the cluster and business needs. This helps identify which clusters are insufficient in terms of computing power or storage, and which clusters have surplus resources that can be utilized, providing direction for subsequent resource allocation and cluster optimization.
[0041] S400 determines the maximum number of smart computing cards that the parameter plane network can support and the maximum number of servers that the sample plane network can support based on the characteristics of network switching equipment and server configuration.
[0042] Among them, the maximum number of smart computing cards that the parameter plane network can support includes a) the maximum number of smart computing cards that can be connected to the parameter plane network as it is currently in operation; b) the maximum number of smart computing cards that can be connected to the parameter plane network with the Spine layer unchanged but the Tor / Leaf layer expanded only; and c) the maximum number of smart computing cards that can be connected to the parameter plane network with the Super Spine layer unchanged but the Spine layer and below network devices expanded only.
[0043] The maximum number of servers that the sample plane network can support includes: d) the maximum number of servers that can be connected to the sample plane network as it is currently in operation; e) the maximum number of servers that can be connected to the sample plane network if the Spine layer is not changed and only the Tor / Leaf layer is expanded; and f) the maximum number of servers that can be connected to the sample plane network if the Super Spine layer is not changed and only the Spine layer and lower network devices are expanded.
[0044] Determining the upper limit of network support for the number of smart computing cards and servers allows for the assessment of the cluster's network capacity. This is crucial for determining whether the cluster can further expand its computing and storage resources. If the network cannot support more device connections, even adding hardware resources will not fully utilize their performance, thus avoiding the waste of resources caused by blind expansion.
[0045] S500 obtains information on the matching capability of computing power cluster scale based on the characteristics of cluster hardware resources, the maximum number of intelligent computing cards that the parameter plane network can support, and the maximum number of servers that the sample plane network can support.
[0046] By comprehensively considering hardware resources and network support capabilities, the obtained information on the matching capability of computing power cluster scale can fully reflect the actual situation of the cluster. This helps to identify potential bottlenecks in the process of cluster scaling, provides a basis for formulating reasonable cluster scaling plans, and ensures that the cluster's hardware and network resources can develop in a coordinated manner.
[0047] S600 determines the scale analysis results for each computing cluster based on its current computing power level, current storage level, and computing power cluster size matching capability information.
[0048] By conducting comprehensive analysis to obtain the scale analysis results of each cluster, we can provide clear decision-making basis for enterprises' computing power resource planning and management. Enterprises can optimize and adjust different clusters in a targeted manner based on the analysis results, improve the operating efficiency and performance of the entire computing power cluster, and meet the computing power needs of business development.
[0049] The method disclosed in this application, through comprehensive analysis of business needs and cluster characteristics, can accurately determine the resource matching of each cluster, avoiding over- or under-configuration of resources, ensuring full utilization of computing and storage resources, and effectively avoiding unnecessary cost expenditures from blindly purchasing hardware equipment and upgrading the network. Based on the scale analysis results, enterprises can make targeted resource allocation and optimization, reduce operating costs, ensure that the computing cluster can meet the computing and storage needs of business development, improve business processing speed and responsiveness, enhance enterprise competitiveness, identify problems and bottlenecks in cluster hardware resources and network configuration, provide a basis for optimizing cluster architecture, and make cluster performance more stable and reliable.
[0050] This method comprehensively considers multiple aspects, including cluster hardware resource dimensions, network switching equipment dimensions, and server configuration dimensions. Different dimensions reflect the performance and potential of the computing power cluster from different perspectives. Multi-dimensional evaluation avoids the one-sidedness of single-dimensional evaluation and can present the actual situation of each computing power cluster more comprehensively and accurately, thereby achieving a precise evaluation of the scale of the computing power cluster.
[0051] Reference Figure 2 The method for S300 to "determine the current computing power level and current storage level of each cluster based on demand information and cluster hardware resource characteristics" includes:
[0052] A100 determines the current computing power of each computing cluster based on the current computing power of a single smart computing card and the number of GPU cards connected to the current parameter plane.
[0053] Currently, it has computing power F c :F c =P u ×N G , where P u N represents the current computing power of a single smart computing card. G This represents the number of GPU cards currently connected to the parameter plane.
[0054] Accurately calculating the current computing power of each computing cluster is the foundation for subsequent assessments of whether the cluster meets business needs; by clarifying the existing computing power scale, we can intuitively understand the current computing power status of the cluster, providing data support for further resource planning and allocation.
[0055] A200 determines the current computing power level of each cluster based on the current computing power demand and the existing computing power.
[0056] If Fu≦Fc, the current computing power level of each cluster is determined to be the first level; otherwise, it is determined to be the second level.
[0057] By comparing the current computing power requirement with the existing computing power, the computing power level can be determined, and it is possible to quickly determine whether the computing power of each cluster is sufficient. This helps enterprises to identify computing power bottlenecks in a timely manner. For clusters with insufficient computing power, measures such as adding hardware equipment and optimizing algorithms can be taken in a timely manner to improve computing power and meet business needs.
[0058] A300 determines the current storage capacity of each computing cluster based on the available storage capacity of a single storage server and the number of storage servers currently connected to the sample plane.
[0059] The current available storage capacity is T. c :T c =B u ×N S2 B u N represents the current available storage capacity of a single storage server. S2 N is the number of storage servers currently connected to the sample surface. S2 .
[0060] Determining the current storage capacity of each computing cluster helps in understanding the cluster's data storage capabilities. This is especially important for business scenarios that require processing large amounts of data, such as big data analytics and artificial intelligence training. Accurately calculating storage capacity provides a basis for data storage and management, preventing data loss or business disruptions due to insufficient storage.
[0061] A400 determines the current storage level for each cluster based on the current storage demand and the existing storage capacity.
[0062] If T u ≦T c Each cluster is assigned a first-level computing power rating, and vice versa.
[0063] By comparing current storage requirements with existing storage capacity, the storage level can be determined, providing a clear understanding of the storage status of each cluster. For clusters with low storage levels, it is possible to plan for adding storage devices or optimizing storage strategies in a timely manner to ensure the secure storage and efficient access of business data.
[0064] The method disclosed in this embodiment, by accurately determining the current computing power and storage level of each cluster, allows enterprises to optimize resource allocation in a targeted manner based on business needs and the actual situation of the cluster. For clusters with insufficient computing power or storage, timely expansion can be implemented; for clusters with excess resources, resource investment can be appropriately reduced, thereby improving resource utilization efficiency and lowering costs. This ensures that the computing power and storage of each cluster can meet business needs, avoiding problems such as slow business processing speeds due to insufficient computing power or data loss due to insufficient storage, thus guaranteeing stable business operation. As business continues to develop, the demand for computing power and storage will also change. This solution can assess the computing power and storage status of the cluster in real time, providing strong support for enterprise business development and enabling enterprises to adjust resource allocation in a timely manner to adapt to business growth and changes.
[0065] The S400 method for "determining the upper limit of the number of smart computing cards that the parameter plane network can support and the maximum number of servers that the sample plane network can support based on the dimensional characteristics of network switching equipment and the dimensional characteristics of server configuration" includes:
[0066] S410, based on the number of TOR / Leaf switches L and the number of ports on a single TOR / Leaf switch P L The maximum number of smart computing cards that can be accessed by the parameter plane network is obtained by using the following formulas: TOR / Leaf switch access side single port rate α1, server side single port rate α2, number of ports used by each server to connect to the sample plane β2, number of smart computing cards per server Γ, and the first preset formula.
[0067] This step can quickly estimate the maximum number of smart computing cards that can be connected to the parameter plane network under the existing network architecture, based on the existing hardware configuration of the TOR / Leaf switch, providing basic data for subsequent network evaluation and planning.
[0068] S420, based on the current network conditions, the maximum number of smart computing cards that can be connected is 'a', the number of Spine switches is 'S', and the number of ports per Spine switch is 'P'. S The parameters are: the single-port rate α1 of the TOR / Leaf switch access side, the single-port rate α2 of the server side, the number of ports β2 used by each server to connect to the sample plane, the number of smart computing cards Γ of each server, and the second preset formula. The parameters are: the network surface remains unchanged, the Spine layer is expanded, and the maximum number of smart computing cards b that can be connected to the Tor / Leaf layer is increased.
[0069] Without changing the Spine layer network equipment, this study evaluates the additional number of smart computing cards that can be added by expanding the Tor / Leaf layer, providing an economical and effective basis for network expansion.
[0070] S430, based on the number of TOR / Leaf switches L and the number of ports per TOR / Leaf switch P L TOR / Leaf switch access side single port rate α1, server side single port rate α2, number of Spine switches S, number of ports per Spine switch P S Number of Super Spine switches (C), Number of ports per Super Spine switch (P) C The parameters are: β2, the number of ports β2 used by each server to connect to the sample plane; Γ, the number of smart computing cards Γ on each server; and the third preset formula. The parameters are obtained by keeping the Super Spine layer unchanged and only expanding the maximum capacity of smart computing cards c for network devices at the Spine layer and below.
[0071] This step considers the maximum smart card access capacity that can be achieved by expanding network devices at the Spine layer and below without changing the Super Spine layer, providing a more comprehensive planning direction for the hierarchical expansion of the network.
[0072] S440, based on the number of TOR / Leaf switches L and the number of ports per TOR / Leaf switch P L The maximum number of servers that can be accessed by the sample plane network is obtained by using the single-port rate α1 of the TOR / Leaf switch access side, the single-port rate α2 of the server side, the number of ports used by each server to connect to the parameter plane β1, and the fourth preset formula.
[0073] This step can determine the maximum number of servers that the sample plane can access under the existing network architecture based on the current hardware configuration of the sample plane network, which helps to understand the carrying capacity of the sample plane network.
[0074] S450, based on the current sample network status, the maximum number of servers that can be connected is d, the number of Spine switches is S, and the number of ports per Spine switch is P. S The maximum number of servers that can be connected to the sample network is obtained by using the following formulas: α1 (single port rate on the access side of the TOR / Leaf switch), α2 (single port rate on the server side), β1 (number of ports used by each server to connect to the parameter plane), and the fifth preset formula. The Spine layer is not changed, but the Tor / Leaf layer is expanded.
[0075] Without changing the Spine layer, this study evaluates the additional number of server connections that the sample plane network can achieve by scaling up the Tor / Leaf layer, providing a feasible solution for scaling up the sample plane network.
[0076] S460, based on the current sample network status, the maximum number of accessible servers (d), the number of Spine switches (S), and the number of ports per Spine switch (P) S , TOR / Leaf switch access side single port rate α1, server side single port rate α2, number of ports per server used for connecting to the parameter plane β1, number of Super Spine switches C, number of ports per Super Spine switch P C And the sixth preset formula, to obtain the maximum number of servers f that can be accessed by the network devices at the Spine layer and below without changing the Super Spine layer.
[0077] This step considers the maximum number of servers that the sample network can connect to after expanding the network devices at the Spine layer and below without changing the Super Spine layer, providing a comprehensive planning basis for the hierarchical expansion of the sample network.
[0078] The method disclosed in this embodiment, through detailed calculation and evaluation of the capacity expansion of the parameter plane and sample plane networks at different network device levels, can accurately plan the network expansion scheme and avoid the problems of over-expansion or under-expansion. Based on the capacity expansion calculation results of different levels, the most economical and effective expansion method can be selected. For example, if the demand is met, priority can be given to expanding the Tor / Leaf layer, thereby reducing network construction costs. It clarifies the number of intelligent computing cards and servers that the network can support under different states, which helps to rationally allocate and optimize network resources and improve network utilization efficiency. It provides a clear direction for future network upgrades and expansions, enabling the network to expand in an orderly manner with the development of services, ensuring network stability and scalability.
[0079] The maximum number of smart computing cards that can be connected to the current parameter plane network is a:
[0080]
[0081] τ=H(L)·[1+H(S)·(1+H(C))].
[0082] Among them, the maximum number of smart computing cards that can be connected is b, where the Spine layer remains unchanged and only the Tor / Leaf layer is expanded:
[0083] in,
[0084] Among them, the parameter plane network remains unchanged, the Super Spine layer is only expanded, and the maximum capacity of the smart computing card that can be connected to network devices at the Spine layer and below is c:
[0085]
[0086] The maximum number of servers that can be connected to the current sample network is d:
[0087] Among them, the maximum number of servers that can be connected to the sample network, with the Spine layer unchanged and only the Tor / Leaf layer expanded, is e:
[0088] Among them, the sample network remains unchanged, the Super Spine layer is not modified, and only the Spine layer and below network devices are expanded. The maximum number of servers that can be connected is f:
[0089]
[0090] Where L represents the number of TOR / Leaf switches, S represents the number of Spine switches, C represents the number of Super Spine switches, and P represents the number of Super Spine switches. L For the number of ports on a single TOR / Leaf switch, P S For the number of ports on a single Spine switch, P C α1 is the number of ports on a single Super Spine switch, α2 is the single port rate on the access side of the TOR / Leaf switch, α3 is the single port rate on the server side, β1 is the number of ports on each server used to connect to the parameter plane, β2 is the number of ports on each server used to connect to the sample plane, Γ is the number of smart computing cards on each server, H() is the step function, τ is the network level, η(τ) is the first network level constant, and γ(τ) is the second network level constant.
[0091] In this embodiment, the computing power cluster scale matching capability information includes the current overall resource supply capability characteristics, computing power demand characteristics within a preset period, storage demand characteristics within a preset period, smart computing card expansion rate demand characteristics, the current network smart computing card expansion ratio characteristics without changing it, server expansion rate demand characteristics, the current network server expansion ratio characteristics without changing it, the expansion ratio of smart computing cards that only expand Tor / Leaf layer networks, the expansion ratio of servers that only expand Tor / Leaf layer networks, the expansion ratio of smart computing cards that only expand Spine layer and below networks, and the expansion ratio of servers that only expand Spine layer and below networks.
[0092] The method for S500 to "obtain computing power cluster scale matching capability information based on cluster hardware resource dimension characteristics, the maximum number of intelligent computing cards supported by the parameter plane network, and the maximum number of servers supported by the sample plane network" specifically includes:
[0093] Based on the characteristics of cluster hardware resources, the maximum number of smart computing cards that the parameter plane network can support, and the maximum number of servers that the sample plane network can support, the following characteristics are obtained: current overall resource supply capacity, computing power demand within a preset period, storage demand within a preset period, smart computing card expansion rate demand, the proportion of smart computing cards that can be expanded without changing the current network, server expansion rate demand, the proportion of servers that can be expanded without changing the current network, the proportion of smart computing cards that can be expanded only in the Tor / Leaf layer network, the proportion of servers that can be expanded only in the Tor / Leaf layer network, the proportion of smart computing cards that can be expanded only in the Spine layer and below network, and the proportion of servers that can be expanded only in the Spine layer and below network.
[0094] Among them, the current overall resource supply capacity characteristics include the current computing power F c And the current storage capacity T c .
[0095] Among them, the computing power demand characteristics within the preset period are Q1: Q1 = F f +F u , of which F f F represents the computing power expansion needs over the next three years. u This represents the current computing power requirement.
[0096] Among them, the storage demand characteristic within the preset period is Q2: Q2 = T f +T u , among which, T f T represents the storage expansion demand over the next three years. u This represents the current storage requirement.
[0097] Among them, the characteristic of the intelligent computing card expansion rate requirement is k. g : Among them, P fTo expand the computing power of a single smart computing card, P u This represents the current computing power of a single smart computing card.
[0098] Among them, the current network intelligent computing card expansion ratio characteristic is k. a : Where, N G 'a' represents the number of GPU cards currently connected to the parameter plane, and 'a' represents the maximum number of intelligent computing cards that can be connected to the parameter plane network at present.
[0099] Among them, the server expansion rate requirement is characterized by k. s :
[0100]
[0101] Where, N f To determine the number of smart computing cards that can be installed on a single smart computing server to be expanded, B f N represents the available storage capacity of a single storage server to be expanded. u B represents the number of smart computing cards installed on a single smart computing server. u This represents the current available storage capacity of a single storage server.
[0102] Among them, the current network server's expandability ratio is characterized by k. d : Where, N S d represents the number of servers currently connected to the sample plane, and d represents the maximum number of servers that can be connected to the sample plane in the current network state.
[0103] Among them, the expansion ratio of only the Tor / Leaf layer network intelligent computing card is k. b : Where b represents the maximum number of smart computing cards that can be connected to while the Spine layer remains unchanged and the Tor / Leaf layer is expanded.
[0104] Among them, the expansion ratio for expanding only the Tor / Leaf layer network servers is k. e : Where e represents the maximum number of servers that can be connected to while the Spine layer remains unchanged and only the Tor / Leaf layer is expanded.
[0105] Among them, the expansion ratio for network intelligent computing cards at the Spine layer and below is k. c : Where c represents the maximum capacity of the smart computing card that can be accessed by network devices at the Spine layer and below, while the Super Spine layer remains unchanged.
[0106] Among them, the expansion ratio for only Spine layer and below network servers is k. f : For sample network, the Super Spine layer remains unchanged; only the Spine layer and below network devices are expanded to accommodate the maximum number of servers that can be connected.
[0107] Reference Figure 3 The S600 method of "determining the scale analysis results of each computing cluster based on the current computing power level, current storage level, and computing power cluster size matching capability information" includes the following methods for determining the scale analysis results of each computing power cluster:
[0108] S610 determines the current overall resource capability characteristics, current computing power expansion capability characteristics, current storage expansion capability characteristics, current overall resource expansion capability characteristics, Tor / Leaf layer network expansion characteristics, Spine layer and below network expansion characteristics, and network architecture expansion characteristics of each cluster based on the current computing power level, current storage level, and computing power cluster scale matching capability information of each cluster.
[0109] A comprehensive understanding of the resource characteristics and scaling capabilities of each cluster provides a detailed data foundation for subsequent scale analysis, which helps to formulate more accurate scaling strategies.
[0110] S620: When the overall capability characteristics of the current resources meet the first preset condition, the first preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster.
[0111] Among them, the current overall resource capability characteristics satisfying the first preset condition include: F f +F u -F c ≦0 and 0≧T f +T u -T c In this condition, both the current computing power level and the current storage level are at level one.
[0112] S630: When the overall capability characteristics of the current resources do not meet the first preset condition and the current computing power expansion capability characteristics meet the second preset condition, the second preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster.
[0113] Among them, the current computing power expansion capability characteristics that meet the second preset condition include: F f +F u -F c ≦0 and 0 <T f +T u -T c In this condition, both the current computing power level and the current storage level are at level one.
[0114] S640: When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, and the current storage expansion capability characteristics meet the third preset condition, the third preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster.
[0115] Among them, the current storage expansion capability characteristics that meet the third preset condition include: 0 <F f +F u -T c And 0≧T f +T u -T c In this condition, both the current computing power level and the current storage level are at level one.
[0116] S650: When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, and the current overall resource expansion capability characteristics meet the fourth preset condition, the fourth preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster.
[0117] Among them, the current overall resource expansion capacity characteristics meet the fourth preset condition, namely: k g ≦k a And k s ≦k d In this scenario, both the current computing power level and the current storage level are at the first level. This embodiment demonstrates that the current requirements can be met, and no network modifications are needed; simply expanding the computing power and storage servers is sufficient to meet the computing power and storage expansion needs for the next three years.
[0118] Specifically, k g ≤k a This indicates that the current business's demand for smart computing card expansion is less than or equal to the expansion ratio achievable with smart computing cards without altering the current network, k. s ≤k d This indicates that the current business's demand for server expansion is less than or equal to the achievable expansion ratio of the server without altering the current network. These two inequalities together reflect the relationship between the expansion needs and scalability of intelligent computing cards and servers without changing the current network architecture. If both inequalities hold true, it means that the expansion needs of intelligent computing cards and servers can be met on the existing network foundation, without requiring large-scale network modifications, thus saving costs and time.
[0119] S660: When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, the current overall resource expansion capability characteristics do not meet the fourth preset condition, and the Tor / Leaf layer network expansion characteristics meet the fifth preset condition, the fifth preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster.
[0120] Among them, the Tor / Leaf layer network expansion characteristics satisfying the fifth preset condition include: k a <k g ≦k b And k d <k s ≦k e In this embodiment, it is shown that the current demand can be met, and that the computing power and storage expansion needs for the next three years can be met by only expanding the Tor / Leaf layer network.
[0121] k a <k g ≦k b This indicates that, without modifying the current network, the existing conditions cannot meet the expansion requirements of the smart computing card (k). a <k g However, if only the Tor / Leaf layer network is expanded, it can meet or just meet the expansion requirements of the smart computing card (k). g ≦k b This means that in order to achieve the expansion target of the smart computing card, the Tor / Leaf layer network needs to be adjusted and upgraded accordingly.
[0122] k d <k s ≦k e This indicates that without changing the current network, it is impossible to meet the server's expansion requirements (k). d <k s However, simply expanding the Tor / Leaf layer network is sufficient to meet or just barely meet the server expansion requirements (k s ≦k e This means that to expand server capacity, the Tor / Leaf layer network needs to be optimized and upgraded.
[0123] These two sets of inequalities, taken together, reflect the relationship between the expansion needs and scalability of intelligent computing cards and servers under existing network conditions. When k appears... a <k g and k d <k s In this case, it indicates that the current network architecture alone cannot meet the expansion requirements. And k g ≤kb and k s ≤k e This suggests that expanding the Tor / Leaf layer network can solve the expansion problem of smart computing cards and servers, providing a specific direction for network planning and upgrades.
[0124] S670: When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, the current overall resource expansion capability characteristics do not meet the fourth preset condition, the Tor / Leaf layer network expansion characteristics do not meet the fifth preset condition, and the Spine layer and below network expansion characteristics meet the sixth preset condition, the sixth preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster.
[0125] Among them, the network expansion characteristics of the Spine layer and below satisfy the sixth preset condition, including: k b <k g ≦k c And k e <k s ≦k f In this condition, both the current computing power level and the current storage level are at the first level. This embodiment illustrates that the current demand can be met, and the computing power and storage expansion needs for the next three years can be met by expanding the Spine layer and below, but not by expanding the Tor / Leaf layer.
[0126] Specifically, k b <k g ≦k c This indicates that simply scaling up the Tor / Leaf layer network is insufficient to meet the scaling requirements of intelligent computing cards (k b <k g However, when scaling up the Spine layer and below (i.e., simultaneously scaling up the Spine layer and the Tor / Leaf layer), the scaling requirements of the smart computing card can be met or just met. g ≦k c ).
[0127] k e <k s ≦k f This indicates that simply expanding the Tor / Leaf layer network is insufficient to meet the server expansion requirements (k e <k s When expanding the network at the Spine layer and below, the expansion requirements of the servers can be met or just met. s ≦k f ).
[0128] These two sets of inequalities comprehensively reflect the relationship between the expansion requirements and scalability of intelligent computing cards and servers under different network expansion levels. When k appears... b <k g and k e <k s In this case, it indicates that simply scaling up the Tor / Leaf layer network is insufficient to meet the scaling requirements. And k g ≦k c and k s ≦k f This suggests that by expanding the network at the Spine layer and below, we can solve the expansion problem of smart computing cards and servers, providing a more specific strategy and direction for network planning and upgrades.
[0129] S680: When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, the current overall resource expansion capability characteristics do not meet the fourth preset condition, the Tor / Leaf layer network expansion characteristics do not meet the fifth preset condition, the Spine layer and below network expansion characteristics do not meet the sixth preset condition, and the network architecture expansion characteristics meet the seventh preset condition, the seventh preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster.
[0130] Among them, the network architecture expansion characteristics that meet the seventh preset condition include: k c <k g And k f <k s In this condition, both the current computing power level and the current storage level are at the first level. This embodiment illustrates that the current demand can be met, but the expansion demand for the next three years cannot be met by scaling up networks below the Spine layer.
[0131] Specifically, k c <k g It was explicitly stated that even after capacity expansion operations were performed on the Spine layer and below, the network's capacity to support smart computing cards was still less than the actual demand for smart computing card expansion. This means that adjustments only to the Spine layer and below are insufficient to meet the business's requirements for smart computing card expansion. f <k s This means that simply expanding the network at the Spine layer and below cannot meet the actual expansion needs of the servers.
[0132] The method disclosed in this embodiment expands computing power, storage, or network capacity in a targeted manner based on different conditions, avoiding blind expansion, improving resource utilization efficiency, and reducing costs. While meeting business needs, it prioritizes lower-cost expansion methods, such as expanding resources without modifying the network, reducing unnecessary network upgrade costs. It provides a clear direction and specific strategies for network planning and upgrades, making the expansion process more orderly and efficient. It ensures that each computing cluster can meet current and future business needs, guaranteeing stable business operation.
[0133] Furthermore, for S620 "when the overall capability characteristics of the current resources meet the first preset condition, the first preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster", it includes:
[0134] If the computing power demand characteristics within the preset period (F) f +F u () Not greater than the current computing power F c And the storage demand characteristics within the preset period (T) f +T u No greater than the current available storage capacity T c At that time, the corresponding computing power cluster size is determined as the first level, and the score corresponding to the first level is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
[0135] In a preferred embodiment, the score corresponding to the first level is Z7.
[0136] When the corresponding computing cluster size is determined to be Level 1, it means that the current computing and storage capacity can meet the expansion needs for the current period and the next three years (no network expansion is required).
[0137] The method for S630 that "when the overall capability characteristics of the current resources do not meet the first preset condition and the current computing power expansion capability characteristics meet the second preset condition, the second preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster" includes:
[0138] If the computing power demand characteristics within the preset period (F) f +F u () Not greater than the current computing power F c And the storage demand characteristics within the preset period (T) f +T u () Greater than the current available storage capacity T c At that time, the corresponding computing power cluster size is determined as the second level, and the score corresponding to the second level is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
[0139] In a preferred embodiment, the score corresponding to the second level is Z6.
[0140] When the corresponding computing cluster size is determined to be Level 2, it means that the current computing power can meet the expansion needs for the current period and the next three years.
[0141] The method for S640, which calls the third preset strategy to obtain the scale analysis results of the corresponding computing cluster when the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics meet the second preset condition, and the current storage expansion capability characteristics meet the third preset condition, includes:
[0142] If the computing power demand characteristics within the preset period (F) f +F u () is greater than the current computing power F c And the storage demand characteristics within the preset period (T) f +T u No greater than the current available storage capacity T c At that time, the corresponding computing power cluster size is determined to be the third level, and the score corresponding to the third level is obtained according to the preset scoring list as the scale analysis score Z6 of the corresponding computing power cluster.
[0143] In a preferred embodiment, the score corresponding to the third level is Z6.
[0144] When the corresponding computing cluster size is determined to be Level 3, it means that the current storage capacity can meet the expansion needs for the present and the next three years.
[0145] The S650 method for "when the overall current resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, and the overall current resource expansion capability characteristics meet the fourth preset condition, calling the fourth preset strategy to obtain the scale analysis results of the corresponding computing power cluster" includes:
[0146] If the smart card expansion rate requires k g Not greater than the current network intelligent computing card expansion ratio k a And the server expansion rate requirement k s Not greater than the current network server's expandable capacity ratio k d At that time, the corresponding computing power cluster size is determined to be the fourth level, and the score corresponding to the fourth level is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
[0147] In a preferred embodiment, the score corresponding to the fourth level is Z5.
[0148] When the corresponding computing power cluster size is determined to be Level 4, it means that the current demand can be met, and: there is no need to change the network, only to expand the computing power and storage servers to meet the computing power and storage expansion needs within three years.
[0149] The method for S660 that "when the overall current resource capacity characteristics do not meet the first preset condition, the current computing power expansion capacity characteristics do not meet the second preset condition, the current storage expansion capacity characteristics do not meet the third preset condition, the overall current resource expansion capacity characteristics meet the fourth preset condition, and the Tor / Leaf layer network expansion characteristics meet the fifth preset condition, calls the fifth preset strategy to obtain the scale analysis results of the corresponding computing power cluster" includes:
[0150] If the current network smart card remains unchanged, the expansion ratio can be increased by k. a The required capacity expansion rate k is less than that of the smart computing card. g Smart computing card expansion rate requirement k g No more than the expansion ratio k of Tor / Leaf layer network intelligent computing cards. b The current network server expansion ratio (k) remains unchanged. d Less than the server expansion rate requirement k s And the server expansion rate requirement k s No more than the scalability ratio k of Tor / Leaf layer network servers that can be expanded only. e At that time, the corresponding computing power cluster size is determined to be Level 5, and the score corresponding to Level 5 is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
[0151] In a preferred embodiment, the score corresponding to the fifth level is Z4.
[0152] When the corresponding computing power cluster size is determined to be Level 5, it means that the current demand can be met, and the computing power and storage expansion demand within three years can be met by only expanding the Tor / Leaf layer network.
[0153] The method for S670 that "when the overall current resource capacity characteristics do not meet the first preset condition, the current computing power expansion capacity characteristics do not meet the second preset condition, the current storage expansion capacity characteristics do not meet the third preset condition, the overall current resource expansion capacity characteristics do not meet the fourth preset condition, the Tor / Leaf layer network expansion characteristics do not meet the fifth preset condition, and the Spine layer and below network expansion characteristics meet the sixth preset condition, calls the sixth preset strategy to obtain the scale analysis results of the corresponding computing power cluster" includes:
[0154] If only the Tor / Leaf layer network intelligent computing card is expanded, the expansion ratio can be k. b The required capacity expansion rate k is less than that of the smart computing card. g Smart computing card expansion rate requirement k g The expansion ratio (k) for network intelligent computing cards at or below the Spine layer is not greater than the expansion ratio for that layer. c Expanding only the Tor / Leaf layer network servers allows for a scalability ratio of k. eLess than the server expansion rate requirement k s And the server expansion rate requirement k s No more than the expansion ratio (k) for network servers at or below the Spine layer. f At that time, the corresponding computing power cluster size is determined to be level six, and the score corresponding to level six is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
[0155] In a preferred embodiment, the score corresponding to the sixth level is Z3.
[0156] When the corresponding computing power cluster size is determined to be level 6, it means that the current demand can be met, and: the computing power and storage expansion demand within three years can be met by expanding the Spine layer and below the network, but: it cannot be met by expanding the Tor / Leaf layer.
[0157] The S680 method for "when the overall current resource capacity characteristics do not meet the first preset condition, the current computing power expansion capacity characteristics do not meet the second preset condition, the current storage expansion capacity characteristics do not meet the third preset condition, the overall current resource expansion capacity characteristics do not meet the fourth preset condition, the Tor / Leaf layer network expansion characteristics do not meet the fifth preset condition, the Spine layer and below network expansion characteristics do not meet the sixth preset condition, and the network architecture expansion characteristics meet the seventh preset condition, calling the seventh preset strategy to obtain the scale analysis results of the corresponding computing power cluster" includes:
[0158] If only Spine layer and below network intelligent computing cards are expanded, the expansion ratio can be k. c The required capacity expansion rate k is less than that of the smart computing card. g And only the Spine layer and below network servers can be expanded up to a capacity ratio of k. f Less than the server expansion rate requirement k s The corresponding computing power cluster size is determined to be level seven, and the score corresponding to level seven is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
[0159] In a preferred embodiment, the score corresponding to the seventh level is Z2.
[0160] When the corresponding computing power cluster size is determined to be level 7, it means that the current demand can be met, but: the expansion demand within three years cannot be met by expanding the network below the Spine layer.
[0161] Furthermore, when F u >F c Or T u >T cWhen the current computing power and storage requirements cannot be met, the corresponding computing power cluster size is determined to be level eight, and the score corresponding to level eight is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
[0162] In a preferred embodiment, the score corresponding to the eighth level is Z1.
[0163] When the corresponding computing power cluster size is determined to be level eight, it means that the current computing power and storage requirements cannot be met.
[0164] Among them, Z7>Z6>Z5>Z4>Z3>Z2>Z1.
[0165] The intelligent computing power cluster scale analysis method disclosed in this application obtains the computing power scale requirements of the target business scenario, which enables the clarification of actual business needs during expansion planning. Secondly, by comprehensively acquiring the cluster hardware resource dimension characteristics, network switching device dimension characteristics, and server configuration dimension characteristics of each computing power cluster, a detailed understanding of the numerous hardware devices involved in the intelligent computing network architecture is achieved. By collecting these multi-faceted characteristics, the specific situation of the intelligent computing network architecture can be fully grasped, providing detailed data support for subsequent analysis and decision-making. Next, based on the demand information and cluster hardware resource dimension characteristics, the current computing power and storage level of the cluster are determined. Based on the network switching device dimension characteristics and server configuration dimension characteristics, the upper limit of network support for the number of intelligent computing cards and servers is determined. By combining hardware resources and network support capabilities, information on the matching capability of computing power cluster scale is obtained. This collaborative analysis method can deeply understand the interrelationship between hardware and network, solving the problem of efficient collaboration between hardware and software. Finally, based on the current computing power level, current storage level, and computing power cluster scale matching capability information of each cluster, the scale analysis results of each computing power cluster are determined. Based on these analysis results, targeted expansion plans can be formulated. If the analysis results show that there is a bottleneck in the network of a certain cluster, the network can be upgraded first, and then the addition of hardware devices can be considered. If the storage level is insufficient, storage devices can be added in a targeted manner. This approach can ensure that the newly added computing power and storage resources are seamlessly integrated with the original system, while also taking into account practical issues such as energy consumption and space occupation, effectively reducing the complexity of expansion.
[0166] This method constructs a quantitative indicator system through a series of steps to evaluate the scalability of intelligent computing systems. Step S300 determines the current computing power level and storage level; step S400 determines the maximum number of intelligent computing cards that the parameter plane network can support and the maximum number of servers that the sample plane network can support; and step S500 obtains information on the matching capability of the computing power cluster scale. These are all quantitative indicators that comprehensively and accurately reflect the scalability of intelligent computing systems in practical applications. In step S600, multiple quantitative indicators are comprehensively considered to determine the scale analysis results of each computing power cluster, thereby providing a comprehensive evaluation of the scalability of the intelligent computing system. This comprehensive evaluation method avoids the limitations of traditional methods that can only evaluate from partial perspectives, and can more comprehensively reflect the ability of intelligent computing systems to effectively expand with reasonable cost and resource consumption when facing ever-increasing computing demands. For example, it considers not only the performance indicators of hardware devices but also network support capabilities and storage conditions, making the evaluation of scalability more accurate and scientific, and providing a reliable reference for planning the construction and development of intelligent computing systems.
[0167] By acquiring information on the computing power requirements of the target business scenario and integrating it into the entire analysis process, this method closely aligns the evaluation results with actual business needs. Based on specific business requirements, it accurately determines whether each computing cluster is suitable for a particular business scenario, improving the relevance and practicality of the evaluation results. In situations where multiple computing clusters coexist, this method facilitates optimized resource allocation. Based on the scale analysis results of each cluster, different business tasks can be assigned to the most suitable cluster.
[0168] From acquiring demand information to finalizing the scale analysis results, this method forms a complete automated computing process. The entire process requires minimal human intervention and judgment, reducing human error and improving the accuracy and consistency of the assessment. Simultaneously, the automated process significantly shortens the assessment time, enabling the analysis of multiple computing clusters in a short period, meeting the high-efficiency assessment requirements of modern computing resource scheduling and trading platforms. This method can perform scale analysis in real time based on changes in business needs and the state of the computing clusters themselves. This dynamic assessment and adjustment mechanism allows computing clusters to better adapt to constantly changing business environments, improving resource responsiveness and flexibility.
[0169] This method provides the intelligent computing industry with a unified approach and quantitative indicator for analyzing the scale of computing clusters. Within the industry, different enterprises and organizations may employ different evaluation methods, leading to a lack of comparability in the results. This method helps standardize industry evaluation criteria, enabling fair and impartial comparisons of different computing clusters on the same platform, thus promoting standardization and normalization within the industry. Accurate scale analysis results allow different enterprises and organizations to better understand each other's computing resources. Regarding resource sharing and cooperation, enterprises can select suitable partners based on the analysis results to achieve resource complementarity and sharing. For example, a company's computing cluster may have strong scale matching capabilities in one aspect but weak capabilities in others. By collaborating with other companies, it can jointly complete more complex business tasks, improving the overall resource utilization efficiency and competitiveness of the industry.
[0170] Secondly, this application discloses a computing power scheduling method for an intelligent computing power cluster, including:
[0171] The analysis results of the corresponding computing power cluster are obtained based on the scale analysis method of the intelligent computing power cluster disclosed in the first aspect of this application;
[0172] All computing power clusters whose analysis results exceed a preset threshold are considered as currently available computing power resources.
[0173] Different computing clusters exhibit varying performance and availability due to factors such as hardware configuration and operational status. By setting preset thresholds, computing clusters with better analysis results and higher performance and reliability can be selected. These preset thresholds can be dynamically adjusted based on actual circumstances.
[0174] Including multiple computing clusters with good analysis results in the available resource pool can enhance the system's fault tolerance. When a cluster fails, computing tasks can be quickly transferred to other available clusters to avoid task interruption and ensure business continuity.
[0175] Specifically, the matching degree between current computing power and storage capacity and demand, as well as the scalability within the same network cluster (up to three layers), are key factors to consider when scheduling computing power (i.e., scheduling computing tasks to computing power clusters). Based on the analysis results (i.e., the scoring results), computing power clusters with higher scores are selected.
[0176] Thirdly, this application discloses a scale analysis system for intelligent computing power clusters, used to execute the scale analysis method for intelligent computing power clusters disclosed in the first aspect of this application. The system includes:
[0177] The demand information acquisition module is used to acquire the computing power scale requirements of the target business scenario;
[0178] The multi-dimensional feature acquisition module is used to acquire the cluster hardware resource dimension features, network switching device dimension features, and server configuration dimension features of each computing power cluster.
[0179] The rating determination module is used to determine the current computing power rating and current storage rating of each cluster based on demand information and cluster hardware resource characteristics.
[0180] The quantity determination module is used to determine the maximum number of smart computing cards that the parameter plane network can support and the maximum number of servers that the sample plane network can support, based on the dimensional characteristics of network switching equipment and the dimensional characteristics of server configuration.
[0181] The computing power cluster scale matching capability information acquisition module is used to obtain computing power cluster scale matching capability information based on the cluster hardware resource dimension characteristics, the maximum number of intelligent computing cards that the parameter plane network can support, and the maximum number of servers that the sample plane network can support.
[0182] The analysis module is used to determine the scale analysis results of each computing cluster based on the current computing power level, current storage level, and computing power cluster scale matching capability information of each cluster.
[0183] A computer device according to an embodiment of this disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.
[0184] The processor may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of this disclosure, the processor is used to run the computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the aforementioned intelligent computing power cluster scale analysis method or the intelligent computing power cluster computing power scheduling method of the various embodiments of this disclosure.
[0185] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.
[0186] like Figure 4This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 4 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0187] like Figure 4 As shown, a computer device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0188] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 4 A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.
[0189] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the intelligent computing power cluster scaling analysis method or the intelligent computing power cluster computing power scheduling method of the embodiments of this disclosure are performed.
[0190] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.
[0191] According to embodiments of the present disclosure, a computer-readable storage medium stores non-transitory computer-readable instructions. When these non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the aforementioned intelligent computing power cluster scale analysis method or the intelligent computing power cluster computing power scheduling method of the various embodiments of the present disclosure are performed.
[0192] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).
[0193] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.
[0194] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0195] In this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, devices, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc., are open-ended terms meaning "including but not limited to," and are used interchangeably with them. The terms "or" and "and" as used herein refer to the terms "and / or," and are used interchangeably with them unless the context clearly indicates otherwise. The term "such as" as used herein refers to the phrase "such as but not limited to," and is used interchangeably with it.
[0196] Additionally, as used herein, the “or” used in a list of items beginning with “at least one” indicates a separate list, such that a list of, for example, “at least one of A, B, or C” means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word “exemplary” does not imply that the described example is preferred or better than other examples.
[0197] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.
[0198] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.
[0199] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0200] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A method for analyzing the scale of an intelligent computing power cluster, characterized in that, include: Obtain information on the computing power requirements of the target business scenario; Obtain the cluster hardware resource dimension characteristics, network switching device dimension characteristics, and server configuration dimension characteristics of each computing power cluster; Based on the demand information and the characteristics of the cluster hardware resources, determine the current computing power level and current storage level of each cluster; Based on the network switching device dimension characteristics and the server configuration dimension characteristics, the maximum number of smart computing cards that the parameter plane network can support and the maximum number of servers that the sample plane network can support are determined. Based on the cluster hardware resource dimension characteristics, the maximum number of smart computing cards that the parameter plane network can support, and the maximum number of servers that the sample plane network can support, information on the computing power cluster scale matching capability is obtained. Based on the current computing power level, current storage level, and computing power cluster size matching capability information of each cluster, the scale analysis results of each computing power cluster are determined.
2. The method for analyzing the scale of intelligent computing power clusters according to claim 1, characterized in that, The step of determining the current computing power level and current storage level of each cluster based on the demand information and the cluster hardware resource dimension characteristics includes: Determine the current computing power available for each computing cluster; wherein, the current computing power F is... c :F c =P u ×N G , where P u N represents the current computing power of a single smart computing card. G This represents the number of GPU cards currently connected to the parameter panel. The current computing power level of each cluster is determined based on the current computing power requirement and the existing computing power. Based on the available storage capacity of a single storage server and the number of storage servers currently connected to the sample plane, determine the current storage capacity of each computing cluster; wherein, the current available storage capacity is T. c :T c =B u ×N S2 B u N represents the current available storage capacity of a single storage server. S2 This represents the number of storage servers currently connected to the sample surface. The current storage level of each cluster is determined based on the current storage demand and the existing storage capacity.
3. The method for analyzing the scale of intelligent computing power clusters according to claim 2, characterized in that, The step of determining the maximum number of smart computing cards that the parameter plane network can support based on the network switching device dimensional characteristics and the server configuration dimensional characteristics includes: Based on the dimensional characteristics of the network switching device and the first preset formula, the maximum number of smart computing cards that can be accessed in the current state of the parameter plane network is obtained. Based on the maximum number of smart computing cards that the parameter plane network can support, the dimensional characteristics of the network switching equipment, the dimensional characteristics of the server configuration, and the second preset formula, the maximum number of smart computing cards that can be connected to the parameter plane network without moving the Spine layer but only expanding the Tor / Leaf layer is obtained. Based on the network switching device dimensional characteristics, the server configuration dimensional characteristics, and the third preset formula, the maximum capacity of the smart computing card that can be accessed by the network devices at the Spine layer and below is obtained while keeping the Super Spine layer unchanged.
4. The method for analyzing the scale of intelligent computing power clusters according to claim 3, characterized in that, The step of determining the maximum number of servers that the sample plane network can support based on the network switching device dimensional characteristics and the server configuration dimensional characteristics includes: Based on the network switching device dimension characteristics, the server configuration dimension characteristics, and the fourth preset formula, the maximum number of servers that can be accessed in the current state of the sample network is obtained. Based on the maximum number of servers that can be connected to the current sample network, the dimensional characteristics of the network switching equipment, the dimensional characteristics of the server configuration, and the fifth preset formula, the maximum number of servers that can be connected to the sample network without moving the Spine layer but only expanding the Tor / Leaf layer is obtained. Based on the maximum number of servers that can be accessed in the current state of the sample network, the dimensional characteristics of the network switching devices, the dimensional characteristics of the server configuration, and the sixth preset formula, the maximum number of servers that can be accessed in the sample network without changing the Super Spine layer, but only by expanding the Spine layer and below network devices.
5. The method for analyzing the scale of intelligent computing power clusters according to claim 4, characterized in that, The maximum number of smart computing cards that can be connected to the current parameter plane network is a: τ=H(L)·[1+H(S)·(1+H(C))]; The parameter plane network, without changing the Spine layer, only expands the Tor / Leaf layer, and the maximum number of smart computing cards that can be connected is b: The parameter plane network remains unchanged at the Super Spine layer; only the capacity of network devices at the Spine layer and below is expanded. The maximum capacity of the smart computing card that can be connected is c: The maximum number of servers that can be connected to the current sample network is d: The sample network, without changing the Spine layer but only expanding the Tor / Leaf layer, can connect to a maximum of e servers: The sample network remains unchanged at the Super Spine layer, only the Spine layer and below network devices are expanded. The maximum number of servers that can be connected is f: Where L represents the number of TOR / Leaf switches, S represents the number of Spine switches, C represents the number of Super Spine switches, and P represents the number of Super Spine switches. L For the number of ports on a single TOR / Leaf switch, P S For the number of ports on a single Spine switch, P C α1 is the number of ports on a single Super Spine switch, α2 is the single port rate on the access side of the TOR / Leaf switch, α3 is the single port rate on the server side, β1 is the number of ports on each server used to connect to the parameter plane, β2 is the number of ports on each server used to connect to the sample plane, Γ is the number of smart computing cards on each server, H() is the step function, τ is the network level, η(τ) is the first network level constant, and γ(τ) is the second network level constant.
6. The method for analyzing the scale of intelligent computing power clusters according to claim 5, characterized in that, The process of obtaining computing power cluster scale matching capability information based on the cluster hardware resource dimension characteristics, the maximum number of intelligent computing cards supported by the parameter plane network, and the maximum number of servers supported by the sample plane network includes: Based on the cluster hardware resource dimension characteristics and the demand information, the current overall resource supply capacity characteristics are obtained, wherein the current overall resource supply capacity characteristics include the current computing power and the current storage capacity. Based on the demand information, obtain the computing power demand characteristics within a preset period; Based on the demand information, obtain the storage demand characteristics within a preset period; Based on the aforementioned demand information, obtain the characteristics of the smart computing card expansion rate demand; Based on the cluster hardware resource dimension characteristics and the maximum number of smart computing cards that the parameter plane network can support, the current network smart computing card expansion ratio characteristics are obtained. Based on the aforementioned demand information, obtain the server expansion rate demand characteristics; Based on the cluster hardware resource dimension characteristics and the maximum number of servers that the sample plane network can support, the current network server expansion ratio characteristics are obtained. Based on the cluster hardware resource dimension characteristics and the maximum number of servers that the sample plane network can support, obtain the expansion ratio of intelligent computing cards that can be expanded by only expanding the Tor / Leaf layer network; Based on the cluster hardware resource dimension characteristics and the maximum number of servers that the sample plane network can support, obtain the expansion ratio that can be achieved by only expanding the Tor / Leaf layer network servers. Based on the cluster hardware resource dimension characteristics and the maximum number of servers that the sample plane network can support, the expansion ratio of intelligent computing cards that can be expanded only at the Spine layer and below is obtained. Based on the cluster hardware resource dimension characteristics and the maximum number of servers that the sample network can support, the expansion ratio of only the Spine layer and below network servers can be obtained.
7. The method for analyzing the scale of intelligent computing power clusters according to claim 6, characterized in that, The computing power demand characteristic within the preset period is Q1: Q1 = F f +F u , of which F f F represents the computing power expansion needs over the next three years. u This represents the current computing power requirement. The storage demand characteristic within the preset period is Q2: Q2 = T f +T u , among which, T f T represents the storage expansion demand over the next three years. u This represents the current storage requirement. The smart card expansion rate requirement characteristic is k. g : Among them, P f To expand the computing power of a single smart computing card, P u This represents the current computing power of a single smart computing card. The current network smart computing card's expandability ratio is characterized by k. a : Where, N G 'a' represents the number of GPU cards currently connected to the parameter plane, and 'a' represents the maximum number of intelligent computing cards that can be connected to the parameter plane network at its current state. The server expansion rate requirement is characterized by k. s : Where, N f To determine the number of smart computing cards that can be installed on a single smart computing server to be expanded, B f N represents the available storage capacity of a single storage server to be expanded. u B represents the number of smart computing cards installed on a single smart computing server. u This represents the current available storage capacity of a single storage server. The current network server's expandability ratio is characterized by k. d : Where, N S d represents the number of servers currently connected to the sample surface, and d represents the maximum number of servers that the sample surface network can connect to. The expansion ratio of the Tor / Leaf layer network smart computing card is k. b : Where b represents the maximum number of smart computing cards that can be connected to the parameter plane network without moving the Spine layer and only expanding the Tor / Leaf layer; The expansion ratio for the Tor / Leaf layer network server is k. e : Where e represents the maximum number of servers that can be connected to by the sample plane network without moving the Spine layer but only expanding the Tor / Leaf layer; The expansion ratio for network intelligent computing cards at only the Spine layer and below is k. c : Wherein, c represents the maximum capacity of the smart computing card that can be accessed by network devices at the Spine layer and below, without changing the Super Spine layer; The expansion ratio for only Spine layer and below network servers is k. f : For the sample network described in f, the Super Spine layer remains unchanged; only the Spine layer and below network devices are expanded to accommodate a maximum number of servers.
8. The method for analyzing the scale of intelligent computing power clusters according to claim 7, characterized in that, The step of determining the scale analysis results for each computing cluster based on its current computing power level, current storage level, and computing cluster scale matching capability information includes: Based on the current computing power level, current storage level, and computing power cluster size matching capability information of each cluster, determine the current overall resource capability characteristics, current computing power expansion capability characteristics, current storage expansion capability characteristics, current overall resource expansion capability characteristics, Tor / Leaf layer network expansion characteristics, Spine layer and below network expansion characteristics, and network architecture expansion characteristics corresponding to each cluster. When the overall capability characteristics of the current resources meet the first preset condition, the first preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster; When the current overall resource capability characteristics do not meet the first preset condition and the current computing power expansion capability characteristics meet the second preset condition, the second preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster. When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics meet the second preset condition, and the current storage expansion capability characteristics meet the third preset condition, the third preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster. When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, and the current overall resource expansion capability characteristics meet the fourth preset condition, the fourth preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster. When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, the current overall resource expansion capability characteristics do not meet the fourth preset condition, and the Tor / Leaf layer network expansion characteristics meet the fifth preset condition, the fifth preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster. When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, the current overall resource expansion capability characteristics do not meet the fourth preset condition, the Tor / Leaf layer network expansion characteristics do not meet the fifth preset condition, and the Spine layer and below network expansion characteristics meet the sixth preset condition, the sixth preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster. When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, the current overall resource expansion capability characteristics do not meet the fourth preset condition, the Tor / Leaf layer network expansion characteristics do not meet the fifth preset condition, the Spine layer and below network expansion characteristics do not meet the sixth preset condition, and the network architecture expansion characteristics meet the seventh preset condition, the seventh preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster.
9. The method for analyzing the scale of intelligent computing power clusters according to claim 8, characterized in that, When the overall capability characteristics of the current resources meet the first preset condition, the first preset strategy is invoked to obtain the scale analysis result of the corresponding computing power cluster, including: if the computing power demand characteristics within the preset period are not greater than the currently available computing power and the storage demand characteristics within the preset period are not greater than the currently available storage capacity, the scale of the corresponding computing power cluster is determined to be the first level, and the score corresponding to the first level is obtained as the scale analysis score of the corresponding computing power cluster according to the preset scoring list.
10. The method for analyzing the scale of intelligent computing power clusters according to claim 9, characterized in that, When the overall current resource capability characteristics do not meet the first preset condition and the current computing power expansion capability characteristics meet the second preset condition, the second preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster, including: If the computing power demand characteristics within the preset period are not greater than the currently available computing power and the storage demand characteristics within the preset period are greater than the currently available storage capacity, the corresponding computing power cluster size is determined to be the second level, and the score corresponding to the second level is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
11. The method for analyzing the scale of intelligent computing power clusters according to claim 10, characterized in that, When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics meet the second preset condition, and the current storage expansion capability characteristics meet the third preset condition, the third preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster, including: If the computing power demand characteristics within the preset period are greater than the currently available computing power and the storage demand characteristics within the preset period are not greater than the currently available storage capacity, the corresponding computing power cluster size is determined to be Level 3, and the score corresponding to Level 3 is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
12. The method for analyzing the scale of intelligent computing power clusters according to claim 11, characterized in that, When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, and the current overall resource expansion capability characteristics meet the fourth preset condition, the fourth preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster, including: If the required expansion rate of the smart computing card is not greater than the current network's expandable smart computing card ratio and the required expansion rate of the server is not greater than the current network's expandable server ratio, the corresponding computing power cluster size is determined to be Level 4, and the score corresponding to Level 4 is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
13. The method for analyzing the scale of intelligent computing power clusters according to claim 12, characterized in that, When the overall current resource capacity characteristics do not meet the first preset condition, the current computing power expansion capacity characteristics do not meet the second preset condition, the current storage expansion capacity characteristics do not meet the third preset condition, the overall current resource expansion capacity characteristics do not meet the fourth preset condition, and the Tor / Leaf layer network expansion characteristics meet the fifth preset condition, the fifth preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster, including: If the current network smart computing card expansion ratio is less than the smart computing card expansion rate requirement, the smart computing card expansion rate requirement is not greater than the expansion ratio of the smart computing card that only expands the Tor / Leaf layer network, and the current network server expansion ratio is less than the server expansion rate requirement and the server expansion rate requirement is not greater than the expansion ratio of the server that only expands the Tor / Leaf layer network, then the corresponding computing power cluster size is determined to be Level 5, and the score corresponding to Level 5 is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
14. The method for analyzing the scale of intelligent computing power clusters according to claim 13, characterized in that, When the current overall resource capability characteristics do not meet the first preset condition, the current computing power expansion capability characteristics do not meet the second preset condition, the current storage expansion capability characteristics do not meet the third preset condition, the current overall resource expansion capability characteristics do not meet the fourth preset condition, the Tor / Leaf layer network expansion characteristics do not meet the fifth preset condition, and the Spine layer and below network expansion characteristics meet the sixth preset condition, the sixth preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster, including: If the expansion ratio of the Tor / Leaf layer network smart computing cards is less than the expansion rate requirement of the smart computing cards, the expansion rate requirement of the smart computing cards is not greater than the expansion ratio of the Spine layer and below network smart computing cards, and the expansion ratio of the Tor / Leaf layer network servers is less than the server expansion rate requirement and the server expansion rate requirement is not greater than the expansion ratio of the Spine layer and below network servers, the corresponding computing power cluster size is determined to be Level 6, and the score corresponding to Level 6 is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
15. The method for analyzing the scale of intelligent computing power clusters according to claim 14, characterized in that, When the overall current resource capacity characteristics do not meet the first preset condition, the current computing power expansion capacity characteristics do not meet the second preset condition, the current storage expansion capacity characteristics do not meet the third preset condition, the overall current resource expansion capacity characteristics do not meet the fourth preset condition, the Tor / Leaf layer network expansion characteristics do not meet the fifth preset condition, the Spine layer and below network expansion characteristics do not meet the sixth preset condition, and the network architecture expansion characteristics meet the seventh preset condition, the seventh preset strategy is invoked to obtain the scale analysis results of the corresponding computing power cluster, including: If the expansion ratio of network smart computing cards at the Spine layer and below is less than the expansion rate requirement of the smart computing cards, and the expansion ratio of network servers at the Spine layer and below is less than the expansion rate requirement of the servers, the corresponding computing power cluster scale is determined to be level seven, and the score corresponding to level seven is obtained according to the preset scoring list as the scale analysis score of the corresponding computing power cluster.
16. A method for scheduling computing power in an intelligent computing cluster, characterized in that, include: The analysis results of the corresponding computing power cluster are obtained based on the scale analysis method of the intelligent computing power cluster according to any one of claims 1-15; All computing power clusters whose analysis results are greater than a preset threshold are considered as currently available computing power resources.
17. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the scale analysis method of the intelligent computing power cluster according to any one of claims 1-15 or the computing power scheduling method of the intelligent computing power cluster according to claim 16.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the scale analysis method of the intelligent computing power cluster according to any one of claims 1-15 or the computing power scheduling method of the intelligent computing power cluster according to claim 16.
19. A computer program product comprising computer instructions, characterized in that, When executed by the processor, the computer instruction implements the steps of the scale analysis method for the intelligent computing power cluster as described in any one of claims 1-15 or the computing power scheduling method for the intelligent computing power cluster as described in claim 16.