Computing device deployment method and electronic device
By training a model in a heterogeneous computing device resource pool to learn the relationship between business operation data and performance data, and predicting the optimal combination of computing devices, the problem of unstable model operation in a heterogeneous computing device resource pool is solved, and performance optimization and cost reduction are achieved.
Patent Information
- Application Number
- CN202511150203.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-14
AI Technical Summary
When deploying models in a heterogeneous computing device resource pool, there are problems such as insufficient performance utilization and poor compatibility leading to unstable model operation and unstable business processing. Existing solutions rely on continuous expansion of hardware resources.
By collecting performance data of different services running on different combinations of computing devices, the first model is trained to learn the relationship between service operation data and performance data, predict the optimal combination of computing devices for each service characteristic period, and optimize model deployment to improve performance.
By predicting the best-performing combination of computing devices, the model optimizes business execution on a heterogeneous computing device resource pool, reduces resource idleness, lowers energy consumption, improves overall inference efficiency and business response speed, and reduces software adaptation and hardware maintenance costs.
Smart Images

Figure CN120950260A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of operation and maintenance technology, and more specifically, to a method for deploying computing devices and electronic devices. Background Technology
[0002] In existing technologies, models and inference environments are typically deployed on a computing resource pool comprising multiple computing devices to provide business interfaces such as multimodal computing and LLM (Large Language Model). These computing devices include, for example, GPUs (Graphics Processing Units) and NPUs (Neural Processing Units). Computing resource pools are categorized into homogeneous and heterogeneous pools. Heterogeneous pools are favored due to their ability to integrate the advantages of different types of computing devices (such as the general computing power of GPUs and the AI inference optimization capabilities of NPUs), balancing the diverse performance, accuracy, and cost requirements of various business scenarios. Furthermore, they can improve overall resource utilization through hybrid deployment, making them increasingly popular in the context of increasingly complex intelligent computing needs across all scenarios.
[0003] However, when deploying the model on a heterogeneous computing device resource pool, due to the different R&D paths of computing devices from different manufacturers, the significant differences in the performance, accuracy, and technological maturity of different chips, it is difficult to meet the needs of full-scenario applications. Furthermore, computing devices from different manufacturers usually rely on the same preset hardware environment rules during the inference process, which can easily lead to problems such as some computing devices having idle resources while others experience bottlenecks and unstable business operations. Currently, the only solution is to continuously expand hardware resources.
[0004] Therefore, it is necessary to address issues such as instability in model operation and business processing caused by insufficient performance utilization and poor compatibility of existing computing devices when deploying models on heterogeneous computing device resource pools.
[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this disclosure is to provide a computing device deployment method and electronic device to solve problems such as unstable model operation and unstable business processing when performing inference business through a heterogeneous computing device resource pool.
[0007] According to a first aspect of the present disclosure, a method for deploying computing devices is provided, comprising: determining P combinations of computing devices corresponding to N computing devices, wherein the N computing devices are not identical, N>1, P>1; collecting performance data corresponding to service operation data when the computing devices process multiple services in different combinations of computing devices; training a first model based on the performance data of the N computing devices corresponding to the service operation data; determining at least one service characteristic period within a target strategy period based on historical data, and predicted service operation data corresponding to each of the service characteristic periods; inputting the predicted service operation data into the first model to determine predicted performance data of the computing devices corresponding to the predicted service operation data in different combinations of computing devices, and determining the computing device combination corresponding to each of the service characteristic periods based on the predicted performance data.
[0008] In one exemplary embodiment of this disclosure, the P computing device combinations include a first computing device combination in operation and / or a second computing device combination not in operation. The performance data collected when the computing devices process multiple services in different computing device combinations, corresponding to the service operation data, includes: collecting performance data corresponding to the computing devices and service operation data when M1 online services are running on the first computing device combination, where M1 ≥ 0; collecting performance data corresponding to the computing devices and service operation data when M1 online services are running on the second computing device combination; and collecting performance data corresponding to the computing devices and service operation data when M2 offline services are running on the P computing device combinations, where M2 ≥ 0.
[0009] In one exemplary embodiment of this disclosure, when the computing device processes multiple services in different combinations of computing devices, the performance data corresponding to the service operation data includes: at a sampling time point, forming operation characteristic data of the computing device at the sampling time point based on the service operation data and the performance data. The operation characteristic data includes the sampling time point, the number of the computing device, the computing device combination in which the computing device is located at the sampling time point, the service characteristics of the service processed by the computing device combination in which the computing device is located at the sampling time point, the service data of the service processed by the computing device combination in which the computing device is located at the sampling time point, the hardware performance parameters of the computing device at the sampling time point, and the hardware specifications of the computing device. The service characteristics include the type of target model deployed on the computing device combination, the model complexity, and the type of service performed by the target model. The type of service includes classification and inference. The service data includes at least one data type, data volume, and number of data dimensions.
[0010] In one exemplary embodiment of this disclosure, determining at least one business characteristic period within a target strategy period, and the predicted business operation data corresponding to each business characteristic period, based on historical data includes: continuously collecting measured business operation data corresponding to the operation of M1 types of online services on a first combination of computing devices in operation, to form the historical data; training a second model using the historical data, so that the second model learns the correlation between the measured business operation data and time; inputting the start and end times of the target strategy period into the second model, and obtaining at least one business characteristic period within the target strategy period, and the predicted business operation data corresponding to each business characteristic period, output by the second model.
[0011] In one exemplary embodiment of this disclosure, the second model and the first model are the same model.
[0012] In an exemplary embodiment of this disclosure, determining the computing device combination corresponding to each service characteristic time period based on the predicted performance data includes: for a target service characteristic time period, determining P1 computing device combinations where any predicted performance data does not meet a preset condition under the predicted service operation data corresponding to the target service characteristic time period, where P1≥0; determining the unavailable computing devices corresponding to the target service characteristic time period, and P2 computing device combinations to which the unavailable computing devices belong, where P2≥0; deduplicating and merging the P1 computing device combinations and the P2 computing device combinations to obtain P3 unavailable computing device combinations, deleting the P3 unavailable computing device combinations from the P computing device combinations to obtain P-P3 available computing device combinations, where P3≥0; determining the total score of the predicted performance data of the computing devices corresponding to each available computing device combination, and determining the available computing device combination with the highest total score as the computing device combination corresponding to the target service characteristic time period.
[0013] In one exemplary embodiment of this disclosure, determining the total score of the predicted performance data of the computing devices corresponding to each available computing device combination includes: determining the target predicted performance data of a target computing device in a target available computing device combination; determining the performance score and accuracy score of the computing device in the target service characteristic period based on the target predicted performance data; determining a first weight corresponding to the performance score and a second weight corresponding to the accuracy score based on the predicted service operation data corresponding to the target service characteristic period; determining the score of the predicted performance data of the target computing device based on the sum of the products of the first weight and the performance score and the second weight and the accuracy score; and determining the total score corresponding to the target available computing device combination based on the sum of the scores of the predicted performance data of the computing devices corresponding to the target available computing device combination.
[0014] According to a second aspect of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the method as described in any one of the preceding embodiments based on instructions stored in the memory.
[0015] According to a third aspect of this disclosure, a computer-readable storage medium is provided having a program stored thereon that, when executed by a processor, implements the computing device deployment method as described in any of the preceding claims.
[0016] According to a fourth aspect of this disclosure, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of the method as described in any of the preceding claims.
[0017] This embodiment of the disclosure collects performance data of computing devices running different services on different combinations of computing devices, and then trains a first model. This enables the first model to learn the correspondence between service operation data, computing device combinations and the performance data of each computing device. Subsequently, it can predict the performance of each computing device combination under the predicted service operation data corresponding to each service characteristic period within the target strategy period, thereby determining the best-performing computing device combination for each service characteristic period in advance. By adjusting the deployment of computing devices according to the service characteristic period, the performance of the model deployed on the heterogeneous computing device resource pool when executing various services is optimized.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0020] Figure 1 This is a flowchart of a computing device deployment method in an exemplary embodiment of this disclosure.
[0021] Figure 2 This is a sub-flowchart of step S2 in an exemplary embodiment of this disclosure.
[0022] Figure 3 This is a schematic diagram illustrating the determination of predicted business operation data for each business characteristic time period in an exemplary embodiment of this disclosure.
[0023] Figure 4This is a schematic diagram illustrating the relationship between the first model and the second model in an exemplary embodiment of this disclosure.
[0024] Figure 5 This is a sub-flowchart of step S5 in an exemplary embodiment of this disclosure.
[0025] Figure 6 This is a sub-flowchart of step S54 in an exemplary embodiment of this disclosure.
[0026] Figure 7 This is a schematic diagram of an application scenario in an exemplary embodiment of this disclosure.
[0027] Figure 8 This is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0028] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0029] Furthermore, the accompanying drawings are merely illustrative of this disclosure, and the same reference numerals in the drawings denote the same or similar parts, thus repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0030] The exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0031] Figure 1 This is a flowchart of a computing device deployment method in an exemplary embodiment of this disclosure.
[0032] refer to Figure 1 The computing device deployment method 100 may include:
[0033] Step S1: Determine P combinations of computing devices corresponding to N computing devices, wherein the N computing devices are not completely identical, N>1, P>1;
[0034] Step S2: Collect performance data corresponding to the service operation data when the computing device processes multiple services in different combinations of the computing device;
[0035] Step S3: Train the first model based on the performance data of the N computing devices corresponding to the service operation data;
[0036] Step S4: Determine at least one business characteristic period within the target strategy period based on historical data, and the predicted business operation data corresponding to each business characteristic period;
[0037] Step S5: Input the predicted service operation data into the first model to determine the prediction performance data of the computing device in different combinations of computing devices corresponding to the predicted service operation data, and determine the combination of computing devices corresponding to each of the service characteristic time periods based on the prediction performance data.
[0038] This embodiment of the disclosure collects performance data of computing devices running different services on different combinations of computing devices, and then trains a first model. This enables the first model to learn the correspondence between service operation data, computing device combinations and the performance data of each computing device. Subsequently, it can predict the performance of each computing device combination under the predicted service operation data corresponding to each service characteristic period within the target strategy period, thereby determining the best-performing computing device combination for each service characteristic period in advance. By adjusting the deployment of computing devices according to the service characteristic period, the performance of the model deployed on the heterogeneous computing device resource pool when executing various services is optimized.
[0039] The following is a detailed description of each step of the computing device deployment method 100.
[0040] In step S1, P combinations of computing devices are determined corresponding to N computing devices. The N computing devices are not completely identical, N>1, P>1.
[0041] In this embodiment of the disclosure, N computing devices refer to all computing devices in the currently available computing device resource pool, and these N computing devices are heterogeneous computing devices. In the field of computer hardware and computing, "heterogeneous" refers to a system or cluster composed of components / devices with different architectures, models, manufacturers, or technical standards. These devices differ significantly in hardware characteristics (such as computing power, precision, and video memory) and software ecosystems (such as drivers, interfaces, and adaptation frameworks), and are not completely unified or compatible.
[0042] For example, a cluster composed of computing hardware from different manufacturers or of different types, such as NVIDIA GPUs (Graphics Processing Units), Huawei Ascend NPUs (Neural Processing Units), and Cambricon Siyuan chips, can be called a "heterogeneous computing cluster." In contrast, "homogeneous" refers to components with the same architecture, model, or manufacturer, and with highly unified characteristics and ecosystem.
[0043] In step S1, from the N not entirely identical computing devices (i.e., heterogeneous devices), all permutations and combinations of computing devices are traversed, and all possible valid combinations (a total of P) are selected. The P valid combinations include a first computing device combination and a second computing device combination. The "first computing device combination" refers to a combination containing computing devices that are already in operation, and the "second computing device combination" refers to other computing device combinations that are not yet in operation. Generally, the computing device combinations in operation will include all computing devices in the computing device resource pool; that is, the first computing device combination contains N computing devices, and the second computing device combination contains 1 to N-1 devices. However, in some cases, the number of computing devices in the first computing device combination may be less than the number of computing devices in the second computing device combination, for example, when several new computing devices are added to the computing resource pool, and these additional computing devices are not yet in operation.
[0044] The screening process requires iterating through all theoretical device combinations, but combinations that are completely incompatible with the ecosystem (such as hardware interfaces or software drivers that cannot work together) must be excluded. The effectiveness of the combinations is ensured through a mutual exclusion strategy (explicitly stipulating that incompatible devices cannot be in the same group).
[0045] Assume N=3, and the 3 computing devices are as follows:
[0046] Device A: Model A GPU (supports CUDA, compatible with LLM framework);
[0047] Device B: Model B NPU (supports CANN, compatible with domestic multimodal framework);
[0048] Device C: C-type GPU (supports proprietary drivers, only compatible with specific image processing frameworks).
[0049] Theoretical combinations include:
[0050] First computing device combination: {A, B, C} (all 3 units included);
[0051] Secondary computing device combinations: {A,B}, {A,C}, {B,C} (2-unit combination); {A}, {B}, {C} (1-unit combination), for a total of 6 secondary computing device combinations.
[0052] Upon testing, it was found that the software ecosystems of devices B and C are completely incompatible (driver interface conflicts, unable to handle any business collaboratively). Therefore, the combination {B,C} was excluded through a mutual exclusion strategy.
[0053] The final P = 1 (first computing device combination) + 5 (effective second computing device combination) = 6, which are as follows:
[0054] First computing device combination: {A,B,C};
[0055] The second computing device combination is: {A,B}, {A,C}, {A}, {B}, {C}.
[0056] Therefore, all P combinations of computing devices obtained are combinations of computing devices that are in normal working order.
[0057] In this embodiment of the disclosure, by finding all reasonable and incompatible effective combinations of computing devices in the current heterogeneous computing device resource pool in step S1, and then predicting the performance of each effective computing device in each business scenario in subsequent steps, it is possible to select the best-performing combination of computing devices to deploy the model and execute inference business for different business scenarios (business characteristic periods), instead of using all computing devices in the heterogeneous computing device resource pool to deploy the model and execute inference business in all scenarios.
[0058] The inventors of this disclosure discovered in their research that, in some scenarios, using a combination of computing devices that only includes a subset of the computing devices to perform inference tasks can significantly outperform using a combination of all computing devices from a heterogeneous computing device resource pool. This is because the total computing devices may include those incompatible with the current business scenario (e.g., the business requires FP8 precision calculations, but a device in the combination only supports FP16, increasing data conversion overhead), or devices with weak ecosystem compatibility (e.g., communication latency between devices from different vendors is much higher than the business calculation time). These devices can become the "bottleneck" in overall performance. By excluding incompatible or poorly compatible computing devices, a partial combination can reduce cross-device communication losses, avoid redundant resource scheduling, and make the characteristics of the computing devices within the combination (e.g., computing power, precision, software stack) more aligned with business needs (e.g., for image processing tasks, only GPUs that excel at image processing are retained, without including NPUs that excel at text inference but have poor adaptability), thereby improving overall inference efficiency.
[0059] Furthermore, using a combination of computing devices to perform inference tasks avoids idle devices consuming computing power and memory, reducing energy consumption—especially when the workload is low, as it does not require starting all devices to meet the demand. At the same time, it reduces the potential risks associated with inter-device collaboration (such as a single device failure causing a complete interruption), and the impact on the business is smaller when a single device malfunctions. It also allows for rapid switching of combinations based on real-time business needs (e.g., using small, high-efficiency combinations for short-term, high-frequency tasks, and highly adaptable combinations for complex tasks), improving response speed. Moreover, it eliminates the need for additional compatibility development to adapt to the full combination of devices, which can greatly reduce software adaptation and hardware maintenance costs.
[0060] In step S2, performance data corresponding to the service operation data is collected when the computing device processes multiple services in different combinations of the computing device.
[0061] The computing device combination refers to the effective combination determined in step S1.
[0062] In an exemplary embodiment, in step S2, a computing device combination can be selected, and one or more target models can be deployed on the computing device combination to process various different services.
[0063] For example, in a computing device combination including GPU A, NPU B, and GPU C, a "Large Language Model (LLM)" (primarily utilizing the computing power of GPUs A and C) and an "Image Generation Model" (primarily utilizing the dedicated acceleration unit of NPUB) can be deployed simultaneously. Different models can choose suitable devices based on their own characteristics (e.g., LLM relies on the high computing power of GPUs, while lightweight classification models can be deployed on low-power NPUs), and the resource pool avoids resource conflicts between models through unified scheduling. Furthermore, since models can be broken down into independent components (e.g., input processing, core inference, and output processing modules), the same computing device can simultaneously support compatible components of different models. For example, computing device A can simultaneously deploy the "attention computing component" (relying on Tensor Core acceleration) of an LLM model and the "feature extraction component" (relying on the parallel computing capabilities of CUDA cores) of an image classification model. As long as the hardware requirements of the two components (e.g., accuracy support, memory usage) do not conflict, the computing device can run these two components simultaneously through time slicing or parallel computing resource allocation (e.g., the multi-streaming mechanism of GPUs).
[0064] Whether it's multi-model deployment or component-shared computing devices, the computing power, precision, and memory of the computing devices must be able to support the needs of multiple models / components simultaneously. Furthermore, virtualization technology or resource slicing (such as the MIG function of GPUs) must be used to avoid resource contention between different models / components. In addition, the drivers and frameworks of the computing devices must be compatible with components of different models (such as supporting operators of PyTorch and TensorFlow at the same time).
[0065] The matching relationship between the computing device combination and the target model can be set according to business needs, and no special restrictions are imposed here.
[0066] Next, the target model is run and various business processes (which can be actual business processes or test cases) are processed, collecting relevant data from both the business and computing device levels.
[0067] From a business perspective, business operation data can include business characteristics, business types, and business data. Business characteristics can include the type of the target model deployed (e.g., LLM, image classification model), model complexity (e.g., parameter size, number of layers), and business type (e.g., classification, inference). Business type refers to the type of task performed by the target model, such as classification, inference, etc. Business data refers to the attributes of the input data processed by the model, such as data type (text, image, video), data volume (e.g., 100 texts, 50 images, 200 tokens), and data dimensions (e.g., text length, image resolution).
[0068] At the computing device level, performance data can include static performance data and dynamic performance data. Static performance data includes, for example, hardware data of the computing device (such as chip model, memory specifications, supported precision, etc.), while dynamic performance data refers to the hardware performance parameters of the computing device when processing the data of this business operation (such as throughput with time stamps, inference latency, accuracy, bandwidth utilization, computing power utilization, memory occupancy, power consumption, etc.).
[0069] In an exemplary embodiment, data can be collected from each computing device in the currently selected computing device combination at the sampling time point. Based on the collected service operation data and performance data, operational characteristic data of a computing device at the sampling time point is formed. That is, the operational characteristic data is comprehensive data collected at the sampling time point describing the operating status of the computing device, including sampling time, device number, combination, characteristics and data of the processed services, hardware performance parameters and specifications, etc. This operational characteristic data can be, for example, a feature vector, which combines the service operation data and performance data into a unified feature vector for the time series and then performs normalization processing.
[0070] For example, the operational characteristic data of a computing device collected at a sampling time point may include: the sampling time point (time stamp), the number of the computing device, the computing device combination in which the computing device is located at the sampling time point, the business characteristics of the business processed by the computing device combination in which the computing device is located at the sampling time point, the business data of the business processed by the computing device combination in which the computing device is located at the sampling time point, the hardware performance parameters of the computing device at the sampling time point, and the hardware specifications of the computing device; wherein, the business characteristics include the type of target model deployed on the computing device combination, the model complexity, and the type of business performed by the target model, and the type of business includes classification and inference; the business data includes at least one type of data, data volume, and number of data dimensions.
[0071] Data collection can be divided into two types: data collection of actual operating data from systems that are already in operation (online), and data collection of a test nature from systems that are not in operation or have ceased operation.
[0072] Figure 2 This is a sub-flowchart of step S2 in an exemplary embodiment of this disclosure.
[0073] refer to Figure 2 In an exemplary embodiment, the effective P computing device combinations can be divided into a first computing device combination that is operational and a second computing device combination that is not operational. Operational refers to having deployed at least one target model and processing business through that target model. Step S2 may include:
[0074] Step S21: Collect the performance data corresponding to the computing device and service operation data when the M1 types of services that have been launched are running on the first computing device combination, M1≥0;
[0075] Step S22: Collect performance data corresponding to the computing device and service operation data when the M1 types of services that have been launched are running on the second computing device combination;
[0076] Step S23: Collect the performance data corresponding to the computing devices and service operation data when the M2 types of services that are not online are running on P combinations of computing devices, where M2≥0.
[0077] The data collection for actual operation, i.e., step S21, is based on the first combination of computing devices put into operation, generally based on all N computing devices (existing deployment methods usually utilize all computing devices in the computing device resource pool). Let the number of online business types be M1. These M1 business types include different task types corresponding to the same target model, the same task types corresponding to different target models, and different task types corresponding to different target models. That is, "model-business type" is considered as one business.
[0078] In some embodiments, there may be no online services or no actual first computing device combination in operation, in which case step S21 may not be performed.
[0079] In an exemplary embodiment, the test-oriented data collection may include steps S22 and S23. Step S22 is used to determine the performance of the M1 type of online service when running on different computing device combinations in a pool of non-operational computing devices. If there is no M1 type of online service, step S22 is not executed. Step S23 is used to determine the performance of the M2 type of offline service when running on all P computing device combinations in a pool of different computing devices. The M2 type of offline service is also categorized according to the logic of treating "model-service type" as a single service. If there is no M2 type of offline service, step S23 may not be executed. However, at least one of steps S21, S22, and S23 must be executed.
[0080] pass Figure 2 The illustrated embodiment can comprehensively collect data from all available computing devices and be compatible with all services. Examples are given below.
[0081] Assuming scenario N=3, the heterogeneous computing device resource pool includes three heterogeneous computing devices: device A (excels in general computing power); device B (excels in AI inference acceleration); and device C (excels in image processing). The effective combinations determined in step S1 include:
[0082] The first combination of computing devices actually put into operation: {A,B,C};
[0083] The second computing unit combination that is not yet in operation consists of: {A,B}, {A,C}, {A}, {B}, and {C} (a total of 5 units).
[0084] When collecting data, we can first collect the performance data of each computing device when running the online services on the first computing device combination. Assuming that Model 1 and Model 2 are deployed on the first computing device combination, Model 1 provides text generation and text translation services, and Model 2 provides image classification and object detection services, then the online services are M1=4, and the service scenarios include four scenarios: Model 1-text generation, Model 1-text translation, Model 2-image classification, and Model 2-object detection services.
[0085] At each sampling time point (e.g., every 5 minutes, with the current sampling time being 10:00), record the service operation data on the first computing device combination and the performance data of each computing device:
[0086] Business characteristics (model type: large language model (parameter scale of 100 billion), high complexity; business type: generation (text continuation, creation));
[0087] Business data (Data type: plain text; Data volume: 500 tokens / minute (each text is a short sentence entered by the user); Data dimension: single text length ≤ 200 tokens):
[0088] Device A: 85% computing power utilization, 70GB of video memory usage, and 300ms inference latency (mainly the attention calculation component for Model 1);
[0089] Device B: Computing power utilization rate 60%, power consumption 120W (input encoding component of auxiliary processing model 1 120W (input encoding component of auxiliary processing model 1);
[0090] Device C: 10% computing power utilization (idle, due to poor text processing capabilities).
[0091] Example of running feature data (or feature vectors):
[0092] “Sampling time 10:00, device A, computing device combination {A,B,C}, business characteristics: model 1 (LLM, 100 billion parameters), business type = generation; business data: text, 500 pieces / minute, 200 tokens; hardware performance: computing power utilization 85%; hardware specifications: NVIDIA A100 GPU.”
[0093] That is, in step S21, by collecting actual operating data and performance data of each computing device at multiple sampling time points (this step can also be called continuous monitoring), the performance data of device A, device B and device C in the above four business scenarios can be sampled.
[0094] Next, we proceed to the testing phase, collecting performance data of the online services (M1=4) on a second set of computing devices that are not actually in operation. Offline testing can be conducted either before a second set of computing devices is put into operation, or after it is put into operation but during maintenance. During testing, input data can be generated by simulating the workload of actual operational data (such as text volume, image resolution).
[0095] For example, a second computing device combination {A,B} is specifically tested to collect performance data of computing device A in the above four business scenarios, as well as the corresponding performance data of computing device B in the above four business scenarios.
[0096] Since the deployment method of the model is generally fixed for the same combination of computing devices, only the type of business of the input data needs to be changed to perform the test step S22.
[0097] Finally, in step S23, performance data of the offline services (M2=1 type) are collected on all P valid computing device combinations.
[0098] Among them, the unlaunched business can include two types: old model-new business type and new model-new business type.
[0099] For the old model-new service type, assuming the unlaunched service is Model 2 - Video Generation, Model 1 and Model 2 can be deployed simultaneously on a single computing device combination. Test data for the video generation service is input, and performance data for each computing device in that combination is collected. Then, the computing device combination is changed, and Model 1 and Model 2 are deployed simultaneously on the new combination again for testing. This testing can also be combined with the testing in step S22, i.e., after testing Model 2 - Object Detection, Model 2 - Video Generation is tested directly to save time.
[0100] For the new model - new business type, assume that the business not yet online is model 3 - video generation (business characteristics: model type is StableDiffusion, high complexity, business type is generation; business data: video type, resolution 1080P, single segment duration 10s, data volume 100 segments / hour).
[0101] In the test environment, Model 3 can be deployed to 6 valid combinations (1 first computing device combination + 5 second computing device combinations), input test data of video generation service, and record the performance data of each device (such as combination {B,C} which has been excluded due to ecosystem incompatibility and does not need to be tested).
[0102] In some embodiments, if it is necessary to test a scenario where Model 3 is deployed simultaneously with Model 1 and Model 2 on a computing device combination, it can be deployed directly, and then the test data of the video generation service can be used for testing. However, since this deployment change also affects the operation of Model 1 and Model 2 on each computing device, if it is necessary to test a scenario where Model 1, Model 2, and Model 3 are deployed simultaneously on a computing device combination, then in the testing phase of step S23, these three models need to be deployed simultaneously on each computing device combination. Then, five service scenarios—the two services provided by Model 1, the two services provided by Model 2, and the one service provided by Model 3—are tested to obtain the performance data of each computing device in each computing device combination under these five service scenarios.
[0103] That is, when different model combinations need to be tested, the deployed model combinations corresponding to the computing device combination at the sampling time point can be added to the running feature data (for example, the computing device combination is {A,B}, and the deployed model combination is {model1, model2, model3}) to identify the impact of different model combinations on the performance data of each computing device.
[0104] In step S3, a first model is trained based on the performance data of the N computing devices corresponding to the service operation data.
[0105] After comprehensive data collection through monitoring and / or testing, a first model can be trained to learn the correlation between business characteristics, business types, business data, computing device combinations, and the performance data of each computing device.
[0106] In an exemplary embodiment, an LSTM regression prediction model (first model) can be constructed. The model architecture and number of layers are determined according to business and hardware complexity and data volume. The first model is trained using preprocessed time-series feature vectors (i.e., the aforementioned operational feature data). During training, the input data is set to business operational data, and the output data is the performance data of each computing device under each computing device combination.
[0107] The first model can be implemented through various models, and the embodiments disclosed herein do not impose any special limitations.
[0108] In step S4, at least one business characteristic period within the target strategy period is determined based on historical data, as well as the predicted business operation data corresponding to each business characteristic period.
[0109] In an exemplary embodiment, a business characteristic period refers to a period of time with typical business characteristics, such as peak business hours, low business hours, or peak periods for text business or image business.
[0110] Assuming the target strategy period is one day, based on historical business data, the day can be divided into two business characteristic periods: 6:00-20:00 (peak period of mixed multimodal business) and 20:00-6:00 (off-peak period dominated by light text business).
[0111] Assuming the target strategy period is one month, based on historical business data, the month can be divided into three business characteristic periods: weekdays (Monday to Friday), weekends (Saturday to Sunday), and the special period at the end of the month (the last 3 days of the month). During weekdays (Monday to Friday), the overall business volume is stable and relatively high. From 9:00 to 18:00, text-based tasks (translation and generation) from Model 1 dominate (accounting for 60%), while image-based tasks (classification and detection) are concentrated from 10:00 to 15:00 (accounting for 30%). During weekends (Saturday to Sunday), the total business volume decreases by 20% compared to weekdays, but the proportion of image-based tasks increases to 50% (mostly processing user-uploaded leisure scene images), while text-based tasks are mainly short-duration, high-frequency generation tasks. During the special period at the end of the month (the last 3 days of the month), due to data aggregation needs, the target detection business volume of Model 2 surges (increasing by 80% compared to weekdays), and this is mostly high-resolution image data, significantly increasing the demand on equipment computing power.
[0112] After determining the characteristic business periods, the predicted business operation data for each period is determined based on historical data. The predicted business operation data for each period can be determined based on factors such as the historical average for the same period.
[0113] For example, when the target strategy period is one day, predictions can be made based on historical data:
[0114] 6:00-20:00: Model 1's text generation volume is 1200 texts / hour (reaching 1800 texts / hour during the morning peak of 7:00-9:00), and its text translation volume is 800 texts / hour; Model 2's image classification volume is 900 images / hour, and its object detection volume is 600 images / hour, with high resolution (1024×1024 pixels) accounting for 40% of the image data.
[0115] 20:00-6:00: The text generation workload of Model 1 is reduced to 200 texts / hour (single text length ≤ 100 tokens), and the text translation workload is 150 texts / hour; the image classification workload of Model 2 is only 150 images / hour (mostly low resolution 512×512 pixels), and the object detection workload is suspended.
[0116] When the target strategy period is one month, it is assumed that the predicted text generation volume during weekdays is 800 texts / minute and the image classification volume is 500 images / minute; the image detection volume during weekends is 600 images / minute; and the high-resolution image data volume in the last 3 days of the month reaches 1000 images / minute.
[0117] In some embodiments, identifying business characteristic periods and determining the predicted business operation data for each business characteristic period can be done manually by staff. In this embodiment, it can be achieved using a model trained with historical data.
[0118] Figure 3 This is a schematic diagram illustrating the determination of predicted business operation data for each business characteristic time period in an exemplary embodiment of this disclosure.
[0119] refer to Figure 3 In an exemplary embodiment, a second model can be trained to perform step S3. The training process of the second model may include:
[0120] Step S31: Continuously collect the actual service operation data corresponding to the M1 types of services that have been launched and are running on the first combination of computing devices put into operation, so as to form historical data;
[0121] Step S32: Train the second model using historical data so that the second model learns the correlation between measured business operation data and time;
[0122] Step S33: Input the start and end times of the target strategy period into the second model, obtain at least one business feature time period corresponding to the target strategy period output by the second model, and the predicted business operation data corresponding to each business feature time period.
[0123] exist Figure 3 In the embodiment shown, the second model can be directly trained using the operational feature data collected through monitoring in step S21, so that the second model learns the business features within a preset strategy period (one day or one month), and can then divide the business feature time period according to the input target strategy period, and predict the predicted business operation data corresponding to each business feature time period in the target strategy period based on historical data, i.e., the operational feature data collected through monitoring in step S21.
[0124] In an exemplary embodiment, the second model is the same model as the first model, or the second model is implemented through the first model, the second model is set as a sub-module of the first model, or the second model is set as the input terminal of the first model.
[0125] Figure 4 This is a schematic diagram illustrating the relationship between the first model and the second model in an exemplary embodiment of this disclosure.
[0126] refer to Figure 4 In an exemplary embodiment, the second model 42 can be set as the input terminal of the first model 41. Thus, during training, the first model 41 and the second model 42 can be jointly trained by monitoring and collecting operational feature data in step S21, and the first model 41 can be specifically trained by testing and collecting operational feature data in steps S22 and S23.
[0127] When applied, the target strategy period can be directly input into the second model 42, and the second model 42 will automatically output at least one business characteristic period within the target strategy period, as well as the predicted business operation data corresponding to each business characteristic period. The output data is in the form of, for example, "business characteristic period - predicted business operation data".
[0128] After receiving the "business characteristic period - predicted business operation data", the first model 41 executes the subsequent step S5.
[0129] In step S5, the predicted service operation data is input into the first model to determine the prediction performance data of the computing device corresponding to the predicted service operation data in different combinations of computing devices, and the combination of computing devices corresponding to each service characteristic time period is determined based on the prediction performance data.
[0130] After training, the first model can provide performance data for each combination of computing devices under different combinations of input business operation data. For example:
[0131] When the input business operation data is "Model 1 - Text Generation, text volume 1500 pieces / hour, single text length 200 tokens", the first model can output the prediction performance data of each computing device combination:
[0132] Computing device combination {A,B,C}: Device A has a computing power utilization of 85% and a video memory usage of 70GB; Device B has a computing power utilization of 60%; Device C has a computing power utilization of 10% and an inference latency of 300ms.
[0133] Computing device combination {A,B}: Device A has a computing power utilization of 92% and a video memory usage of 75GB; Device B has a computing power utilization of 75% and an inference latency of 280ms (no communication overhead from Device C).
[0134] Computing device combination {A}: Device A has a computing power utilization of 98%, a video memory usage of 80GB, and an inference latency of 350ms (the latency increases due to excessive load on a single device).
[0135] The predicted performance data output by the first model can be output in the format of "Computing Device Combination-Computing Device Number-Performance Data". Referring to the example above, the first model can output a total of 6 predicted performance data in the format of "Computing Device Combination {A,B,C}-Device A-Computing Power Utilization 85%, Graphics Memory Usage 70GB".
[0136] for Figure 4In the illustrated embodiment, after receiving data in the format of "business characteristic period - predicted business operation data" corresponding to x business characteristic periods of the target strategy period output by the second model 42, the first model 41 can output predicted performance data in the format of "business characteristic period - computing device combination - computing device A - performance data", and output predicted performance data corresponding to each computing device combination for different business characteristic periods. Figure 4 Let y be the number of available computing device combinations. Following the example above, assuming that for a set of business operation data, the first model outputs 6 predicted performance data points, while the second model provides three sets of predicted business operation data corresponding to 3 business characteristic time periods, then the first model can output 3 × 6 = 18 predicted performance data points.
[0137] Next, based on the predicted performance data provided by the first model, the optimal combination of computing devices corresponding to the operational data of the business can be determined.
[0138] Figure 5 This is a sub-flowchart of step S5 in an exemplary embodiment of this disclosure.
[0139] refer to Figure 5 In an exemplary embodiment, step S5 may include:
[0140] Step S51: For a target business characteristic period, determine P1 combinations of computing devices where any prediction performance data does not meet the preset conditions under the prediction business operation data corresponding to the target business characteristic period, where P1≥0.
[0141] Step S52: Determine the unavailable computing devices corresponding to the target service characteristic time period, and the P2 combinations of computing devices to which the unavailable computing devices belong, where P2≥0;
[0142] Step S53: After deduplicating P1 computing device combinations and P2 computing device combinations, merge them to obtain P3 unusable computing device combinations. Delete P3 unusable computing device combinations from the P computing device combinations to obtain P-P3 usable computing device combinations, where P3≥0.
[0143] Step S54: Determine the total score of the predicted performance data of the computing devices corresponding to each available computing device combination, and determine the available computing device combination with the highest total score as the computing device combination corresponding to the target business characteristic time period.
[0144] Continuing from the previous example, after the first model outputs 18 prediction performance data points, these 18 data points can be filtered first. If it is found that in the data point "Business characteristic period 1 - Computing device combination {A} - Device A - Computing power utilization 98%, GPU memory usage 80GB, inference latency 350ms", the performance data part "Computing power utilization 98%, GPU memory usage 80GB, inference latency 350ms" does not meet the preset condition (inference latency ≤ 300ms), then the computing device combination {A} is directly listed as an unavailable computing device combination, P1 = 1.
[0145] If, during business characteristic period 1, any one of the predicted performance data of devices A, B, and C in the computing device combination {A,B,C} fails to meet the preset conditions, then the computing device combination {A,B,C} will also be listed as an unavailable computing device combination, and P1 will be incremented by 1.
[0146] Furthermore, if device C is predicted to be unavailable during business characteristic period 1, then the computing device group {A,B,C} containing device C is directly added to the unavailable computing device group, and P2 is incremented by one. In the example above, P2 = 1. The reasons why computing devices may be unavailable during a certain business characteristic period include, but are not limited to, planned maintenance (such as firmware upgrades, hardware repairs), resource usage conflicts (being occupied for a long time by a higher-priority dedicated task), or temporary interruptions in ecosystem compatibility (such as driver crashes that cannot be repaired immediately).
[0147] Predicting the unavailability of a device during a specific business period in the future can be done by analyzing historical failure records and patterns, as well as trends in device operating status. For example, if device C experienced 1-2 instances of memory overflow-induced downtime during peak periods (high-load scenarios) at the end of each month over the past three months, combined with a predicted surge in target detection traffic during those periods, it can be inferred that it faces a high risk of unavailability at the end of this month. Alternatively, by monitoring device C's hardware parameters in real time (e.g., temperature consistently exceeding the 90°C threshold and showing an upward trend during the peak image traffic period from 10:00 to 15:00 for three consecutive days), combined with maintenance records of the aging cooling system, it can be predicted that it may become unavailable due to overheat protection during similar peak periods in the future. Furthermore, predictions can also be made based on planned maintenance schedules (e.g., firmware upgrade times notified in advance by the device manufacturer) or resource reservation plans (e.g., scheduling of the device for other dedicated tasks), ensuring a more comprehensive assessment of unavailability risks.
[0148] Predicting which devices are unavailable during a specific business period can be done manually, by training a third model, or by training a first model. This disclosure does not impose any special limitations on this method.
[0149] Finally, after removing duplicates from P1 and P2 unavailable computing device combinations (computing device combination {A,B,C}), they are merged to determine the final P3 unavailable computing device combinations [computing device combination {A,B,C}, computing device combination {A}], P3 = 2.
[0150] After deleting P3 unavailable computing device combinations from all computing device combinations, the available computing device combinations corresponding to the service characteristic period are obtained, such as the computing device combination {A,B} in the example above.
[0151] Generally, the computing device resource pool contains a large number of computing devices (at least 8), resulting in numerous combinations of computing devices. The final selected usable computing device combinations are usually multiple. The examples described above in this disclosure are for illustrative purposes only and represent simplified setups.
[0152] In an exemplary embodiment, after multiple available computing device combinations are selected for a business characteristic period, it is necessary to evaluate the performance of each computing device combination in order to determine the best performing computing device combination for that business characteristic period.
[0153] Figure 6 This is a sub-flowchart of step S54 in an exemplary embodiment of this disclosure.
[0154] refer to Figure 6 In an exemplary embodiment, step S54 may include:
[0155] Step S541: Determine the target prediction performance data corresponding to a target computing device in a target available computing device combination;
[0156] Step S542: Determine the performance score and accuracy score of the computing device during the target service characteristic period based on the target prediction performance data;
[0157] Step S543: Based on the predicted business operation data corresponding to the target business characteristic time period, determine the first weight corresponding to the performance score and the second weight corresponding to the accuracy score.
[0158] Step S544: Determine the score of the predicted performance data of the target computing device based on the sum of the product of the first weight and the performance score, and the product of the second weight and the accuracy score.
[0159] Step S545: Determine the total score corresponding to the target available computing device combination based on the sum of the scores of the predicted performance data of the computing devices corresponding to the target available computing device combination.
[0160] In this embodiment of the disclosure, the scoring of the predicted performance data of each computing device can be divided into two aspects: performance and accuracy. The final score is obtained by a weighted sum of the two, with the weights adjusted according to the service type. That is, the predicted performance data score = w1 * performance score + w2 * accuracy score, with the first weight w1 and the second weight w2 dynamically adjusted according to the service.
[0161] For example, during business characteristic period 1 (Model 1 - mainly text translation business, predicted business volume 1000 messages / hour), devices A and B in the available combination {A,B} are evaluated:
[0162] Device A: Target prediction performance data: computing power utilization 85%, translation latency 410ms.
[0163] Performance rating: Based on latency and computing load, it scores 80 points (the lower the latency and the more balanced the load, the higher the score);
[0164] Accuracy rating: Translation accuracy rate 95%, score 90 (the higher the accuracy rate, the higher the score);
[0165] Because text translation requires higher accuracy (erroneous translations have a greater impact), the first weight w1 = 0.4 and the second weight w2 = 0.6 are set.
[0166] The score for device A is 0.4 × 80 + 0.6 × 90 = 86 points.
[0167] Device B: Target prediction performance data: rule matching efficiency 85%, output fluency 92 points.
[0168] Performance rating: 75 points based on matching efficiency;
[0169] Accuracy score: Based on fluency (reflecting grammatical accuracy), 88 points;
[0170] For the same business weight, w1 = 0.4, w2 = 0.6;
[0171] The score for device B is 0.4 × 75 + 0.6 × 88 = 82.8 points.
[0172] During periods when image classification is the primary business (where "performance," or processing speed, is more critical), the first and second weights mentioned above can be adjusted to w1 = 0.7 and w2 = 0.3, respectively, to ensure that combinations with higher throughput are prioritized.
[0173] Finally, based on the sum of the scores for device A and device B, the total score corresponding to the target combination of available computing devices is determined. For example, in the above example, the total score for the combination of available computing devices {A,B} is 86 + 82.8 = 168.8 points.
[0174] Assuming that computing device combination {A} (device A only) is also an available computing device combination, then according to the above scoring logic, device A in this computing device combination has a performance score of 70 points (high single device load leads to increased latency) and an accuracy score of 92 points. Calculated with equal weights, this results in a total score of 83.2 points. The total score of this computing device combination is 83.2 points, which is lower than the total score of computing device combination {A,B}. Therefore, combination {A,B} performs better during this business characteristic period.
[0175] By measuring the performance of computing devices through both performance and accuracy, and adjusting the weights of these two aspects in real time according to business needs, a refined assessment of the suitability of computing device combinations can be achieved. This avoids the loss of accuracy caused by a single dimension (such as focusing solely on speed) (e.g., translation services that prioritize speed and result in mistranslations), while also preventing the sacrifice of efficiency for the sake of overemphasizing accuracy (e.g., image classification services that prioritize accuracy and slow down processing speed). At the same time, the dynamic adjustment of weights can match business needs (e.g., reasoning tasks in financial scenarios prioritize accuracy, while short video generation scenarios prioritize performance). The final selected computing device combination can maximize resource utilization efficiency while meeting the bottom line of business quality, thereby improving the overall service adaptability and stability.
[0176] Finally, after determining the computing device combination corresponding to each business characteristic time period within the target strategy period, multiple computing device deployment strategies in the form of "target strategy period - business characteristic time period - computing device combination" can be output, corresponding to the number of business characteristic time periods. This allows the system to automatically switch computing device deployment strategies when it reaches the starting point of the target strategy period and the corresponding business characteristic time period, and allocate the business to the computing device combination corresponding to that business characteristic time period.
[0177] Therefore, the embodiments of this disclosure enable dynamic and scenario-based scheduling of computing device resources: the system does not need to maintain full device operation at all times, but can call the optimal combination of computing devices during peak business periods and switch to a lightweight combination during off-peak periods to save energy; at the same time, it matches dedicated combinations of computing devices for different business characteristics (such as text peaks, image peaks) to avoid efficiency losses caused by resource mismatch. This strategy not only ensures the processing performance and accuracy of business at different times, but also maximizes the utilization rate of device resources, reduces operation and maintenance costs, and ultimately achieves intelligent and efficient operation of heterogeneous computing resource pools.
[0178] The following describes an application scenario of an embodiment of this disclosure.
[0179] In practical applications, embodiments of this disclosure can provide monitoring capabilities, service and performance prediction capabilities, and policy distribution module capabilities for remote heterogeneous GPU clusters. It automatically sends interactive monitoring data between the remote GPU cluster and service requirements to the control terminal for predictive learning, thereby forming scheduling strategies for heterogeneous GPUs under different service requirements.
[0180] Figure 7 This is a schematic diagram of an application scenario in an exemplary embodiment of this disclosure.
[0181] refer to Figure 7 In an exemplary embodiment, the heterogeneous computing device resource pool scheduling system 700 may include a monitoring module 71, a vector processing module 72, a prediction and regression module 73, and a scheduling policy distribution module 74. Wherein:
[0182] Monitoring module 71 provides relatively accurate preliminary performance accuracy data through a small cluster environment. Combined with performance accuracy data from actual business operations, monitoring module 72 extracts standard output business metrics and GPU metrics. Business metrics include data type, data size, and data dimension with time stamps. GPU inference metrics include throughput, latency, accuracy, and bandwidth utilization with time stamps. In other words, monitoring module 71 extracts multi-dimensional business operation data, including actual business operation data and standard test data, and combines this with heterogeneous GPUs and business characteristics (classification, recognition, LLM, etc.) to form feature output.
[0183] The vector processing module 72 combines the business characteristics (model type, model complexity), monitoring metrics (business metrics, GPU metrics), and hardware characteristics (heterogeneous chip identifier, memory specifications, supported accuracy) output by the monitoring module 71 into a unified feature vector for the time series, and performs normalization processing. The normalized data uses a weighted average as the evaluation function to form a performance-accuracy label (performance-accuracy metric = w1 * performance metric + w2 * accuracy metric, with weights dynamically adjusted according to business needs). For example, the vector processing module 72 combines multi-dimensional features with time features to form a time series feature vector, and then scales the time series feature vector to a uniform range using the StandardScaler normalization algorithm to ensure data consistency and model training stability. Furthermore, the vector processing module 72 can use a weighted average method to combine performance metrics and accuracy metrics into a comprehensive performance-accuracy metric as a label, facilitating subsequent prediction model training.
[0184] The prediction and regression module 73 constructs an LSTM regression prediction model. Based on business and hardware complexity and data volume, the model architecture and number of layers are determined. The input is a preprocessed time-series feature vector, and the output is the performance-accuracy prediction value matched to different businesses on heterogeneous hardware. A periodic update mechanism can be established for the LSTM-based regression prediction model. Through a fully connected layer, the output is converted into a single prediction value representing the performance-accuracy output. That is, the first module can directly output the prediction performance data scores of each computing device in each computing device combination corresponding to the business characteristic time period, rather than outputting lower-level performance data. In this embodiment, the scores are output first, followed by filtering and comparison, which is more intuitive.
[0185] The scheduling policy distribution module 74 compares the scores of the LSTM prediction outputs, sets performance-accuracy thresholds that conform to business characteristics, and triggers the scheduling mechanism for business states below the threshold. For example, if the score is below the threshold, the heterogeneous GPU solution is adjusted, including scaling up or down heterogeneous instance resources and updating resource configurations. If the score is above the threshold, the stable solution is maintained until the next scheduled LSTM training update. The scheduling policy distribution module 74 completes resource adjustments such as instance and configuration based on GPU hardware metrics.
[0186] In an exemplary embodiment, the system 700 can monitor and collect overall business system operation metrics and heterogeneous GPU cluster performance and accuracy metrics through components such as Prometheus and Grafana. The metrics are then categorized according to different business types and integrated into data vectors for unified vector processing. The vectors contain time series features, business features (model type, model complexity), business monitoring metrics (data type, data size, data dimension), GPU monitoring metrics (inference throughput, latency, accuracy, memory access bandwidth usage), and hardware characteristics (heterogeneous chip identifier, memory specifications, supported accuracy). After unified vector processing, the labeled data is fed into LSTM for regression prediction, and the cluster deployment of heterogeneous GPU-supported services is adjusted based on the prediction results and preset thresholds.
[0187] For situations where the system has not yet been officially launched or where new heterogeneous GPU hardware has been introduced into the deployment plan, a standard test request is sent to the small cluster to generate standard business performance and accuracy test data.
[0188] The computing device deployment scheduling method for heterogeneous computing device resource pools proposed in this disclosure extracts data including business data and computing device data to form a unified feature vector, which is then normalized and labeled according to time series to construct an LSTM-based regression model. After training, the performance-accuracy prediction results for heterogeneous GPUs in multivariate inference tasks are obtained. Based on the prediction model and appropriate performance-accuracy thresholds, inference task scheduling among heterogeneous hardware is performed. By scheduling heterogeneous GPU chips, the strengths of each chip in business operations can be fully utilized to improve resource utilization efficiency. The design considers both performance and accuracy, balancing performance and accuracy according to different scenario requirements to provide superior business service quality. By considering the scalability of new heterogeneous GPUs, feature vectors can be provided according to standard rules, enabling deployment scheduling of services in new environments. To a certain extent, considering differences in metrics such as memory access optimization, scheduling suggestions can be provided for the stability of inference services. For services with high stability requirements, hardware with strong memory access optimization capabilities is recommended.
[0189] This disclosure enables accurate prediction of the applicability of different computing devices in different combinations of computing devices for different business scenarios, and dynamically completes the scheduling and configuration of heterogeneous resources. For example, in LLM inference scenarios, there are different requirements for accuracy and complexity. Different domestic GPUs can work together heterogeneously to avoid the problem that a single hardware device cannot fully meet the needs of inference business. Through performance-accuracy prediction, a better heterogeneous hardware resource configuration is provided. In exemplary application scenarios, by continuously collecting and analyzing business and hardware GPU data, LSTM is used to dynamically predict and adjust the scheduling strategy to achieve better resource allocation. At the same time, it can quickly connect to new heterogeneous GPUs to complete the allocation of inference tasks and provide dynamic adaptation capabilities for heterogeneous hardware. By integrating performance and accuracy as evaluation indicators, scheduling is performed according to different business requirements for performance and accuracy. By utilizing the predictive capabilities of LSTM for time series, peak inference requests in different time periods can be effectively supported, and effective hardware resource support can be achieved.
[0190] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0191] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely hardware implementations, entirely software implementations (including firmware, microcode, etc.), or implementations combining hardware and software aspects, collectively referred to herein as “circuits,” “modules,” or “systems.”
[0192] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided. Referring below... Figure 8 To describe an electronic device 800 according to this embodiment of the present invention.
[0193] The electronic device 800 can be used to implement any step of the computing device deployment method in the embodiments of this disclosure, or it can be used to implement at least one of the first model and the second model in the embodiments of this disclosure, or it can be used to implement the training method for at least one of the first model and the second model.
[0194] Figure 8 The electronic device 800 shown is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the present invention. Figure 8 As shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one processor 810, at least one memory 820, and a bus 830 connecting different system components (including memory 820 and processor 810).
[0195] The memory stores program code that can be executed by the processor 810, causing the processor 810 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of the present invention. For example, the processor 810 can perform the methods shown in the embodiments of this disclosure.
[0196] The memory 820 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 8201 and / or cache 8202, and may further include read-only memory (ROM) 8203.
[0197] The memory 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0198] Bus 830 can represent one or more of several types of bus structures, including a memory bus or memory controller, peripheral bus, graphics acceleration port, processor, or a local bus using any of the various bus structures.
[0199] Electronic device 800 can also communicate with one or more external devices 900 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 800, and / or with any device that enables electronic device 800 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 850. Furthermore, electronic device 800 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 860. As shown, network adapter 860 communicates with other modules of electronic device 800 via bus 830. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 800, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0200] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0201] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of the invention may also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of the invention described in the "Exemplary Methods" section of this specification.
[0202] The program product for implementing the above-described method according to embodiments of the present invention may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0203] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0204] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable signal medium may also be any readable medium other than a readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0205] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0206] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and concept of this disclosure are indicated by the claims.
Claims
1. A method for deploying computing devices, characterized in that, include: Determine P combinations of computing devices corresponding to N computing devices, wherein the N computing devices are not identical, N>1, P>1; Collect performance data corresponding to the service operation data when the computing device processes multiple services in different combinations of the computing device; The first model is trained based on the performance data of the N computing devices corresponding to the business operation data; Based on historical data, at least one business characteristic period within the target strategy cycle is determined, along with the predicted business operation data corresponding to each business characteristic period. The predicted service operation data is input into the first model to determine the prediction performance data of the computing device in different combinations of computing devices corresponding to the predicted service operation data, and the combination of computing devices corresponding to each of the service characteristic time periods is determined based on the prediction performance data.
2. The computing device deployment method as described in claim 1, characterized in that, The P computing device combinations include a first computing device combination in operation and / or a second computing device combination not in operation. Performance data corresponding to the service operation data is collected when the computing devices process multiple services in different computing device combinations, including: When M1 types of services that have been launched are running on the first combination of computing devices, the performance data corresponding to the computing devices and the service operation data is collected, and M1≥0; The performance data of the computing devices and the service operation data are collected when the M1 types of services that have been launched are running on the second computing device combination. When M2 types of services that are not yet online are running on the P combinations of computing devices, the performance data corresponding to the running data of the computing devices and the services is collected, and M2≥0.
3. The computing device deployment method as described in claim 1 or 2, characterized in that, When the computing device processes multiple services in different combinations of the computing device, the performance data corresponding to the service operation data includes: At the sampling time point, based on the business operation data and the performance data, operational characteristic data of the computing device at the sampling time point is generated. The operational characteristic data includes the sampling time point, the number of the computing device, the computing device combination in which the computing device is located at the sampling time point, the business characteristics of the business processed by the computing device combination in which the computing device is located at the sampling time point, the business data of the business processed by the computing device combination in which the computing device is located at the sampling time point, the hardware performance parameters of the computing device at the sampling time point, and the hardware specifications of the computing device. The business characteristics include the type of target model deployed on the computing device combination, the model complexity, and the types of business performed by the target model, wherein the types of business include classification and reasoning; the business data includes at least one type of data, data volume, and number of data dimensions.
4. The computing device deployment method as described in claim 1, characterized in that, Based on historical data, at least one business characteristic period within the target strategy cycle is determined, and the predicted business operation data corresponding to each business characteristic period includes: Continuously collect actual service operation data corresponding to the M1 types of services that have been launched and are running on the first combination of computing devices put into operation, so as to form the historical data; The second model is trained using the historical data so that it learns the correlation between the measured business operation data and time. The second model is input with the start and end times of the target strategy period, and at least one business characteristic time period corresponding to the target strategy period is obtained from the output of the second model, as well as the predicted business operation data corresponding to each business characteristic time period.
5. The computing device deployment method as described in claim 4, characterized in that, The second model is the same as the first model.
6. The computing device deployment method as described in claim 1, characterized in that, The computing device combination corresponding to each business characteristic time period is determined based on the predicted performance data, including: For a target business characteristic period, identify P1 combinations of computing devices where any predicted performance data does not meet the preset conditions under the predicted business operation data corresponding to the target business characteristic period, where P1≥0; Identify the unavailable computing devices corresponding to the target service characteristic time period, and the P2 combinations of computing devices to which the unavailable computing devices belong, where P2≥0; After deduplication of the P1 computing device combinations and the P2 computing device combinations, they are merged to obtain P3 unusable computing device combinations. The P3 unusable computing device combinations are then deleted from the P computing device combinations to obtain P-P3 usable computing device combinations, where P3≥0. Determine the total score of the predicted performance data of the computing devices corresponding to each available computing device combination, and determine the available computing device combination with the highest total score as the computing device combination corresponding to the target business characteristic time period.
7. The computing device deployment method as described in claim 6, characterized in that, The total score for determining the predicted performance data of the computing devices corresponding to each of the available computing device combinations includes: Determine the target predicted performance data for a target computing device within a target available computing device portfolio; Based on the target prediction performance data, determine the performance score and accuracy score of the computing device during the target service characteristic period; Based on the predicted business operation data corresponding to the target business characteristic time period, determine the first weight corresponding to the performance score and the second weight corresponding to the accuracy score; The score of the predicted performance data of the target computing device is determined by the sum of the product of the first weight and the performance score, and the product of the second weight and the accuracy score. The total score corresponding to the target available computing device combination is determined by summing the scores of the predicted performance data of the computing devices corresponding to the target available computing device combination.
8. An electronic device, characterized in that, include: Memory; as well as A processor coupled to the memory, the processor being configured to perform the method as described in any one of claims 1-7 based on instructions stored in the memory.
9. A computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the method as described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1-7.