Dynamic load balancing optimization method of AI intelligent computing system server equipment

By combining LSTM, ARIMA and wavelet transform functions, the load changes of the server equipment of the AI ​​intelligent computing system are predicted, the equipment is prioritized dynamically and the optimal resource allocation strategy is generated, which solves the problems of insufficient load prediction accuracy and unreasonable resource allocation in the existing technology, and significantly improves resource utilization efficiency and service quality.

CN119938339AInactive Publication Date: 2025-05-06CHANGSHA SHAOGUANG SEMICONDUCTOR CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510216213.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When facing the rapidly changing workload and complex task resource requirements, the load balancing method of existing AI intelligent computing system server equipment has insufficient prediction accuracy and unreasonable resource allocation, resulting in low resource utilization efficiency.

Method used

Long-term short-term memory network (LSTM) is used to combine the ARIMA model and wavelet transformation function to capture long-term dependencies, short-term fluctuations and multi-scale periodic changes in the state characteristics of the device, predict future load changes, dynamically divide equipment priorities, and generate optimal resource allocation strategies based on device priorities and resource requirements through resource matching algorithms.

Benefits of technology

It significantly improves the response speed and service quality of server equipment, maximizes resource utilization efficiency, and ensures that resource allocation is always in the optimal state.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938339A_ABST
    Figure CN119938339A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic load balancing optimization method for AI intelligent computing system server equipment, which relates to the technical field of load balancing optimization, and comprises the following steps: acquiring equipment performance data, and preprocessing the acquired equipment performance data; based on the preprocessed equipment performance data, extracting equipment state features through time sequence analysis and periodic fluctuation detection; predicting the future load change of the equipment based on the state characteristics of the equipment, dividing the priority of the equipment according to the prediction result, and analyzing the resource demand; generating a resource allocation strategy by applying a resource matching algorithm according to the device priority and the resource demand; and executing the resource allocation strategy, receiving feedback information in real time in the execution process, and optimizing the resource allocation strategy according to the feedback information. According to the method, the response speed and the service quality of the server equipment are remarkably improved, and the resource utilization efficiency is maximized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of load balancing optimization, and in particular to a dynamic load balancing optimization method for an AI intelligent computing system server device. Background Art

[0002] In recent years, with the rapid growth of artificial intelligence (AI) and big data processing needs, the performance and efficiency of AI intelligent computing system server equipment has become a research hotspot. Traditional load balancing methods mainly rely on static rules or simple dynamic adjustment mechanisms, which are incapable of dealing with complex and changing workloads. To solve this problem, researchers have introduced a variety of advanced technical means, such as machine learning, time series analysis, and periodic fluctuation detection, to improve the load balancing effect of server equipment. In particular, the method of combining the long short-term memory network (LSTM) with the ARIMA model performs well in capturing long-term dependencies and short-term fluctuations in device state characteristics, greatly improving the prediction accuracy. In addition, wavelet transform, as a multi-scale analysis tool, can effectively extract the periodic characteristics of device status, further enhancing the accuracy of load prediction.

[0003] However, despite the significant progress made in the above technologies, existing methods still have some shortcomings, especially in dealing with rapidly changing workloads and complex task resource requirements. Traditional methods often focus on single-dimensional data analysis, which makes it difficult to fully reflect the actual operating status of the equipment; at the same time, in the process of generating resource allocation strategies, there is a lack of full consideration of the differences in resource consumption of equipment priorities and task types, resulting in inefficient resource utilization. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a dynamic load balancing optimization method for an AI intelligent computing system server device to solve the problems of insufficient prediction accuracy and unreasonable resource allocation in the prior art.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for dynamic load balancing optimization of an AI intelligent computing system server device, which includes collecting device performance data and preprocessing the collected device performance data; Based on the pre-processed equipment performance data, equipment status characteristics are extracted through time series analysis and periodic fluctuation detection; Based on the device status characteristics, predict the future load changes of the device, prioritize the devices according to the prediction results, and analyze resource requirements; Based on device priorities and resource requirements, a resource matching algorithm is applied to generate a resource allocation strategy; Execute resource allocation strategies, receive feedback information in real time during execution, and optimize resource allocation strategies based on the feedback information.

[0007] As a preferred solution of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device described in the present invention, the device performance data includes CPU utilization, memory usage, disk I / O rate, network traffic and response time.

[0008] As a preferred solution of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device of the present invention, the preprocessing of the collected device performance data is performed in the following specific steps: Time series alignment of equipment performance data is performed through interpolation method; Standardize equipment performance data through data filtering and conversion rules; Use data cleaning to remove invalid and abnormal data from performance data; Through normalization processing, the performance data of equipment of different magnitudes are unified.

[0009] As a preferred solution of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device of the present invention, wherein: based on the pre-processed device performance data, the device status characteristics are extracted through time series analysis and periodic fluctuation detection, and the specific steps are as follows: The time series decomposition method is used to decompose the preprocessed performance data into long-term trend data and short-term fluctuation data; Using Fourier transform, identify periodic and non-periodic components in performance data; The long-term trend characteristics of equipment status are extracted from the long-term trend data through the moving average method; Extract the short-term fluctuation characteristics of equipment status from short-term fluctuation data through exponential smoothing method; The periodic characteristics of the equipment status are extracted from the periodic components through wavelet transform.

[0010] As a preferred solution of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device of the present invention, wherein: based on the device state characteristics, the future load changes of the device are predicted, and the device priority is divided according to the prediction results. The specific steps are as follows: The long short-term memory network combined with the ARIMA method is used to capture the long-term dependency and short-term fluctuations in the state characteristics of the equipment. At the same time, the wavelet transformation function based on different wavelet bases is used to capture the periodic changes of the periodic characteristics of the equipment state at multiple scales, and through nonlinear transformation, the future load changes of the equipment are predicted. The expression is: in, It's time point The load change value of the equipment when Indicates a time interval. Indicates the length of the historical data considered when predicting future load changes on the equipment. Indicates the current time point. represents the number of wavelet bases used, is the index variable of the wavelet basis, It represents the periodic variation intensity of the device within the frequency range of time t, Indicates the device status at a future time point The long-term trend characteristics of Indicates the device status at a future time point The periodic characteristics of represents the wavelet transform function based on the i-th wavelet basis, represents the state characteristics of the device at time point t, Represents the state characteristics of the device captured by LSTM long-term dependencies in Based on the load distribution in historical data, define a low load threshold L and a high load threshold H; when ≥H, it is considered that the device will When the device is under high load, the device is classified as low priority; When L< <H, it is considered that the device will When the device is in medium load state, the device is classified as medium priority; when ≤L, it is considered that the device will When the load is low, the device is classified as high priority.

[0011] As a preferred solution of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device of the present invention, wherein: the analysis of resource requirements, the specific steps are as follows: Classify task types based on the CPU time, memory usage, disk read / write volume, and network traffic actually consumed by historical tasks; By monitoring the current resource usage of the device in real time and analyzing the actual resource consumption of each task type in combination with the device log; Based on the actual resource consumption of each task type, the average and maximum resource consumption of each task type is analyzed.

[0012] As a preferred solution of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device of the present invention, wherein: the resource matching algorithm is applied to generate a resource allocation strategy according to the device priority and resource requirements, and the specific steps are as follows: The matching score between the task and the device is calculated based on the device priority, device status characteristics, average resource consumption of each task type, and maximum resource consumption. The expression is: ; in, Indicates the task type at time t The resource matching score with the device, A scaling factor that characterizes the state of the device, A scaling factor indicating the device priority, represents the priority score of the device at time t, A sensitive factor indicating the current workload of the device. represents the actual workload of the equipment at time t, Indicates the critical point of device busyness. Indicates the task type Scaling factor for average resource consumption, Indicates the task type The average resource consumption at time t, Indicates the task type The scaling factor for maximum resource consumption, Indicates the task type Maximum resource consumption at time t; For each task in the time window, the corresponding resource matching score is calculated. Sort from high to low; According to the sorting results, the resource matching scores are The highest tasks are assigned to high-priority devices first and used as a resource allocation strategy.

[0013] As a preferred solution of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device of the present invention, wherein: the feedback information includes real-time device performance data, task execution progress and resource consumption; The specific steps of optimizing the resource allocation strategy according to the feedback information are as follows: Update the load change value of the device based on the latest device performance data and task resource consumption and the average resource consumption of task types and maximum resource consumption , and recalculate the resource matching score ; The resource matching score for the new calculation Re-sort and optimize the resource allocation strategy based on the new sorting results.

[0014] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device as described in the first aspect of the present invention.

[0015] The beneficial effects of the present invention are as follows: by adopting a long short-term memory network (LSTM) in combination with an ARIMA model and a wavelet transform function, the long-term dependencies, short-term fluctuations and multi-scale periodic changes in the device state characteristics can be accurately captured, and future load changes can be predicted through nonlinear transformations. The device priorities can be dynamically divided, and the actual resource consumption of the task types can be analyzed in detail. The optimal resource allocation strategy based on device priorities and resource requirements is generated, which ultimately significantly improves the response speed and service quality of the server device and maximizes the resource utilization efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0017] Figure 1 This is a flow chart of the dynamic load balancing optimization method for the AI ​​intelligent computing system server device in Example 1.

[0018] Figure 2 This is a flowchart for extracting device status features in Example 1. DETAILED DESCRIPTION

[0019] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0020] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0022] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, and provides a dynamic load balancing optimization method for an AI intelligent computing system server device, comprising the following steps: S1: Collect equipment performance data and pre-process the collected equipment performance data.

[0023] S1.1: Device performance data includes CPU utilization, memory usage, disk I / O rate, network traffic, and response time.

[0024] Furthermore, device performance data is collected in real time through monitoring agents integrated in the server or dedicated hardware sensors.

[0025] S1.2: Perform time series alignment on equipment performance data by interpolation.

[0026] For example, for missing CPU utilization data in certain time periods, linear interpolation is used to fill these gaps to ensure that the data at all time points are continuous and consistent, thereby achieving time series alignment.

[0027] S1.3: Standardize equipment performance data through data filtering and transformation rules.

[0028] For example, the Z-score standardization method is applied to convert the memory usage data into a standard normal distribution with a mean of 0 and a standard deviation of 1 to eliminate the differences between data of different magnitudes and facilitate subsequent analysis.

[0029] S1.4: Use data cleaning to remove invalid and abnormal data from performance data.

[0030] For example, we identified and removed outliers in the disk I / O rate data that were clearly outside the normal range (such as values ​​exceeding the 99% confidence interval), as well as invalid data with incorrect or repeated records, to ensure the purity and reliability of the data set.

[0031] S1.5: Unify the performance data of devices of different magnitudes through normalization.

[0032] For example, converting network traffic data from byte units to a standardized [0, 1] interval allows data of different magnitudes to be compared and analyzed on the same scale, improving the consistency and accuracy of multi-dimensional data analysis.

[0033] S2: Based on the preprocessed equipment performance data, the equipment status characteristics are extracted through time series analysis and periodic fluctuation detection.

[0034] S2.1: Use the time series decomposition method to decompose the preprocessed performance data into long-term trend data and short-term fluctuation data.

[0035] Furthermore, using classic time series decomposition techniques (such as STL decomposition), the pre-processed device performance data (such as CPU utilization) is decomposed into two main components: a data set that reflects long-term trend changes and a data set that captures short-term rapid changes. This method can clearly distinguish long-term evolution patterns and short-term disturbances in the data, laying the foundation for further feature extraction.

[0036] S2.2: Use Fourier transform to identify periodic and non-periodic components in performance data.

[0037] Furthermore, by applying Fourier transform, the performance data is converted from the time domain to the frequency domain, so that the periodic components (such as daily or weekly repetitive patterns) and non-periodic components (such as random noise) in the data can be accurately identified. This process helps to separate regular periodic signals and irregular transient changes, providing a more refined data structure for subsequent analysis.

[0038] S2.3: Extract the long-term trend characteristics of the equipment status from the long-term trend data through the moving average method.

[0039] Furthermore, first, according to the characteristics of the equipment performance data and the analysis requirements, determine a suitable moving window size; then, put the first M data points in the time series into the initial window, and calculate the average of these data points as the first smoothed data point; then, slide the window one by one, remove the earliest data point each time and add the next data point, recalculate the average of the data in the window, and so on, until the entire time series is traversed. This process effectively smooths short-term fluctuations and highlights long-term trend changes, thereby accurately capturing the long-term evolution characteristics of the equipment status.

[0040] S2.4: Extract the short-term fluctuation characteristics of equipment status from short-term fluctuation data through exponential smoothing method.

[0041] Furthermore, when processing short-term fluctuating data, an appropriate smoothing method is first selected to capture recent trends. Next, the smoothing process is initialized using the first data point of the time series. Each subsequent data point is then processed step by step, with each update giving more weight to the most recent data while retaining a portion of the influence of the previous smoothing results. This method can effectively highlight short-term fluctuation characteristics and reduce the impact of random noise, thereby sensitively reflecting immediate changes in device status.

[0042] S2.5: Extract the periodic characteristics of the device status from the periodic components through wavelet transform.

[0043] Further, first, select appropriate wavelet basis functions, such as Morlet or Mexican Hat wavelets, to adapt to different periodic characteristics; then, perform multi-scale wavelet transform on the periodic components to decompose the data into components in different frequency ranges; then, calculate the wavelet coefficients at each scale, which reflect the intensity of periodic changes at different time scales; finally, analyze the wavelet coefficients at each scale, identify significant periodic patterns, and extract the periodic characteristics of the equipment status from them. This method can capture periodic changes at different time scales and provide detailed periodic characteristic information.

[0044] S3: Based on the device status characteristics, predict the future load changes of the device, prioritize the devices according to the prediction results, and analyze resource requirements.

[0045] S3.1: The long short-term memory network combined with the ARIMA method is used to capture the long-term dependency and short-term fluctuations in the state characteristics of the equipment. At the same time, the wavelet transformation function based on different wavelet bases is used to capture the periodic changes of the periodic characteristics of the equipment state at multiple scales, and through nonlinear transformation, the future load changes of the equipment are predicted. The expression is: in, It's time point The load change value of the equipment when Indicates a time interval. Indicates the length of the historical data considered when predicting future load changes on the equipment. Indicates the current time point. represents the number of wavelet bases used, is the index variable of the wavelet basis, It represents the periodic variation intensity of the device within the frequency range of time t, Indicates the device status at a future time point The long-term trend characteristics of Indicates the device status at a future time point The periodic characteristics of represents the wavelet transform function based on the i-th wavelet basis, represents the state characteristics of the device at time point t, Represents the state characteristics of the device captured by LSTM long-term dependencies in .

[0046] Furthermore, the long short-term memory network combined with the ARIMA method is used to first analyze the state characteristics of the equipment to capture the long-term dependencies and short-term fluctuations. Then, the wavelet transform function of different wavelet bases is used to deeply explore the periodic change characteristics of the equipment state at multiple scales. Then, these features are integrated through nonlinear transformation technology to build a prediction model. This process can fully reflect the historical operation mode of the equipment and its potential periodic laws, so as to accurately predict the future load changes of the equipment. Finally, based on the prediction results, resource allocation can be planned in advance to ensure that the equipment is in the optimal working state at a future time point.

[0047] S3.2: Based on the load distribution in the historical data, define a low load threshold L and a high load threshold H.

[0048] when ≥H, it is considered that the device will When the device is under high load, it will be classified as low priority.

[0049] For example, if the forecast results show the load change value of a server in the future Exceeding the set high load threshold H indicates that the server is expected to face a very high workload. Based on this prediction, the server is classified as low priority, which means that other less loaded devices will be given priority when allocating resources to avoid further burdening it.

[0050] When L< <H, it is considered that the device will When the device is in medium load state, the device is classified as medium priority.

[0051] For example, if the forecast results show the load change value of a server in the future Between the low load threshold L and the high load threshold H, it means that the server is expected to be in a relatively stable working state, neither too busy nor too idle. Therefore, the server is classified as medium priority, and tasks are reasonably arranged according to actual needs during resource allocation to ensure effective use of resources.

[0052] when ≤L, it is considered that the device will When the load is low, the device is classified as high priority.

[0053] For example, if the forecast results show the load change value of a server in the future If the value is lower than the set low load threshold L, it indicates that the server is expected to have a low workload. Based on this prediction, the server is classified as high priority, which means that this device can be given priority when allocating resources, making full use of its remaining computing power and improving overall resource utilization.

[0054] S3.3: Classify task types based on the CPU time, memory usage, disk read / write volume, and network traffic actually consumed by historical tasks.

[0055] S3.4: Monitor the current resource usage of the device in real time and analyze the actual resource consumption of each task type in combination with the device log.

[0056] Current resource usage includes CPU utilization, memory utilization, disk I / O rate, and network traffic.

[0057] Furthermore, first, the real-time monitoring tool continuously collects the performance data of the device, including key indicators such as CPU utilization, memory usage, disk I / O rate, network traffic and response time. Then, these real-time data are transmitted to the central processing unit for analysis. At the same time, the device logs are also parsed regularly to extract detailed resource consumption records for each task type. Then, the real-time monitoring data is combined with the historical records in the device log to compare and analyze the CPU time, memory usage, disk read and write volume, and network traffic actually consumed by each task type during execution. Finally, based on this comprehensive analysis result, the average and maximum resource consumption of each task type can be accurately evaluated, providing a reliable basis for subsequent resource allocation and optimization.

[0058] S3.5: Analyze the average and maximum resource consumption of each task type based on the actual resource consumption of each task type.

[0059] Furthermore, first, collect and summarize all performance data generated by each task type during execution, including CPU utilization, memory usage, disk I / O rate, network traffic, and response time. Then, based on these data, calculate the resource consumption of each task type in different time periods, and calculate the average resource consumption of each task type, that is, the average CPU time, memory usage, disk read and write volume, and network traffic required by the task type during a typical run. Then, identify and record the maximum resource consumption of each task type, which represents the highest resource demand of the task type under peak load. Finally, by comparing the average and maximum resource consumption of different task types, we can gain an in-depth understanding of the resource demand characteristics of each task type and provide accurate data support for subsequent resource allocation strategies.

[0060] S4: Based on device priorities and resource requirements, a resource matching algorithm is applied to generate a resource allocation strategy.

[0061] S4.1: Calculate the matching score between the task and the device based on the device priority, device status characteristics, average resource consumption of each task type, and maximum resource consumption. The expression is: ; in, Indicates the task type at time t The resource matching score with the device, A scaling factor that characterizes the state of the device, A scaling factor indicating the device priority, represents the priority score of the device at time t, A sensitive factor indicating the current workload of the device. represents the actual workload of the equipment at time t, Indicates the critical point of device busyness. Indicates the task type Scaling factor for average resource consumption, Indicates the task type The average resource consumption at time t, Indicates the task type The scaling factor for maximum resource consumption, Indicates the task type Maximum resource consumption at time t.

[0062] Furthermore, first, the load situation of the equipment at the current time point and its future change trend are evaluated by combining the status characteristics and priority information of the equipment. Next, the historical resource consumption data of each task type is analyzed to determine its average resource consumption and maximum resource consumption to understand the actual needs of the task. Then, the priority and status characteristics of the equipment are combined with the resource requirements of the task type, and the matching score between the task and the equipment is calculated through a comprehensive evaluation method. This score reflects the adaptability of the task when it is assigned to a specific device, and a high score indicates that the device is more suitable for processing this type of task. Finally, based on these matching scores, the optimal resource allocation strategy can be formulated for each task in the time window to ensure that critical tasks are given priority and maximize the use of existing resources.

[0063] S4.2: For each task in the time window, score the task according to the corresponding resource matching score. Sort from high to low.

[0064] Furthermore, the time window is set according to the task scheduling requirements. It can be a fixed period (such as every hour, every day) or dynamically adjusted according to the actual load. In this way, it ensures that task scheduling and resource allocation can be optimized within a reasonable and flexible time period, so as to better adapt to different workloads.

[0065] S4.3: According to the sorting results, match the resources with scores The highest tasks are assigned to high-priority devices first and used as a resource allocation strategy.

[0066] For example, within a specific time window, after calculation and sorting, it is found that Task A has the highest resource matching score. Based on this result, Task A will be preferentially assigned to Server Device 1, which currently has the highest priority, because this device not only has a higher processing capacity, but also has a lower current load, and can efficiently complete Task A. This allocation method ensures that critical tasks can obtain the best resource support while maximizing the overall resource utilization efficiency.

[0067] S5: Execute the resource allocation strategy, receive feedback information in real time during the execution process, and optimize the resource allocation strategy according to the feedback information.

[0068] S5.1: Feedback information includes real-time device performance data, task execution progress, and resource consumption.

[0069] Furthermore, during the execution of the resource allocation strategy, the monitoring agent deployed on the server is used to continuously collect real-time performance data such as CPU utilization and memory usage. Secondly, the execution progress of each task is recorded from the start to the completion. Finally, the actual resource consumption of each task, such as CPU time and memory usage, is analyzed in combination with the device log. All data is centrally transmitted to the central management system for processing to ensure the accuracy and timeliness of the feedback information.

[0070] S5.2: Optimize resource allocation strategy based on feedback information. The specific steps are as follows: S5.2.1: Update the load change value of the device based on the latest device performance data and task resource consumption and the average resource consumption of task types and maximum resource consumption , and recalculate the resource matching score .

[0071] Furthermore, by combining real-time monitoring tools and device logs, we analyze the actual resource consumption of each task type and update its average resource consumption. and maximum resource consumption Then, based on the updated device performance data and task resource consumption, re-predict the future load change value of the device Finally, the resource matching score between each task and device is recalculated using the updated load change value and task resource consumption data. , ensuring that the scoring results can accurately reflect the current status of server devices and provide the latest basis for subsequent resource allocation. This process ensures that the resource allocation strategy is always optimized and adjusted based on the most accurate data.

[0072] S5.2.2: Newly calculated resource matching score Re-sort and optimize the resource allocation strategy based on the new sorting results.

[0073] It should be noted that based on the new sorting results, the resource allocation strategy is adjusted to prioritize tasks with high matching scores to high-priority devices, thereby ensuring that key tasks can receive optimal resource support. This process not only improves resource utilization efficiency, but also enhances response speed and service quality, ensuring that resource allocation is always in the optimal state.

[0074] This embodiment also provides a computer device, which is suitable for the dynamic load balancing optimization method of the AI ​​intelligent computing system server device, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the dynamic load balancing optimization method of the AI ​​intelligent computing system server device proposed in the above embodiment.

[0075] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0076] The present embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the dynamic load balancing optimization method for implementing an AI intelligent computing system server device as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage device, a flash memory, a disk or an optical disk.

[0077] In summary, the present invention adopts a long short-term memory network (LSTM) combined with an ARIMA model and a wavelet transform function to accurately capture long-term dependencies, short-term fluctuations and multi-scale periodic changes in device state characteristics, and predicts future load changes through nonlinear transformations, dynamically divides device priorities, and analyzes the actual resource consumption of task types in detail, generating an optimal resource allocation strategy based on device priority and resource requirements, ultimately significantly improving the response speed and service quality of server devices and maximizing resource utilization efficiency.

[0078] Example 2, referring to Table 1, is the second example of the present invention. To further verify the technical solution of the present invention, experimental simulation data of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device are provided.

[0079] In order to verify the effectiveness of the proposed dynamic load balancing optimization method for server equipment in the AI ​​intelligent computing system, a comparative experiment was designed and implemented. Two groups of server clusters were selected in the experiment, one group adopted the traditional static load balancing strategy (called the control group), and the other group applied the method of the present invention (called the experimental group). The experimental environment simulated a typical cloud computing data center scenario, including multiple virtual machine instances, database services, and web applications to ensure that the test conditions were as close to the real-world application environment as possible.

[0080] First, before the experiment started, a monitoring agent was deployed to all servers participating in the experiment to collect key performance indicators in real time, including CPU utilization, memory usage, disk I / O rate, network traffic, and response time. To ensure the quality and consistency of the data, the raw data was preprocessed, including time series alignment, standardization, data cleaning, and normalization.

[0081] Secondly, we use interpolation to fill in missing data points in certain time periods to ensure that the data at all time points are continuous and consistent. We use the Z-score normalization method to convert the memory usage into a standard normal distribution to eliminate the differences between data of different magnitudes. We identify and remove outliers and invalid data in the disk I / O rate to ensure the purity of the data set. We also standardize the network traffic to the [0, 1] interval so that data of different magnitudes can be compared on the same scale.

[0082] Then, based on the preprocessed equipment performance data, the equipment status characteristics are extracted through time series analysis and periodic fluctuation detection. The specific steps include using STL decomposition technology to decompose the performance data into long-term trend and short-term fluctuation data; using Fourier transform to identify periodic and non-periodic components; smoothing short-term fluctuations in long-term trend data through moving average method; capturing instant changes in short-term fluctuation data through exponential smoothing method; and finally, extracting multi-scale periodic features from periodic components through wavelet transform.

[0083] Finally, based on the extracted equipment status features, the future load changes of the equipment are predicted, and the equipment priorities are divided according to the prediction results. The long short-term memory network is combined with the ARIMA model to capture long-term dependencies and short-term fluctuations, and the wavelet transform function is used to mine multi-scale periodic change characteristics. The low load threshold L and the high load threshold H are defined, and the equipment is divided into different priorities based on them. The resource allocation strategy is generated through the resource matching algorithm to ensure that key tasks obtain the best resource support and maximize resource utilization efficiency.

[0084] Traditional static load balancing strategy refers to a resource allocation method that does not rely on real-time performance data and prediction models, including simple round robin or proportional allocation of tasks to each server.

[0085] The details are shown in Table 1 below: Table 1 Server equipment performance data comparison table Through the analysis of the data in the above table, it can be clearly seen that the present invention significantly improves the operating efficiency of the server and reduces resource consumption by optimizing the preprocessing of device performance data and resource allocation strategies. For example, after applying the present invention, the CPU utilization of the server in the experimental group was reduced from an average of 75% to 60%, and the memory utilization was also reduced from an average of 80% to 70%. At the same time, the disk I / O rate and network traffic were reduced by 20MB / s and 200Mbps respectively. These improvements show that the present invention not only improves the effective utilization of computing and storage resources, but also reduces the demand for external network bandwidth, thereby enhancing the overall operational efficiency and service quality of the data center.

[0086] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A dynamic load balancing optimization method for an AI intelligent computing system server device, characterized in that: include, Collect equipment performance data and pre-process the collected equipment performance data; Based on the pre-processed equipment performance data, equipment status characteristics are extracted through time series analysis and periodic fluctuation detection; Based on the device status characteristics, predict the future load changes of the device, prioritize the devices according to the prediction results, and analyze resource requirements; Based on device priorities and resource requirements, a resource matching algorithm is applied to generate a resource allocation strategy; Execute resource allocation strategies, receive feedback information in real time during execution, and optimize resource allocation strategies based on the feedback information.

2. The method for dynamic load balancing optimization of an AI intelligent computing system server device according to claim 1, characterized in that: The device performance data includes CPU utilization, memory usage, disk I / O rate, network traffic, and response time.

3. The dynamic load balancing optimization method for the AI ​​intelligent computing system server device according to claim 2, characterized in that: The specific steps of preprocessing the collected equipment performance data are as follows: Time series alignment of equipment performance data is performed through interpolation method; Standardize equipment performance data through data filtering and conversion rules; Use data cleaning to remove invalid and abnormal data from performance data; Through normalization processing, the performance data of equipment of different magnitudes are unified.

4. The dynamic load balancing optimization method for the AI ​​intelligent computing system server device according to claim 3, characterized in that: The device performance data after preprocessing is used to extract the device status characteristics through time series analysis and periodic fluctuation detection. The specific steps are as follows: The time series decomposition method is used to decompose the preprocessed performance data into long-term trend data and short-term fluctuation data; Using Fourier transform, identify periodic and non-periodic components in performance data; The long-term trend characteristics of equipment status are extracted from the long-term trend data through the moving average method; Extract the short-term fluctuation characteristics of equipment status from short-term fluctuation data through exponential smoothing method; The periodic characteristics of the equipment status are extracted from the periodic components through wavelet transform.

5. The dynamic load balancing optimization method for the AI ​​intelligent computing system server device according to claim 4, characterized in that: The method predicts the future load change of the equipment based on the equipment status characteristics and divides the equipment priority according to the prediction results. The specific steps are as follows: The long short-term memory network combined with the ARIMA method is used to capture the long-term dependency and short-term fluctuations in the state characteristics of the equipment. At the same time, the wavelet transformation function based on different wavelet bases is used to capture the periodic changes of the periodic characteristics of the equipment state at multiple scales, and through nonlinear transformation, the future load changes of the equipment are predicted. The expression is: in, It's time point The load change value of the equipment when Indicates a time interval. Indicates the length of the historical data considered when predicting future load changes on the equipment. Indicates the current time point. represents the number of wavelet bases used, is the index variable of the wavelet basis, It represents the periodic variation intensity of the device within the frequency range of time t, Indicates the device status at a future time point The long-term trend characteristics of Indicates the device status at a future time point The periodic characteristics of represents the wavelet transform function based on the i-th wavelet basis, represents the state characteristics of the device at time point t, Represents the state characteristics of the device captured by LSTM long-term dependencies in Based on the load distribution in historical data, define a low load threshold L and a high load threshold H; when ≥H, it is considered that the device will When the device is under high load, the device is classified as low priority; When L< <H, it is considered that the device will When the device is in medium load state, the device is classified as medium priority; when ≤L, it is considered that the device will When the load is low, the device is classified as high priority.

6. The method for dynamic load balancing optimization of an AI intelligent computing system server device according to claim 5, characterized in that: The specific steps of analyzing resource requirements are as follows: Classify task types based on the CPU time, memory usage, disk read / write volume, and network traffic actually consumed by historical tasks; By monitoring the current resource usage of the device in real time and analyzing the actual resource consumption of each task type in combination with the device log; Based on the actual resource consumption of each task type, the average and maximum resource consumption of each task type is analyzed.

7. The dynamic load balancing optimization method for the AI ​​intelligent computing system server device according to claim 6, characterized in that: The resource allocation strategy is generated by applying a resource matching algorithm according to device priority and resource requirements. The specific steps are as follows: The matching score between the task and the device is calculated based on the device priority, device status characteristics, average resource consumption of each task type, and maximum resource consumption. The expression is: ; in, Indicates the task type at time t The resource matching score with the device, A scaling factor that characterizes the state of the device, A scaling factor indicating the device priority, represents the priority score of the device at time t, A sensitive factor indicating the current workload of the device. represents the actual workload of the equipment at time t, Indicates the critical point of device busyness. Indicates the task type Scaling factor for average resource consumption, Indicates the task type The average resource consumption at time t, Indicates the task type The scaling factor for maximum resource consumption, Indicates the task type Maximum resource consumption at time t; For each task in each time window, the corresponding resource matching score is calculated. Sort from high to low; According to the sorting results, the resource matching scores are The highest tasks are assigned to high-priority devices first and used as a resource allocation strategy.

8. The dynamic load balancing optimization method for the AI ​​intelligent computing system server device according to claim 7, characterized in that: The feedback information includes real-time device performance data, task execution progress and resource consumption; The specific steps of optimizing the resource allocation strategy according to the feedback information are as follows: Update the load change value of the device based on the latest device performance data and task resource consumption and the average resource consumption of task types and maximum resource consumption , and recalculate the resource matching score ; The resource matching score for the new calculation Re-sort and optimize the resource allocation strategy based on the new sorting results.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps of the dynamic load balancing optimization method of the AI ​​intelligent computing system server device described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Dynamic configuration computing resource management method and system based on power wireless terminal

    CN120602992A

  • Dynamic configuration computing resource management method and system based on power wireless terminal

    CN120602992B

  • Server operation and maintenance adaptive management method and system

    CN120909872A

  • Data center intelligent operation and maintenance method and system based on Internet of Things

    CN120929273A

  • Load balancing method and device, electronic equipment, program product and storage medium

    CN121603498A