Server optimization method, system and device based on AI big data and readable storage medium

By employing multi-dimensional data processing and a dual-channel decision-making mechanism, the problem of inaccurate resource assessment caused by single-dimensional monitoring has been solved, enabling accurate assessment and efficient utilization of server resources.

CN120929276BActive Publication Date: 2025-12-30CHANGSHA SHAOGUANG SEMICONDUCTOR CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511453207.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2025-12-30
Estimated Expiration
2045-10-13

AI Technical Summary

Technical Problem

Existing technologies rely on single-dimensional monitoring, which leads to inaccurate resource assessment and makes it impossible to allocate resources effectively under complex loads.

Method used

By acquiring multi-dimensional data from the hardware, software, and network layers, and performing time-aligned processing, the server stress index and topology affinity matrix are calculated. This data is then processed in a dual-channel manner, and resource scheduling instructions are executed.

Benefits of technology

It achieves multi-dimensional feature fusion analysis, accurately assesses resources, and reduces energy consumption, shortens response time, and improves resource utilization through a dual-channel decision-making mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929276B_ABST
    Figure CN120929276B_ABST
Patent Text Reader

Abstract

The application discloses an AI big data-based server optimization method, system and device and a readable storage medium. Hardware layer collection data, software layer collection data and network layer collection data are acquired, and then the hardware layer collection data, the software layer collection data and the network layer collection data are subjected to time alignment processing to obtain a heterogeneous data set. A server stress index is calculated based on the heterogeneous data set, and a topological affinity matrix is generated. Double-channel processing is performed based on the server stress index and the topological affinity matrix, and the processed results are subjected to decision fusion. Finally, a resource scheduling instruction is executed. The application can accurately evaluate resources through multi-dimensional feature fusion analysis, and can reduce energy consumption, shorten response time and effectively improve resource utilization through a double-channel decision mechanism and adaptive resource scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cloud computing server management, and in particular to a server optimization method, system, device, and readable storage medium based on AI big data. Background Technology

[0002] Server optimization refers to improving server performance and efficiency by adjusting various server configurations and parameters. Server optimization includes multiple aspects such as hardware optimization, operating system optimization, and security optimization.

[0003] Currently, server optimization mainly involves monitoring CPU / memory to assess resources. However, monitoring only one dimension can lead to inaccurate resource assessment, making it impossible to allocate resources under complex loads.

[0004] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this invention is to provide a server optimization method, system, device, and readable storage medium based on AI big data, in order to solve the problem that current single-dimensional monitoring leads to inaccurate resource assessment, thus making it impossible to allocate resources under complex loads.

[0006] To achieve the above objectives, the present invention provides a server optimization method based on AI big data, the server optimization method based on AI big data comprising:

[0007] Acquire hardware layer data, software layer data, and network layer data. The hardware layer data includes CPU temperature data and PDU real-time power consumption data. The software layer data includes container memory usage and API gateway request response time. The network layer data includes flow table statistics and server communication latency matrix.

[0008] The hardware layer data, the software layer data, and the network layer data are time-aligned to obtain a heterogeneous dataset, wherein the heterogeneous dataset is the dataset obtained after time alignment of the hardware layer data, the software layer data, and the network layer data.

[0009] The server stress index is calculated based on the heterogeneous dataset, and a topological affinity matrix is ​​generated.

[0010] Dual-channel processing is performed based on the server stress index and the topology affinity matrix, wherein short-term prediction is performed based on the server stress index and topology analysis is performed based on the topology affinity matrix.

[0011] The results of short-term prediction channel processing and topology analysis channel processing are fused together for decision-making, and resource scheduling instructions are executed.

[0012] Further, the step of performing time alignment processing on the hardware layer data, the software layer data, and the network layer data to obtain a heterogeneous dataset includes:

[0013] The reference time source is extracted from the CPU temperature data, wherein the original timestamp of the hardware layer sensor clock is selected as the reference time source.

[0014] Determine whether the PDU is a high-precision PDU. If the PDU is a high-precision PDU, directly obtain the power consumption data timestamp. If the PDU is a normal PDU, apply dynamic delay compensation to obtain the power consumption data timestamp. Here, the high-precision PDU refers to a PDU device equipped with a dedicated clock chip, and the normal PDU refers to a PDU device that relies on a software clock.

[0015] Match the reference time source with the power consumption data timestamp to unify the hardware layer time axis;

[0016] The target time point values ​​are calculated using cubic spline interpolation on the data collected by the software layer.

[0017] Network layer data is aggregated using an exponentially weighted moving average.

[0018] Furthermore, if the PDU is a regular PDU, the step of applying dynamic delay compensation to obtain the power consumption data timestamp includes:

[0019] If the PDU is a regular PDU, then the corrected timestamp is calculated using a first preset formula, and the corrected PDU timestamp is used as the power consumption data timestamp, wherein the first preset formula is:

[0020] , The corrected PDU timestamp As a smoothing factor, For the precise local time of the data acquisition server, The network time obtained by the PDU device via the NTP protocol. This is the original timestamp of the PDU.

[0021] Furthermore, the step of matching the reference time source with the power consumption data timestamp to unify the hardware layer time axis includes:

[0022] Based on the reference time source, the time residual is calculated on the power consumption data timestamp according to the second preset formula, wherein the second preset formula is: , The time residual is K, where K is the number of reference time points, and K is greater than or equal to 2. This refers to the time point of the power consumption data of the j-th PDU after the timestamp of the j-th PDU has undergone dynamic delay compensation. This is the base timestamp of the i-th CPU;

[0023] Based on the time residual, a unified timestamp is calculated using a preset time mapping function to obtain a unified timestamp hardware layer dataset. This hardware layer dataset includes the CPU temperature data and the real-time power consumption data of the PDU after the unified timestamp. The time mapping function is: , The unified timestamp after alignment For time residuals, For dynamic compensation coefficients, The timestamp of the PDU to be mapped The timestamp of the current PDU The most recent CPU benchmark time point, This represents the local time deviation between the current PDU time point and the nearest CPU time point.

[0024] Furthermore, the step of calculating the target time point value using cubic spline interpolation on the data collected by the software layer includes:

[0025] Obtain the set of data points and the target time point from the data collected by the software layer;

[0026] The target time point value is calculated using a cubic spline function, wherein the cubic spline function is: , The target time point value, the For constant terms, The coefficient of the linear term, The coefficient of the quadratic term, The coefficient of the cubic term, As the independent variable, This represents the k-th data acquisition time.

[0027] Furthermore, the step of aggregating network layer data using an exponentially weighted moving average includes:

[0028] Define the time window centered on the hardware reference time;

[0029] Calculate the exponential weight for each network data point within the time window;

[0030] The weighted aggregate value is calculated based on the exponential weights to obtain the time-aligned network layer acquisition data.

[0031] Furthermore, the step of calculating the server stress index based on the heterogeneous dataset and generating the topological affinity matrix includes:

[0032] The server stress index , To calculate the weighting coefficients, , For memory weighting coefficients, , These are the network weight coefficients. , The CPU utilization is represented by (memory / 100), and the bus bandwidth usage percentage is represented by (memory / 100). For the number of packets retransmitted over the network, + + =1;

[0033] The topological affinity matrix ,|delay i -Delay j | represents the communication delay from server i to j.

[0034] Furthermore, to achieve the above objectives, the present invention also provides a server optimization system based on AI big data, comprising:

[0035] The acquisition module is used to acquire hardware layer data, software layer data, and network layer data. The hardware layer data includes CPU temperature data and PDU real-time power consumption data. The software layer data includes container memory usage and API gateway request response time. The network layer data includes flow table statistics and server communication latency matrix.

[0036] The time alignment processing module is used to perform time alignment processing on the data collected by the hardware layer, the data collected by the software layer, and the data collected by the network layer to obtain a heterogeneous dataset, wherein the heterogeneous dataset is the dataset obtained after time alignment processing of the data collected by the hardware layer, the data collected by the software layer, and the data collected by the network layer.

[0037] The calculation module is used to calculate the server stress index based on heterogeneous datasets and generate a topological affinity matrix;

[0038] A dual-channel processing module is used to perform dual-channel processing based on the server stress index and the topology affinity matrix, wherein the short-term prediction channel is performed based on the server stress index, and the topology analysis channel is performed based on the topology affinity matrix.

[0039] The decision fusion module combines the results of the short-term prediction channel with the results of the topology analysis channel to make decisions and execute resource scheduling instructions.

[0040] In addition, to achieve the above objectives, the present invention also provides a server optimization device based on AI big data. The server optimization device based on AI big data includes: a memory, a processor, and an AI big data-based server optimization program stored in the memory and executable on the processor. When the AI ​​big data-based server optimization program is executed by the processor, it implements the steps of the server optimization method based on AI big data as described above.

[0041] In addition, to achieve the above objectives, the present invention also provides a readable storage medium storing a server optimization program based on AI big data, wherein when the server optimization program based on AI big data is executed by a processor, it implements the steps of the server optimization method based on AI big data as described above.

[0042] This application acquires data from the hardware layer, software layer, and network layer, then performs time alignment processing on these data to obtain a heterogeneous dataset. Based on this heterogeneous dataset, a server stress index is calculated, and a topology affinity matrix is ​​generated. Dual-channel processing is then performed on the server stress index and the topology affinity matrix. The results of the short-term prediction channel processing and the topology analysis channel processing are fused for decision-making, and finally, resource scheduling instructions are executed. This application can accurately assess resources through multi-dimensional feature fusion analysis, and its dual-channel decision-making mechanism and adaptive resource scheduling can reduce energy consumption, shorten response time, and effectively improve resource utilization. Attached Figure Description

[0043] Figure 1 This is a flowchart illustrating an embodiment of server optimization based on AI big data in this application;

[0044] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0046] This invention further provides a server optimization method based on AI big data. (Refer to...) Figure 1 , Figure 1This is a flowchart illustrating an embodiment of the server optimization method based on AI big data according to the present invention.

[0047] In this embodiment, the execution entity of the AI-based big data server optimization method is an AI-based big data server optimization system. This system includes an AI-based big data server optimization device or equipment, which can be a PC, PDA, or other terminal device. This invention acquires hardware-layer, software-layer, and network-layer data, then performs time-alignment processing on these data to obtain a heterogeneous dataset. Based on this heterogeneous dataset, a server stress index is calculated, and a topology affinity matrix is ​​generated. Dual-channel processing is then performed on the server stress index and the topology affinity matrix. The results of the short-term prediction channel processing and the topology analysis channel processing are fused for decision-making, and finally, resource scheduling instructions are executed. This application can accurately assess resources through multi-dimensional feature fusion analysis, and through a dual-channel decision-making mechanism and adaptive resource scheduling, it can reduce energy consumption, shorten response time, and effectively improve resource utilization.

[0048] The steps of this AI-based big data-driven server optimization method include:

[0049] Step S10: Acquire hardware layer data, software layer data, and network layer data;

[0050] In this embodiment, the hardware layer collects data including CPU (central processing unit) temperature data and PDU (Power Distribution Unit) real-time power consumption data; the software layer collects data including container memory usage and API gateway request response time; and the network layer collects data including flow table statistics and server communication latency matrix. CPU temperature data can be obtained via the IPMI (Intelligent Platform Management Interface) protocol, and PDU real-time power consumption data can be directly read. The Prometheus exporter captures container memory usage. A container refers to an isolated environment created based on operating system-level virtualization technology, achieving hardware resource isolation through cgroups and process, network, and file system isolation through namespaces. Each container shares the host kernel but has an independent user space. The AI ​​big data server optimization system also records the API gateway request response time. By parsing the flow table, flow table statistics, such as those for OpenFlow flow tables, can be obtained, and a server communication latency matrix (N×N array) can be constructed.

[0051] In the above data acquisition sampling frequencies, the hardware layer sampling frequency can be 1Hz, the software layer sampling frequency can be 0.2Hz, and the network layer sampling frequency can be 5Hz.

[0052] Step S20: Perform time alignment processing on the data collected from the hardware layer, the data collected from the software layer, and the data collected from the network layer to obtain a heterogeneous dataset;

[0053] In this embodiment, the heterogeneous dataset is a dataset obtained by time-aligning data collected from the hardware layer, software layer, and network layer. The hardware layer data can be given a unified time axis, the software layer data can be interpolated using cubic spline interpolation, and the network layer data can be processed using a moving average.

[0054] Step S20 includes:

[0055] Step S21: Extract the reference time source from the CPU temperature data, wherein the original timestamp of the hardware layer sensor clock is selected as the reference time source.

[0056] Step S22: Determine whether the PDU is a high-precision PDU. If the PDU is a high-precision PDU, directly obtain the power consumption data timestamp. If the PDU is a normal PDU, apply dynamic delay compensation to obtain the power consumption data timestamp.

[0057] In one embodiment, the method of unifying the time axis for data acquisition at the hardware layer is as follows: the hardware layer sensor clock signal (raw timestamp) is selected as the reference time source from the CPU temperature data, and its clock error is less than ±1ms; it is determined whether the PDU is a high-precision PDU. If the PDU is a high-precision PDU, the power consumption data timestamp is directly obtained. If the PDU is a normal PDU, dynamic delay compensation is applied to obtain the power consumption data timestamp. Here, a high-precision PDU refers to a PDU device equipped with a dedicated clock chip, and a normal PDU refers to a PDU device that relies on a software clock.

[0058] The aforementioned dynamic delay compensation is performed by calculating the corrected timestamp using a first preset formula, and then using the corrected PDU timestamp as the power consumption data timestamp. The first preset formula is: , The corrected PDU timestamp As a smoothing factor, For the precise local time of the data acquisition server, The network time obtained by the PDU device via the NTP protocol. This is the original timestamp of the PDU.

[0059] Step S23: Match the reference time source with the power consumption data timestamp to unify the hardware layer time axis;

[0060] The reference time source extracted from the CPU temperature data is matched with the power consumption timestamp of the PDU (corrected timestamp) to ultimately unify the hardware layer timeline.

[0061] Specifically, firstly, based on the reference time source, the time residual is calculated on the power consumption data timestamp according to the second preset formula, whereby: , The time residual is K, where K is the number of reference time points, and K is greater than or equal to 2. This refers to the time point of the power consumption data of the j-th PDU after the timestamp of the j-th PDU has undergone dynamic delay compensation. This is the base inter-stamp of the i-th CPU;

[0062] Then, based on the time residual, a unified timestamp is calculated using a preset time mapping function to obtain a unified timestamp hardware layer dataset. The hardware layer dataset includes the CPU temperature data after the unified timestamp and the real-time power consumption data of the PDU after the unified timestamp.

[0063] The time mapping function is: , For the aligned unified timestamp, For time residuals, For dynamic compensation coefficients, The timestamp of the PDU to be mapped The timestamp of the current PDU The most recent CPU benchmark time point, This represents the local time deviation between the current PDU time point and the nearest CPU time point.

[0064] Step S24: Calculate the target time point value using cubic spline interpolation on the data collected by the software layer;

[0065] In one embodiment, the method for time alignment processing of data acquired by the software layer includes cubic spline interpolation, which is used to calculate the target time point value.

[0066] Specifically, first, the data point set and target time point from the software layer's collected data are obtained. Then, the value of the target time point is calculated using a cubic spline function. The data point set from the software layer's collected data is used to construct the cubic spline function as follows: , Values ​​at the target time point. This is a constant term (starting index value). It is the coefficient of the linear term (the instantaneous rate of change at the starting point). The coefficient of the quadratic term (initial curvature intensity). The coefficient of the cubic term (rate of change of curvature). The independent variable (the moment when interpolation needs to be calculated). This represents the k-th data acquisition time.

[0067] The data point set in the software layer data collection is , Let i be the time of data collection for the i-th data point. For a moment Monitoring metrics (such as container memory usage), and the target time point for data collection at the software layer. This is the hardware layer reference time. Construct the above cubic spline function over the interval.

[0068] To eliminate endpoint oscillations, natural spline boundary conditions are used:

[0069]

[0070] in, The time of the first node in the data sequence. The time of the last node in the data sequence. Let be the second derivative at the left endpoint, and be the curvature at the starting point, indicating that the function at... The unevenness of the surface; Let be the second derivative at the right endpoint, and be the curvature at the endpoint, indicating that the function at... The unevenness of the surface.

[0071] The specific steps for constructing and solving tridiagonal equations are as follows:

[0072] First, calculate the step size: , This represents the time interval between data points.

[0073] Then construct a system of equations: , ;in, , for nodes The second derivative at that point; , where is the weighting coefficient for the left interval; , where is the weighting coefficient of the right interval. , is the right-hand term.

[0074] Boundary condition injection: transforming natural boundary conditions into...

[0075] Then, the chasing method is used to solve the problem: forward elimination:

[0076]

[0077]

[0078]

[0079]

[0080] Back-substitution solution:

[0081]

[0082]

[0083] Then, the interpolation coefficients are calculated: [The results are incomplete in the original text.] Then calculate the coefficients for each interval.

[0084]

[0085]

[0086]

[0087]

[0088] in, , For monitoring values ​​at the interval endpoints, The interval length is... , is the second derivative of the endpoints of the interval.

[0089] Step S25: Aggregate the network layer data using an exponentially weighted moving average.

[0090] In one embodiment, the method for time-aligning network layer acquisition data is exponentially weighted moving average aggregation. First, a time window is defined with the hardware reference time as the center. Then, an exponential weight is calculated for each network data point within the time window. Based on the exponential weight, a weighted aggregation value is calculated to obtain the time-aligned network layer acquisition data.

[0091] Specifically, firstly, based on hardware reference time Construct a symmetrical time window centered on the time window. ;

[0092] Then, dynamic adjustments are made, with the following rules: base value. =0.5ms, if there is no data in the window: expand to =1.0ms, if network jitter rate >30%, shrink to =0.3ms, where, The hardware reference time is obtained through the CPU temperature sensor clock. The radius of the time window.

[0093] Next, time distance calculation is performed, calculating the absolute time distance for each data point within the window. , For absolute time distance, The data is timestamped on the network; conditions are imposed as follows: If the constraints are not met, then that point is excluded.

[0094] Then, the exponential weights are calculated using the exponential decay function: , The attenuation coefficient is [0.5, 1.0]; a weighted truncation mechanism is used to remove far-end noise points with a contribution of <1%, increasing the effective signal ratio by 23%. ,in The weights are the processed weights.

[0095] Finally, weighted aggregation is performed to output the aligned network layer acquisition data. This includes normalizing and weighting the valid data points within the window. , This is the aggregated output, representing the aligned network metrics. For valid points, satisfying The number of points, These are network metric values, which are the original network data.

[0096] Step S30: Calculate the server stress index based on the heterogeneous dataset and generate the topological affinity matrix;

[0097] In this embodiment, the server stress index , To calculate the weighting coefficients, , For memory weighting coefficients, , These are the network weight coefficients. (Memory / 100) represents CPU utilization. For bus bandwidth usage ratio, For the number of packets retransmitted over the network, + + =1;

[0098] Topological affinity matrix ,|delay i -Delay j | represents the communication delay from server i to j.

[0099] Step S40: Perform dual-channel processing based on the server stress index and topology affinity matrix;

[0100] In this embodiment, short-term prediction channel processing is performed based on the server stress index, and topology analysis channel processing is performed based on the topology affinity matrix.

[0101] The short-term forecast channel processing based on the server stress index specifically includes: using the server stress index as data input, and the server stress index sequence within the time window is as follows:

[0102] The window size is 6, the span is 5 minutes, and LSTM is used for modeling.

[0103]

[0104]

[0105]

[0106]

[0107]

[0108]

[0109] Temperature sensing gating is introduced in the above process (the weight of the forget gate is increased when the CPU is >80℃).

[0110] Employing an attention mechanism:

[0111]

[0112] Predicted output:

[0113]

[0114] For example, the load probability distribution for the next 30 seconds (e.g., [CPU:0.85, Mem:0.72, Net:0.15]).

[0115] Topology analysis channel processing based on the topology affinity matrix specifically includes: constructing the affinity matrix. ;

[0116] Then, the weights are calculated: ,in, This is the network distance weight (default 0.5). Weight for service call frequency (default 0.3). The weight for resource complementarity (default 0.2). .

[0117] Step S50: The results of the short-term prediction channel processing and the results of the topology analysis channel processing are fused together for decision-making, and resource scheduling instructions are executed.

[0118] In this embodiment, the decision-making process is determined by fusion decision triggering conditions, which are as follows:

[0119] Enable topology decision-making, where, The threshold is dynamic (default 0.15).

[0120] Resource scheduling is performed using resource scheduling rules, which include expansion decisions, migration decisions, and energy-saving decisions.

[0121] In the expansion decision, priority is given to "backup clusters" in the same community, and the next priority is neighboring communities with an affinity of >0.6.

[0122] In the migration decision, high-pressure nodes are migrated to low-pressure communities. The migration delay must be less than the interruption time × 0.2. For example, the migration delay must be < 5ms × 0.2 = 1ms.

[0123] In energy-saving decision-making, the shutdown order is as follows: servers within the same PDU power supply unit, then in descending order of affinity (retaining highly connected nodes), ensuring that at least two nodes are retained in each community, including nodes that can be shut down. ∩{non-boundary nodes}, where For server node i, Let be the pressure index of node i.

[0124] Furthermore, this invention also proposes a server optimization system based on AI big data, which includes:

[0125] The acquisition module is used to acquire hardware layer data, software layer data, and network layer data. The hardware layer data includes CPU temperature data and PDU real-time power consumption data. The software layer data includes container memory usage and API gateway request response time. The network layer data includes flow table statistics and server communication latency matrix.

[0126] The time alignment processing module is used to perform time alignment processing on the data collected by the hardware layer, the data collected by the software layer, and the data collected by the network layer to obtain a heterogeneous dataset, wherein the heterogeneous dataset is the dataset obtained after time alignment processing of the data collected by the hardware layer, the data collected by the software layer, and the data collected by the network layer.

[0127] The calculation module is used to calculate the server stress index based on heterogeneous datasets and generate a topological affinity matrix;

[0128] A dual-channel processing module is used to perform dual-channel processing based on the server stress index and the topology affinity matrix, wherein the short-term prediction channel is performed based on the server stress index, and the topology analysis channel is performed based on the topology affinity matrix.

[0129] The decision fusion module combines the results of the short-term prediction channel with the results of the topology analysis channel to make decisions and execute resource scheduling instructions.

[0130] Furthermore, this embodiment of the invention also provides a server optimization device based on AI big data. The server optimization device based on AI big data includes: a memory, a processor, and an AI big data-based server optimization program stored in the memory and executable on the processor. When the AI ​​big data-based server optimization program is executed by the processor, it implements the steps of the above-described AI big data-based server optimization method.

[0131] In addition, this embodiment also provides a readable storage medium storing a server optimization program based on AI big data. When the server optimization program based on AI big data is executed by the processor, it implements the above-mentioned steps of the server optimization method based on AI big data.

[0132] The processor and memory can be connected via a bus or other means.

[0133] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0134] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0135] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0136] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0137] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. An AI big data-based server optimization method, characterized by, The AI big data-based server optimization method comprises the following steps: acquiring hardware layer collected data, software layer collected data and network layer collected data, wherein the hardware layer collected data comprises CPU temperature data and real-time power consumption data of a PDU, the software layer collected data comprises container memory occupancy rate and API gateway request response time, and the network layer collected data comprises flow table statistical information and server communication delay matrix; performing time alignment processing on the hardware layer collected data, the software layer collected data and the network layer collected data to obtain a heterogeneous data set; calculating a server stress index based on the heterogeneous data set and generating a topology affinity matrix; performing double-channel processing based on the server stress index and the topology affinity matrix, wherein short-term prediction channel processing is performed according to the server stress index, and topology analysis channel processing is performed according to the topology affinity matrix; performing decision fusion on the results of the short-term prediction channel processing and the topology analysis channel processing, and executing resource scheduling instructions, wherein the decision fusion refers to determining whether to enable topology decision by using a fusion decision trigger condition, and if the fusion decision trigger condition is met, the topology decision is enabled, and the execution of the resource scheduling instructions refers to resource scheduling by using resource scheduling rules, and the resource scheduling rules comprise expansion decision, migration decision and energy saving decision.

2. The AI big data-based server optimization method of claim 1, wherein, The step of performing time alignment processing on the hardware layer collected data, the software layer collected data and the network layer collected data to obtain a heterogeneous data set comprises the following steps: extracting a reference time source from the CPU temperature data, wherein an original time stamp of a hardware layer sensor clock is selected as the reference time source; determining whether the PDU is a high-precision PDU, if the PDU is a high-precision PDU, directly acquiring power consumption data time stamps, if the PDU is a normal PDU, applying dynamic delay compensation to obtain the power consumption data time stamps, wherein the high-precision PDU refers to a PDU device equipped with a dedicated clock chip, and the normal PDU refers to a PDU device relying on a software clock; matching the reference time source and the power consumption data time stamps to unify the hardware layer time axis; calculating target time point values for the software layer collected data by using a cubic spline interpolation method; aggregating network layer data by using an exponential weighted moving average. 3.The AI big data-based server optimization method of claim 2, wherein, If the PDU is a normal PDU, the step of applying dynamic delay compensation to obtain the power consumption data time stamps comprises the following steps: if the PDU is a normal PDU, calculating a corrected time stamp by using a first preset formula, and taking the corrected PDU time stamp as the power consumption data time stamp, wherein the first preset formula is: , is the corrected PDU timestamp, is the smoothing factor, is the accurate local time of the data collection server, is the network time acquired by the PDU device through the NTP protocol, is the original timestamp of the PDU.

4. The AI big data-based server optimization method of claim 3, wherein The step of matching the reference time source and the power consumption data time stamps to unify the hardware layer time axis comprises the following steps: The time residual is calculated according to a second preset formula on the basis of the reference time source and the power consumption data timestamp, wherein the second preset formula is: , is the time residual, K is the number of reference time points, K is greater than or equal to 2, is the jth PDU power consumption data time point after dynamic delay compensation of the jth PDU timestamp, is the ith CPU reference timestamp; According to the time residual, a statistical uniform timestamp is calculated by a preset time mapping function to obtain a hardware layer dataset of the uniform timestamp, wherein the hardware layer dataset comprises the CPU temperature data after the uniform timestamp and real-time power consumption data of the PDU after the uniform timestamp, and the time mapping function is: , is the aligned uniform timestamp, is the time residual, is a dynamic compensation coefficient, is a PDU timestamp to be mapped, is a local time deviation of a current PDU time point from a nearest CPU point. is a nearest CPU reference time point, is a local time deviation of a current PDU time point from a nearest CPU point. 5.The AI big data-based server optimization method of claim 2, wherein, The step of calculating target time point values for the software layer collected data by using a cubic spline interpolation method comprises the following steps: acquire a set of data points in the software layer collected data and a target time point; The target time point value is calculated by using a cubic spline function, wherein the cubic spline function is: , is the target time point value, the target time point value is calculated by using a cubic spline function, wherein the cubic spline function is: is a constant term, is a linear term coefficient, is a quadratic term coefficient, is a cubic term coefficient, is an independent variable, is the kth data acquisition time.

6. The AI big data-based server optimization method of claim 2, wherein, the step of aggregating the network layer data by exponential weighted moving average comprises: defining a time window centered on a hardware reference time; calculating an exponential weight for each network data point in the time window; calculating a weighted aggregation value according to the exponential weight to obtain the network layer collected data after time alignment processing.

7. The AI big data-based server optimization method of claim 2, wherein, the step of calculating a server stress index based on the heterogeneous data set and generating a topology affinity matrix comprises: The server stress index , For calculating the weight coefficient, , For memory weight coefficient, , For network weight coefficient, , CPU utilization rate, (memory / 100) is the bus bandwidth usage ratio, Network retransmission packet number, + + =1; The topological affinity matrix | delay i - delay j | is the communication delay from server i to j.

8. An AI big data-based server optimization system, characterized by, the AI big data based server optimization system comprises: an acquisition module for acquiring hardware layer collected data, software layer collected data and network layer collected data, wherein the hardware layer collected data includes CPU temperature data and real-time power consumption data of PDU, the software layer collected data includes container memory occupancy rate and API gateway request response time, and the network layer collected data includes flow table statistical information and server communication delay matrix; a time alignment processing module for performing time alignment processing on the hardware layer collected data, the software layer collected data and the network layer collected data to obtain a heterogeneous data set; a calculation module for calculating a server stress index based on the heterogeneous data set and generating a topology affinity matrix; a dual-channel processing module for performing dual-channel processing based on the server stress index and the topology affinity matrix, wherein short-term prediction channel processing is performed according to the server stress index, and topology analysis channel processing is performed according to the topology affinity matrix; a decision fusion module for performing decision fusion on the results of short-term prediction channel processing and topology analysis channel processing, and executing resource scheduling instructions, wherein the decision fusion refers to determining whether to enable topology decision by using a fusion decision trigger condition, and if the fusion decision trigger condition is met, the topology decision is enabled, and the execution of resource scheduling instructions refers to resource scheduling by using resource scheduling rules, and the resource scheduling rules include expansion decision, migration decision and energy saving decision. 9.A server optimization apparatus based on AI big data, characterized by, The AI big data based server optimization device comprises a memory, a processor and an AI big data based server optimization program stored on the memory and executable on the processor, and the AI big data based server optimization program implements the steps of the method according to any one of claims 1 to 7 when executed by the processor.

10. A readable storage medium, characterized by, The readable storage medium stores an AI big data based server optimization program, and the AI big data based server optimization program implements the steps of the method according to any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • Multi-source computing power data integration and intelligent scheduling system and method

    CN118916147A

  • Decision-making large model-oriented multi-level heterogeneous memory collaborative scheduling method

    CN119576555A