Method for analyzing affinity of application load and computing storage system

By combining deep learning methods of computing storage systems and application load characteristics, neural network models are trained, and the problem of insufficient platform adaptability in the existing technology is solved, automated analysis and accurate prediction of computing storage systems and application loads are realized, and system resource utilization and application performance are improved.

CN120386698APending Publication Date: 2025-07-29XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510432675.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing technology is difficult to take into account various platforms, cannot adapt to the dynamic needs of different storage systems, and the model cannot fully reflect the complexity of the system and application load, resulting in inaccurate prediction results and insufficient capture of details that may affect performance.

Method used

Combining the characteristics of the computing storage system and application load characteristics, deep learning methods are used to characterize the resource consumption mode of the computing storage system and application load, and use neural network models for automated analysis, and train neural network models to predict the affinity between the application load and the computing storage system.

Benefits of technology

It realizes automated analysis of various types of computing storage systems, comprehensively characterizes system resources and application behavior, improves prediction accuracy and adaptability, and can better match application load and computing storage systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386698A_ABST
    Figure CN120386698A_ABST
Patent Text Reader

Abstract

The invention discloses a method for analyzing the affinity of an application load and a computing storage system. The method comprises the following implementation steps of: 1, describing the characteristics of the computing storage system, quantifying and storing the characteristics of the computing storage system by a characteristic module of the computing storage system; 2, the application load feature description module quantifies the features of the universal application load by acquiring an application runtime index from a time sequence database Prometheus; 3, a neural network model training and predictive analysis module trains the neural network model through the application load features and the calculation storage system features; and 4, the trained model calculates the affinity evaluation result between the application load and the calculation storage system. According to the invention, through the affinity analysis of the application load and the calculation storage system, the capability of the calculation and storage system is maximized, and the overall performance of the application load is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electronic digital data processing, and further relates to a method for analyzing the affinity between an application load and a computing and storage system in the technical field of computer hardware performance testing. The present invention can be used to analyze the degree of goodness or badness of the affinity between the two according to the characteristics of the computing and storage system and the characteristics of the application load, and provide a basis for more reasonably allocating the application load to the computing and storage system. Background Art

[0002] Currently, with the rapid development of computer technology, the types of computing and storage are becoming increasingly rich, and the differences in computing and storage requirements between computer application loads and within application loads are gradually becoming significant. This difference has led to a huge gap between dynamic application loads and static computing and storage infrastructures, making it a difficult problem for application loads to operate efficiently on existing systems. Generally, application loads need to go through a process of small-scale trial runs and continuous optimization before gradually improving the operation scale and efficiency. However, this process is inefficient and the optimization cost is extremely high. For example, the numerical simulation software in a certain hydrometeorological field took more than half a year to be ported and optimized on specific computing and storage facilities before achieving large-scale operation. The fundamental problem is that the computing and storage model of this application load does not match the system design, resulting in challenges in subsequent performance optimization. The existing technology scheduling decisions still rely on manual judgment; either a single indicator or a static threshold is used, which is difficult to capture the complex relationships between multi-dimensional data, insufficiently characterizes system resources, and does not fully consider the specific requirements and operating characteristics of the user application load program. As a result, the complexity of the system and application load may not be fully reflected ultimately, leading to inaccurate results.

[0003] Dell Products Limited proposed a processing device and method for storage workload allocation based on input / output (IO) mode affinity calculation in its patent document "Storage Workload Allocation Based on Input / Output Mode Affinity Calculation" (Application No. 202210719865.8, Publication No. CN 117331484 A). The processing device is configured to identify storage workloads to be run on a storage system and determine an input / output (IO) mode mix associated with the identified storage workloads. The IO mode mix includes: a first set of IO modes that characterize the types of IO operations performed by a first storage workload; and at least a second set of IO modes that characterize the types of IO operations performed by a second storage workload. The processing device is further configured to calculate an affinity metric for the IO mode mix, and the calculated affinity metric characterizes the difference between (i) the IO mode mix running concurrently and (ii) the performance metrics of the first set of IO modes and the second set of IO modes running separately. The processing device is further configured to allocate the identified storage workloads to storage devices of the storage system based on the calculated affinity metric. The problem solved by this technical solution is to optimize the input / output (IO) performance of the storage system and evaluate the affinity of various storage workloads with input / output modes. The deficiency still existing in this technical solution is that it is difficult for a fixed threshold to adapt to the dynamic requirements of different storage systems, and there is a lack of automated processing, that is, the adaptability to a dynamic environment is insufficient.

[0004] Shanghai Feiqi Network Technology Co., Ltd. disclosed an application performance prediction model based on multi-dimensional resource utilization and its setting method in its patent application document "A User SLO Modeling Method and Device Based on Application Performance Prediction Model" (Application No. CN 202211662665.X Application Publication No. CN 115934343A). This prediction model uses virtualization resource utilization to construct an application performance prediction model based on multi-dimensional resource utilization, predicts the predicted operating performance of the user application, and obtains at least one candidate resource configuration scheme that meets the predicted operating performance. In this way, virtualization resources are reasonably allocated to user business applications, while ensuring the operating performance of user applications and improving the resource utilization of the virtualization system. However, the shortcomings of this technical solution are: first, relying solely on multi-dimensional virtualization resource utilization data to train the application performance prediction model, resulting in insufficient characterization of virtualization system resources, and also failing to fully consider the specific needs and operating characteristics of user application loads; second, the use of a virtual environment makes it difficult to adapt to different platforms (especially domestic platforms). Therefore, this invention is difficult to take into account various platforms, and the model cannot fully reflect the complexity of the system and application load, resulting in inaccurate prediction results and failure to fully capture details that may affect performance. Summary of the Invention

[0005] The purpose of the present invention is to address the defects of the above-mentioned existing technologies and propose an analysis method for the affinity between application loads and computing storage systems. The method aims to solve the problems existing in the existing technologies, such as the use of fixed thresholds, which makes it difficult to take into account various platforms and cannot adapt to the dynamic needs of different storage systems; and the model cannot fully reflect the complexity of the system and application loads, resulting in inaccurate prediction results and the inability to fully capture details that may affect performance.

[0006] The technical idea for achieving the purpose of the present invention is: the method of the present invention combines the characteristics of the computing storage system and the application load characteristics, based on the affinity calculation requirements between the application load and the computing storage system, adopts a deep learning method for affinity analysis between the computing storage system and the application load, by characterizing the characteristics of the computing storage system and the application load characteristics, describing the resource consumption pattern exhibited during the operation of the application load, and combining with the neural network model, it can automatically analyze the affinity between the application load and the computing storage system, and solve the problem of comprehensively characterizing the characteristics of the computing storage system and the application load, relying on manual experience rules, and adapting to various platforms. The present invention regards various application loads and computing storage systems as a black box. There is no need to understand their internal implementation details, only to focus on their external behavior, so that the present invention can take into account various types of computing storage systems.

[0007] To achieve the above object, the steps of the present invention include the following:

[0008] Step 1: normalize the performance score and metadata information of the computing storage system to obtain the computing storage system characteristics;

[0009] Step 2: Use performance monitoring tools to monitor and collect resource usage while running common application workloads. Then, normalize and extract important features from the collected data to obtain application workload characteristics.

[0010] Step 3: Generate a training set based on the characteristics of the computing storage system and the application load, train a neural network model, and use the trained network model as an affinity prediction model;

[0011] In step 4, the characteristics of the computing storage system and the characteristics of the application load collected and processed in real time are input into the affinity prediction model in the same manner as in steps 1 and 2, and the affinity score is calculated and output.

[0012] Furthermore, obtaining the performance score of the computing storage system refers to obtaining computing performance test score data, memory performance test score data, disk performance test score data, and network performance test score data through all computing loads, memory loads, disk loads, and network loads in the benchmark program deployed on the computing storage system under test.

[0013] Furthermore, obtaining metadata information of the computing storage system refers to recording four metadata information of the computing storage system based on the output of executing the lshw command on the computing storage system under test: first, recording CPU information including CPU architecture, CPU cache size, and instruction set; second, recording memory information including memory size, memory architecture, and clock cycle; third, recording disk information including disk size and interface type; fourth, recording network card information including network card online bandwidth.

[0014] The normalization process adopts any one of minimum-maximum normalization, Z-Score normalization, decimal scaling normalization, Log normalization and exponential normalization.

[0015] Furthermore, the resource usage data obtained when the application load is running includes CPU resource usage data, memory resource usage data, disk resource usage data, and network resource usage data.

[0016] Furthermore, the steps for obtaining resource usage data during application load operation are as follows:

[0017] The first step is to deploy monitoring tools on the computing and storage systems under test.

[0018] Second, run the application payload on the computing storage system under test, and obtain four types of metric information through monitoring tools: First, CPU metric data, including the utilization rate of each core of the CPU, CPU system time, CPU user time, CPU wait time, and CPU idle time; Second, memory metric data, including the occupied size of memory, the occupied size of the file system cache, and the occupied size of the file system buffer; Third, disk metric data, including disk IOPS, disk IO queue length, disk read bytes, and disk write bytes; Fourth, network metric data, including network upload rate, network download rate, total number of bytes sent over the network, and total number of bytes received over the network;

[0019] Third, after the application payload finishes running, count the running time of the application payload, and use the time elapsed from the start of running to the end of running as the application performance.

[0020] Furthermore, the training set refers to a data set formed by combining the computing storage system characteristics and the application payload characteristics, and after performing the following processing on the data set, a generated training set;

[0021] First, data preprocessing: regard the computing storage system characteristics as hardware characteristics, regard the application payload characteristics as application characteristics, regard the running time on the application payload hardware combination as the application performance, and regard the price of the application payload hardware combination as the hardware cost of the computing storage system, obtaining the data set format as {hardware characteristics, application characteristics, application performance, hardware cost of the computing storage system};

[0022] Second, randomly shuffle the data set to obtain the shuffled data set;

[0023] Third, data encoding: limit the CPU architecture types in the shuffled data set, and create a vector with a length equal to the number of categories for each category, set the index position of this category to 1, and the rest to 0, and replace the original value of each CPU architecture with its corresponding [0,1] encoded vector.

[0024] Furthermore, the structure of the neural network model is an input layer, a linear transformation layer, an activation layer, a hidden layer, a non-linear transformation layer, and an output layer connected in series in sequence; set the dimensions of the linear transformation to 64 dimensions, 32 dimensions, and 1 dimension in sequence, and the activation layer is implemented with the Relu activation function.

[0025] Furthermore, the steps for training the neural network model are as follows:

[0026] First, set the hyperparameters for training the neural network model as: the number of iterations is 500, the optimizer is Adam, the learning rate is 0.001, and the size of each batch is 32;

[0027] In the second step, the gradient descent method is used and the Adam optimizer is used to iteratively update the network parameters until the cross entropy loss function of the network converges to obtain a trained neural network model.

[0028] The cross entropy loss function is as follows:

[0029] loss(y,y_hat)=-∑ylog(y_hat)

[0030] Where y represents the actual value of the application performance input into the neural network model, y_hat represents the predicted value of the application performance predicted by the neural network model, and log represents the logarithm operation with base 10.

[0031] Furthermore, the affinity score is obtained by the following formula:

[0032] affinity = α × application performance + (1-α) × computing storage system hardware cost

[0033] Among them, affinity represents the affinity score. The closer the value is to 0, the stronger the affinity is. α is a coefficient of 0.5.

[0034] Compared with the prior art, the present invention has the following advantages:

[0035] First, the present invention combines the characteristics of computing and storage systems with those of application loads, employing deep learning methods to automatically analyze the affinity between application loads and computing and storage systems. This overcomes the existing limitations of platform adaptability and the inadequate characterization of system resources and application behavior. This allows the present invention to analyze the affinity between application loads and computing and storage systems through a deep learning model and predict their matching. Consequently, the present invention can comprehensively characterize computing and storage system resources and application load behavior, taking into account a variety of application loads and computing and storage systems.

[0036] Second, the present invention extracts the characteristics of the computing storage system and the application load characteristics for training the constructed neural network model, and finally obtains a trained network model for calculating affinity. This overcomes the existing problems of relying on manual experience rules and lacking automated analysis in the prior art, allowing the present invention to comprehensively and automatically analyze the affinity between the computing storage system and the application load. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flow chart of affinity analysis of the present invention;

[0038] Figure 2 It is a flow chart of characterization of the computing storage system of the present invention;

[0039] Figure 3 It is a schematic diagram depicting the characteristics of the computing storage system of the present invention;

[0040] Figure 4 This is a flow chart of the application load characterization of the present invention;

[0041] Figure 5 It is a data acquisition scheme diagram of the present invention;

[0042] Figure 6 Schematic diagram of affinity analysis based on deep learning of the present invention;

[0043] Figure 7 This is a flowchart of affinity analysis based on deep learning of the present invention;

[0044] Figure 8 This is a diagram of the deep learning training model of the present invention;

[0045] Figure 9 This is the deep learning prediction model diagram of the present invention. DETAILED DESCRIPTION

[0046] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0047] The affinity mentioned in the present invention refers to a computing storage system that can enable an application load to operate with higher performance and lower cost, which is called a computing storage system with higher affinity for the application load.

[0048] Reference Figure 1 , the steps of affinity analysis of the present invention are further described.

[0049] Step 1: Characterize the computing and storage system. This involves using various computing and storage system benchmarks and metadata to characterize the system's performance. This analysis focuses on analyzing the system's performance under specific workloads, including computing power, storage capacity, and network bandwidth. Due to the limited number of devices, this paper uses Docker to simulate different hardware combinations, expand the dataset, and run the Sysbench and iperf benchmarks and typical application loads for each hardware combination.

[0050] Step 2: For the characterization of the application, use performance monitoring tools to monitor and collect resource usage when running general application loads. Then, normalize the collected data and extract important features to obtain the characteristics of the application load.

[0051] The performance characteristics of a compute-storage system are characterized using the performance scores of benchmark testing tools for various compute-storage systems and the metadata of the compute-storage system. The focus is on analyzing the performance capabilities of the compute-storage system under specific workloads, including computing power, storage capacity, and network bandwidth, etc. At the same time, due to the limited number of devices, in the present invention, Docker is used to simulate different hardware combinations, expand the dataset, and run the Sysbench and iperf benchmark testing programs and typical application workloads of the compute-storage system for each hardware combination.

[0052] Step 3, predicting affinity through a neural network model. The application characteristics and the compute-storage system characteristics are processed and then input into the neural network model. Through continuous iteration, a prediction model is obtained. This prediction model completes the affinity calculation between the application and the compute-storage system, and obtains the affinity score between the application and the compute-storage system.

[0053] Refer to Figure 2 , and a further description is given to the process of characterizing the compute-storage system features of the present invention.

[0054] The process of characterizing the compute-storage system features is data collection, data preprocessing, and outputting the compute-storage system features. In the embodiments of the present invention, the features of the compute-storage system are characterized by the performance scores of the Sysbench and iperf benchmark testing programs of various compute-storage systems and the metadata of the compute-storage system. Because just looking at the performance scores may not comprehensively reflect the characteristics of the compute-storage system, therefore, the present invention also adds metadata to further enrich the feature description.

[0055] Step 1, data collection.

[0056] For the performance information of the compute-storage system, including CPU performance information, memory performance information, disk performance information, and network performance information, the CPU performance information, memory performance information, disk performance information, and network performance information are collected by running the performance test workloads of the compute-storage system.

[0057] The performance test workloads include compute load, memory load, disk load, and network load. The compute load refers to the resources consumed by the system when executing various computing tasks, including the CPU usage rate, processing speed, and execution efficiency. The memory load refers to the memory resources used by the system during operation, including the memory usage rate, read-write speed, and storage efficiency. The disk load refers to the resources consumed by the system when reading and writing disk data, including the disk read-write speed, I / O operation times, and latency. The network load refers to the resources consumed by the system during network communication, including the network bandwidth usage rate, data transmission speed, and network latency. The main score information collected in the present invention is shown in Table 1.

[0058] Table 1 List of Benchmark Information Mainly Collected by the Present Invention

[0059] CPU single-core integer performance test score CPU single-core floating-point performance test score CPU multi-core integer performance test score CPU multi-core floating-point performance test score Memory random read test score Memory random write test score Memory sequential read test score Memory sequential write test score Disk random read test score Random disk write test scores Disk sequential read test score Disk sequential write test scores Network upload test score Network download test score

[0060] For the metadata information of the computing storage system, the present invention obtains it through system built-in commands (such as the lshw command in Linux), as shown in Table 2.

[0061] Table 2 List of Metadata Information Mainly Collected by the Present Invention

[0062] CPU CPU architecture, cache, instruction set, etc. Memory Memory size, architecture, clock cycle, etc. disk Disk size, interface type, etc. Network Card Network card bandwidth limit

[0063] Step 2, data processing.

[0064] Since there are significant differences in the benchmark scores of different computing storage system performance test loads, the present invention scales these scores to the same order of magnitude to retain the relative differences between the computing storage systems.

[0065] To this end, the present invention uses the Min-Max normalization method to scale the data to a fixed range (usually 0 to 1), and its formula is as follows:

[0066]

[0067] Among them, x' represents the normalized performance score of the computing storage system, x represents the performance score of the computing storage system before normalization, x min represents the minimum score of the performance score of the computing storage system, x max represents the maximum score of the performance score of the computing storage system.

[0068] The specific steps of data processing are as follows:

[0069] The first step is to collect the original data and obtain the original scores of all benchmark programs in Sysbench and iperf. Each benchmark will have scores of multiple hardware devices.

[0070] The second step is that the embodiment of the present invention selects Min-Max normalization.

[0071] The third step is to calculate the normalization parameters and obtain the minimum and maximum values of each benchmark.

[0072] The fourth step is to apply the Min-Max normalization formula to normalize the scores of each benchmark.

[0073] Step 3, output the characteristics of the computing storage system as shown in Table 3.

[0074] Table 3 List of Characteristics of the Computing Storage System

[0075]

[0076]

[0077] Referring to Figure 3 , the entire process of characterizing the computing storage system features of the present invention will be further described.

[0078] Since the present invention needs to accommodate various application workloads, the present invention regards it as a black box, that is, a binary program. Therefore, when characterizing the application performance features, there is no need to understand its internal implementation details, and only the external behavior and resource consumption of the application workload need to be concerned about, so as to accommodate various application workloads. In this case, the embodiments of the present invention utilize system resources to execute their functions.

[0079] For this reason, the present invention uses the Prometheus monitoring tool to monitor and collect its resource usage when running the application workload, and uses this as an indicator for feature characterization. In order to ensure that the selected indicators are representative, the present invention adopts the core indicators that most monitoring systems will focus on.

[0080] Referring to Figure 4 , the application workload characterization process of the present invention will be further described.

[0081] Step A, indicator collection.

[0082] The application workload indicators selected by the present invention mainly include the runtime data of the application workload and some meta-information. The present invention uses Prometheus combined with a custom Exporter for data collection. The specific application workload indicators collected are shown in Table 4.

[0083] Table 4 List of application workload indicators collected in the embodiments of the present invention

[0084]

[0085]

[0086] The reasons for selecting these indicators to characterize the application workload features are as follows:

[0087] Utilization rate of each core of the CPU: High utilization rate may indicate that the application workload is fully utilizing all available cores.

[0088] CPU system state time: Measures the time consumed by the operating system kernel for processing system calls and other kernel-level tasks. If this time accounts for too large a proportion, it may indicate frequent system calls.

[0089] CPU user state time: Measures the time consumed by the operating system for executing user code.

[0090] CPU wait time: Measures the time spent waiting for I / O or other events to occur. If the wait time is too long, it may indicate poor I / O performance and may require better I / O devices or reducing I / O requests.

[0091] CPU Idle Time: Measures the amount of time the processor is idle and not performing any useful work. If the idle time is too long, it may indicate that the CPU resources are not being used efficiently.

[0092] Memory usage: Displays memory usage to understand whether the application load uses a lot of memory.

[0093] File system cache usage: A high cache usage may indicate frequent file read I / O operations.

[0094] File system buffer occupancy: A high buffer occupancy may indicate frequent file write I / O operations.

[0095] Disk IOPS: High IOPS may indicate that the disk is busy and the application load is placing high demands on the disk.

[0096] Disk IO queue length: A long queue length may indicate a backlog of I / O requests and may require more disk.

[0097] Disk Read Bytes and Disk Write Bytes: High read and write volumes might indicate heavy disk usage and the need for more disk space or increased disk capacity.

[0098] Network upload rate and network download rate: can indicate the demand of application load on the network.

[0099] Total Network Bytes Sent and Total Network Bytes Received: High byte counts may indicate high network activity and a need for increased bandwidth.

[0100] Impact of multiple processors on program performance: Small performance gains may indicate insufficient multi-threading optimization, requiring optimization of the multi-threaded code or increasing the number of processors.

[0101] Impact of thread synchronization, lock contention, and context switching on performance: Frequent context switching may indicate improper use of thread synchronization and locks, requiring concurrent code optimization.

[0102] Step B: data preprocessing.

[0103] During the monitoring process, various factors may cause noise and outliers in the data. Therefore, these inaccurate data need to be removed before analysis to ensure the reliability of the data. First, the data needs to be comprehensively reviewed and data exploratory analysis methods, such as drawing line graphs, histograms, box plots, etc., are used to identify potential data quality issues. Secondly, there may be missing values in the data, which is usually caused by equipment failure or other reasons during the data collection process. If the proportion of missing values is small, then delete the records containing missing values. For numerical data, the mean, median or mode can be used to fill in the missing values. Here, in an embodiment of the present invention, the mean value is selected to fill in the missing data.

[0104] After completing the missing value processing, it is necessary to normalize the data of different indicators. In the embodiment of the present invention, this step is the same as the data preprocessing for computing the storage system feature characterization, and adopts the Min-Max normalization method.

[0105] After the above processing, the final application load feature format is shown in Table 5.

[0106] Table 5 Application load feature format list

[0107] serial number Indicator unit 1 CPU core utilization percentage 2 CPU system state time Second 3 CPU user mode time Second 4 CPU wait time Second 5 CPU idle time Second 6 Memory usage GB 7 File system cache size GB 8 File system buffer size GB 9 Disk IOPS none 10 Disk IO queue length none 11 Disk read bytes GB 12 Disk write bytes GB 13 Network upload rate GB / s 14 Network download speed GB / s 15 Total bytes sent over the network GB 16 The total number of bytes received by the network GB

[0108] Reference Figure 5 , further describes the data collection scheme of the embodiment of the present invention.

[0109] This embodiment of the present invention uses Docker to deploy and run application workloads, and combines Prometheus with a custom exporter to collect relevant data. Due to the limited number of devices, this embodiment of the present invention uses Docker to simulate different hardware combinations, expand the data set, and run the Sysbench and iperf benchmarks for the computing and storage system, as well as typical application workloads, on each hardware combination.

[0110] Step S1: Prepare a typical application load as shown in Table 6.

[0111] Table 6 List of application loads in embodiments of the present invention

[0112]

[0113] Step S2: Run a typical application load and collect data.

[0114] In the embodiments of the present invention, Prometheus and MySQL are deployed on the control node, and Docker and node Exporter are deployed on the measured computing and storage system to monitor and collect the runtime data of the application in real time. The Sysbench, iperf benchmark test programs for the computing and storage system and Docker images of typical application loads are prepared in advance.

[0115] Overall operation process of the control node side control script:

[0116] Step S2.1, traverse all available measured computing and storage systems.

[0117] Step S2.2, obtain the hardware meta-information of the measured computing and storage system.

[0118] Step S2.3, determine the resource limit combination of Docker and calculate the hardware cost according to the ratio.

[0119] Step S3.1, traverse the number of CPU cores, from 1 to N (assuming there are N cores).

[0120] Step S3.2, traverse the memory size, from 0.25 to 1, increasing by 0.25 each time, i.e., [0.25, 0.5, 0.75, 1] (divide the memory into 4 segments to simulate different memory situations).

[0121] Step S3.3, traverse the IOPS of the disk, from 0.25 times to 1 time, increasing by 0.25 times each time, divide the performance of the disk into four segments and gradually improve to simulate different disk situations.

[0122] Step S4, traverse all typical application loads.

[0123] Step S5, generate test tasks, insert the tasks into the database and send them to the measured computing and storage system for execution.

[0124] Finally, the output data format of the test tasks is shown in Table 7

[0125] Table 7 List of Output Data Formats of Test Tasks

[0126]

[0127] The overall operation process of the measured computing and storage system includes the following steps:

[0128] Step 1) Run the Sysbench and iperf benchmark test programs, generate hardware characteristics and insert them into the database;

[0129] Step 2) Run typical application loads, generate application load characteristics and the running time of the application load, and insert them into the database.

[0130] Reference Figure 6 , further describes the affinity analysis process based on regression prediction in an embodiment of the present invention.

[0131] Deep learning-based affinity analysis uses a deep learning model to calculate the affinity between application loads and computing and storage systems, making it suitable for more complex scenarios. The model inputs the collected application load and computing and storage system characteristics, and outputs an affinity score. Application load characteristics are extracted using the application load characterization method, while computing and storage system characteristics are extracted using the computing and storage system characterization method. Affinity is characterized using the application load's performance objectives and costs.

[0132] Reference Figure 7 , further describing the affinity analysis process based on deep learning of the present invention.

[0133] Step 1) Data preprocessing.

[0134] In the previous steps, the present invention has completed data collection, and the final data format obtained is (hardware characteristics, application characteristics, application performance, and hardware cost in the computing storage system). Among them, hardware characteristics: the computing storage system characteristics mentioned above, which are a vector; application characteristics: the application load characteristics mentioned above, which are a vector; application performance: the time it takes for the application to run on this hardware combination, which is a scalar; hardware cost in the computing storage system: the price of the hardware combination.

[0135] Step 2) Split the dataset.

[0136] The dataset needs to be randomly shuffled first and then split into a training set, a validation set, and a test set, with the training set accounting for 70%, the validation set accounting for 15%, and the test set accounting for 15%.

[0137] Step 3) Encode the input data:

[0138] For CPU architecture data, it is not feasible to directly input the neural network, so one-hot encoding is used for processing.

[0139] (3a) Determining the number of categories: The embodiment of the present invention limits the CPU architecture to five types: [x86, arm, sw, riscv, misp];

[0140] (3b) Construct the encoding: Create a vector for each category with a length equal to the number of categories, set the index position of the category to 1, and the remaining positions to 0.

[0141] (3c) Data encoding: Replace the original value of each CPU architecture with its corresponding [0,1] encoding vector.

[0142] The remaining data can be directly input into the neural network as scalar data.

[0143] Step 4) Train the model:

[0144] This method uses a DNN (Deep Neural Network) to build a model, aiming to capture the characteristics of the application load from the operation monitoring data of the application load and combine the characteristic data of the hardware to fit the affinity target.

[0145] Refer to Figure 8 , for a further description of the deep learning training mode of the present invention.

[0146] The basic unit of the DNN model is a neuron. Each neuron receives inputs from other neurons and changes the influence of the inputs on the neuron by adjusting the weights. In the DNN, data starts from the input layer, goes through successive calculations in multiple hidden layers, and finally reaches the output layer. At the same time, before the output of each layer of neurons serves as the input of the next layer of neurons, a non-linear transformation is achieved through an activation function. The entire training process relies on the backpropagation algorithm and the gradient descent algorithm. First, the error between the output layer and the true label is calculated, and then the error is backpropagated to each layer of neurons to update the weights and bias terms of the neurons to minimize the prediction error. The DNN structure applied to the present invention is as follows:

[0147] An input layer: Receives the application load characteristic data and the computing storage system characteristic data, performs a linear transformation on them through a linear layer Linear, converts the dimension of the input data to 64 dimensions, and achieves a non-linear transformation through the Relu activation function.

[0148] A hidden layer: Takes the output after the non-linear transformation of the input layer as the input, performs a linear transformation through a linear layer Linear, converts the 64-dimensional input data to 32 dimensions, and achieves a non-linear transformation through the Relu activation function to capture the complex relationships between the input data.

[0149] An output layer: Performs a linear transformation through a linear layer Linear, converts the 32-dimensional input data to a one-dimensional data, that is, the application performance is obtained.

[0150] The hyperparameters for model training are shown in Table 8.

[0151] Table 8 List of Hyperparameters for Model Training

[0152] Training parameters value Iterations 500 Optimizer Adam Loss Function Cross Entropy Loss Function Learning rate 0.001 The size of each batch 32

[0153] Based on the above hyperparameters, a neural network model is built. At each training epoch, the cross-entropy loss function and the Adam optimizer are combined to complete the neural network training. The trained model is saved to a file for calculating subsequent affinity results, ultimately obtaining the affinity between the application load and the computing storage system.

[0154] Step 5) Affinity prediction:

[0155] After completing affinity model training, affinity prediction can be performed for application workloads and computing and storage systems. This involves inputting the compute and storage system characteristics and application workload characteristics, which are collected and processed in real time, into the affinity prediction model to output a calculated affinity score.

[0156] Reference Figure 9 , further describing the deep learning prediction model of the present invention.

[0157] First, the target application load is run on the benchmark computing and storage system, and the application load characteristics are obtained through the application load characterization method. Second, the Sysbench and iperf benchmark programs are run on the computing and storage system to obtain the performance scores of each hardware. The hardware characteristics are then obtained through the computing and storage system characterization method. Finally, the application load characteristics and hardware characteristics are used as inputs to the neural network to obtain an initial affinity score between the application load and the computing and storage system, that is, the application performance in the formula. To further and more accurately characterize the affinity, the application performance and cost are weighted and summed:

[0158] affinity = α × application performance + (1-α) × cost

[0159] The final affinity score is obtained by the above formula. The closer the value is to 0, the stronger the affinity. α is 0.5.

Claims

1. An analysis method for the affinity between application load and a compute storage system, characterized in that Calculate the characteristics of the storage system and the application workload respectively, and use deep learning methods to analyze the affinity between the application workload and the computing storage system; the steps of this method are as follows: Step 1, normalize the performance scores and metadata information of the computing storage system obtained to obtain the characteristics of the computing storage system; Step 2, through a performance monitoring tool, monitor and collect the resource usage of the general application workload during operation, and then normalize and extract important features from the collected data to obtain the characteristics of the application workload; Step 3, generate a training set with the characteristics of the computing storage system and the characteristics of the application workload, train a neural network model, and use the trained network model as an affinity prediction model; Step 4, in the same way as in Step 1 and Step 2, input the characteristics of the computing storage system and the application workload after real-time collection and post-processing into the affinity prediction model, and output the calculated affinity score.

2. The affinity analysis method according to claim 1, characterized in that, The performance scores of the computing storage system obtained in Step 1 refer to obtaining the computing performance test score data, memory performance test score data, disk performance test score data, and network performance test score data through all the computing loads, memory loads, disk loads, and network loads in the benchmark test program deployed on the computing storage system under test.

3. The affinity analysis method according to claim 1, wherein The metadata information of the computing storage system obtained in Step 1 refers to recording four metadata information of the computing storage system according to the output of executing the lshw command on the computing storage system under test: First, record the CPU information including CPU architecture, CPU cache size, and instruction set; Second, record the memory information including memory size, memory architecture, and clock cycle; Third, record the disk information including disk size and interface type; Fourth, record the network card information of the online bandwidth of the network card.

4. The analysis method of affinity according to claim 1, characterized in that, The data of the resource usage of the application workload during operation obtained in Step 2 includes the data of the CPU resource usage, the data of the memory resource usage, the data of the disk resource usage, and the data of the network resource usage.

5. The affinity analysis method according to claim 4, characterized in that, The steps of obtaining the data of the resource usage of the application workload during operation in Step 2 are as follows: The first step is to deploy a monitoring tool on the computing storage system under test; The second step is to run the application workload on the computing storage system under test, and obtain four types of index information through the monitoring tool: First, CPU index data, including the utilization rate of each core of the CPU, the system time of the CPU, the user time of the CPU, the waiting time of the CPU, and the idle time of the CPU; Second, memory index data, including the occupied size of the memory, the occupied size of the file system cache, and the occupied size of the file system buffer; Third, disk index data, including disk IOPS, disk IO queue length, disk read bytes, and disk write bytes; Fourth, network index data, including network upload rate, network download rate, total number of network sent bytes, and total number of network received bytes; The third step is to count the running time of the application workload after it finishes running, and use the time elapsed from the start to the end of the run as the application performance.

6. The affinity analysis method according to claim 1, characterized in that, The training set mentioned in step 3 refers to the training set generated after processing the dataset composed of computing storage system features and application workload features as follows: The first step is data preprocessing: taking the computing storage system features as hardware features, the application workload features as application features, the time consumed when running on the application workload-hardware combination as application performance, and the price of the application workload-hardware combination as the hardware cost of the computing storage system, obtaining the dataset in the format of {hardware features, application features, application performance, hardware cost of the computing storage system}; The second step is to randomly shuffle the dataset to obtain the shuffled dataset; The third step is data encoding: limiting the architecture types of CPUs in the shuffled dataset, creating a vector with a length equal to the number of categories for each category, setting the index position of this category to 1 and the rest to 0, and replacing the original value of each CPU architecture with its corresponding [0,1] encoded vector.

7. The method for analyzing affinity according to claim 1, characterized in that, The structure of the neural network model mentioned in step 3 is an input layer, a linear transformation layer, an activation layer, a hidden layer, a non-linear transformation layer, and an output layer connected in series in sequence; the dimensions of the linear transformation are set to 64 dimensions, 32 dimensions, and 1 dimension in sequence, and the activation layer is implemented with the Relu activation function.

8. The analysis method of affinity according to claim 1, characterized in that, The steps for training the neural network model mentioned in step 3 are as follows: The first step is to set the hyperparameters for training the neural network model as: the number of iterations is 500, the optimizer is Adam, the learning rate is 0.001, and the size of each batch is 32; The second step is to use the gradient descent method, utilize the Adam optimizer, and iteratively update the network parameters until the cross-entropy loss function of the network converges, obtaining the trained neural network model.

9. The method for analyzing affinity according to claim 8, characterized in that, The cross-entropy loss function is as follows: loss(y,y_hat)=-∑ylog(y_hat) where y represents the true value of the application performance input into the neural network model, y_hat represents the predicted value of the application performance predicted by the neural network model, and log represents the logarithmic operation with base 10.

10. The analysis method of affinity according to claim 6, characterized in that, The affinity score mentioned in step 4 is obtained by the following formula: affinity=α×application performance+(1-α)×hardware cost of the computing storage system where affinity represents the affinity score, the closer this value is to 0, the stronger the affinity, and α is a coefficient of 0.5.

Citation Information

Patent Citations

  • User SLO modeling method and device based on application performance prediction model

    CN115934343A

  • Storage workload allocation based on input / output mode affinity calculations

    CN117331484A