Load generation method, computer program product and network attached storage system
By loading the scene load template in the scene library and adjusting the load using reinforcement learning models, the problem that the NAS system load generation method cannot simulate the dynamic changes in traffic in real business scenarios is solved, and more accurate performance testing and better resource dynamic adjustment capabilities are achieved.
Patent Information
- Application Number
- CN202510490596.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing NAS system load generation method cannot effectively simulate the dynamic changes in traffic in real business scenarios, resulting in a large deviation from the actual system performance.
By loading the scene load template in the scene library of the network attached storage system, the target test load of the target test scenario is generated, and the reinforcement learning model is used to obtain load adjustment instructions based on real-time performance indicators and dynamically adjust the load.
It accurately simulates the dynamic changes in traffic in real business scenarios, improves the accuracy of test results, and verifies the system's ability to dynamically adjust resources through dynamic load-triggered NAS system's elastic expansion and scaling mechanism.
Smart Images

Figure CN120029870A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of load generation, and in particular to a load generation method, a computer program product, and a network attached storage system. Background Art
[0002] With the rapid development of cloud computing, edge computing, and artificial intelligence technologies, Network Attached Storage (NAS) systems are facing increasingly complex business scenarios, and the performance requirements for NAS systems are also getting higher and higher.
[0003] At present, in NAS system performance testing, loads are mainly generated through fixed modes, such as pre-defined read-write ratios and fixed data block sizes, to test NAS system throughput, latency and other indicators. This type of method cannot effectively simulate the dynamic characteristics of traffic changes in real business scenarios, resulting in a large deviation between the test results and the actual system performance. Summary of the invention
[0004] The present application provides a load generation method, a computer program product, and a network attached storage system to at least solve the problem that the NAS system load generation method in the related art cannot effectively simulate the dynamic change characteristics of traffic in real business scenarios, resulting in a large deviation between the test results and the actual system performance.
[0005] The present application provides a load generation method, which is applied to a network attached storage system, and includes: Loading a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system to generate a target test load corresponding to the target test scenario, wherein the scenario library includes a business scenario tag and the scenario load template corresponding to the business scenario tag; Based on the real-time performance indicators of the network attached storage system under the target test load, obtaining load adjustment instructions through a reinforcement learning model; The load adjustment instruction is executed to adjust the target test load.
[0006] The present application also provides a computer program product, which is applied to a network attached storage system, and the computer program product includes: A first processing module is used to load a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system to generate a target test load corresponding to the target test scenario, wherein the scenario library includes a business scenario tag and the scenario load template corresponding to the business scenario tag; A second processing module, configured to obtain a load adjustment instruction through a reinforcement learning model based on the real-time performance indicator of the network attached storage system under the target test load; The third processing module is used to execute the load adjustment instruction to adjust the target test load.
[0007] The present application also provides a network attached storage system for executing the load generation method as described above.
[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned load generation methods when executing the computer program.
[0009] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned load generation methods are implemented.
[0010] Through this application, a target test load corresponding to a target test scenario is generated by loading a scenario load template, the load environment of the target test scenario is quickly constructed, and real-time performance indicators are obtained and input into a reinforcement learning model. Since the reinforcement learning model can perform load adjustment actions in the target test scenario, and use real-time performance indicators as feedback, output load adjustment instructions to guide the network attached storage system to adjust the load, it can solve the technical problem that the load generation method cannot effectively simulate the dynamic change characteristics of traffic in real business scenarios, resulting in a large deviation between the test results and the actual system performance. The load is dynamically adjusted based on the real-time performance of the system to accurately simulate the dynamic change characteristics of traffic in real business scenarios and improve the accuracy of the test results. At the same time, the dynamically changing load can effectively trigger the elastic expansion and contraction mechanism of the network attached storage system, thereby fully verifying the system's ability to dynamically adjust resources, so that the system can automatically allocate resources according to actual needs and achieve a balance between performance and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0012] Figure 1 One of the flow charts of the load generation method provided in the embodiment of the present application; Figure 2 A schematic diagram of the structure of a computer program product provided in an embodiment of the present application; Figure 3A schematic diagram of the structure of a network attached storage system provided in an embodiment of the present application; Figure 4 This is the second flow chart of the load generation method provided in the embodiment of the present application; Figure 5 This is the third flow chart of the load generation method provided in the embodiment of the present application; Figure 6 This is the fourth flow chart of the load generation method provided in the embodiment of the present application. DETAILED DESCRIPTION
[0013] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0014] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0015] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0016] An embodiment of the present application provides a load generation method, and the method is described in detail in conjunction with the execution flow of the load generation method.
[0017] The load generation method can be applied to a network attached storage (NAS) system. The network attached storage system is used for data storage and sharing and can be connected through a local area network or the Internet to provide centralized file storage and access services for multiple users or clients.
[0018] like Figure 1 As shown, the load generation method of the embodiment of the present application includes step 110, step 120 and step 130.
[0019] Step 110: Load a scenario load template corresponding to the target test scenario from a scenario library of the network attached storage system, and generate a target test load corresponding to the target test scenario.
[0020] Among them, the scenario library includes business scenario tags and scenario load templates corresponding to the business scenario tags.
[0021] It can be understood that the business scenario label is a mark for classifying the business scenarios faced by the network attached storage system. For the network attached storage system, the classification of its business scenarios may include steady-state scenarios, peak scenarios, and edge scenarios.
[0022] Each business scenario tag in the scenario library has a corresponding scenario load template. The scenario load template corresponding to a business scenario tag may include setting information of load parameters under the business scenario corresponding to the business scenario tag, such as the number of reads and writes per second (IOPS), data block size, protocol ratio and other parameters.
[0023] For example, the business scenario of daily office document collaboration can be classified as a steady-state scenario. The corresponding scenario load template can be set with a read-write ratio of 6:4, mainly processing small files of 4KB-64KB.
[0024] For another example, the business scenario of image content distribution network (CDN) backhaul during e-commerce promotions can be classified as a peak scenario. The corresponding scenario load template can be set to a read ratio of > 95% and sequential reading of large files larger than 1 MB.
[0025] For another example, business scenarios with high-frequency writes on IoT devices can be classified as edge scenarios. The corresponding scenario load template can be set to write 95% of files below 4KB, with large fluctuations in network latency.
[0026] In this step, according to the target test scenario to be tested, a label search is performed in the scenario library of the network attached storage system, the business scenario label corresponding to the target test scenario is determined, the scenario load template corresponding to the business scenario label is loaded, and the target test load corresponding to the target test scenario is generated according to the load parameter settings of the scenario load template.
[0027] For example, if the target test scenario is an edge scenario, search for the business scenario tag corresponding to the edge scenario in the scenario library, load the scenario load template corresponding to the business scenario tag, and generate the target test load corresponding to the edge scenario.
[0028] In actual execution, the target test load corresponding to the target test scenario can be generated through the load generation engine, which converts the scenario load template into an executable test load and quickly builds a load environment close to the actual business scenario.
[0029] Step 120: Based on the real-time performance indicators of the network attached storage system under the target test load, a load adjustment instruction is obtained through a reinforcement learning model.
[0030] Generate a target test load corresponding to the target test scenario, monitor the real-time performance indicators of the network attached storage system under the target test load in real time, and the real-time performance indicators are used to characterize the performance status of the network attached storage system.
[0031] In actual implementation, the real-time performance indicators of the network-attached storage system can be monitored and stored in real time through tools such as the data visualization and monitoring platform Grafana and the time series database InfluxDB.
[0032] The reinforcement learning model is a machine learning method that learns the optimal decision-making strategy by interacting with the environment. The agent performs actions in a given state and continuously adjusts the strategy for executing actions based on the rewards fed back by the environment to maximize the long-term cumulative rewards and achieve autonomous learning and adaptive optimization.
[0033] The target test scenario and real-time performance indicators can be input into the reinforcement learning model. The reinforcement learning model performs load adjustment actions under the target test scenario, provides real-time feedback with real-time performance indicators, dynamically adjusts the load, and outputs load adjustment instructions for guiding the network attached storage system to perform load adjustment.
[0034] Step 130: Execute the load adjustment instruction to adjust the target test load.
[0035] In this step, the load adjustment instructions output by the reinforcement learning model are obtained, the target test load is adjusted according to the load adjustment instructions, and the real-time performance indicators of the network attached storage system are monitored at the same time. The simulation and verification of the target test scenario are completed, and a test report of the target test scenario is generated.
[0036] It should be noted that the load adjustment instructions output by the reinforcement learning model can take effect in real time.
[0037] It is understandable that with the rapid development of cloud computing, edge computing and artificial intelligence technologies, the business scenarios faced by network-attached storage systems have expanded from traditional file sharing to artificial intelligence (AI) training, Internet of Things (IoT) data lakes, real-time video analysis, etc., and the load patterns have become dynamic and hybrid.
[0038] Related technologies mainly use fixed-mode load generation tools such as FIO and IOzone to generate loads by pre-defining read-write ratios and fixed data block sizes to test throughput, latency and other indicators. This type of static load mode is divorced from actual business scenarios, and fixed load parameters cannot simulate the dynamic changes of business traffic.
[0039] At the same time, static loads make it difficult to trigger the elastic expansion and contraction mechanism of the network-attached storage system, and it is impossible to verify the dynamic adjustment capability of resources.
[0040] In an embodiment of the present application, a target test load corresponding to a target test scenario is generated by loading a scenario load template, the load environment of the target test scenario is quickly constructed, and real-time performance indicators are obtained and input into a reinforcement learning model. Since the reinforcement learning model can perform load adjustment actions in the target test scenario, and use real-time performance indicators as feedback, output load adjustment instructions to guide the network attached storage system to perform load adjustment, the technical problem that the load generation method cannot effectively simulate the dynamic change characteristics of traffic in real business scenarios, resulting in a large deviation between the test results and the actual system performance, can be solved. The load is dynamically adjusted based on the real-time performance of the system to accurately simulate the dynamic change characteristics of traffic in real business scenarios and improve the accuracy of the test results. At the same time, the elastic expansion and contraction mechanism of the network attached storage system can be effectively triggered by the dynamically changing load, thereby fully verifying the system's ability to dynamically adjust resources, so that the system can automatically allocate resources according to actual needs and achieve a balance between performance and efficiency.
[0041] According to the load generation method provided in the embodiment of the present application, a target test load corresponding to a target test scenario is generated by loading a scenario load template, the load environment of the target test scenario is quickly constructed, real-time performance indicators are obtained and input into a reinforcement learning model, and the reinforcement learning model outputs load adjustment instructions to guide the network attached storage system to perform load adjustment. The load is dynamically adjusted based on the real-time performance of the system, and the dynamic change characteristics of traffic in real business scenarios are accurately simulated, which helps to improve the accuracy of the test results of the network attached storage system.
[0042] In some embodiments, the scene library is constructed by the following steps: Obtain system log data from a network attached storage system; Build a scenario library based on system log data.
[0043] Based on the system log data of the network attached storage system, different business scenarios faced by the network attached storage system (such as video streaming, collaborative office, etc.) can be identified, and a scenario library including business scenario tags and scenario load templates corresponding to the business scenario tags can be constructed.
[0044] In actual implementation, Fluentd can be used to collect system log data from network-attached storage systems and buffer the data through Kafka message queues.
[0045] Among them, Fluentd is an open source data collector that can build a unified log collection layer for users to better use and understand log data; Kafka is an open source distributed stream processing platform.
[0046] It is understandable that the collected system log data can be cleaned to improve data quality, providing a reliable foundation for subsequent data analysis and construction of a scenario library.
[0047] For example, use data cleaning tools such as Apache Spark to clean system log data, such as abnormal breakpoint logs, duplicate records, standardize timestamp formats, and complete missing fields.
[0048] In actual implementation, business scenarios are identified based on system log data, and business scenario labels are marked for different business scenarios. Business scenario labels can be manually labeled or automatically generated.
[0049] For example, label system log data according to business types (such as video streaming and database backup), and mark business scenario labels corresponding to typical business scenarios.
[0050] For another example, by linking the timestamp with the system, the business scenario labels corresponding to different business scenarios can be automatically marked.
[0051] In the related technology, specific business scenarios are simulated by writing scripts, which rely on manually predefined scenarios. Manual script writing is time-consuming and has limited scenario coverage, making it difficult to simulate business scenarios with complex load characteristics such as edge computing and AI training.
[0052] During the implementation of this application, based on system log data, the driver identifies the business scenarios of the network attached storage system, creates business scenario tags and scenario load templates corresponding to the business scenario tags, builds a scenario library for the network attached storage system, and generates target test loads corresponding to the target test scenarios by loading the scenario load templates of the scenario library, accurately simulating real business scenarios. Since the system log data is real data and is constantly updated, business scenario tags and scenario load templates corresponding to the business scenario tags are created in a dynamic generation manner, without the need for manual script writing, and can cover various business scenarios.
[0053] It should be noted that the scenario library is dynamically updated based on the incremental learning mechanism. When a new business scenario is identified based on the system log data, an incremental update is triggered. Based on scenario similarity matching, the cosine similarity is used to calculate the matching degree between the system log data and the existing scenario category. If the similarity is less than the preset threshold, a new scenario category is created.
[0054] In some embodiments, the system log data includes at least one of a protocol level log, a file metadata log, and a user behavior log of the network attached storage system.
[0055] The protocol-level log may refer to the operation details of the file transfer protocol of the network-attached storage system, including information such as read and write request timestamps, file paths, data block sizes, and client addresses.
[0056] It should be noted that when the network attached storage system includes multiple file transfer protocols, the protocol level log includes the operation details of the multiple file transfer protocols.
[0057] File metadata logs may include log data describing file background, content, structure, and management processes, such as file creation, deletion, modification events, and permission changes in a network-attached storage system.
[0058] User behavior logs may include user operation trajectories and interaction data such as user login frequency, number of concurrent sessions, and directory access patterns.
[0059] In actual implementation, one or more combinations of protocol-level logs, file metadata logs, and user behavior logs of the network-attached storage system may be obtained to identify business scenarios of the network-attached storage system and build a scenario library.
[0060] In some embodiments, a scenario library is constructed based on system log data, including: Extract features from system log data to obtain business scenario features of the network attached storage system; Cluster business scenario features and build a scenario library.
[0061] The business scenario features are extracted based on system log data, and the business scenario features can reflect the characteristics of the business scenarios faced by the network attached storage system.
[0062] Through unsupervised learning clustering processing, the business scenario features extracted from system log data are divided into several groups. The data in the same group are similar to each other (that is, these data belong to the same type of business scenario), and the data between different groups are obviously different. Different business scenarios of network-attached storage systems can be accurately identified, providing a reliable basis for building a scenario library.
[0063] In some embodiments, feature extraction is performed on system log data to obtain business scenario features of the network attached storage system, including: Based on the system log data, at least one of the time distribution features, input and output mode features, protocol and user behavior features, and resource occupancy features is extracted to obtain business scenario features.
[0064] Among them, the time distribution feature is used to characterize the regular pattern of system log data changing over time, which can be expressed in a periodic or sudden form.
[0065] For example, the time distribution characteristics may include request frequency periodic characteristics (such as daily peak hours, weekend low loads) and burst traffic characteristics.
[0066] Input-output (IO) pattern characteristics are used to characterize the characteristics of data flow in a network-attached storage system.
[0067] In actual implementation, the input and output pattern characteristics may include data flow characteristic parameters such as dynamic distribution of read-write ratio, data block size distribution, and random access to sequential access ratio.
[0068] Protocol and user behavior characteristics are used to characterize the file transfer protocol and user behavior characteristics in network attached storage systems.
[0069] In actual implementation, protocol and user behavior characteristics may include the proportion of multi-protocol traffic (such as 60% for NFS, 30% for SMB, and 10% for FTP), the number of concurrent user sessions, and the frequency of file lock contention (such as lock timeout events under the SMB protocol).
[0070] Resource occupancy characteristics are used to characterize the consumption patterns and extent of various computing resources of a network-attached storage system during operation. They can be expressed as dynamic changes in the time dimension (such as periodic peaks, sustained high loads, or sudden occupancy) and can be quantified through indicators such as processor (CPU) utilization, memory usage, response delay, etc.
[0071] In actual implementation, resource occupancy characteristics may include the correlation between CPU or memory utilization and the number of read / write operations per second (Input / Output Operations Per Second, IOPS). Resource occupancy characteristics may also include the linear relationship between network bandwidth and throughput (such as the maximum throughput under a gigabit network card).
[0072] Perform feature extraction on system log data to obtain a combination of one or more of time distribution features, input and output pattern features, protocol and user behavior features, and resource occupancy features to obtain business scenario features.
[0073] In actual implementation, feature engineering such as sliding window statistics, principal component analysis or time series feature encoding can be used to extract at least one of the time distribution characteristics, input and output pattern characteristics, protocol and user behavior characteristics and resource occupancy characteristics of system log data.
[0074] In some embodiments, clustering is performed on business scenario features to construct a scenario library, including: According to the density clustering algorithm or the hierarchical clustering algorithm, the business scenario features are clustered to obtain the feature clustering results; Map the feature clustering results into feature vector templates to build a scenario library.
[0075] Among them, the density clustering algorithm (DBSCAN) is suitable for non-uniformly distributed log data and automatically identifies noise points (such as temporary test traffic).
[0076] In actual implementation, the neighborhood radius (Eps) of the density clustering algorithm can be determined by the k-distance curve, and the minimum number of samples (MinPts) can be set to 10% of the requests in the time window.
[0077] The hierarchical clustering algorithm is a clustering method based on hierarchical decomposition. By recursively dividing the data into different levels, it can assist in verifying the hierarchical structure of business scenarios (e.g., "video streaming" can be subdivided into live and on-demand subcategories).
[0078] The business scenario features are clustered by a density clustering algorithm or a hierarchical clustering algorithm, and the feature clustering results obtained by clustering are mapped into a feature vector template to build a scenario library, wherein the feature vector template includes information about the business scenario label and the scenario load template corresponding to the business scenario label.
[0079] For example, the feature vector template may include information such as scenario ID, read-write ratio, data block distribution, protocol proportion, number of concurrent threads, network conditions, etc. The scenario ID corresponds to the business scenario label, and the scenario load template corresponding to the business scenario label includes information such as read-write ratio, data block distribution, protocol proportion, number of concurrent threads and network conditions.
[0080] It should be noted that the constructed scenario library can be verified and tuned.
[0081] In actual implementation, the business scenario tags and scenario load templates of the scenario library can be verified through cross-validation, business semantic consistency check, load playback test, etc.
[0082] For example, the historical data of system log data is divided into training set and test set according to time, and the stability of clustering results is verified through cross-validation.
[0083] For example, by manually reviewing the clustering results, we can ensure that the scenario definitions of business scenario labels such as "video streaming" and "database backup" are consistent with the actual business logic, thereby achieving business semantic consistency checks.
[0084] For another example, perform a load playback test, drive the load generation engine based on the scenario load template, compare the test results with historical performance data, and detect whether the error rate is within the preset range.
[0085] In actual implementation, tuning strategies include but are not limited to feature weight optimization, noise filtering and other strategies.
[0086] For example, the importance of features is calculated through the random forest model, and the feature weights of the clustering algorithm are adjusted (such as increasing the weight of the "burst traffic indicator") to optimize data clustering.
[0087] For another example, temporary loads that last less than 5 minutes (such as temporary backup operations by maintenance personnel) can be eliminated to filter out noise and optimize data clustering.
[0088] In the embodiment of the present application, based on the multi-dimensional feature extraction of system log data, a clustering algorithm is used to automatically generate a scenario library, support incremental updates and similarity matching, and convert unstructured log data into a drivable load template, breaking through the limitations of manually predefined scenarios, accurately simulating real business scenarios, and providing more comprehensive business scenario coverage.
[0089] In some embodiments, the real-time performance indicator includes at least one of a system resource indicator, a network status indicator, and a protocol level indicator.
[0090] The system resource indicators may include CPU utilization, memory usage, disk IOPS, and other indicators that characterize the resource status of the network attached storage system.
[0091] In actual implementation, system resource indicators can be collected through tools such as Prometheus Agent.
[0092] System resource indicators may also include indicators that characterize the status of the storage layer of the network attached storage system, such as cache hit rate, solid state drive (SSD) wear leveling status, etc.
[0093] The network status indicators may include network bandwidth, delay, packet loss rate, and other indicators that characterize the network status of the network attached storage system.
[0094] In actual implementation, network bandwidth can be collected through Prometheus, latency can be collected through PingMesh, and packet loss rate can be obtained through NetFlow analysis.
[0095] Protocol-level metrics are indicators that characterize the file transfer protocol performance of network-attached storage systems.
[0096] For example, for the NFS file transfer protocol, you can count the remote procedure call (RPC) latency, and for the SMB file transfer protocol, you can monitor the file lock contention frequency.
[0097] In actual implementation, the real-time performance indicators of the network-attached storage system can be monitored and stored in real time through tools such as the data visualization and monitoring platform Grafana and the time series database InfluxDB.
[0098] In some embodiments, the state space of the reinforcement learning model includes performance metrics of the network-attached storage system, protocol characteristics, scenario context, and historical trends; The action space of the reinforcement learning model includes load intensity adjustment of the network-attached storage system, protocol ratio adjustment, and network impairment injection; The reward function of the reinforcement learning model is constructed based on latency, throughput, overrun penalty terms, and protocol conflict penalty terms.
[0099] It can be understood that the reinforcement learning model belongs to the state-action-reward model structure. Its state is determined by real-time performance metrics and target test scenarios. Dynamically adjusting load parameters is an action that the reinforcement learning model can execute, and the model optimization is guided according to the reward function.
[0100] Defining the state space of the reinforcement learning model can include configuring the performance metrics, protocol characteristics, scenario context, and historical trends of the network-attached storage system.
[0101] Among them, the performance metrics include CPU utilization, memory occupancy, average input / output latency, throughput, etc.; the protocol characteristics include the request proportion of different file transfer protocols (such as NFS / SMB / FTP), file lock competition frequency, RPC retransmission rate, etc.; the scenario context includes the ID of the current test scenario, network impairment mode (such as "high latency + packet loss"), etc.; the historical trends can include the metric variance and maximum value of the system data within a sliding window (5 minutes) (used to detect burst traffic).
[0102] The actions that the reinforcement learning model can execute include load intensity adjustment, protocol ratio adjustment, and network impairment injection of the network-attached storage system.
[0103] Load intensity adjustment can include adjusting the number of concurrent threads and the IOPS target value. For example, the number of concurrent threads is adjusted from 200 to 300, and the IOPS target value is increased from 50000 to 60000.
[0104] Protocol ratio adjustment can be to dynamically allocate traffic weights for different file transfer protocols. For example, the traffic weight of the NFS file transfer protocol is reduced from 60% to 55%, and the traffic weight of the SMB file transfer protocol is increased from 30% to 35%.
[0105] Network impairment injection can include operations such as increasing latency, reducing latency, increasing packet loss rate, and reducing packet loss rate.
[0106] The reward function of the reinforcement learning model is built based on delay, throughput, over-limit penalty and protocol conflict penalty. Corresponding weight coefficients (dynamically adjustable) can be set for delay, throughput, over-limit penalty and protocol conflict penalty, and the weighted sum of delay, throughput, over-limit penalty and protocol conflict penalty is used as the reward function.
[0107] For example, the formula for the reward function R of the reinforcement learning model is as follows: R=w 1 ×Delay+w 2 ×Throughput-w 3 ×Overlimit Penalty Item-w 4 ×Protocol violation penalty item; Among them, w 1 、w 2 、w 3 、w 4 They are the weight coefficients of delay, throughput, over-limit penalty and protocol conflict penalty respectively.
[0108] For example, 1 =0.6, w 2 =0.2, w 3 =0.1, w 4 =0.1.
[0109] In actual implementation, the over-limit penalty item can be triggered based on the CPU utilization of the network attached storage system. For example, if the CPU utilization is >90%, the over-limit penalty item is triggered and 10 points are deducted.
[0110] The protocol conflict penalty item can be determined based on the number of conflicts between different file transfer protocols of the network attached storage system. For example, if the number of SMB protocol lock timeouts is >100 times / minute, the protocol conflict penalty item is triggered and 5 points are deducted.
[0111] It can be understood that latency refers to the time from when data is sent to when it is received. The smaller the latency, the faster the data transmission speed, and the more points the reward function adds. The larger the latency, the slower the data transmission speed, and the fewer points the reward function adds.
[0112] For example, a delay of 1 second will give 1 point, and a delay of 0.5 seconds will give 2 points.
[0113] For target test scenarios that focus on data transmission speed, the weight coefficient corresponding to the delay can be increased so that the score caused by the delay accounts for a higher proportion.
[0114] Throughput is the amount of data processed per unit time. The greater the throughput, the more points you get.
[0115] For example, a throughput of 100 MB / s is worth 100 points, and a throughput of 200 MB / s is worth 200 points.
[0116] For target test scenarios that focus on data processing volume, the weight coefficient corresponding to throughput can be increased to make the score brought by throughput account for a higher proportion.
[0117] The over-limit penalty item can be triggered based on the CPU utilization of the network-attached storage system. If the CPU utilization exceeds the safety line, the network-attached storage system is at risk of freezing. Each time the CPU utilization exceeds the safety line, an over-limit penalty item is triggered and a point is deducted (a minus sign is placed before the over-limit penalty item). The greater the excess ratio, the more points are deducted.
[0118] For example, if the CPU utilization is 90%, 1 point will be deducted, and if the CPU utilization is 95%, 3 points will be deducted.
[0119] For target test scenarios that focus on whether the system is stuck, the weight coefficient corresponding to the over-limit penalty item can be increased.
[0120] The number of protocol conflicts can be the number of times a file is processed simultaneously by different file transfer protocols. One point will be deducted for each protocol conflict.
[0121] For example, if there is one conflict, one point will be deducted; if there are five conflicts, five points will be deducted.
[0122] For target test scenarios that focus on avoiding protocol conflicts, the weight coefficient corresponding to the protocol conflict penalty item can be increased.
[0123] The reinforcement learning model is built based on the deep Q network. The training phase of the reinforcement learning model includes an offline pre-training phase and an online learning phase.
[0124] Among them, Deep Q-Network (DQN) is a reinforcement learning network that combines deep learning and Q learning. It approximates the Q-value function (action value function) through a neural network. The reinforcement learning model is built based on the deep Q-network and is suitable for the high-dimensional state space and load-related continuous action decisions of network-attached storage systems.
[0125] It can be understood that the training phase of the reinforcement learning model includes an offline pre-training phase, in which preliminary training is performed before the reinforcement learning model is put into the load adjustment of the network attached storage system so that the reinforcement learning model can learn a general feature representation.
[0126] In actual implementation, historical data sets (e.g., containing 100,000 state-action-reward records) can be used to pre-train the reinforcement learning model and initialize the model. During the pre-training process, experience replay can be used to reduce data correlation and improve convergence speed.
[0127] The training phase of the reinforcement learning model includes an online learning phase, which continuously optimizes the load strategy based on the online learning mechanism to adapt to the dynamic changes of the network-attached storage system (such as node expansion and cache strategy update).
[0128] In actual implementation, the reinforcement learning model can collect data from the network-attached storage system at regular intervals (such as 5 minutes), generate actions and execute them, and record reward values to update model parameters.
[0129] In some embodiments, loading a scenario load template corresponding to a target test scenario from a scenario library of a network attached storage system to generate a target test load corresponding to the target test scenario may include: Parse the scenario load template corresponding to the target test scenario, perform initial load parameter configuration, protocol engine configuration, network environment pre-simulation and resource baseline calibration, and generate the target test load.
[0130] Generating a target test load may include sub-processes such as initial load parameter configuration, protocol engine configuration, network environment pre-simulation, and resource baseline calibration.
[0131] The initial load parameter configuration may refer to generating basic load parameters (eg, IOPS, data block size, protocol ratio, etc.) of the network attached storage system according to the scenario load template definition.
[0132] The protocol engine configuration may refer to initializing the file transfer protocol client of the network attached storage system and setting basic client parameters such as authentication and mount points.
[0133] In some embodiments, the protocol engine configuration may include: Initialize multiple protocol clients based on different protocol types and set parameters for each protocol client.
[0134] The network attached storage system may include multiple protocol clients of different protocol types. For example, the network attached storage system may include a client based on the NFS file transfer protocol, a client based on the SMB file transfer protocol, a client based on the FTP file transfer protocol, and the like.
[0135] Configure the protocol engine, initialize the NFS, SMB, and FTP protocol clients, and set basic client parameters such as authentication and mount points for each protocol client.
[0136] It should be noted that when loading the scenario load template and generating the target test load, the protocol engine configuration is performed, and the load adjustment instructions output by the deep learning model may include protocol engine control.
[0137] The deep learning model can automatically allocate the traffic ratios of the NFS, SMB, and FTP protocols according to the requirements of the target test scenario (such as 60% for NFS, 30% for SMB, and 10% for FTP), and simulate resource competition between protocols (such as lock conflicts and metadata operations) to achieve dynamic allocation of protocol traffic.
[0138] The deep learning model can generate differentiated loads based on different protocol characteristics (for example, NFS prefers sequential reading and writing of large files, while SMB prefers random access to small files), thus achieving protocol-sensitive load generation.
[0139] Most of the related technologies conduct separate tests for different protocols, without considering resource competition and performance interference in multi-protocol concurrent scenarios (such as the interactive impact of SMB locks and NFS cache failures), rely on manual experience to configure parameters, lack the ability to dynamically allocate cross-protocol traffic, cannot adapt to changes in business scenarios, and cannot verify the stability of NAS under multi-protocol mixed access.
[0140] In the embodiments of the present application, full consideration is given to resource competition and performance interference in multi-protocol concurrent scenarios, target test complexity is generated through protocol engine configuration, a multi-protocol hybrid access environment is constructed, dynamic traffic allocation between different protocols is achieved through a deep learning model, and differentiated loads are generated for different protocol characteristics. This can adapt to changes in different business scenarios and effectively verify the stability and reliability of the network attached storage system in a multi-protocol hybrid access environment.
[0141] Network environment pre-simulation can refer to configuring corresponding network damage rules according to the target test scenario and simulating the network environment. For example, configuring network damage rules related to delay and packet loss for edge scenarios.
[0142] Resource baseline calibration can refer to collecting the initial state of the network-attached storage system (e.g., CPU utilization, memory occupancy, etc.) as a benchmark for the subsequent dynamic adjustment of the reinforcement learning model.
[0143] A specific embodiment is described below.
[0144] A scenario load template corresponding to a target test scenario is loaded from a scenario library of a network attached storage system, and the scenario load template corresponding to the target test scenario is parsed.
[0145] The contents of the scene load template include: Scenario ID: Edge_AI_Training, Read-Write Ratio: 70:30, Data Block Distribution: {4KB: 40%, 64KB: 50%, 1MB: 10%}, Protocol Ratio: {4KB: 40%, 64KB: 50%, 1MB: 10%}, Number of Concurrent Threads: 200, Network Conditions: {Benchmark Delay: 50ms, Jitter Range: ±30ms, Packet Loss Rate: 2%}, Special Rules: Peak write is triggered every 10 minutes during the training cycle.
[0146] Initial load parameter configuration is performed according to parameter mapping rules, and IOPS is calculated by scaling the peak load of historical scenarios (such as 100KIOPS) and the current hardware configuration (such as the number of CPU cores of the test machine). A layered random sampling algorithm is used to ensure that the data block size distribution conforms to the template definition (such as 40% of 4KB requests) and calculate the data block distribution. A weighted round robin algorithm is used to schedule multi-protocol traffic. For example, 60 out of every 100 requests are allocated to NFS, 30 to SMB, and 10 to FTP, and protocol weights are allocated.
[0147] Configure the protocol engine, configure the mount parameters and authentication mode for the NFS client, configure the Samba connection parameters and file lock policy for the SMB client, and configure the passive mode (PASV) and large file transfer optimization for the FTP client.
[0148] Network environment pre-simulation includes tool chain integration and dynamic damage rules. Tool chain integration includes using TC (Traffic Control) and NetEm to simulate network delay and packet loss, and building a virtual network topology through Mininet to simulate cross-domain communication between edge nodes and cloud NAS (delay of 10ms).
[0149] Dynamic damage rules can include periodic fluctuations, such as randomly adjusting the delay (±20ms) and packet loss rate (±1%) every 5 minutes to simulate the instability of mobile networks. Dynamic damage rules can also include sudden interruptions, simulating edge nodes to reconnect after 10 seconds of disconnection, and testing the NAS automatic retransmission mechanism.
[0150] Resource baseline calibration can include taking an initial status snapshot of the network-attached storage system, collecting initial device indicators such as CPU utilization and memory occupancy through SNMP or REST API, and recording storage layer status such as RAID group health status and SSD remaining life.
[0151] Resource baseline calibration can also include baseline threshold settings, such as setting a dynamic scaling trigger condition where the load generator automatically reduces the number of concurrent threads by 10% when the CPU utilization is > 80% for three consecutive minutes.
[0152] In some embodiments, adjusting the target test load may include: Perform at least one of protocol engine control, network impairment real-time update, and load strength adjustment.
[0153] The reinforcement learning model outputs load adjustment instructions, which can be used to instruct protocol engine control, real-time update of network damage and / or load intensity adjustment.
[0154] Among them, the protocol engine control can include operations such as dynamic expansion and contraction of the client and protocol ratio switching.
[0155] For example, for dynamic scaling of NFS and SMB clients, the number of Pod replicas can be dynamically increased or decreased through the Kubernetes API (such as expanding from 10 Pods to 15), or the thread pool size of a single client can be adjusted to generate differentiated loads for different protocol characteristics, thereby achieving protocol-sensitive load generation.
[0156] For another example, a weighted random algorithm is used to switch the ratio of different protocols, and every 100 or 200 requests are allocated according to the new ratio. This can simulate resource competition between protocols and realize dynamic allocation of protocol traffic.
[0157] It can be understood that network impairment refers to the phenomenon that data transmission quality or service performance is reduced due to various factors in the network transmission process (such as bandwidth limitation, delay, packet loss, jitter, noise, etc.).
[0158] Real-time update of network damage refers to adjusting the network's bandwidth limit, delay, packet loss and other parameters to simulate different network states and accurately test the performance of the network-attached storage system under different network states.
[0159] For example, you can call the TC command to dynamically modify the network rules, adjust the delay from 50ms to 60ms, and update the network impairment of the network attached storage system.
[0160] Load intensity adjustment may refer to adjusting the load that a network attached storage system can bear within a certain period of time.
[0161] For example, through the programming interface (API) or command line interaction provided by the Flexible I / O Tester (FIO) tool, load-related parameters such as the number of reads and writes per second (IOPS), data block size, and number of concurrent threads can be adjusted in real time without restarting the process.
[0162] It should be noted that the reinforcement learning model can achieve closed-loop optimization and adaptive learning, and continuously optimize the load adjustment strategy.
[0163] Among them, closed-loop optimization can be divided into short-term feedback and long-term optimization.
[0164] Short-term feedback can be used to perform model reasoning at regular intervals (such as 5 minutes), generate adjustment actions, update the Q value table based on the real-time reward value, and optimize the load adjustment strategy.
[0165] Long-term optimization can generate test reports at regular intervals (such as 24 hours), mark high-frequency adjustment actions, and specifically enhance the sample weights of relevant state-action pairs when training the model offline.
[0166] When a new protocol is detected, the protocol dimension of the state space of the reinforcement learning model is automatically expanded, and the exploration mode is started to adapt to the new protocol.
[0167] When the capacity of the network-attached storage system nodes is expanded (such as adding an SSD storage pool), the reinforcement learning model automatically identifies the resource margin and increases the load intensity limit.
[0168] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0169] An embodiment of the present application also provides a computer program product, which can be applied to a network attached storage system.
[0170] like Figure 2 As shown, the computer program product comprises: A first processing module 210 is used to load a scenario load template corresponding to a target test scenario from a scenario library of a network attached storage system to generate a target test load corresponding to the target test scenario, wherein the scenario library includes a business scenario tag and a scenario load template corresponding to the business scenario tag; A second processing module 220, configured to obtain a load adjustment instruction through a reinforcement learning model based on the real-time performance index of the network attached storage system under the target test load; The third processing module 230 is used to execute the load adjustment instruction to adjust the target test load.
[0171] It should be noted that the description of the features in the embodiments corresponding to the computer program product can refer to the relevant description of the embodiments corresponding to the load generation method, which will not be repeated here.
[0172] An embodiment of the present application also provides a network attached storage system for executing the load generation method as described above.
[0173] like Figure 3As shown, the network attached storage system can be divided into application layer, control layer, processing layer, data layer and infrastructure layer.
[0174] The infrastructure layer of the network-attached storage system includes NAS device clusters, edge nodes, and network damage devices. A NAS device cluster refers to a device cluster formed by multiple interconnected network-attached storage devices, which has high availability, load balancing, and horizontal expansion capabilities. Edge nodes are lightweight computing storage units close to data sources or end users, and are used in edge computing scenarios such as the Internet of Things.
[0175] Among them, the test execution console can be a management tool for the network attached storage system to execute different business scenario tests. The visual report dashboard can monitor and store the load, performance indicators and other data of the network attached storage system in real time, and intuitively display the test process.
[0176] The network impairment controller can control the network impairment device to inject network impairment into the network attached storage system to simulate different network environments.
[0177] The load generation engine can load the scenario load template corresponding to the business scenario tag, generate the target test load corresponding to the target test scenario, convert the scenario load template into an executable test load, and quickly build a load environment close to the actual business scenario.
[0178] The input of the reinforcement learning model may include the target test scenario and the real-time performance indicators. The load adjustment action is performed under the target test scenario, and the real-time performance indicators are used for real-time feedback, the load is dynamically adjusted, and the load adjustment instructions are output. The load generation engine executes the load adjustment instructions.
[0179] Among them, the real-time performance indicators input by the reinforcement learning model can be provided by the performance monitoring database and protocol traffic capture.
[0180] The system log data of the network attached storage system is collected by NAS logs. Through the data analysis engine, the different business scenarios faced by the network attached storage system are identified, and a scenario library including business scenario tags and scenario load templates corresponding to the business scenario tags is constructed to provide a data foundation for the scenario modeling engine and the load generation engine.
[0181] A specific embodiment is described below.
[0182] like Figure 4 As shown, NAS logs collect system log data, the scenario library provides scenario load templates, the scenario modeling engine and the load generation engine build the target test load under the target test scenario, perform real-time performance monitoring of the network attached storage system, and use the load adjustment strategy output by the reinforcement learning model to dynamically adjust the load generated by the load generation engine.
[0183] The load adjustment strategy output by the reinforcement learning model can also include network damage control, which can realize the performance testing of edge nodes such as the Internet of Things and the Internet of Vehicles.
[0184] like Figure 5 As shown, taking the target test scenario as an edge scenario as an example, data collection and scenario modeling are performed, system log data is collected for feature extraction, and a scenario library is constructed.
[0185] The load generation engine is initialized, the scenario load template corresponding to the edge scenario in the scenario library is parsed, initial load parameter configuration, protocol engine configuration, network environment pre-simulation and resource baseline calibration are performed, and the target test load is generated.
[0186] Monitor the real-time performance indicators of the network-attached storage system, adjust the load parameters through the reinforcement learning model, perform edge scenario simulation and verification (such as network damage injection, fault recovery test, data consistency verification), etc., and generate test reports corresponding to edge scenarios.
[0187] like Figure 6 As shown in the figure, real-time performance indicators such as monitoring delay, throughput, and CPU utilization are input into the reinforcement learning model, and the reinforcement learning model outputs action instructions such as adjusting the number of concurrent threads, switching protocol ratios, and modifying network damage parameters, and performs corresponding load adjustments. The effect of the adjustment is verified according to the real-time performance indicators, and the rollback to the output action instructions is triggered according to the adjustment effect, so as to continuously optimize the load adjustment instructions.
[0188] In the embodiment of the present application, a scenario library is built based on log data, and load adjustment is performed through a reinforcement learning model. The dynamic load generation strategy can effectively improve the coverage of test scenarios, make edge scenario simulation more realistic, accurately simulate the actual business pressure of the network attached storage system, automate protocol traffic distribution, and improve the efficiency of multi-protocol concurrent testing. The dynamic load can also trigger the elastic expansion and contraction of the network attached storage system to optimize resources and reduce costs.
[0189] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned load generation method embodiments.
[0190] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned load generation method embodiments when running.
[0191] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0192] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0193] The above is a detailed introduction to a load generation method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A load generation method, characterized in that: The method is applied to a network attached storage system, and the method comprises: Loading a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system to generate a target test load corresponding to the target test scenario, wherein the scenario library includes a business scenario tag and the scenario load template corresponding to the business scenario tag; Based on the real-time performance indicators of the network attached storage system under the target test load, obtaining load adjustment instructions through a reinforcement learning model; The load adjustment instruction is executed to adjust the target test load.
2. The load generation method according to claim 1, characterized in that: The scene library is constructed through the following steps: Obtaining system log data of the network attached storage system; The scenario library is constructed based on the system log data.
3. The load generation method according to claim 2, characterized in that: The system log data includes at least one of a protocol level log, a file metadata log, and a user behavior log of the network attached storage system.
4. The load generation method according to claim 2, characterized in that: The step of constructing the scenario library based on the system log data includes: Extracting features from the system log data to obtain business scenario features of the network attached storage system; Clustering is performed on the business scenario features to construct the scenario library.
5. The load generation method according to claim 4, characterized in that: The extracting features of the system log data to obtain the business scenario features of the network attached storage system includes: Based on the system log data, at least one of time distribution features, input and output mode features, protocol and user behavior features, and resource occupancy features is extracted to obtain the business scenario features.
6. The load generation method according to claim 4, characterized in that: The clustering process of the business scenario features to construct the scenario library includes: Clustering the business scenario features according to a density clustering algorithm or a hierarchical clustering algorithm to obtain a feature clustering result; The feature clustering result is mapped into a feature vector template to construct the scene library.
7. The load generation method according to any one of claims 1 to 6, characterized in that: The real-time performance indicator includes at least one of a system resource indicator, a network status indicator and a protocol level indicator.
8. The load generation method according to any one of claims 1 to 6, characterized in that: The state space of the reinforcement learning model includes performance indicators, protocol characteristics, scenario context and historical trends of the network attached storage system; The action space of the reinforcement learning model includes load intensity adjustment, protocol ratio adjustment and network damage injection of the network attached storage system; The reward function of the reinforcement learning model is constructed based on delay, throughput, over-limit penalty and protocol conflict penalty.
9. The load generation method according to any one of claims 1 to 6, characterized in that: The reinforcement learning model is constructed based on a deep Q network, and the training phase of the reinforcement learning model includes an offline pre-training phase and an online learning phase.
10. The load generation method according to any one of claims 1 to 6, characterized in that: The step of loading a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system to generate a target test load corresponding to the target test scenario includes: The scenario load template corresponding to the target test scenario is parsed, initial load parameter configuration, protocol engine configuration, network environment pre-simulation and resource baseline calibration are performed, and the target test load is generated.
11. The load generation method according to claim 10, characterized in that: The protocol engine configuration includes: Initialize multiple protocol clients based on different protocol types, and set parameters of each of the protocol clients.
12. The load generation method according to any one of claims 1 to 6, characterized in that: The adjusting the target test load includes: Perform at least one of protocol engine control, network impairment real-time update, and load strength adjustment.
13. A computer program product, characterized in that The computer program product is applied to a network attached storage system, and the computer program product comprises: A first processing module is used to load a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system to generate a target test load corresponding to the target test scenario, wherein the scenario library includes a business scenario tag and the scenario load template corresponding to the business scenario tag; A second processing module, configured to obtain a load adjustment instruction through a reinforcement learning model based on the real-time performance indicator of the network attached storage system under the target test load; The third processing module is used to execute the load adjustment instruction to adjust the target test load.
14. A network attached storage system, characterized in that: Used to execute the load generation method as described in any one of claims 1-12.
15. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the load generation method according to any one of claims 1 to 12 when executing the computer program.
16. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the load generation method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Performance test method in server cluster and related equipment
CN112162891A
Parameter adjustment method and related equipment
CN117667227A
Cloud computing platform real-time CPU load balancing method and system based on reinforcement learning
CN118885291A
Cloud storage system performance test method and system based on large model
CN119088628A
Cited By
Network disk suite integration method, device and system of network attached storage device
CN120849320A
Server POC test method based on AI intelligent agent cooperation
CN122489443A