Load generation method, computer program product, and network attached storage system
By loading the scene load template in the NAS system and using the reinforcement learning model to dynamically adjust the load, the problem that the load generation method cannot simulate the dynamic changes in traffic in real business scenarios is solved, which improves the accuracy of the test results and realizes the dynamic adjustment ability of system resources.
Patent Information
- Application Number
- CN202510490596.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing NAS system load generation method cannot effectively simulate the dynamic changes in traffic in real business scenarios, resulting in a large deviation from the actual system performance.
By loading the scene load template corresponding to the target test scenario from the scene library of the network attached storage system, the target test load of the target test scenario is generated, and the reinforcement learning model is used to obtain load adjustment instructions based on real-time performance indicators, and dynamically adjust the load to accurately simulate dynamic traffic changes.
It improves the accuracy of the test results and can effectively trigger the elastic scaling mechanism of the network additional storage system to achieve a balance of performance and efficiency.
Smart Images

Figure CN120029870B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of load generation, and in particular to a load generation method, a computer program product, and a network attached storage system. Background Art
[0002] With the rapid development of cloud computing, edge computing, and artificial intelligence technologies, Network Attached Storage (NAS) systems are facing increasingly complex business scenarios, and the performance requirements for NAS systems are also getting higher and higher.
[0003] Currently, NAS system performance testing primarily uses fixed load patterns, such as predefined read / write ratios and fixed data block sizes, to test metrics like throughput and latency. These methods fail to effectively simulate the dynamic traffic dynamics of real-world scenarios, leading to significant discrepancies between test results and actual system performance. Summary of the Invention
[0004] The present application provides a load generation method, a computer program product, and a network attached storage system to at least address the problem in the related art that the load generation method of the NAS system cannot effectively simulate the dynamic change characteristics of traffic in real business scenarios, resulting in a large deviation between the test results and the actual system performance.
[0005] The present application provides a load generation method, which is applied to a network attached storage system, and includes:
[0006] Loading a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system to generate a target test load corresponding to the target test scenario, wherein the scenario library includes a business scenario tag and the scenario load template corresponding to the business scenario tag;
[0007] Based on the real-time performance indicators of the network attached storage system under the target test load, obtaining load adjustment instructions through a reinforcement learning model;
[0008] The load adjustment instruction is executed to adjust the target test load.
[0009] The present application also provides a computer program product, which is applied to a network attached storage system, and includes:
[0010] A first processing module is configured to load a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system, and generate a target test load corresponding to the target test scenario, wherein the scenario library includes a business scenario tag and the scenario load template corresponding to the business scenario tag;
[0011] A second processing module is configured to obtain a load adjustment instruction through a reinforcement learning model based on the real-time performance indicator of the network attached storage system under the target test load;
[0012] The third processing module is configured to execute the load adjustment instruction to adjust the target test load.
[0013] The present application also provides a network attached storage system for executing the load generation method as described above.
[0014] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned load generation methods when executing the computer program.
[0015] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned load generation methods are implemented.
[0016] Through this application, a target test load corresponding to a target test scenario is generated by loading a scenario load template, the load environment of the target test scenario is quickly constructed, and real-time performance indicators are obtained and input into a reinforcement learning model. Since the reinforcement learning model can perform load adjustment actions under the target test scenario, and use real-time performance indicators as feedback, it outputs load adjustment instructions to guide the network attached storage system to perform load adjustment. This can solve the technical problem that the load generation method cannot effectively simulate the dynamic change characteristics of traffic in real business scenarios, resulting in a large deviation between the test results and the actual system performance. The load is dynamically adjusted based on the real-time performance of the system to accurately simulate the dynamic change characteristics of traffic in real business scenarios and improve the accuracy of the test results. At the same time, the dynamically changing load can effectively trigger the elastic expansion and contraction mechanism of the network attached storage system, thereby fully verifying the system's ability to dynamically adjust resources, so that the system can automatically allocate resources according to actual needs and achieve a balance between performance and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 This is a flow chart of a load generation method according to an embodiment of the present application;
[0019] Figure 2 A schematic diagram of the structure of a computer program product provided in an embodiment of the present application;
[0020] Figure 3 A schematic diagram of the structure of a network attached storage system provided in an embodiment of the present application;
[0021] Figure 4 This is the second flow chart of the load generation method provided in the embodiment of the present application;
[0022] Figure 5 This is the third flow chart of the load generation method provided in the embodiment of the present application;
[0023] Figure 6 This is the fourth flow chart of the load generation method provided in the embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0026] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0027] An embodiment of the present application provides a load generation method, and the method is described in detail in conjunction with the execution process of the load generation method.
[0028] The load generation method can be applied to a network attached storage (NAS) system. A network attached storage system is used for data storage and sharing and can be connected through a local area network or the Internet to provide centralized file storage and access services for multiple users or clients.
[0029] like Figure 1 As shown, the load generation method of the embodiment of the present application includes step 110, step 120 and step 130.
[0030] Step 110: Load a scenario load template corresponding to the target test scenario from a scenario library of the network attached storage system, and generate a target test load corresponding to the target test scenario.
[0031] The scenario library includes business scenario tags and scenario load templates corresponding to the business scenario tags.
[0032] It can be understood that the business scenario label is a mark for classifying the business scenarios faced by the network attached storage system. For the network attached storage system, the classification of its business scenarios may include steady-state scenarios, peak scenarios, and edge scenarios.
[0033] Each business scenario tag in the scenario library has a corresponding scenario load template. The scenario load template corresponding to a business scenario tag may include the setting information of the load parameters under the business scenario corresponding to the business scenario tag, such as the settings of parameters such as the number of reads and writes per second (IOPS), data block size, and protocol ratio.
[0034] For example, the business scenario of daily office document collaboration can be classified as a steady-state scenario. The corresponding scenario load template can be set to a read-write ratio of 6:4, mainly processing small files of 4KB-64KB.
[0035] For example, the business scenario of image content delivery network (CDN) back-to-source during e-commerce promotions can be classified as a peak scenario. The corresponding scenario load template can be set to a read ratio of > 95%, with large files larger than 1 MB read sequentially.
[0036] For another example, business scenarios with high-frequency writes on IoT devices can be classified as edge scenarios. The corresponding scenario load template can be set to write 95% of files under 4KB, with large fluctuations in network latency.
[0037] In this step, according to the target test scenario to be tested, a tag search is performed in the scenario library of the network attached storage system to determine the business scenario tag corresponding to the target test scenario, and the scenario load template corresponding to the business scenario tag is loaded. According to the load parameter settings of the scenario load template, the target test load corresponding to the target test scenario is generated.
[0038] For example, if the target test scenario is an edge scenario, search for the business scenario tag corresponding to the edge scenario in the scenario library, load the scenario load template corresponding to the business scenario tag, and generate the target test load corresponding to the edge scenario.
[0039] In actual execution, the target test load corresponding to the target test scenario can be generated through the load generation engine, which converts the scenario load template into an executable test load and quickly builds a load environment close to the real business scenario.
[0040] Step 120: Based on the real-time performance indicators of the network attached storage system under the target test load, obtain load adjustment instructions through the reinforcement learning model.
[0041] Generate a target test load corresponding to the target test scenario, and monitor the real-time performance indicators of the network attached storage system under the target test load in real time. The real-time performance indicators are used to characterize the performance status of the network attached storage system.
[0042] In actual implementation, the real-time performance indicators of the network-attached storage system can be monitored and stored in real time through tools such as the data visualization and monitoring platform Grafana and the time series database InfluxDB.
[0043] The reinforcement learning model is a machine learning method that learns optimal decision-making strategies by interacting with the environment. The agent performs actions in a given state and continuously adjusts its action execution strategy based on the rewards fed back by the environment to maximize long-term cumulative rewards, achieving autonomous learning and adaptive optimization.
[0044] The target test scenario and real-time performance indicators can be input into the reinforcement learning model. The reinforcement learning model performs load adjustment actions under the target test scenario, provides real-time feedback with real-time performance indicators, dynamically adjusts the load, and outputs load adjustment instructions for guiding the network attached storage system to perform load adjustment.
[0045] Step 130: Execute the load adjustment instruction to adjust the target test load.
[0046] In this step, the load adjustment instructions output by the reinforcement learning model are obtained, the target test load is adjusted according to the load adjustment instructions, and the real-time performance indicators of the network attached storage system are monitored at the same time. The target test scenario is simulated and verified, and a test report for the target test scenario is generated.
[0047] It should be noted that the load adjustment instructions output by the reinforcement learning model can take effect in real time.
[0048] It is understandable that with the rapid development of cloud computing, edge computing and artificial intelligence technologies, the business scenarios faced by network-attached storage systems have expanded from traditional file sharing to artificial intelligence (AI) training, Internet of Things (IoT) data lakes, real-time video analysis, etc., and the load patterns have become dynamic and hybrid.
[0049] Related technologies mainly use fixed-mode load generation tools such as FIO and IOzone to generate loads by pre-defining read-write ratios and fixed data block sizes to test indicators such as throughput and latency. This type of static load mode is divorced from actual business scenarios, and fixed load parameters cannot simulate the dynamic changes of business traffic.
[0050] At the same time, static loads make it difficult to trigger the elastic scaling mechanism of the network-attached storage system, and the dynamic adjustment capability of resources cannot be verified.
[0051] In an embodiment of the present application, a target test load corresponding to a target test scenario is generated by loading a scenario load template, a load environment of the target test scenario is quickly constructed, and real-time performance indicators are obtained and input into a reinforcement learning model. Since the reinforcement learning model can perform load adjustment actions under the target test scenario, and uses real-time performance indicators as feedback, it outputs load adjustment instructions to guide the network attached storage system to perform load adjustment. This can solve the technical problem that the load generation method cannot effectively simulate the dynamic change characteristics of traffic in real business scenarios, resulting in a large deviation between the test results and the actual system performance. The load is dynamically adjusted based on the real-time performance of the system to accurately simulate the dynamic change characteristics of traffic in real business scenarios and improve the accuracy of the test results. At the same time, the dynamically changing load can effectively trigger the elastic expansion and contraction mechanism of the network attached storage system, thereby fully verifying the system's ability to dynamically adjust resources, so that the system can automatically allocate resources according to actual needs and achieve a balance between performance and efficiency.
[0052] According to the load generation method provided in the embodiment of the present application, a target test load corresponding to a target test scenario is generated by loading a scenario load template, a load environment for the target test scenario is quickly constructed, real-time performance indicators are obtained and input into a reinforcement learning model, and the reinforcement learning model outputs load adjustment instructions to guide the network attached storage system to perform load adjustment. The load is dynamically adjusted based on the real-time performance of the system, accurately simulating the dynamic change characteristics of traffic in real business scenarios, which helps to improve the accuracy of the test results of the network attached storage system.
[0053] In some embodiments, the scene library is constructed by the following steps:
[0054] Obtain system log data from network attached storage systems;
[0055] Build a scenario library based on system log data.
[0056] Based on the system log data of the network-attached storage system, we can identify different business scenarios faced by the network-attached storage system (such as video streaming, collaborative office, etc.) and build a scenario library that includes business scenario tags and scenario load templates corresponding to the business scenario tags.
[0057] In actual implementation, Fluentd can be used to collect system log data from network-attached storage systems and buffer the data through Kafka message queues.
[0058] Among them, Fluentd is an open source data collector that can build a unified log collection layer for users to better use and understand log data; Kafka is an open source distributed stream processing platform.
[0059] It is understandable that the collected system log data can be cleaned to improve data quality and provide a reliable basis for subsequent data analysis and construction of scenario libraries.
[0060] For example, use data cleaning tools such as Apache Spark to clean system log data, such as abnormal breakpoint logs, duplicate records, standardize timestamp formats, and complete missing fields.
[0061] In actual implementation, business scenarios are identified based on system log data, and business scenario labels are added to different business scenarios. Business scenario labels can be manually labeled or automatically generated.
[0062] For example, system log data can be tagged according to business types (such as video streaming and database backup), and business scenario tags corresponding to typical business scenarios can be added.
[0063] For another example, by linking the timestamp with the system, the business scenario labels corresponding to different business scenarios can be automatically marked.
[0064] In related technologies, scripting is used to simulate specific business scenarios, which relies on manually predefined scenarios. Manual script writing is time-consuming and has limited scenario coverage, making it difficult to simulate business scenarios with complex load characteristics such as edge computing and AI training.
[0065] During the implementation of this application, based on system log data, the business scenarios of the network attached storage system are driven to be identified, business scenario tags and scenario load templates corresponding to the business scenario tags are created, and a scenario library of the network attached storage system is constructed. By loading the scenario load templates of the scenario library, the target test load corresponding to the target test scenario is generated to accurately simulate the real business scenario. Since the system log data is real data and is constantly updated, business scenario tags and scenario load templates corresponding to the business scenario tags are created in a dynamic generation manner, without the need for manual script writing, and various business scenarios can be covered.
[0066] It should be noted that the scenario library is dynamically updated based on the incremental learning mechanism. When a new business scenario is identified based on system log data, an incremental update is triggered. Based on scenario similarity matching, the cosine similarity is used to calculate the matching degree between the system log data and the existing scenario category. If the similarity is less than the preset threshold, a new scenario category is created.
[0067] In some embodiments, the system log data includes at least one of a protocol level log, a file metadata log, and a user behavior log of the network attached storage system.
[0068] The protocol-level log may refer to the operation details of the file transfer protocol of the network-attached storage system, including information such as read and write request timestamps, file paths, data block sizes, and client addresses.
[0069] It should be noted that when the network attached storage system includes multiple file transfer protocols, the protocol-level log includes operation details of the multiple file transfer protocols.
[0070] File metadata logs can include log data describing the file background, content, structure, and management process, such as file creation, deletion, modification events, and permission changes in the network attached storage system.
[0071] User behavior logs may include user operation traces and interaction data such as user login frequency, number of concurrent sessions, and directory access patterns.
[0072] In actual implementation, one or more combinations of protocol-level logs, file metadata logs, and user behavior logs of the network-attached storage system can be obtained to identify business scenarios of the network-attached storage system and build a scenario library.
[0073] In some embodiments, a scenario library is constructed based on system log data, including:
[0074] Extract features from system log data to obtain business scenario features of the network-attached storage system;
[0075] Cluster business scenario features and build a scenario library.
[0076] The business scenario features are extracted based on system log data, and can reflect the characteristics of the business scenarios faced by the network attached storage system.
[0077] Through unsupervised learning clustering processing, the business scenario features extracted from system log data are divided into several groups. The data within the same group are similar to each other (that is, these data belong to the same type of business scenario), and the data between different groups are significantly different. This allows accurate identification of different business scenarios of network-attached storage systems, providing a reliable basis for building a scenario library.
[0078] In some embodiments, feature extraction is performed on system log data to obtain business scenario features of the network attached storage system, including:
[0079] Based on system log data, at least one of time distribution features, input and output pattern features, protocol and user behavior features, and resource occupancy features is extracted to obtain business scenario features.
[0080] Among them, the time distribution feature is used to characterize the regular pattern of system log data changing over time, which can be expressed as periodic or sudden forms.
[0081] For example, time distribution characteristics may include request frequency periodicity characteristics (such as daily peak hours and low load on weekends) and burst traffic characteristics.
[0082] Input-output (IO) pattern characteristics are used to characterize the characteristics of data flow in a network-attached storage system.
[0083] In actual implementation, the input and output pattern characteristics may include data flow characteristic parameters such as the dynamic distribution of read-write ratio, data block size distribution, and the ratio of random access to sequential access.
[0084] Protocol and user behavior characteristics are used to characterize file transfer protocols and user behavior characteristics in network attached storage systems.
[0085] In actual implementation, protocol and user behavior characteristics may include the proportion of multi-protocol traffic (such as 60% for NFS, 30% for SMB, and 10% for FTP), the number of concurrent user sessions, and the frequency of file lock contention (such as lock timeout events under the SMB protocol).
[0086] Resource usage characteristics are used to characterize the consumption patterns and extent of various computing resources of a network-attached storage system during operation. They can be expressed as dynamic changes in the time dimension (such as periodic peaks, sustained high loads, or sudden usage) and can be quantified through indicators such as processor (CPU) utilization, memory usage, and response latency.
[0087] In actual implementation, resource usage characteristics may include the correlation between CPU or memory utilization and the number of read / write operations per second (IOPS). Resource usage characteristics may also include the linear relationship between network bandwidth and throughput (such as the maximum throughput of a gigabit network card).
[0088] Perform feature extraction on system log data to obtain a combination of one or more of time distribution features, input and output pattern features, protocol and user behavior features, and resource occupancy features to obtain business scenario features.
[0089] In actual implementation, feature engineering such as sliding window statistics, principal component analysis, or time series feature coding can be used to extract at least one of the time distribution characteristics, input and output pattern characteristics, protocol and user behavior characteristics, and resource occupancy characteristics of system log data.
[0090] In some embodiments, clustering business scenario features to build a scenario library includes:
[0091] Cluster the business scenario features using a density clustering algorithm or a hierarchical clustering algorithm to obtain feature clustering results;
[0092] Map the feature clustering results into feature vector templates to build a scenario library.
[0093] Among them, the density clustering algorithm (DBSCAN) is suitable for non-uniformly distributed log data and automatically identifies noise points (such as temporary test traffic).
[0094] In actual implementation, the neighborhood radius (Eps) of the density clustering algorithm can be determined by the k-distance curve, and the minimum number of samples (MinPts) can be set to 10% of the request volume in the time window.
[0095] The hierarchical clustering algorithm is a clustering method based on hierarchical decomposition. By recursively dividing data into different levels, it can help verify the hierarchical structure of business scenarios (for example, "video streaming" can be subdivided into live and on-demand subcategories).
[0096] The business scenario features are clustered using a density clustering algorithm or a hierarchical clustering algorithm, and the feature clustering results obtained by clustering are mapped into a feature vector template to build a scenario library, where the feature vector template includes information about the business scenario label and the scenario load template corresponding to the business scenario label.
[0097] For example, the feature vector template may include information such as scenario ID, read-write ratio, data block distribution, protocol proportion, number of concurrent threads, network conditions, etc. The scenario ID corresponds to the business scenario label, and the scenario load template corresponding to the business scenario label includes information such as read-write ratio, data block distribution, protocol proportion, number of concurrent threads and network conditions.
[0098] It should be noted that the constructed scenario library can be verified and tuned.
[0099] In actual implementation, the business scenario tags and scenario load templates of the scenario library can be verified through cross-validation, business semantic consistency check, load playback test, etc.
[0100] For example, the historical data of system log data is divided into training set and test set according to time, and the stability of clustering results is verified through cross-validation.
[0101] For example, by manually reviewing the clustering results, we can ensure that the scenario definitions of business scenario labels such as "video streaming" and "database backup" are consistent with the actual business logic, thereby achieving business semantic consistency checks.
[0102] For another example, perform a load playback test, drive the load generation engine based on the scenario load template, compare the test results with historical performance data, and detect whether the error rate is within the preset range.
[0103] In actual implementation, tuning strategies include but are not limited to feature weight optimization, noise filtering and other strategies.
[0104] For example, the importance of features can be calculated through the random forest model, and the feature weights of the clustering algorithm can be adjusted (such as increasing the weight of the "burst traffic indicator") to optimize data clustering.
[0105] For another example, temporary loads lasting less than 5 minutes (such as temporary backup operations by maintenance personnel) can be eliminated to filter out noise and optimize data clustering.
[0106] In the embodiment of the present application, based on the multi-dimensional feature extraction of system log data, a clustering algorithm is used to automatically generate a scenario library, support incremental updates and similarity matching, and convert unstructured log data into a drivable load template, breaking through the limitations of manually predefined scenarios, accurately simulating real business scenarios, and providing more comprehensive business scenario coverage.
[0107] In some embodiments, the real-time performance indicator includes at least one of a system resource indicator, a network status indicator, and a protocol level indicator.
[0108] The system resource indicators may include CPU utilization, memory usage, disk IOPS, and other indicators that represent the resource status of the network attached storage system.
[0109] In actual implementation, system resource indicators can be collected through tools such as Prometheus Agent.
[0110] System resource indicators may also include indicators that characterize the status of the storage layer of the network-attached storage system, such as cache hit rate, solid-state drive (SSD) wear leveling status, etc.
[0111] Network status indicators may include network bandwidth, delay, packet loss rate, and other indicators that characterize the network status of the network attached storage system.
[0112] In actual implementation, network bandwidth can be collected through Prometheus, latency can be collected through PingMesh, and packet loss rate can be obtained through NetFlow analysis.
[0113] Protocol-level metrics are indicators that characterize the file transfer protocol performance of network-attached storage systems.
[0114] For example, for the NFS file transfer protocol, you can count the remote procedure call (RPC) latency, and for the SMB file transfer protocol, you can monitor the file lock contention frequency.
[0115] In actual implementation, the real-time performance indicators of the network-attached storage system can be monitored and stored in real time through tools such as the data visualization and monitoring platform Grafana and the time series database InfluxDB.
[0116] In some embodiments, the state space of the reinforcement learning model includes performance metrics, protocol characteristics, scenario context, and historical trends of the network attached storage system;
[0117] The action space of the reinforcement learning model includes load intensity adjustment of network-attached storage systems, protocol ratio adjustment, and network impairment injection;
[0118] The reward function of the reinforcement learning model is constructed based on latency, throughput, over-limit penalty, and protocol violation penalty.
[0119] It can be understood that the reinforcement learning model belongs to the state-action-reward model structure, where its state is determined by real-time performance indicators and target test scenarios. Dynamic adjustment of load parameters is an action that the reinforcement learning model can perform, guiding model optimization according to the reward function.
[0120] Defining the state space of the reinforcement learning model may include configuring the performance indicators, protocol characteristics, scenario context, and historical trends of the network-attached storage system.
[0121] Among them, performance indicators include CPU utilization, memory usage, average input and output delays, throughput, etc.; protocol characteristics include the request ratio of different file transfer protocols (such as NFS / SMB / FTP), file lock competition frequency, RPC retransmission rate, etc.; scenario context includes the ID of the current test scenario, network damage mode (such as "high latency + packet loss"), etc.; historical trends can include the indicator variance and maximum value (used to detect burst traffic) of system data within a sliding window (5 minutes).
[0122] Actions that the reinforcement learning model can perform include load scaling for network-attached storage systems, protocol scaling, and network impairment injection.
[0123] Load intensity adjustment can include adjusting the number of concurrent threads and the IOPS target value. For example, the number of concurrent threads can be adjusted from 200 to 300, and the IOPS target value can be increased from 50,000 to 60,000.
[0124] Protocol ratio adjustment can dynamically allocate traffic weights for different file transfer protocols. For example, the traffic weight of the NFS file transfer protocol can be reduced from 60% to 55%, while the traffic weight of the SMB file transfer protocol can be increased from 30% to 35%.
[0125] Network impairment injection can include operations such as increasing delay, decreasing delay, increasing packet loss rate, and decreasing packet loss rate.
[0126] The reward function of the reinforcement learning model is constructed based on delay, throughput, over-limit penalty, and protocol conflict penalty. Corresponding weight coefficients (dynamically adjustable) can be set for delay, throughput, over-limit penalty, and protocol conflict penalty, and the weighted sum of delay, throughput, over-limit penalty, and protocol conflict penalty is used as the reward function.
[0127] For example, the reward function R of the reinforcement learning model is formulated as follows:
[0128] R = w1 × delay + w2 × throughput - w3 × over-limit penalty - w4 × protocol conflict penalty;
[0129] Among them, w1, w2, w3, and w4 are the weight coefficients of delay, throughput, over-limit penalty, and protocol conflict penalty, respectively.
[0130] For example, w1=0.6, w2=0.2, w3=0.1, w4=0.1.
[0131] In actual implementation, the over-limit penalty item can be triggered based on the CPU utilization of the network-attached storage system. For example, if the CPU utilization is greater than 90%, the over-limit penalty item is triggered and 10 points are deducted.
[0132] The protocol conflict penalty item can be determined based on the number of conflicts between different file transfer protocols of the network attached storage system. For example, if the number of SMB protocol lock timeouts is greater than 100 times / minute, the protocol conflict penalty item will be triggered and 5 points will be deducted.
[0133] It can be understood that latency refers to the time from when data is sent to when it is received. The smaller the latency, the faster the data transmission speed, and the more points the reward function adds. The larger the latency, the slower the data transmission speed, and the fewer points the reward function adds.
[0134] For example, a delay of 1 second will give 1 point, and a delay of 0.5 seconds will give 2 points.
[0135] For target test scenarios that focus on data transmission speed, the weight coefficient corresponding to the delay can be increased so that the score caused by the delay accounts for a higher proportion.
[0136] Throughput is the amount of data processed per unit time. The higher the throughput, the more bonus points you get.
[0137] For example, a throughput of 100 MB / s is worth 100 points, and a throughput of 200 MB / s is worth 200 points.
[0138] For target test scenarios that focus on data processing volume, the weight coefficient corresponding to throughput can be increased to increase the score contribution from throughput.
[0139] Over-limit penalties can be triggered based on the CPU utilization of the NAS system. If the CPU utilization exceeds the safety limit, the NAS system is at risk of freezing. Each time the CPU utilization exceeds the safety limit, an over-limit penalty is triggered, and points are deducted (a minus sign appears before the over-limit penalty item). The greater the excess percentage, the more points are deducted.
[0140] For example, if the CPU utilization is 90%, 1 point will be deducted; if the CPU utilization is 95%, 3 points will be deducted.
[0141] For target test scenarios that focus on whether the system is stuck, the weight coefficient corresponding to the over-limit penalty item can be increased.
[0142] The number of protocol conflicts can be the number of times a file is processed simultaneously by different file transfer protocols. One point will be deducted for each protocol conflict.
[0143] For example, if there is one conflict, 1 point will be deducted; if there are five conflicts, 5 points will be deducted.
[0144] For target test scenarios that focus on avoiding protocol conflicts, the weight coefficient corresponding to the protocol conflict penalty item can be increased.
[0145] The reinforcement learning model is built based on the deep Q network. The training phase of the reinforcement learning model includes the offline pre-training phase and the online learning phase.
[0146] Among them, the Deep Q-Network (DQN) is a reinforcement learning network that combines deep learning and Q-learning. It approximates the Q-value function (action-value function) through a neural network. The reinforcement learning model is built based on the Deep Q-Network and is suitable for the high-dimensional state space and load-related continuous action decision-making of network-attached storage systems.
[0147] It is understandable that the training phase of the reinforcement learning model includes an offline pre-training phase, in which preliminary training is performed before the reinforcement learning model is put into the load adjustment of the network attached storage system so that the reinforcement learning model can learn a general feature representation.
[0148] In actual implementation, a historical dataset (e.g., containing 100,000 state-action-reward records) can be used to pre-train the reinforcement learning model and initialize the model. During the pre-training process, experience replay can be used to reduce data correlation and improve convergence speed.
[0149] The training phase of the reinforcement learning model includes an online learning phase, which continuously optimizes the load strategy based on the online learning mechanism to adapt to the dynamic changes of the network-attached storage system (such as node expansion and cache strategy update).
[0150] In actual implementation, the reinforcement learning model can collect data from the network-attached storage system at regular intervals (e.g., 5 minutes), generate actions and execute them, and record the reward values to update the model parameters.
[0151] In some embodiments, loading a scenario load template corresponding to a target test scenario from a scenario library of a network attached storage system and generating a target test load corresponding to the target test scenario may include:
[0152] Parse the scenario load template corresponding to the target test scenario, perform initial load parameter configuration, protocol engine configuration, network environment pre-simulation and resource baseline calibration, and generate the target test load.
[0153] Generating target test load may include sub-processes such as initial load parameter configuration, protocol engine configuration, network environment pre-simulation, and resource baseline calibration.
[0154] The initial load parameter configuration may refer to generating basic load parameters (eg, IOPS, data block size, protocol ratio, etc.) of the network attached storage system according to the scenario load template definition.
[0155] The protocol engine configuration may refer to initializing the file transfer protocol client of the network attached storage system and setting basic client parameters such as authentication and mount points.
[0156] In some embodiments, the protocol engine configuration may include:
[0157] Initialize multiple protocol clients based on different protocol types and set parameters for each protocol client.
[0158] The network attached storage system may include multiple protocol clients of different protocol types. For example, the network attached storage system may include clients based on the NFS file transfer protocol, clients based on the SMB file transfer protocol, clients based on the FTP file transfer protocol, and the like.
[0159] Configure the protocol engine, initialize the NFS, SMB, and FTP protocol clients, and set basic client parameters such as authentication and mount points for each protocol client.
[0160] It should be noted that when loading the scenario load template and generating the target test load, the protocol engine configuration is performed, and the load adjustment instructions output by the deep learning model may include protocol engine control.
[0161] The deep learning model can automatically allocate the traffic ratios of NFS, SMB, and FTP protocols according to the requirements of the target test scenario (for example, 60% for NFS, 30% for SMB, and 10% for FTP), and simulate resource competition between protocols (such as lock conflicts and metadata operations) to achieve dynamic allocation of protocol traffic.
[0162] The deep learning model can generate differentiated loads based on different protocol characteristics (for example, NFS tends to read and write large files sequentially, while SMB tends to access small files randomly), achieving protocol-sensitive load generation.
[0163] Most of the related technologies conduct separate tests for different protocols, without considering resource competition and performance interference in multi-protocol concurrent scenarios (such as the interactive impact of SMB locks and NFS cache failures). They rely on manual experience to configure parameters, lack the ability to dynamically allocate cross-protocol traffic, cannot adapt to changes in business scenarios, and cannot verify the stability of NAS under multi-protocol mixed access.
[0164] In the embodiments of the present application, full consideration is given to resource competition and performance interference in multi-protocol concurrent scenarios, target test complexity is generated through protocol engine configuration, a multi-protocol hybrid access environment is constructed, dynamic traffic distribution between different protocols is achieved through a deep learning model, and differentiated loads are generated for different protocol characteristics. This can adapt to changes in different business scenarios and effectively verify the stability and reliability of the network attached storage system in a multi-protocol hybrid access environment.
[0165] Network environment pre-simulation can refer to configuring corresponding network impairment rules according to the target test scenario to simulate the network environment. For example, configuring network impairment rules related to delay and packet loss for edge scenarios.
[0166] Resource baseline calibration can refer to collecting the initial state of the network-attached storage system (such as CPU utilization, memory occupancy, etc.) as a benchmark for subsequent dynamic adjustments of the reinforcement learning model.
[0167] A specific embodiment is described below.
[0168] A scenario load template corresponding to a target test scenario is loaded from a scenario library of a network attached storage system, and the scenario load template corresponding to the target test scenario is parsed.
[0169] The contents of the scenario load template include:
[0170] Scenario ID: Edge_AI_Training, Read-Write Ratio: 70:30, Data Block Distribution: {4KB: 40%, 64KB: 50%, 1MB: 10%}, Protocol Ratio: {4KB: 40%, 64KB: 50%, 1MB: 10%}, Number of Concurrent Threads: 200, Network Conditions: {Baseline Latency: 50ms, Jitter Range: ±30ms, Packet Loss Rate: 2%}, Special Rules: Trigger a peak write every 10 minutes during the training cycle.
[0171] Initial load parameter configuration is performed according to parameter mapping rules. IOPS is calculated by scaling the load proportionally based on historical peak loads (e.g., 100K IOPS) and current hardware configuration (e.g., number of CPU cores on the test machine). A stratified random sampling algorithm is used to ensure that the data block size distribution conforms to the template definition (e.g., 40% 4KB requests) and calculate the data block distribution. A weighted round-robin algorithm is used to schedule multi-protocol traffic. For example, out of every 100 requests, 60 are allocated to NFS, 30 to SMB, and 10 to FTP, thereby weighting the protocols.
[0172] Configure the protocol engine, configure the mount parameters and authentication mode for the NFS client, configure the Samba connection parameters and file lock policy for the SMB client, and configure the passive mode (PASV) and large file transfer optimization for the FTP client.
[0173] Network environment pre-simulation includes toolchain integration and dynamic damage rules. Toolchain integration uses TC (Traffic Control) and NetEm to simulate network delay and packet loss, and builds a virtual network topology through Mininet to simulate cross-domain communication between edge nodes and cloud NAS (delay of 10ms).
[0174] Dynamic damage rules can include periodic fluctuations, such as randomly adjusting the delay (±20ms) and packet loss rate (±1%) every 5 minutes to simulate mobile network instability. Dynamic damage rules can also include sudden interruptions, simulating edge nodes to reconnect after 10 seconds of disconnection, and testing the NAS automatic retransmission mechanism.
[0175] Resource baseline calibration can include taking an initial status snapshot of the network-attached storage system, collecting initial device indicators such as CPU utilization and memory occupancy through SNMP or REST API, and recording storage layer status such as RAID group health status and SSD remaining life.
[0176] Resource baseline calibration can also include baseline threshold settings, such as setting a dynamic scaling trigger condition where the load generator automatically reduces the number of concurrent threads by 10% when the CPU utilization is >80% for three consecutive minutes.
[0177] In some embodiments, adjusting the target test load may include:
[0178] Perform at least one of protocol engine control, network impairment real-time update, and load intensity adjustment.
[0179] The reinforcement learning model outputs load adjustment instructions, which can be used to instruct protocol engine control, real-time network damage updates and / or load intensity adjustment.
[0180] Among them, protocol engine control can include operations such as dynamic client expansion and contraction and protocol ratio switching.
[0181] For example, for dynamic scaling of NFS and SMB clients, the number of Pod replicas can be dynamically increased or decreased through the Kubernetes API (for example, from 10 Pods to 15). The thread pool size of a single client can also be adjusted to generate differentiated loads for different protocol characteristics, thus achieving protocol-sensitive load generation.
[0182] For example, using a weighted random algorithm to switch the ratio of different protocols and allocating every 100 or 200 requests according to the new ratio can simulate resource competition between protocols and achieve dynamic allocation of protocol traffic.
[0183] It is understandable that network impairment refers to the phenomenon of reduced data transmission quality or service performance due to various factors in the network transmission process (such as bandwidth limitations, delays, packet loss, jitter, noise, etc.).
[0184] Real-time network damage updates refer to adjusting network parameters such as bandwidth limitations, latency, and packet loss to simulate different network states and accurately test the performance of network-attached storage systems under different network conditions.
[0185] For example, you can call the TC command to dynamically modify the network rules, adjust the delay from 50ms to 60ms, and update the network impairment of the network attached storage system.
[0186] Load intensity adjustment refers to adjusting the load that a network attached storage system can bear within a certain period of time.
[0187] For example, the Flexible I / O Tester (FIO) tool provides an API or command-line interface to adjust load-related parameters such as the number of IOPS, data block size, and number of concurrent threads in real time without restarting the process.
[0188] It should be noted that the reinforcement learning model can achieve closed-loop optimization and adaptive learning, and continuously optimize the load adjustment strategy.
[0189] Among them, closed-loop optimization can be divided into short-term feedback and long-term optimization.
[0190] Short-term feedback can be used to perform model inference at regular intervals (such as 5 minutes), generate adjustment actions, update the Q-value table based on real-time reward values, and optimize the load adjustment strategy.
[0191] Long-term optimization can generate test reports at regular intervals (such as 24 hours), mark high-frequency adjustment actions, and specifically enhance the sample weights of relevant state-action pairs during offline model training.
[0192] When a new protocol is detected, the protocol dimension of the reinforcement learning model state space is automatically expanded, and the exploration mode is started to adapt to the new protocol.
[0193] When the network-attached storage system node capacity is expanded (such as adding an SSD storage pool), the reinforcement learning model automatically identifies resource margins and increases the load intensity limit.
[0194] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0195] An embodiment of the present application further provides a computer program product, which can be applied to a network attached storage system.
[0196] like Figure 2 As shown, the computer program product comprises:
[0197] A first processing module 210 is configured to load a scenario load template corresponding to a target test scenario from a scenario library of a network attached storage system to generate a target test load corresponding to the target test scenario, wherein the scenario library includes business scenario tags and scenario load templates corresponding to the business scenario tags;
[0198] A second processing module 220 is configured to obtain load adjustment instructions through a reinforcement learning model based on the real-time performance indicators of the network attached storage system under the target test load;
[0199] The third processing module 230 is configured to execute the load adjustment instruction to adjust the target test load.
[0200] It should be noted that, for the description of the features in the embodiments corresponding to the computer program product, reference can be made to the relevant description of the embodiments corresponding to the load generation method, which will not be repeated here.
[0201] An embodiment of the present application also provides a network attached storage system for executing the load generation method as described above.
[0202] like Figure 3 As shown in Figure 1, the network attached storage system can be divided into application layer, control layer, processing layer, data layer and infrastructure layer.
[0203] The infrastructure layer of the network-attached storage system includes NAS device clusters, edge nodes, and network damage devices. A NAS device cluster refers to a device cluster formed by multiple interconnected network-attached storage devices, which has high availability, load balancing, and horizontal expansion capabilities. Edge nodes are lightweight computing storage units close to data sources or end users, and are used in edge computing scenarios such as the Internet of Things.
[0204] Among them, the test execution console can be used as a management tool for network-attached storage systems to perform tests in different business scenarios. The visual report dashboard can monitor and store data such as the load and performance indicators of the network-attached storage system in real time, and intuitively display the test process.
[0205] The network impairment controller can control the network impairment device to inject network impairments into the network attached storage system to simulate different network environments.
[0206] The load generation engine can load the scenario load template corresponding to the business scenario tag, generate the target test load corresponding to the target test scenario, convert the scenario load template into an executable test load, and quickly build a load environment close to the real business scenario.
[0207] The input of the reinforcement learning model can include the target test scenario and real-time performance indicators. The load adjustment action is performed under the target test scenario, and real-time feedback is provided with the real-time performance indicators. The load is dynamically adjusted and load adjustment instructions are output. The load generation engine executes the load adjustment instructions.
[0208] Among them, the real-time performance indicators input by the reinforcement learning model can be provided by the performance monitoring database and protocol traffic capture.
[0209] The system log data of the network-attached storage system is collected by NAS logs. Through the data analysis engine, the different business scenarios faced by the network-attached storage system are identified. A scenario library is constructed, including business scenario tags and scenario load templates corresponding to the business scenario tags, providing a data foundation for the scenario modeling engine and load generation engine.
[0210] A specific embodiment is described below.
[0211] like Figure 4As shown, NAS logs collect system log data, the scenario library provides scenario load templates, and the scenario modeling engine and load generation engine construct the target test load under the target test scenario, perform real-time performance monitoring of the network attached storage system, and use the load adjustment strategy output by the reinforcement learning model to dynamically adjust the load generated by the load generation engine.
[0212] The load adjustment strategy output by the reinforcement learning model can also include network damage control, which can realize the performance testing of edge nodes such as the Internet of Things and the Internet of Vehicles.
[0213] like Figure 5 As shown, taking the target test scenario as an edge scenario as an example, data collection and scenario modeling are performed, system log data is collected for feature extraction, and a scenario library is constructed.
[0214] The load generation engine is initialized, the scenario load template corresponding to the edge scenario in the scenario library is parsed, initial load parameter configuration, protocol engine configuration, network environment pre-simulation and resource baseline calibration are performed, and the target test load is generated.
[0215] Monitor the real-time performance indicators of network-attached storage systems, adjust load parameters through reinforcement learning models, perform edge scenario simulation and verification (e.g., network damage injection, fault recovery testing, data consistency verification), and generate test reports corresponding to edge scenarios.
[0216] like Figure 6 As shown in the figure, real-time performance indicators such as monitoring delay, throughput, and CPU utilization are input into the reinforcement learning model, and the reinforcement learning model outputs action instructions such as adjusting the number of concurrent threads, switching protocol ratios, and modifying network damage parameters. The corresponding load adjustment is performed, and the effect of the adjustment is verified according to the real-time performance indicators. Based on the adjustment effect, the rollback to the output action instruction is triggered, and the load adjustment instruction is continuously optimized.
[0217] In an embodiment of the present application, a scenario library is constructed based on log data, and load adjustment is performed through a reinforcement learning model. The dynamic load generation strategy can effectively improve the coverage of test scenarios, make edge scenario simulation more realistic, accurately simulate the actual business pressure of the network-attached storage system, automate protocol traffic distribution, and improve the efficiency of multi-protocol concurrent testing. Dynamic load can also trigger the elastic expansion and contraction of the network-attached storage system to optimize resources and reduce costs.
[0218] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above load generation method embodiments.
[0219] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned load generation method embodiments when running.
[0220] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0221] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0222] The above is a detailed introduction to a load generation method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A load generation method, characterized in that: The method is applied to a network attached storage system, and the method comprises: Loading a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system to generate a target test load corresponding to the target test scenario, wherein the scenario library includes a business scenario tag and the scenario load template corresponding to the business scenario tag; Based on the real-time performance indicators of the network attached storage system under the target test load, obtaining load adjustment instructions through a reinforcement learning model; Executing the load adjustment instruction to adjust the target test load; Inputting the target test scenario and the real-time performance indicator into the reinforcement learning model, the reinforcement learning model performing load adjustment actions under the target test scenario, providing real-time feedback based on the real-time performance indicator, dynamically adjusting the load, and outputting the load adjustment instruction for guiding the network attached storage system to perform load adjustment; The scene library is constructed through the following steps: Obtaining system log data of the network attached storage system; The scenario library is constructed based on the system log data.
2. The load generation method according to claim 1, wherein: The system log data includes at least one of a protocol-level log, a file metadata log, and a user behavior log of the network attached storage system.
3. The load generation method according to claim 1, wherein: The step of constructing the scenario library based on the system log data includes: Performing feature extraction on the system log data to obtain business scenario features of the network attached storage system; Clustering is performed on the business scenario features to construct the scenario library.
4. The load generation method according to claim 3, characterized in that: The extracting features from the system log data to obtain business scenario features of the network attached storage system includes: Based on the system log data, at least one of time distribution features, input and output mode features, protocol and user behavior features, and resource occupancy features is extracted to obtain the business scenario features.
5. The load generation method according to claim 3, characterized in that: The clustering process of the business scenario features to construct the scenario library includes: Clustering the business scenario features according to a density clustering algorithm or a hierarchical clustering algorithm to obtain a feature clustering result; The feature clustering result is mapped into a feature vector template to construct the scene library.
6. The load generation method according to any one of claims 1 to 5, characterized in that: The real-time performance indicator includes at least one of a system resource indicator, a network status indicator, and a protocol level indicator.
7. The load generation method according to any one of claims 1 to 5, characterized in that: The state space of the reinforcement learning model includes performance indicators, protocol characteristics, scenario context and historical trends of the network attached storage system; The action space of the reinforcement learning model includes load intensity adjustment, protocol ratio adjustment and network damage injection of the network attached storage system; The reward function of the reinforcement learning model is constructed based on delay, throughput, over-limit penalty and protocol violation penalty.
8. The load generation method according to any one of claims 1 to 5, characterized in that: The reinforcement learning model is constructed based on a deep Q network, and the training phase of the reinforcement learning model includes an offline pre-training phase and an online learning phase.
9. The load generation method according to any one of claims 1 to 5, characterized in that: The step of loading a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system to generate a target test load corresponding to the target test scenario includes: The scenario load template corresponding to the target test scenario is parsed, and initial load parameter configuration, protocol engine configuration, network environment pre-simulation and resource baseline calibration are performed to generate the target test load.
10. The load generation method according to claim 9, characterized in that: The protocol engine configuration includes: Initialize multiple protocol clients based on different protocol types and set parameters for each of the protocol clients.
11. The load generation method according to any one of claims 1 to 5, characterized in that: The adjusting the target test load includes: Perform at least one of protocol engine control, network impairment real-time update, and load intensity adjustment.
12. A computer program product, characterized in that The computer program product is applied to a network attached storage system, and the computer program product includes: A first processing module is configured to load a scenario load template corresponding to a target test scenario from a scenario library of the network attached storage system, and generate a target test load corresponding to the target test scenario, wherein the scenario library includes a business scenario tag and the scenario load template corresponding to the business scenario tag; A second processing module is configured to obtain a load adjustment instruction through a reinforcement learning model based on the real-time performance indicator of the network attached storage system under the target test load; A third processing module is configured to execute the load adjustment instruction to adjust the target test load; The second processing module is configured to input the target test scenario and the real-time performance indicator into the reinforcement learning model, wherein the reinforcement learning model performs a load adjustment action under the target test scenario, provides real-time feedback based on the real-time performance indicator, dynamically adjusts the load, and outputs the load adjustment instruction for guiding the network attached storage system to perform load adjustment; The scene library is constructed through the following steps: Obtaining system log data of the network attached storage system; The scenario library is constructed based on the system log data.
13. A network attached storage system, characterized in that: Used to execute the load generation method according to any one of claims 1 to 11.
14. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the load generation method according to any one of claims 1 to 11 when executing the computer program.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the load generation method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Performance test method in server cluster and related equipment
CN112162891A
Cloud computing platform real-time CPU load balancing method and system based on reinforcement learning
CN118885291A
Cloud storage system performance test method and system based on large model
CN119088628A