A Solid State Drive Load Testing Method, System, Electronic Device and Medium
By acquiring and processing the IO request feature data of solid-state drives in multiple application scenarios, generating data access thermal maps and separating hot and cold data streams, the problem that existing testing methods are difficult to accurately simulate application scenarios is solved, and the accuracy of solid-state drive load testing is significantly improved.
Patent Information
- Application Number
- CN202510333025.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-20
AI Technical Summary
The existing solid-state drive load testing methods are difficult to accurately simulate data access modes in different application scenarios, resulting in large differences in the test results and actual application scenarios, reducing the accuracy of the test.
By obtaining the IO request feature data of solid-state drives in multiple application scenarios, performing feature classification and de-redundancy processing, and establishing a load feature fingerprint library. Then, a data access thermal map is generated based on the target fingerprint of the target application scenario, and the benchmark load is divided into cold data streams and hot data streams, a mixed load sequence is generated, and it is loaded in parallel to multiple physical areas of the solid-state drive to be tested for performance testing.
This method can accurately capture the characteristics of data access in different application scenarios, generate test loads that are closer to actual application scenarios, and significantly improve the accuracy of solid-state drive load testing.
Smart Images

Figure CN119847900B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly relates to a method, system, electronic device, and medium for solid-state drive load testing. Background Art
[0002] With the continuous growth of data processing requirements, solid-state drives (SSDs) have been widely used in various scenarios such as personal computers, enterprise-level servers, and big data centers due to their high-speed read and write performance and reliability. In different application scenarios, the data access patterns and workload characteristics faced by solid-state drives are significantly different, which directly affects the actual performance of solid-state drives. Therefore, effective performance evaluation of solid-state drives for specific application scenarios is of great significance for system optimization and hardware selection.
[0003] Currently, existing solid-state drive load testing methods mainly use standard benchmark testing tools to simulate workloads by setting fixed input and output, that is, fixed I / O for testing. However, in actual applications, testing with only fixed workloads often fails to effectively capture the changes in data access in different application scenarios, and there are often significant differences between the test results and the actual application scenarios, thus reducing the accuracy of solid-state drive load testing. Summary of the Invention
[0004] This application provides a method, system, electronic device, and medium for solid-state drive load testing, which has the effect of improving the accuracy of solid-state drive load testing.
[0005] In a first aspect, this application provides a method for solid-state drive load testing, including:
[0006] Obtaining the I / O request feature data of the solid-state drive in multiple application scenarios;
[0007] Classifying the I / O request feature data, and performing redundancy removal processing on various data features after feature classification through a segmented compression algorithm to obtain a load feature fingerprint library;
[0008] Extracting the target fingerprint of the target application scenario from the load feature fingerprint library, and generating a data access heat map according to the target fingerprint;
[0009] Based on the data access heat map, dividing the benchmark test load into cold data streams and hot data streams, and performing read and write operations on the cold data streams and the hot data streams to generate a mixed load sequence;
[0010] Parallelly loading the mixed load sequence into multiple physical regions of the solid-state drive to be tested, and performing performance testing to obtain the test result of the solid-state drive to be tested.
[0011] In a second aspect of the present application, a solid-state drive load testing system is provided. The system includes:
[0012] A data acquisition module for acquiring IO request feature data of a solid-state drive under multiple application scenarios;
[0013] A heat map determination module for classifying the features of the IO request feature data, and performing redundancy removal processing on various data features after feature classification through a segmented compression algorithm to obtain a load feature fingerprint library; extracting a target fingerprint of a target application scenario from the load feature fingerprint library, and generating a data access heat map according to the target fingerprint;
[0014] A load sequence determination module for dividing a benchmark test load into a cold data stream and a hot data stream based on the data access heat map, and performing read and write operations on the cold data stream and the hot data stream to generate a mixed load sequence;
[0015] A hard disk test module for parallelly loading the mixed load sequence into multiple physical regions of a solid-state drive to be tested, and performing a performance test to obtain a test result of the solid-state drive to be tested.
[0016] In a third aspect of the present application, an electronic device is provided, including a memory, a processor, and a program stored on the memory and executable on the processor. When the program is loaded and executed by the processor, it can implement a solid-state drive load testing method.
[0017] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is enabled to implement a solid-state drive load testing method.
[0018] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0019] By adopting the above technical solutions, the IO request feature data of solid-state drives in multiple application scenarios are obtained, and these feature data are classified and redundancy-eliminated to establish a load feature fingerprint library, so as to accurately capture the feature differences in data access under different application scenarios. Based on the target fingerprints of the target application scenarios extracted, this solution generates a data access heat map to visually present the data access distribution, and accordingly divides the benchmark test load into cold data streams and hot data streams that conform to the actual usage characteristics. By performing read and write operations on the cold data streams and hot data streams, a mixed load sequence closer to the real application scenario is generated, and then this mixed load sequence is parallelly loaded into multiple physical regions of the solid-state drive to be tested for performance testing, making full use of the physical parallel characteristics of the solid-state drive. This test method based on the characteristics of the actual application scenario effectively solves the problem that the existing test methods are difficult to accurately simulate the data access mode under specific application scenarios, and the test results are closer to the performance performance of the solid-state drive in the actual application environment, thus significantly improving the accuracy of the solid-state drive load test. Description of the Drawings
[0020] Figure 1 is a schematic flowchart of a method for testing the load of a solid-state drive provided by an embodiment of the present application;
[0021] Figure 2 is a schematic structural diagram of a system for testing the load of a solid-state drive provided by an embodiment of the present application;
[0022] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0023] Description of the reference numerals: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Embodiments
[0024] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.
[0025] In the description of the embodiments of the present application, words such as "for example" or "for instance" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "for example" or "for instance" is intended to present relevant concepts in a specific manner.
[0026] In the description of the embodiments of the present application, the term "plurality" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0027] The embodiments of the present application provide a method for solid-state drive load testing. In one embodiment, please refer to Figure 1 , Figure 1 which is a schematic flowchart of the solid-state drive load testing method provided by the embodiments of the present application. This method can be implemented relying on a computer program, which can be integrated into an application or run as an independent tool-type application. This method can also be implemented relying on a single-chip microcomputer and can run on a solid-state drive load testing system based on the von Neumann architecture. Specifically, this method may include the following steps:
[0028] Step 101: Obtain the IO request feature data of the solid-state drive under multiple application scenarios.
[0029] Among them, the application scenario refers to different working environments and uses of the solid-state drive during actual use. In this embodiment, it may specifically include the personal computer application scenario, where the solid-state drive is mainly used for the operation of the operating system and daily software; the enterprise-level server application scenario, where the solid-state drive mainly bears services such as databases and network storage; and the big data center application scenario, where the solid-state drive needs to handle the storage and calculation tasks of massive data. Under different application scenarios, there are significant differences in the data access requirements and workload characteristics faced by the solid-state drive.
[0030] The IO request feature data refers to the feature information when the solid-state drive performs input and output operations in the above application scenarios. Specifically, it may include the read / write type, indicating whether the data access is a read operation or a write operation; the access address, indicating the physical storage location of the data in the solid-state drive; the request size, indicating the data volume size of a single IO operation; the access timestamp, indicating the specific time point when the IO request occurs; the access mode, indicating whether it is a random access or a sequential access; and the concurrency of the IO requests, indicating the number of simultaneously occurring IO requests. These feature data comprehensively reflect the load characteristics of the solid-state drive during actual operation.
[0031] Specifically, first, obtain the IO request feature data of the solid-state drive under multiple application scenarios. Since there are significant differences in the data access patterns and workload characteristics of solid-state drives under different application scenarios, it is necessary to collect the IO request feature data under multiple typical application scenarios to build a more comprehensive test basis. Specifically, the IO request information during the actual operation of the solid-state drive can be collected by deploying an IO feature collection module in the target application scenario. The IO request information includes, but is not limited to, feature data such as the type (read / write) of the IO request, access address, request size, access timestamp, etc. For example, in the database application scenario, the IO feature collection module can record the random read and write requests generated during the execution of the database; in the video storage application scenario, the sequential read and write characteristics of continuous large chunks of data can be collected; in the virtualization environment, the mixed IO characteristics generated by concurrent access of multiple virtual machines can be obtained. By continuously collecting the IO request feature data in different application scenarios, a relatively complete IO feature dataset can be accumulated. These feature data will provide a basis for subsequent feature classification and redundancy removal processing, contribute to building a more accurate load feature fingerprint library, and further improve the authenticity and reliability of the solid-state drive load test. The richer the collected IO request feature data, the more comprehensively it can reflect the actual workload characteristics of the solid-state drive under various application scenarios, thus providing strong support for generating test loads closer to the actual application scenarios in the future.
[0032] Step 102: Classify the features of the IO request feature data, and perform redundancy removal processing on the various data features after feature classification through a segmented compression algorithm to obtain a load feature fingerprint library.
[0033] Among them, the segmented compression algorithm in this embodiment refers to a data compression method based on a sliding window, such as the existing PAA segmented compression algorithm or PLA segmented compression algorithm, etc., which is used to compress the frequency distribution histogram. This algorithm continuously moves on the frequency distribution histogram by setting a sliding window of a fixed size. When the change amplitude of the frequency data within the window is less than a preset threshold, multiple frequency values within the range of this window are compressed into a representative value. This compression method can effectively reduce data redundancy while retaining the data access characteristics.
[0034] The load feature fingerprint library in this embodiment refers to a database that stores the load characteristics of the solid-state drive under multiple application scenarios. Each fingerprint feature in this fingerprint library corresponds to the IO access characteristics after feature classification and redundancy removal under a specific application scenario, including read / write type information, data access frequency distribution information, and frequency distribution segment information after segmented compression.
[0035] Specifically, to extract the typical characteristics of the solid-state drive load in different application scenarios, it is necessary to perform feature classification and redundancy removal on the obtained IO request feature data. First, the IO request feature data is divided into initial data of random read type, random write type, sequential read type, and sequential write type according to the read-write type. This classification method can effectively distinguish the IO characteristics under different access modes. For the initial data of each type, determine its corresponding data volume, and based on a preset data volume distribution table, determine the distribution intervals corresponding to each sub-data volume. By obtaining the data access frequencies within each distribution interval, a frequency distribution histogram that can intuitively represent the IO access characteristics is generated. Since there may be data redundancy in the original frequency distribution histogram, which affects the extraction efficiency of subsequent load characteristics, a sliding window-based segmented compression algorithm is used to compress the frequency distribution histogram. Specifically, the sliding window moves on the frequency distribution histogram. When the amplitude of the frequency change within the window is less than a preset threshold, multiple frequency values in this interval are compressed into a representative value, thereby obtaining the fingerprint characteristics corresponding to multiple preset scenario frequency distribution segments. Finally, each fingerprint characteristic is stored in the initial fingerprint library to form a load characteristic fingerprint library. Through this feature classification and redundancy removal process, data redundancy can be significantly reduced, and the most representative load characteristics can be extracted, providing more refined and effective feature data support for subsequent generation of data access heat maps and construction of test loads. This processing method not only improves the efficiency of feature extraction but also ensures the representativeness of the extracted features, which helps to generate more accurate test loads subsequently.
[0036] Based on the above embodiments, as an optional embodiment, in step 102: performing feature classification on the IO request feature data, and performing redundancy removal on the various data features after feature classification through a segmented compression algorithm to obtain a load characteristic fingerprint library. This step may further include the following steps:
[0037] Step 201: Divide the IO request feature data into multiple types of initial data according to the read-write type, and the types of the initial data are at least one of random read type, random write type, sequential read type, and sequential write type.
[0038] Specifically, to distinguish the data access characteristics of solid-state drives in different application scenarios, it is necessary to first classify the IO request feature data. In specific implementation, by analyzing the access address continuity and data block size in the IO request feature data, the IO request feature data is divided into different types of initial data. When the access addresses of two adjacent IO requests are discontinuous and the data block size is less than a preset threshold (such as 4KB), it is classified as initial data of the random read type or the random write type; when the access addresses of two adjacent IO requests are continuous and the data block size is greater than the preset threshold, it is classified as initial data of the sequential read type or the sequential write type. This classification method can effectively distinguish the IO access patterns in different application scenarios.
[0039] Step 202: Determine the data volume corresponding to each initial data, and based on a preset data volume distribution table, determine the distribution interval corresponding to each data volume.
[0040] Specifically, to quantitatively analyze the distribution characteristics of each type of initial data, first count the total data volume of each type of initial data. Then, based on the pre-established data volume distribution table, divide the total data volume of each type of initial data. The data volume distribution table presets the distribution intervals corresponding to different data volume ranges. For example, 0 - 4KB can be set as the first interval, 4KB - 16KB as the second interval, and so on. In this way, the IO access characteristics at different data volume levels can be described more precisely.
[0041] Step 203: Obtain the data access frequencies within each distribution interval, and generate a frequency distribution histogram based on each data access frequency.
[0042] Specifically, to visually represent the data access characteristics of each distribution interval, it is necessary to calculate the data access frequency within each distribution interval. In specific implementation, count the number of occurrences of IO requests within each distribution interval, and divide by the total number of IO requests to obtain the data access frequency of this interval. Based on the calculated access frequencies of each interval, generate a frequency distribution histogram, where the horizontal axis represents the data volume distribution interval and the vertical axis represents the corresponding access frequency. This visual representation method can clearly display the IO access patterns at different data volume levels.
[0043] Step 204: Use a sliding window-based segmented compression algorithm to compress the frequency distribution histogram to obtain fingerprint features corresponding to multiple preset scenario frequency distribution segments; store each fingerprint feature in an initial fingerprint library to obtain the loaded feature fingerprint library after storage.
[0044] Specifically, to reduce data redundancy and extract typical features, a segment compression algorithm based on a sliding window is used to compress the frequency distribution histogram. In specific implementation, a sliding window with a fixed size (such as containing 3 adjacent distribution intervals) is set. When the change amplitude of the frequency values within the window is less than a preset threshold (such as 5%), the average value of the frequency values within the window is used to replace the original multiple frequency values, thereby obtaining the compressed frequency distribution segments. These compressed frequency distribution segments constitute the fingerprint features of a specific application scenario. Finally, the obtained fingerprint features are stored in the initial fingerprint database to form a load feature fingerprint database containing the features of multiple application scenarios. In this way, both the key load feature information is retained, and the data redundancy is significantly reduced, providing reliable feature data support for subsequent load testing.
[0045] Step 103: Extract the target fingerprint of the target application scenario from the load feature fingerprint database, and generate a data access heat map based on the target fingerprint.
[0046] Among them, the target application scenario in this embodiment refers to a specific working environment that requires solid-state drive performance testing, and this scenario has specific IO access requirements and load characteristics.
[0047] The target fingerprint in this embodiment refers to a data structure that can represent the IO access characteristics of the target application scenario and is extracted from the load feature fingerprint database.
[0048] The data access heat map in this embodiment refers to a visual representation method used to display the heat distribution of data access in the solid-state drive storage space. This heat map divides the storage space of the solid-state drive into multiple consecutive address blocks, and uses different shades of color to represent the access frequency of each block.
[0049] Specifically, to accurately reflect the data access characteristics of the solid-state drive in the target application scenario, it is necessary to generate a data access heat map based on the load feature fingerprint library. First, retrieve the target fingerprint that matches the target application scenario from the load feature fingerprint library. This target fingerprint contains the frequency distribution characteristics of different types of IO requests in the target application scenario. Based on the obtained target fingerprint, divide the storage space of the solid-state drive into multiple data blocks of a preset size. Each data block corresponds to an access counter, which is used to count the number of times the data block is accessed within a specific time window. By analyzing the access address information and access frequency information in the target fingerprint, update the access counter values of each data block. Specifically, when the access count of a certain data block exceeds the preset hot data threshold, mark it as a hot data area; when the access count is lower than the preset cold data threshold, mark it as a cold data area; and mark it as a warm data area if it is between the two thresholds. Subsequently, based on the access heat of different data areas, use different color depths for visual representation. For example, use red to represent the hot data area and blue to represent the cold data area to generate a two-dimensional data access heat map. The horizontal axis of this heat map represents the address range of the storage space, and the vertical axis represents the time dimension. The color depth intuitively reflects the access frequency of different areas. The data access heat map generated in this way can clearly display the access hot spot distribution of the storage space of the solid-state drive in the target application scenario, providing the spatial and time characteristics of the data distribution for subsequent construction of the test load, and helping to generate a more realistic test data stream.
[0050] Based on the above embodiments, as an alternative embodiment, in step 103: Generating a data access heat map according to the target fingerprint, this step may further include the following steps:
[0051] Step 301: Extract the data access location information and access count information of the target application scenario according to the target fingerprint.
[0052] Specifically, to construct a data access distribution model in the target application scenario, it is necessary to extract key information from the target fingerprint. In specific implementation, parse the data structure in the target fingerprint and extract the physical address corresponding to each IO request as the data access location information. This physical address is in bytes and records the actual storage location of the data in the solid-state drive. At the same time, establish an array of counters, where each element of the array corresponds to a physical address, and is used to count the cumulative number of times the address is accessed as the access count information. During the counting process, whenever an access request to a specific physical address occurs, the value of the corresponding counter is incremented by 1. Through this precise information extraction method, both the spatial location information of data access and the statistical characteristics of access intensity are retained, providing a complete data basis for subsequent heat map generation.
[0053] Step 302: Establish a physical address index table based on the data access location information, and map the access count information to the physical address index table.
[0054] Specifically, to establish the spatial mapping relationship of data access, it is necessary to organize the discrete access location information into a structured index table. In specific implementation, first, the entire physical storage space of the solid-state drive is divided into continuous address blocks at a granularity of 4KB, and each address block is assigned an increasing index number starting from 0. For example, for a 1TB solid-state drive, approximately 268,435,456 address blocks can be divided to establish the mapping relationship between the physical address and the index number. Then, create a two-dimensional array as the physical address index table, where the row number represents the number of the time window, and the column number represents the index number of the address block. Based on the access count information statistically obtained in Step 301, calculate the corresponding address block index number according to the physical address of each access request, and accumulate the access count to the corresponding position in the index table. This mapping method based on address blocks not only reduces the discreteness of data but also maintains the spatial locality characteristic.
[0055] Step 303: Normalize the access count information in the physical address index table to obtain an access frequency matrix, and divide the access frequencies of the access frequency matrix into multiple hierarchical intervals.
[0056] Specifically, to eliminate the magnitude difference of access counts in different application scenarios, it is necessary to standardize the access count information. In specific implementation, first, convert the physical address index table into a two-dimensional matrix of M×N, where M represents the number of divided time windows (such as 1000), and N represents the number of address blocks. Normalize each element in the matrix using the following formula: Normalized value = (Original access count - Minimum access count) / (Maximum access count - Minimum access count), ensuring that all values are mapped to the interval [0, 1]. Then, divide the normalized access frequencies into 5 hierarchical intervals: [0 - 0.2] represents extremely cold data, [0.2 - 0.4] represents cold data, [0.4 - 0.6] represents warm data, [0.6 - 0.8] represents hot data, and [0.8 - 1.0] represents extremely hot data. This fine interval division method can more accurately distinguish data regions with different access intensities.
[0057] Step 304: Determine the preset color identifiers corresponding to each hierarchical interval, and render the access frequency matrix based on each color identifier to generate a data access heat map.
[0058] Specifically, to display the hotspot distribution of data access, it is necessary to convert the access frequency information into a visual effect. In the specific implementation, the RGB color model is used to set color identifiers for each level interval: the extremely cold data interval uses dark blue (RGB: 0, 0, 255), the cold data interval uses light blue (RGB: 0, 127, 255), the warm data interval uses green (RGB: 0, 255, 0), the hot data interval uses orange (RGB: 255, 127, 0), and the extremely hot data interval uses red (RGB: 255, 0, 0). Subsequently, each element in the access frequency matrix is traversed, and according to the level interval to which its value belongs, the corresponding color identifier is filled into the corresponding position of the two-dimensional heat map. In the finally generated heat map, the horizontal axis represents the spatial distribution of the physical address, with the address block index number as the scale; the vertical axis represents the number of the time window, and the change in the depth of the color intuitively shows the access heat of different regions. This multi-level color mapping scheme can clearly display the distribution characteristics of data access in the time and space dimensions.
[0059] Step 104: Based on the data access heat map, divide the benchmark load into cold data streams and hot data streams, and perform read and write operations on the cold data streams and hot data streams to generate a mixed load sequence.
[0060] Among them, the benchmark load in this embodiment refers to the standard data set used for the performance test of the solid-state drive, which includes test data files of a preset size and basic read and write instruction sequences.
[0061] The cold data stream and the hot data stream in this embodiment refer to two types of data streams obtained by separating the benchmark load according to the access frequency characteristics of the data access heat map. Among them, the hot data stream corresponds to the data access sequence in the hot data area, with a high access frequency and a short access interval, and the cold data stream corresponds to the data access sequence in the cold data area, with a low access frequency and a long access interval.
[0062] The mixed load sequence in this embodiment refers to a complete test sequence formed by alternately executing the cold data stream and the hot data stream at a preset time interval through a time slice round-robin scheduling method. This sequence contains information such as the physical address of the data block, the access type, and the access timestamp, and dynamically adjusts the scheduling ratio of the cold data stream and the hot data stream according to the time characteristics of the data access heat map.
[0063] Specifically, to generate a test load closer to the actual application scenario, it is necessary to construct cold data streams and hot data streams based on the data access heat map and generate a mixed load sequence. When specifically implementing, first, according to the color depth in the data access heat map, the physical address space is divided into a cold data area and a hot data area. Among them, the area with a color depth greater than the preset hot data threshold (such as 0.7) is marked as the hot data area, and the area with a color depth less than the preset cold data threshold (such as 0.3) is marked as the cold data area. Then, test data is extracted from the benchmark test load, and according to the address ranges of the hot data area and the cold data area, the test data is written into the corresponding data streams respectively to form a hot data stream and a cold data stream. For the hot data stream, a higher access priority and a shorter access time interval (such as 10 ms) are set; for the cold data stream, a lower access priority and a longer access time interval (such as 100 ms) are set. When performing the test, a time slice rotation scheduling method is adopted. Within each preset time slice, the hot data stream is preferentially scheduled to perform read and write operations, and the remaining time is used to perform the read and write operations of the cold data stream. At the same time, based on the time dimension information in the data access heat map, the read and write ratios and access intervals of the cold data stream and the hot data stream are dynamically adjusted. For example, the scheduling weight of the hot data stream is increased during the access intensive period, and the scheduling weight of the cold data stream is increased during the access sparse period. The mixed load sequence generated in this way not only maintains the cold and hot distribution characteristics of data access but also reflects the load change characteristics in the time dimension, and can more realistically simulate the data access mode in the target application scenario. This load generation method based on cold data streams and hot data streams significantly improves the simulation accuracy of the test load for the actual application scenario and provides a more reliable test basis for the performance evaluation of solid-state drives.
[0064] Based on the above embodiments, as an alternative embodiment, in step 104: Based on the data access heat map, dividing the benchmark test load into a cold data stream and a hot data stream, this step may further include the following steps:
[0065] Step 401: Scan the data access heat map to obtain a heat distribution curve; according to the peak and valley information of the heat distribution curve, determine the location information of the hot spot area and the location information of the cold spot area.
[0066] Specifically, to accurately identify the distribution characteristics of hot and cold regions in the data access heat map, quantitative analysis of the heat map is required. In specific implementation, first scan along the spatial dimension of the heat map, and count the access heat values of different physical addresses within each time window to generate a heat distribution curve. The abscissa of this curve represents the physical address range, and the ordinate represents the access heat value. Then, perform waveform analysis on the heat distribution curve, identify the peak positions as hot regions by setting a heat threshold (such as 0.7), and identify the trough positions as cold regions by setting a coldness threshold (such as 0.3). Further record the physical address ranges corresponding to these hot and cold regions to form a position information table containing the start address and end address. This waveform analysis-based method can accurately locate the boundaries of hot and cold regions of data access, providing a basis for spatial division for subsequent data stream generation.
[0067] Step 402: Based on the position information of the hot region, extract the corresponding first data access pattern and first load characteristics to generate a hot data stream template.
[0068] Specifically, to extract the access characteristics of the hot region, in-depth analysis of the hot region is required. In specific implementation, based on the obtained position information of the hot region, extract the data access behavior characteristics within this region. First, analyze the access pattern, including the sequential access ratio, random access ratio, read / write ratio, etc. These statistical information constitute the first data access pattern; then extract the load characteristics, including parameters such as the average access interval (such as 10 ms), burst access length (such as 64 KB), access alignment method (such as 4 KB alignment), etc. These parameters constitute the first load characteristics. Organize the extracted access pattern and load characteristics into a hot data stream template, which defines the typical access behavior and load parameters of the hot data region, providing a behavior pattern reference for subsequent load generation.
[0069] Step 403: Based on the position information of the cold region, extract the corresponding second data access pattern and second load characteristics to generate a cold data stream template.
[0070] Specifically, to characterize the access characteristics of the cold region, an analysis method similar to that of the hot region is adopted. In specific implementation, based on the position information of the cold region, extract the access behavior characteristics within this region. First, count the second data access pattern, including characteristics such as access type distribution and access regularity; then extract the second load characteristics, including parameters such as a longer average access interval (such as 100 ms) and a smaller burst access length (such as 4 KB). These characteristic information are organized into a cold data stream template to guide the access behavior generation of the cold data region. In this way, a complete description of the access behavior of the cold region is established, providing a pattern basis for subsequent load division.
[0071] Step 404: Divide the benchmark load into corresponding cold data streams and hot data streams according to the cold data stream template and the hot data stream template.
[0072] Specifically, first scan the data blocks in the benchmark load, and determine whether they belong to the hot area or the cold area according to the physical addresses of the data blocks. For the data blocks falling in the hot area, set the corresponding access parameters, such as a higher access priority, a shorter access interval, etc., according to the first data access mode and the first load characteristic in the hot data stream template, and allocate these data blocks to the hot data stream; for the data blocks falling in the cold area, set the corresponding access parameters, such as a lower access priority, a longer access interval, etc., according to the second data access mode and the second load characteristic in the cold data stream template, and allocate these data blocks to the cold data stream. Through this template-based division method, the conversion of the benchmark load into cold data streams and hot data streams that conform to the actual access characteristics is realized, laying a foundation for generating a real mixed load sequence.
[0073] Based on the above embodiments, as an alternative embodiment, in step 104: performing read and write operations on the cold data stream and the hot data stream to generate a mixed load sequence, this step may further include the following steps:
[0074] Step 405: Determine the switching period of the cold data stream and the hot data stream and the execution time ratio within a single period; according to the execution time ratio, construct a time slice allocation template, and allocate corresponding execution time slices to the cold data stream and the hot data stream according to the time slice allocation template.
[0075] Specifically, to achieve reasonable scheduling of the cold data stream and the hot data stream, a time slice allocation mechanism needs to be established. When specifically implemented, first, based on the time characteristics of the data access heat map, set the switching period of the cold data stream and the hot data stream, for example, set each period to 1000 ms. Then analyze the time distribution of the access intensity in the heat map to determine the execution time ratio of the cold data stream and the hot data stream within a single period. For example, set the proportion of the hot data stream to 80% (800 ms) and the proportion of the cold data stream to 20% (200 ms) during the access peak period; adjust the proportion of the hot data stream to 40% (400 ms) and the proportion of the cold data stream to 60% (600 ms) during the access trough period. Based on these time allocation parameters, construct a time slice allocation template, which defines the specific execution time slice sizes of the cold data stream and the hot data stream in different time periods. Through this dynamic time slice allocation mechanism, not only the timely processing of hot data is ensured, but also the excessive delay of cold data is avoided, realizing the balanced scheduling of the cold data stream and the hot data stream.
[0076] Based on the above embodiments, as an alternative embodiment, in step 405: allocating corresponding execution time slices to the cold data stream and the hot data stream according to the time slice allocation template, this step may further include the following steps:
[0077] Step 415: Obtain the cycle duration and time slice size in the time slice allocation template; divide the execution time of the cold data stream and the hot data stream into multiple execution cycles according to the cycle duration.
[0078] Specifically, to achieve the periodic scheduling of the cold data stream and the hot data stream, time division needs to be performed based on the time slice allocation template. In specific implementation, first extract the cycle duration parameter (such as 1000 ms) and the time slice size parameter (such as 800 ms for the hot data stream and 200 ms for the cold data stream) from the time slice allocation template. Then, divide the total execution time of the cold data stream and the hot data stream equally according to the cycle duration. For example, divide the total execution time of 10000 ms into 10 execution cycles, and the duration of each execution cycle is 1000 ms. This time division method based on a fixed cycle provides a basic time framework for the scheduling of the cold data stream and the hot data stream, ensuring the regularity and controllability of the execution process.
[0079] Step 425: Determine the start time point and end time point of the cold data stream and the hot data stream within a single execution cycle based on the time slice size of each execution cycle.
[0080] Specifically, to precisely control the switching timing of the cold data stream and the hot data stream within each execution cycle, specific time nodes need to be calculated. In specific implementation, for each execution cycle, calculate the start time point and end time point of the cold data stream and the hot data stream based on the time slice size parameter. For example, within a certain execution cycle, set the start time point of the hot data stream as the start moment of the cycle, and the end time point as the start time point plus the hot data time slice size (800 ms); set the start time point of the cold data stream as the end time point of the hot data stream, and the end time point as the start time point plus the cold data time slice size (200 ms). Through this precise calculation of time points, strict time boundary division of the cold data stream and the hot data stream within each cycle is achieved.
[0081] Step 435: Map the start time point and end time point to the execution time axis to generate the execution time slice sequence of the cold data stream and the hot data stream.
[0082] Specifically, to form a complete execution time plan, it is necessary to map the calculated time points onto the execution time axis. In specific implementation, a unified execution time axis is created, and the time scale of this axis is marked in milliseconds. Then, the calculated start time point and end time point within each execution cycle are mapped onto the time axis in sequence, forming an execution time slice sequence that contains all the time slice information. This sequence clearly marks the switching moments of the cold data stream and the hot data stream during the entire execution process, providing a complete scheduling basis for subsequent time slice allocation.
[0083] Step 445: Allocate corresponding execution time slices to the cold data stream and the hot data stream according to the execution time slice sequence.
[0084] Specifically, to allocate the execution time slices to the cold data stream and the hot data stream, specific allocation needs to be carried out based on the execution time slice sequence. In specific implementation, each time slice information in the execution time slice sequence is traversed, and according to the start time point and end time point of the time slice, the corresponding time period is allocated to the cold data stream or the hot data stream. For example, when reaching a certain time point, the data stream type to be executed currently is determined by querying the execution time slice sequence, and the corresponding time slice is allocated to this data stream. This sequence-based allocation method realizes the precise scheduling of the cold data stream and the hot data stream, ensuring that each data stream can obtain an execution opportunity within the specified time slice, thereby guaranteeing the balance and regularity of the load execution.
[0085] Step 406: Within each execution time slice, perform read and write operations according to the IO request characteristics of the corresponding cold data stream and hot data stream.
[0086] Specifically, to ensure that the cold data stream and the hot data stream are executed according to the actual access characteristics, it is necessary to precisely control the IO operations within the allocated time slice. In specific implementation, within the execution time slice of the hot data stream, according to its IO request characteristics (such as a data block size of 4KB, an access interval of 10ms, a read operation ratio of 70%, etc.), read and write operations are performed according to the preset access mode. Similarly, within the execution time slice of the cold data stream, according to its IO request characteristics (such as a data block size of 16KB, an access interval of 100ms, a read operation ratio of 30%, etc.), corresponding read and write operations are performed. During the execution process, strictly follow the access parameter settings of each data stream, including access granularity, access interval, read and write ratio, etc., to ensure that the generated IO request sequence conforms to the actual access behavior. This precise control method based on characteristics can truly restore the data access behavior of the target application scenario.
[0087] Step 407: Obtain the IO request information for the read and write operations within each execution time slice, and merge the IO request information in sequence to generate a mixed load sequence that includes the alternating execution of read and write operations by the cold data stream and the hot data stream.
[0088] Specifically, to generate a complete hybrid load sequence, it is necessary to perform a merging process on the execution results. In specific implementation, first, record the IO request information generated within each execution time slice, including attributes such as the request initiation timestamp, physical address, operation type (read / write), data size, etc. Then, sort and merge these discrete IO request information in timestamp order to form a continuous request sequence. During the merging process, keep the original attributes of each IO request unchanged and only reorganize them in chronological order. The finally generated hybrid load sequence contains all the IO request information of the alternating execution of cold data streams and hot data streams. These requests are continuously distributed in time and follow the cold and hot distribution characteristics in space, and can accurately simulate the actual load situation of the target application scenario. Through this time-sequential merging method, both the access characteristics of cold data streams and hot data streams are maintained, and the continuity and integrity of the load sequence are ensured, providing real and reliable test data for subsequent performance tests.
[0089] Step 105: Parallelly load the hybrid load sequence into multiple physical regions of the solid state drive to be tested and perform a performance test to obtain the test results of the solid state drive to be tested.
[0090] Among them, the physical region in this embodiment refers to the continuous physical address range in the storage space of the solid state drive to be tested, and each physical region is an independent data storage and access unit.
[0091] The test results in this embodiment refer to the set of performance metrics obtained by parallelly loading the hybrid load sequence.
[0092] The performance test in this embodiment refers to the process of loading the hybrid load sequence into multiple physical regions of the solid state drive to be tested and performing IO operations in a multi-threaded parallel manner.
[0093] Specifically, to comprehensively evaluate the performance of the solid-state drive (SSD) under test in an actual application scenario, it is necessary to conduct a parallel loading test on the generated mixed workload sequence. In specific implementation, first, the storage space of the SSD under test is divided into multiple physical regions, and the size of each physical region is set according to the data volume of the mixed workload sequence. For example, a 256GB storage space is evenly divided into 4 physical regions of 64GB each. Then, multiple execution threads are created, and each thread is responsible for loading the mixed workload sequence into a specified physical region and performing I / O operations. During the execution process, each thread strictly executes read and write operations in the time sequence and access parameters defined in the mixed workload sequence, and simultaneously records performance metrics such as the response time and throughput of each I / O request. For the hot data stream part, the focus is on random read and write performance and access latency; for the cold data stream part, the focus is on sequential read and write performance and bandwidth utilization. Through the parallel loading test, the performance data collected includes key metrics such as the average response time, IOPS (number of I / O operations per second), read and write bandwidth, and queue depth. This parallel test method not only improves the test efficiency but also can truly simulate the performance of the SSD under a multi-application concurrent access scenario, providing reliable test data for evaluating the performance characteristics of the SSD in an actual application environment. The final test results comprehensively reflect the performance level of the SSD under test when processing complex workloads and can be used as an important reference for product performance optimization and application scenario matching.
[0094] Based on the above embodiments, as an alternative embodiment, in step 105: Parallelly loading the mixed workload sequence into multiple physical regions of the SSD under test and performing a performance test to obtain the test results of the SSD under test, this step may further include the following steps:
[0095] Step 501: Determine the physical parallelism parameter based on the number of chips and channels of the SSD under test.
[0096] Specifically, to make full use of the hardware resources of the SSD for parallel testing, it is necessary to first determine the physical parallelism parameter. In specific implementation, obtain the hardware configuration information of the SSD under test, including the number of NAND flash chips (such as 16) and the number of physical channels supported by the controller (such as 8). Based on these hardware parameters, calculate the maximum physical parallelism parameter, which represents the ability of the SSD to process I / O requests simultaneously. For example, when each physical channel is connected to 2 NAND flash chips, the physical parallelism parameter can be set to 8, indicating that up to 8 concurrent I / O requests are supported. This parallelism setting based on the hardware configuration ensures that the test load can make full use of the hardware resources of the SSD.
[0097] Step 502: Distribute the mixed workload sequence to multiple physical channels of the SSD under test according to the physical parallelism parameter.
[0098] Specifically, to achieve the balanced distribution of the hybrid load sequence, load allocation needs to be performed according to the physical parallelism. In specific implementation, first, the hybrid load sequence is divided into multiple subsequences according to the physical address range, and the data volume of each subsequence is equivalent. Then, according to the physical parallelism parameter (such as 8), these subsequences are allocated to each physical channel of the solid-state drive to ensure that each physical channel is allocated an appropriate amount of load. During the allocation process, the bandwidth limit and processing capacity of each physical channel need to be considered to avoid uneven load distribution. Through this load distribution method, the reasonable allocation of the test load on physical resources is achieved, laying a foundation for concurrent testing.
[0099] Step 503: Through the load generators deployed on each physical channel, perform the test concurrently and obtain the load execution data of each physical channel.
[0100] Specifically, to perform the concurrent test, load generators need to be deployed on each physical channel. In specific implementation, an independent load generator instance is created for each physical channel, and these load generators can generate corresponding IO requests according to the allocated subsequences. Each load generator strictly follows the access pattern and timing requirements defined in the hybrid load sequence and performs read and write operations simultaneously. During the test, each load generator works concurrently and records the execution status of each IO request in real time, including information such as the request initiation time, completion time, data size, operation type, etc. This information constitutes the load execution data of the physical channel. This concurrent test mechanism gives full play to the parallel processing ability of the solid-state drive and improves the test efficiency.
[0101] Step 504: Determine the completion time and execution status of each IO request from the load execution data of each physical channel; according to the completion time and execution status, calculate the average response latency and maximum throughput of the solid-state drive to be tested in the target application scenario as the corresponding test results.
[0102] Specifically, to evaluate the actual performance of a solid-state drive, statistical analysis of the test data is required. In specific implementation, first, load execution data is collected from each physical channel, and the completion time of each I / O request (such as the request end time minus the start time) and the execution status (such as successful completion or failure) are extracted. Then, based on these raw data, performance metrics are calculated: by statistically analyzing the completion times of all successfully completed I / O requests, the average response latency (such as 50 microseconds) is obtained; by calculating the amount of data transferred successfully per unit time, the maximum throughput (such as 500 MB / s) is obtained. These performance metrics together constitute the test results of the solid-state drive to be tested in the target application scenario, comprehensively reflecting the ability of the solid-state drive to handle complex loads. Through this performance analysis method, not only quantifiable performance metrics are obtained, but also potential performance bottlenecks can be discovered, providing a basis for product optimization.
[0103] Referring to Figure 2 , a solid-state drive load test system provided by an embodiment of the present application, the system includes: a data acquisition module, a heat map determination module, a load sequence determination module, and a hard disk test module, where:
[0104] The data acquisition module is used to acquire I / O request feature data of the solid-state drive under multiple application scenarios;
[0105] The heat map determination module is used to perform feature classification on the I / O request feature data, and perform redundancy removal processing on various data features after feature classification through a segmented compression algorithm to obtain a load feature fingerprint library; extract the target fingerprint of the target application scenario from the load feature fingerprint library, and generate a data access heat map according to the target fingerprint;
[0106] The load sequence determination module is used to divide the benchmark test load into a cold data stream and a hot data stream based on the data access heat map, and perform read and write operations on the cold data stream and the hot data stream to generate a mixed load sequence;
[0107] The hard disk test module is used to parallelly load the mixed load sequence into multiple physical regions of the solid-state drive to be tested, and perform a performance test to obtain the test results of the solid-state drive to be tested.
[0108] On the basis of the above embodiments, the heat map determination module is further configured to classify the IO request feature data into multiple types of initial data according to the read / write type, and the types of the initial data are at least one of a random read type, a random write type, a sequential read type, and a sequential write type; determine the data volume corresponding to each initial data, and based on a preset data volume distribution table, determine the distribution interval corresponding to each data volume; obtain the data access frequency within each distribution interval, and generate a frequency distribution histogram according to each data access frequency; perform compression processing on the frequency distribution histogram by using a segmented compression algorithm with a sliding window to obtain fingerprint features corresponding to multiple preset scenario frequency distribution segments; store each fingerprint feature in an initial fingerprint library to obtain a loaded feature fingerprint library after storage.
[0109] On the basis of the above embodiments, the heat map determination module is further configured to extract data access location information and access count information of a target application scenario according to a target fingerprint; establish a physical address index table based on the data access location information, and map the access count information to the physical address index table; perform normalization processing on the access count information in the physical address index table to obtain an access frequency matrix, and divide the access frequencies of the access frequency matrix into multiple level intervals; determine preset color identifiers corresponding to each level interval, and render the access frequency matrix based on each color identifier to generate a data access heat map.
[0110] On the basis of the above embodiments, the load sequence determination module is further configured to scan the data access heat map to obtain a heat distribution curve; determine hot spot area location information and cold spot area location information according to the peak and valley information of the heat distribution curve; based on the hot spot area location information, extract corresponding first data access patterns and first load characteristics to generate a hot data flow template; based on the cold spot area location information, extract corresponding second data access patterns and second load characteristics to generate a cold data flow template; divide the benchmark test load into corresponding cold data flows and hot data flows according to the cold data flow template and the hot data flow template.
[0111] On the basis of the above embodiments, the load sequence determination module is further configured to determine the switching period of the cold data flow and the hot data flow and the execution time ratio within a single period; construct a time slice allocation template according to the execution time ratio, and allocate corresponding execution time slices to the cold data flow and the hot data flow according to the time slice allocation template; within each execution time slice, perform read / write operations according to the IO request characteristics of the corresponding cold data flow and hot data flow; obtain the IO request information for performing read / write operations within each execution time slice, and merge the IO request information in sequence to generate a mixed load sequence including alternating read / write operations of the cold data flow and the hot data flow.
[0112] Based on the above embodiments, the load sequence determination module is further configured to obtain the cycle duration and time slice size in the time slice allocation template; divide the execution times of the cold data stream and the hot data stream into multiple execution cycles according to the cycle duration; determine the start time point and end time point of the cold data stream and the hot data stream within a single execution cycle based on the time slice size of each execution cycle; map the start time point and end time point to the execution time axis to generate the execution time slice sequences of the cold data stream and the hot data stream; and allocate corresponding execution time slices to the cold data stream and the hot data stream according to the execution time slice sequences.
[0113] Based on the above embodiments, the hard disk test module is further configured to determine a physical parallelism parameter based on the number of chips and channels of the solid state hard disk to be tested; distribute the mixed load sequence to multiple physical channels of the solid state hard disk to be tested according to the physical parallelism parameter; concurrently execute tests through the load generators deployed on each physical channel, and obtain the load execution data of each physical channel; determine the completion time and execution status of each IO request from the load execution data of each physical channel; and calculate the average response latency and maximum throughput of the solid state hard disk to be tested in the target application scenario as the corresponding test results according to the completion time and execution status.
[0114] It should be noted that when the device provided in the above embodiments implements its functions, only the division of the above function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0115] This application also discloses an electronic device. Refer to Figure 3 , Figure 3 is a schematic structural diagram of an electronic device disclosed in an embodiment of this application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0116] Among them, the communication bus 302 is used to implement connection communication between these components.
[0117] Among them, the user interface 303 may include a display (Display) interface and a camera (Camera) interface. Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.
[0118] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0119] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines, and executes various functions of the server and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling the data stored in the memory 305. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface graphics, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.
[0120] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. Refer to Figure 3 , the memory 305, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a solid-state drive load test method.
[0121] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 301 can be used to call the application program for a solid-state drive load test method stored in the memory 305. When executed by one or more processors 301, the electronic device 300 is caused to execute the method in one or more of the above embodiments. It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0122] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0123] In several implementation manners provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some service interfaces. The indirect couplings or communication connections of the devices or units can be in electrical or other forms.
[0124] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0125] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0126] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.
[0127] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will readily think of other implementation manners of the present disclosure after considering the specification and the practice of the disclosure.
[0128] The present application aims to cover any variations, uses, or adaptive changes of the present disclosure. These variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary.
Claims
1. A solid state hard disk load testing method, characterized in that: include: Obtain IO request characteristic data of solid-state drives in multiple application scenarios; The IO request feature data is feature classified, and redundancy processing is performed on various data features after feature classification through a segmented compression algorithm to obtain a load feature fingerprint library; Extracting a target fingerprint of a target application scenario from the load feature fingerprint library, and generating a data access heat map according to the target fingerprint; Based on the data access heat map, the benchmark test load is divided into a cold data stream and a hot data stream, and read and write operations are performed on the cold data stream and the hot data stream to generate a mixed load sequence; Loading the mixed load sequence in parallel to multiple physical areas of the solid state drive to be tested, and performing a performance test to obtain a test result of the solid state drive to be tested; The performing read and write operations on the cold data stream and the hot data stream to generate a mixed load sequence includes: Determine a switching cycle of the cold data stream and the hot data stream and a ratio of execution time in a single cycle; According to the execution time ratio, a time slice allocation template is constructed, and corresponding execution time slices are allocated to the cold data stream and the hot data stream according to the time slice allocation template; In each of the execution time slices, read and write operations are performed according to the IO request characteristics of the corresponding cold data stream and hot data stream; Acquire IO request information for performing read and write operations in each execution time slice, and merge each IO request information in time sequence to generate a mixed load sequence including the cold data stream and the hot data stream alternately performing read and write operations; The allocating corresponding execution time slices to the cold data stream and the hot data stream according to the time slice allocation template includes: Obtaining the cycle duration and time slice size in the time slice allocation template; According to the cycle length, dividing the execution time of the cold data stream and the hot data stream into a plurality of execution cycles; Determine the start time point and the end time point of the cold data stream and the hot data stream in a single execution cycle based on the time slice size of each execution cycle; Mapping the start time point and the end time point to an execution time axis to generate an execution time slice sequence of the cold data stream and the hot data stream; Corresponding execution time slices are allocated to the cold data stream and the hot data stream according to the execution time slice sequence.
2. The solid state drive load testing method according to claim 1, characterized in that: The IO request feature data is feature classified, and various data features after feature classification are de-redundanted by using a segmented compression algorithm to obtain a load feature fingerprint library, including: Dividing the IO request characteristic data into multiple types of initial data according to the read and write types, wherein the type of the initial data is at least one of a random read type, a random write type, a sequential read type, and a sequential write type; Determine the data volume corresponding to each of the initial data, and determine the distribution interval corresponding to each of the data volume based on a preset data volume distribution table; Acquire the data access frequency within each of the distribution intervals, and generate a frequency distribution histogram according to each of the data access frequencies; The frequency distribution histogram is compressed using a sliding window segmentation compression algorithm to obtain fingerprint features corresponding to a plurality of preset scene frequency distribution segments; Each of the fingerprint features is stored in an initial fingerprint library to obtain a stored load feature fingerprint library.
3. The solid state drive load testing method according to claim 1, characterized in that: Generating a data access heat map according to the target fingerprint includes: Extracting data access location information and access frequency information of the target application scenario according to the target fingerprint; Establishing a physical address index table based on the data access location information, and mapping the access count information to the physical address index table; Normalizing the access count information in the physical address index table to obtain an access frequency matrix, and dividing the access frequency of the access frequency matrix into multiple level intervals; The preset color identifier corresponding to each of the level intervals is determined, and the access frequency matrix is rendered based on each of the color identifiers to generate a data access heat map.
4. The solid state drive load testing method according to claim 1, characterized in that: The step of dividing the benchmark test load into cold data streams and hot data streams based on the data access heat map includes: Scanning the data access heat map to obtain a heat distribution curve; Determine the hot spot area location information and the cold spot area location information according to the peak and trough information of the heat distribution curve; Based on the hot spot area location information, extract the corresponding first data access mode and the first load feature, and generate a hot data flow template; Based on the cold spot area location information, extract the corresponding second data access mode and second load characteristics to generate a cold data flow template; According to the cold data flow template and the hot data flow template, the benchmark test load is divided into corresponding cold data flows and hot data flows.
5. The solid state drive load testing method according to claim 1, characterized in that: The step of loading the mixed load sequence in parallel to multiple physical areas of the solid state drive to be tested, and performing a performance test to obtain a test result of the solid state drive to be tested includes: Determining a physical parallelism parameter based on the number of chips and the number of channels of the solid state drive to be tested; Distributing the mixed load sequence to multiple physical channels of the solid state drive to be tested according to the physical parallelism parameter; Concurrently executing the test through the load generators deployed on each of the physical channels, and obtaining the load execution data of each of the physical channels; Determining the completion time and execution status of each IO request from each of the load execution data; According to the completion time and the execution status, the average response delay and the maximum throughput of the solid state drive to be tested in the target application scenario are calculated as corresponding test results.
6. A solid state hard disk load testing system, characterized in that: The system comprises: A data acquisition module is used to obtain IO request characteristic data of the solid-state drive in multiple application scenarios; The heat map determination module is used to perform feature classification on the IO request feature data, and perform redundancy processing on various data features after feature classification through a segmented compression algorithm to obtain a load feature fingerprint library; extract a target fingerprint of a target application scenario from the load feature fingerprint library, and generate a data access heat map based on the target fingerprint; A load sequence determination module, configured to divide the benchmark test load into a cold data stream and a hot data stream based on the data access heat map, and perform read and write operations on the cold data stream and the hot data stream to generate a mixed load sequence; A hard disk test module, used for loading the mixed load sequence in parallel to multiple physical areas of the solid state disk to be tested, and performing performance testing to obtain a test result of the solid state disk to be tested; The performing read and write operations on the cold data stream and the hot data stream to generate a mixed load sequence includes: Determine a switching cycle of the cold data stream and the hot data stream and a ratio of execution time in a single cycle; According to the execution time ratio, a time slice allocation template is constructed, and corresponding execution time slices are allocated to the cold data stream and the hot data stream according to the time slice allocation template; In each of the execution time slices, read and write operations are performed according to the IO request characteristics of the corresponding cold data stream and hot data stream; Acquire IO request information for performing read and write operations in each execution time slice, and merge each IO request information in time sequence to generate a mixed load sequence including the cold data stream and the hot data stream alternately performing read and write operations; The allocating corresponding execution time slices to the cold data stream and the hot data stream according to the time slice allocation template includes: Obtaining the cycle duration and time slice size in the time slice allocation template; According to the cycle length, dividing the execution time of the cold data stream and the hot data stream into a plurality of execution cycles; Determine the start time point and the end time point of the cold data stream and the hot data stream in a single execution cycle based on the time slice size of each execution cycle; Mapping the start time point and the end time point to an execution time axis to generate an execution time slice sequence of the cold data stream and the hot data stream; Corresponding execution time slices are allocated to the cold data stream and the hot data stream according to the execution time slice sequence.
7. An electronic device, characterized in that: It includes a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the solid-state hard disk load testing method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the solid-state hard disk load testing method according to any one of claims 1 to 5 is executed.
Citation Information
Patent Citations
Cold and hot data identification test method and device based on solid state disk
CN117234879A
Solid state disk performance test data processing method and device, equipment and storage medium
CN118132403A
Hard disk data protection method and device based on artificial intelligence
CN119646906A