Test case generation method and device, storage medium and electronic equipment

By extracting and classifying features from the storage system's real-time log data and generating targeted test cases, this solves the problem in traditional automated testing methods where test cases cannot perceive the storage system status in real time. This enables real-time monitoring of the storage system and rapid problem identification, improving test efficiency and system reliability.

CN120803938APending Publication Date: 2025-10-17JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510932852.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional automated testing methods are unable to generate targeted test cases to test storage systems, resulting in a disconnect between test case execution and storage status, an inability to perceive the status of the storage system in real time, and a lag.

Method used

By extracting features from the log data generated in real time by the storage system, the first feature vector is generated, and then input into the target model for classification to generate targeted test cases.

Benefits of technology

It achieves real-time monitoring of the storage system's operating status and rapid identification of problems, improves the pertinence and efficiency of testing, and ensures the reliability and stability of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803938A_ABST
    Figure CN120803938A_ABST
Patent Text Reader

Abstract

The invention discloses a test case generation method and device, a storage medium and electronic equipment, and relates to the technical field of automatic testing, and the method comprises the steps: determining a first feature vector through first log data of a storage system, and inputting the first feature vector into a target model, the target model outputs a classification result corresponding to the running state of the storage system based on the first feature vector, the classification result comprises a normal running state and abnormal running states of different abnormal types, and a first test case corresponding to the storage system is generated according to the classification result. That is to say, the classification result corresponding to the running state of the storage system is determined in a targeted manner in combination with the first log data, and then the first test case corresponding to the storage system is generated in a targeted manner according to different classification results. Through the method and the device, the problem that a targeted test case cannot be generated to test the storage system due to the fact that the test case in the related technology is developed on the basis of a set test plan can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic testing, in particular to a test case generation method and device, a storage medium and an electronic device. BACKGROUND

[0002] In addition to the necessary hardware devices, the supporting software is also essential for distributed storage. At present, users basically rely on the supporting software to set functions, query, and read and write business, etc. With the growing complexity and scale of distributed storage software, automated testing has become an important means to ensure software quality and reliability. However, the traditional automated testing method has many limitations, for example: the current automated test cases are converted according to the specific single test cases that have been written, and the logs generated by the distributed storage are not concerned during the automated testing process, which leads to the complete separation of the two, and the real-time perception of various indicators and states of the storage system during the automated testing process is impossible. Therefore, the execution of the test case and the storage state are disconnected, and there is a certain lag in analyzing the relevant logs and taking targeted tests when the storage system has a problem.

[0003] That is, in the related art, the test case is developed based on the existing test plan, and it is impossible to generate a targeted test case to test the storage system.

[0004] Therefore, the problem that the test case in the related art is developed based on the existing test plan and it is impossible to generate a targeted test case to test the storage system has not been effectively solved. SUMMARY

[0005] The present application provides a test case generation method and device, a storage medium and an electronic device to at least solve the problem that the test case in the related art is developed based on the existing test plan and it is impossible to generate a targeted test case to test the storage system.

[0006] The present application provides a test case generation method, comprising: performing feature extraction on first log data generated in real time by a storage system to obtain a first feature vector corresponding to the first log data, wherein the first log data at least includes one of the following: performance indicator data, storage state, operation record and storage error record of the storage system at a current time; inputting the first feature vector into a target model to enable the target model to classify the running state of the storage system at the current time based on the first feature vector, and generate a classification result, wherein the classification result at least includes one of the following: normal running state, different abnormal running states of different abnormal categories; and generating a first test case corresponding to the storage system according to the classification result.

[0007] The application further provides a test case generation device, comprising: an extraction module configured to perform feature extraction on first log data generated in real time by a storage system to obtain a first feature vector corresponding to the first log data, wherein the first log data comprises at least one of the following: performance index data of the storage system at a current time, a storage state, an operation record, and a storage error record; an input module configured to input the first feature vector into a target model, so that the target model classifies a running state of the storage system at the current time based on the first feature vector and generates a classification result, wherein the classification result comprises at least one of the following: a normal running state and an abnormal running state of different abnormal categories; and a generation module configured to generate a first test case corresponding to the storage system according to the classification result.

[0008] The application further provides an electronic device, comprising: a memory configured to store a computer program; and a processor configured to implement the steps of any of the test case generation methods described above when executing the computer program.

[0009] The application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of any of the test case generation methods described above when executed by a processor.

[0010] The application further provides a computer program product, comprising a computer program, and the computer program is configured to implement the steps of any of the test case generation methods described above when executed by a processor.

[0011] According to the application, the first feature vector is determined based on the first log data of the storage system, and then the first feature vector is input into the target model, so that the target model outputs a classification result corresponding to the running state of the storage system based on the first feature vector, and the classification result comprises: a normal running state and an abnormal running state of different abnormal categories, and the first test case corresponding to the storage system is generated according to the classification result. That is, the application determines the classification result corresponding to the running state of the storage system in combination with the first log data, and then generates the first test case corresponding to the storage system according to different classification results. Through the application, the problem that the test case in the related art is developed based on an established test plan and cannot generate a targeted test case to test the storage system can be solved, and then a targeted test case is generated according to the first log data of the storage system to test the storage system. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.

[0013] Figure 1 is a hardware structure block diagram of a computer terminal of a test case generation method according to an embodiment of the present application;

[0014] Figure 2 is a flow chart of a test case generation method according to an embodiment of the present application;

[0015] Figure 3 is an architecture diagram of a method for forming dynamic generation of automatic test cases by analyzing storage real-time logs based on an online support vector machine according to an optional embodiment of the present application;

[0016] Figure 4 is a structure block diagram of a test case generation device according to an embodiment of the present application. DETAILED DESCRIPTION

[0017] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort fall within the protection scope of the present application.

[0018] It should be noted that, in the description of the present application, the terms "comprise", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0019] In order to make those skilled in the art better understand the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments.

[0020] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the test case generation method depends, the specific application environment architecture or specific hardware architecture is described here.

[0021] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure diagram of a computer terminal for a method of generating a test case according to an embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MPU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0022] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method for determining the interactive state in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0023] The transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a computer terminal's communications provider. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0024] The embodiment of the present application provides a method for generating a test case. Figure 2 This is a flow chart of a method for generating a test case according to an embodiment of the present application, which can be applied toFigure 1 in a computer terminal as shown in Figure 2 The flow includes the following steps:

[0025] In step S202, feature extraction is performed on first log data generated in real time by the storage system to obtain a first feature vector corresponding to the first log data, wherein the first log data includes at least one of the following: performance indicator data of the storage system at a current time, a storage state, an operation record, and a storage error record.

[0026] In step S204, the first feature vector is input into a target model, so that the target model classifies the running state of the storage system at the current time based on the first feature vector, and generates a classification result, wherein the classification result includes at least one of the following: a normal running state, and an abnormal running state of different abnormal categories.

[0027] The target model can be an online support vector machine model.

[0028] In step S206, a first test case corresponding to the storage system is generated according to the classification result.

[0029] The test case generation method provided in the present application determines a first feature vector based on first log data of a storage system, and then inputs the first feature vector into a target model, so that the target model outputs a classification result corresponding to the running state of the storage system based on the first feature vector. The classification result includes: a normal running state, and an abnormal running state of different abnormal categories. A first test case corresponding to the storage system is generated according to the classification result. That is, the present application determines a classification result corresponding to the running state of the storage system in combination with the first log data, and then generates a first test case corresponding to the storage system according to different classification results. Through the present application, the problem that test cases in related technologies are developed based on an established test plan and cannot generate targeted test cases for testing the storage system can be solved, and then targeted test cases can be generated according to the first log data of the storage system to test the storage system.

[0030] Optionally, the step S202 of extracting features from the first log data generated in real time by the storage system to obtain a first feature vector corresponding to the first log data comprises: performing statistics on scalar data in the first log data based on a sliding window to obtain a multi-dimensional scalar feature vector corresponding to the first log data, wherein the scalar data comprises performance index data of the storage system; encoding sequence data in the first log data by using a long short-term memory network to obtain a multi-dimensional sequence feature vector corresponding to the first log data, wherein the sequence data at least comprises one of the following: storage state, operation record and storage error record of the storage system; and performing merging processing on the multi-dimensional scalar feature vector and the multi-dimensional sequence feature vector to generate the first feature vector.

[0031] It can be understood that the step of generating the first feature vector in the present application can be:

[0032] Step one, performing statistics on scalar data in the first log data based on a sliding window;

[0033] The sliding window is a technique for extracting local features from time series data. In the present application, the sliding window is used to process scalar data (such as performance indicators such as CPU usage rate, disk I / O rate, network delay, etc.) in real-time logs to capture the change trend of these indicators in a short period of time. For example, assuming that the storage system records the CPU usage rate per second when collecting log data. When using the sliding window technique, for example, the window size is 10 seconds, the average value, maximum value, minimum value and standard deviation of the CPU usage rate in the last 10 seconds can be continuously calculated. These statistics constitute the feature vector of the scalar data, for example, the CPU usage rate average value, the CPU usage rate maximum value, the CPU usage rate minimum value, the CPU usage rate standard deviation, etc. a total of 12-dimensional scalar feature vector, each dimension represents a specific statistical indicator.

[0034] Step two, encoding sequence data in the first log data by using a long short-term memory network:

[0035] Long Short-Term Memory (LSTM) is used to process and predict time series data, which can remember information across time span and solve long-term dependency problem. In this application, LSTM is used to encode sequence data in logs, such as storage state, operation records and storage error records, so as to capture the patterns and correlations of these data over time. For example, for the sequence data of operation records, LSTM can encode the following operation event sequence: file read -> file write -> file delete -> file read; LSTM will encode these event sequences into a fixed-length vector (e.g. 128 dimensions).

[0036] Step three, merging the multi-dimensional scalar feature vector and the multi-dimensional sequence feature vector:

[0037] The scalar feature vector and the sequence feature vector are merged to generate a more comprehensive first feature vector. For example, the 12-dimensional scalar feature vector and the 128-dimensional sequence feature vector are merged to finally form a first feature vector with 140 dimensions. The first feature vector contains performance statistics of the storage system in the last 10 seconds and state encoding of the last series of operations, which can be used as input of the online support vector machine model for real-time monitoring and anomaly detection.

[0038] Through the above feature extraction and merging process, more accurate state classification and anomaly identification can be made based on real-time data of the storage system, so as to dynamically generate and execute more effective automated test cases, ensuring the comprehensiveness and accuracy of the test.

[0039] Optionally, before inputting the first feature vector into the target model in the above step S204, the method further includes: obtaining target log data of the storage system in a target time period, and performing feature extraction on the target log data to obtain a second feature vector corresponding to the target log data, wherein the target time period is a time period before the current time; inputting the second feature vector into a kernel function to re-characterize the second feature vector in a high-dimensional feature space through the kernel function, to obtain a re-characterized second feature vector; inputting the re-characterized second feature vector into a target algorithm corresponding to the target model, so that the target algorithm calculates first model parameters corresponding to the target model based on the re-characterized second feature vector; modifying the model parameters corresponding to the target model to the first model parameters to train the target model.

[0040] It can be understood that the target model needs to be trained, and the specific training process is as follows:

[0041] Collecting log data of the storage system in a target time period: During the operation of the storage system, logs are continuously generated, recording its performance indicators (such as CPU usage, I / O operation rate), storage status (online, offline, warning, etc.), operation records, and error information. In order to train the target model, log data in the target time period can be collected. For example, we can focus on the log data generated every 10 seconds in the past 1 hour (i.e., the first log data) of the storage system.

[0042] Feature extraction generates a second feature vector: From the collected 1-hour log data, use the sliding window technique (such as a 10-second window size) to statistically analyze performance indicator data (such as CPU usage, disk I / O rate, etc.), forming a multi-dimensional scalar feature vector reflecting the system state. At the same time, sequence data such as operation records and storage status are encoded by Long Short-Term Memory Network (LSTM) to generate a fixed-length multi-dimensional sequence feature vector. The scalar feature vector and the sequence feature vector are combined to obtain the second feature vector. For example, the second feature vector may include: the average, maximum, minimum, and standard deviation of CPU usage (assuming a dimension of 4), statistical features of disk I / O rate (assuming a dimension of 4), and 128-dimensional LSTM encoded sequence features, which comprehensively reflect the patterns of storage status, operation records, and error information, forming a total of 136-dimensional second feature vector.

[0043] High-dimensional space mapping of the second feature vector by kernel function: The second feature vector is input into the kernel function, which (for example: radial basis function kernel) maps the low-dimensional feature vector to a high-dimensional space to capture more complex data patterns. In high-dimensional space, data that was originally non-linearly separable may become linearly separable, helping to improve the classification performance of the target model.

[0044] Optimizing model parameters: The feature vector mapped by the kernel function is sent to the training algorithm of the target model - online support vector machine. For example, using the LASVM algorithm, dynamically adjust the model parameters based on the feature vector in the high-dimensional space to optimize the decision boundary.

[0045] Training the target model: After the model parameter optimization is completed, the target model (online support vector machine) will use these updated parameters for training, so that it can more accurately identify patterns in real-time logs based on historical data, especially abnormal patterns. For example, if it is found in historical data that disk read errors are related to high CPU usage, the target model will be able to learn this association and identify similar patterns in real-time analysis, and then dynamically generate targeted test cases.

[0046] Through the above technical solutions, real-time and historical log data of the storage system can be continuously utilized to optimize the classification ability of the model and improve the coverage and efficiency of automated testing.

[0047] The method for obtaining the target log data can comprise:

[0048] It can be understood that the method for obtaining the target log data can comprise:

[0049] Obtaining a plurality of historical log data of the storage system at a plurality of time nodes within a target time period: Assuming that the running status of the storage system in the past week is analyzed, especially whether an abnormal event occurs during high load. The system log of each minute in this week can be collected, including but not limited to: CPU utilization, disk I / O, network throughput, storage pool state change, system error record, adding a timestamp in each historical log data. For each collected log data, a specific timestamp is marked, for example: 2023-04-10, 15:30:00.

[0050] Adding a target mark to indicate the business flow of each historical log data: further marking the historical log data, for example, marking all operations related to user A with the mark “UserA”.

[0051] Correlating a plurality of historical log data based on the timestamp and the business flow to generate a full-link view: by integrating time and business marks, isolated log records can be linked together to form a complete business processing path. For example, it is observed that during the high load period (such as from 10 am to 2 pm on weekdays), user A performs a large number of file read-write operations, and then the storage pool state changes from “normal” to “warning”, followed by a disk read error. By correlating these events in chronological order and business flow, a full-link view is constructed.

[0052] Determining the target log data for recording abnormal events in the storage system based on the full-link view: based on the full-link view, log data accompanying the storage system entering an abnormal state can be analyzed. For example, log data recorded when the storage system is in the “warning” state, or log data when abnormal conditions such as disk read error, significant increase in network delay, etc. occur.

[0053] The detection of the abnormal event corresponding to the abnormal running state: by monitoring the abnormal event, the running state of the storage system can be located to be an abnormal running state of different abnormal categories, such as performance degradation caused by high load, data read error caused by disk failure, communication delay caused by network fluctuation, etc. The log data corresponding to these abnormal events is the target log data.

[0054] Through the above technical solution, the real-time log of the storage system can be closely combined with the automated test process, not only improving the pertinence and efficiency of the test, but also enhancing the real-time monitoring of the running state of the storage system and the rapid identification ability of the problem.

[0055] Optionally, the step S204 of generating the first test case corresponding to the storage system according to the classification result comprises: in the case that the classification result indicates that the running state of the storage system at the current time is an abnormal running state of any abnormal category, determining a test template having a corresponding relationship with the any abnormal category in a preset rule library, wherein the preset rule library comprises a corresponding relationship between each abnormal category and each test template; obtaining a running parameter corresponding to the current running state of the storage system; supplementing the test template based on the running parameter; and generating the first test case according to the supplemented test template.

[0056] It can be understood that the technical solution of generating the first test case can be:

[0057] When the classification result indicates that the running state of the storage system at the current time is an abnormal running state of any abnormal category: after the online support vector machine model analyzes the log data of the storage system in real time, a classification result is given to indicate whether the running state of the current storage system is normal or belongs to a certain abnormal category. If the model determines that the system is in an abnormal state, it means that the storage system may have some kind of failure or performance bottleneck. For example: based on the latest collected log data (such as abnormal increase of disk I / O rate, significant increase of network delay, etc.), the online support vector machine model determines that the current storage system may be facing the abnormal running state of "high load".

[0058] Identify the test template that corresponds to the anomaly type in the preset rule library: The preset rule library contains a series of predefined rules that establish a connection between the abnormal operating state type and the automated test case template. Once the abnormal state is identified, the corresponding test template is searched in the preset rule library based on this state. For example, suppose the preset rule library contains a rule: "If the storage system is in the 'high load' abnormal state, perform a disk stress test." Then, when the online support vector machine model outputs the classification result of "high load", the automation framework automatically loads the disk stress test template in this rule and prepares to generate the corresponding first test case.

[0059] Obtaining the operating parameters corresponding to the storage system's current operating status: Before generating the first test case, it is necessary to capture specific operating parameters directly related to the current abnormal state from the storage system. For example, if the storage system is detected to be in a "high load" state, the automation framework will capture real-time parameters such as CPU usage, disk I / O rate, and network traffic. These operating parameters will be used to populate subsequent test templates, ensuring that the test environment is as close to the actual abnormal state as possible.

[0060] Supplement test templates based on operating parameters: Fill the selected test template with operating parameters obtained from the storage system in real time to ensure that the test cases can accurately reflect the current abnormal operating environment.

[0061] Generate a first test case based on the supplemented test template: Finally, based on the test template supplemented with specific operating parameters, automatically generate an executable first test case to simulate and detect the behavior and response of the storage system under the current abnormal operating state.

[0062] Through the above technical solution, not only can abnormal operating conditions in the storage system be discovered in a timely manner, but targeted test cases can also be quickly generated to conduct white box testing, and potential problems can be evaluated and repaired in a timely manner, thereby improving the reliability and stability of the system.

[0063] Optionally, after generating the first test case corresponding to the storage system based on the classification result in the above step S206, the method also includes: executing the first test case based on the target device and determining the execution result corresponding to the first test case; generating second log data corresponding to the first test case based on the execution result, and performing feature extraction on the second log data to obtain a third feature vector corresponding to the second log data; and updating the model parameters in the target model according to the third feature vector.

[0064] It is understandable that the online support vector machine model can be optimized based on the execution results of the automated test cases. Specifically:

[0065] Performing the first test case and determining the execution result: In the automation testing framework, the target device is executed with the preset first test case. For example, the first test case can be a test simulating whether the storage system can work stably and the response time is within the expected range under high concurrent read-write operations. After the test is executed, the detailed execution result can be recorded, including: system response time, error code, I / O operation success rate and other key indicators.

[0066] Generating second log data according to the execution result: Based on the execution result of the automation test, the second log data is generated. The second log data not only records the specific execution of the first test case, but also may include: the running state, resource consumption and performance indicators of the target device during the test. For example, if the storage system has a delay increase or disk I / O failure during the execution of the first test case, this information will be recorded in the second log data.

[0067] Feature extraction on the second log data to obtain the third feature vector: The generated second log data is subjected to feature extraction to form a third feature vector. It can include statistical and encoding of key information in the log, such as the average, maximum and minimum values of system response time, the number and success rate of I / O operations, and the frequency of occurrence of any system errors or warnings, etc. For example, use the sliding window technique to count the response time, and use the LSTM network to encode the operation sequence, finally form a multi-dimensional feature vector containing system state and test result.

[0068] Updating the model parameters in the target model: The third feature vector is input into the online support vector machine model, and the model parameters of the online support vector machine are dynamically adjusted according to these new feature vectors. For example, if the storage system delay increase is recorded in the second log data, the online support vector machine model will learn this abnormal state and update its weights and biases through incremental learning mechanism to better identify and predict abnormal behavior in similar future situations.

[0069] Through the above technical solutions, the results of the automation test can be used to optimize the online support vector machine model in real time, improving the accuracy and efficiency of anomaly detection.

[0070] Optionally, after generating the first test case corresponding to the storage system according to the classification result in step S206, the method further comprises: in a case where it is determined that a plurality of second test cases corresponding to the storage system have been generated, executing the plurality of second test cases in one of the following manners: manner one: acquiring a first emergency degree corresponding to each second test case, and determining a first execution order of the plurality of second test cases according to the first emergency degree; executing the plurality of second test cases in sequence according to the first execution order; manner two: constructing a plurality of parallel nodes, and executing the plurality of second test cases in parallel based on the plurality of parallel nodes; manner three: determining whether there are a plurality of third test cases with identical test steps in the plurality of second test cases; performing merging processing on the plurality of third test cases; processing the merging-processed test cases and other test cases except the merging-processed test cases in the plurality of second test cases in parallel; or, determining a second emergency degree of the merging-processed test cases and the other test cases; determining a second execution order of the merging-processed test cases and the other test cases based on the second emergency degree; and executing the merging-processed test cases and the other test cases in sequence based on the second execution order.

[0071] It can be understood that in a case where there are a plurality of test cases, the plurality of second test cases can be executed in the following multiple manners:

[0072] Manner one: determining the execution order according to the emergency degree: the emergency degree determines the priority of the test case, and the test case with a high emergency degree can be executed preferentially. For example, it is assumed that five second test cases are dynamically generated, which correspond to different abnormal events, such as storage pool overload, disk error, network delay, file system conflict and data consistency problem. Each second test case has an emergency degree score associated therewith, for example, the storage pool overload is marked as the highest emergency degree (5 points), and the data consistency problem is slightly lower (3 points). In this way, the test case with the highest score is executed first, that is, the storage pool overload is tested first, followed by the disk error, and so on, until all test cases are executed.

[0073] Manner two: constructing parallel nodes and executing test cases in parallel: parallel execution allows multiple test cases to run simultaneously, shortening the overall test time, and is particularly suitable for independent and mutually independent test cases. For example, in order to speed up the test process, three parallel nodes can be configured. This allows three test cases to run simultaneously, for example, storage pool overload test, disk error test and network delay test can start simultaneously on three different nodes. As long as there is no mutual dependence between these test cases, they can be executed simultaneously without waiting for a test case to complete before proceeding to the next one.

[0074] Way three: merge use cases with the same test steps and sort by urgency: If multiple test cases share the same test steps, merging these test cases can reduce repetitive work and improve efficiency. Furthermore, sorting the merged use cases and other independent use cases by urgency ensures that urgent problems are detected first. For example: suppose there are two dynamically generated use cases that focus on network latency and data consistency problems, respectively, but both of them contain network request sending and receiving steps. The network request parts of these two use cases can be merged into a common test step to form a third test case, reducing repetitive testing. Then, according to the urgency of each remaining use case (for example: storage pool overload is 5 points, disk error is 4 points), the execution order is determined, such as first executing the storage pool overload test, then the disk error test, and then testing the combined use case of network latency and data consistency problems.

[0075] Through the above three ways, not only can the dynamically generated test cases be effectively managed, but also the priority of solving key problems can be ensured, and the efficiency and resource utilization of testing can be improved.

[0076] In order to better understand the process of the above test case generation method, the implementation process of the test case generation method will be described in combination with optional embodiments below, but it is not used to limit the technical solutions of the embodiments of the present application.

[0077] The efficient operation of a distributed storage system not only depends on hardware, but also needs the support of supporting software to realize functions such as setting, data query, and business read-write operations. With the expansion of the complexity and scale of such software, it becomes increasingly important to ensure its high quality and high reliability, and automated testing has become an indispensable part. However, the automated testing method in the related art exposes a series of inherent limitations, which are particularly prominent when facing distributed storage software:

[0078] Static test cases: Traditional automated testing relies on predefined test scripts, which are usually developed for a single function point and lack consideration of the overall state of the system. Once the software is updated or the system state changes, the generation, transformation, and maintenance of test cases will become heavy and costly.

[0079] Fixed test execution order: Use case execution follows a pre-set order. Although this serial execution method is orderly, it ignores the randomness and contingency of user operations, making it difficult to simulate real-world usage scenarios and limiting test coverage.

[0080] Isolation of logs and tests: The log data generated by the storage system, which contains valuable information about the storage state, operation details and their results, is not fully utilized in the automated testing process. The separation of testing and log analysis means that the test cannot provide real-time feedback on the system health, reducing the effectiveness of the test and increasing the delay in discovering and fixing storage problems after they occur.

[0081] Therefore, the traditional automated testing method not only has high cost and is difficult to maintain, but also fails to achieve ideal test coverage due to the neglect of real-time analysis of log data, especially when the system is complex and dynamic.

[0082] To solve the above problems, the optional embodiment of the present application provides a technical scheme of forming dynamic generation of automated test cases based on online support vector machine storage real-time log analysis. The logs generated by the automated testing and real-time analysis of the storage are combined, and executable automated test cases are dynamically generated based on the classification results of the online support vector machine. The purpose is to associate the single function point test with the storage log, to master the storage indicators and states in real time during the automated testing process, to dynamically generate relevant test cases for abnormal indicators and error prompts in the log, to form white box testing, to improve the automated test coverage, to improve the reliability and security of the system, to reduce the labor and test cost investment, and to improve the test efficiency.

[0083] Figure 3 According to the optional embodiment of the present application, a method for forming dynamic generation of automated test cases based on online support vector machine storage real-time log analysis is shown in the architectural diagram as Figure 3

[0084] The optional embodiment of the present application uses the online support vector machine to analyze the logs generated by the storage in real time and combines the automated testing, dynamically generates executable automated test cases based on the classification results of the online support vector machine, and forms white box testing. The purpose is to automatically associate the single function point test with the storage log, to master the storage indicators and states in real time during the testing process, to dynamically generate relevant test cases for abnormal indicators and error prompts in the log, to improve the automated test coverage, to improve the reliability and security of the system, to reduce the labor and test cost investment, and to improve the test efficiency. The optional embodiment of the present application includes two parts: a log collection and processing part and an online support vector machine model online updating part, specifically:

[0085] 1. Log collection and processing part:

[0086] ​The log collection processing part is mainly divided into multi-mode log collection (i.e. log collection) and log stream processing. The multi-mode log collection is used to adapt to various stored logs, such as text logs, kernel logs, etc. Because the optional embodiment of the present application adopts an online support vector machine model based on incremental training, a log stream processing scheme is adopted to solve the problem that the batch processing model cannot use dynamic logs. Specifically:

[0087] (1) Log collection:

[0088] The log collection part can collect text logs and kernel logs of the storage system, and then associate the cross-node logs based on context marking time to realize full-link tracking of storage operations.

[0089] The logs generated by the storage system record various indicators of the storage system (such as load, disk pressure, input / output (IO), etc.), storage operation records (such as adding, deleting, modifying, and querying, etc.), storage status (such as health, alarm, unavailability, storage pool status, etc.), and storage errors (such as software code errors, pointer value null, etc.), so log analysis has great value.

[0090] However, the related technical solution uses converted use cases to perform single function point testing, which is completely isolated from storage log analysis, resulting in the inability to automatically cover test cases for specific exceptions and errors.

[0091] The optional embodiment of the present application introduces an online support vector machine model with incremental learning in the automatic test framework to analyze storage logs in real time, and then dynamically generates targeted test cases according to the output classification results, avoiding the lagging nature of automatic test case updates and the problem of insufficient coverage. At the same time, it can also improve test efficiency, save manpower, and reduce the cost of maintaining test cases.

[0092] The above log collection supports text logs and kernel logs of the storage system, and then associates the cross-node logs based on context marking time to realize full-link tracking of storage operations.

[0093] (2) Log stream processing:

[0094] 1) Preprocessing and feature extraction: The original log data (i.e. first log data) first goes through a preprocessing stage to clean and standardize the data, ensuring the accuracy of subsequent analysis.

[0095] 2) Scalar feature statistics: In order to capture the performance indicator changes of the storage system in a short period of time, a sliding window mechanism is used to calculate scalar features every 10 seconds, such as load, disk pressure, I / O rate, etc. The statistical results in each sliding window are encoded as a dimension, with a total of 12 dimensions, forming a feature subset (i.e. multi-dimensional scalar feature vector) reflecting the instantaneous system state.

[0096] 3) Event sequence encoding: In addition to the scalar features, the logs also contain a series of operation and event records, which may exhibit temporal sequence correlation. To capture this sequence information, the log event sequences are encoded using a long short-term memory network (LSTM), and each event sequence is converted into a fixed-length vector representation with a dimension of 128 (i.e., a multi-dimensional sequence feature vector). LSTM is capable of remembering long-term dependencies and is very suitable for processing log data with temporal dependencies.

[0097] Based on the above steps, a comprehensive feature vector with a total dimension of 134 (i.e., a first feature vector) is formed. The instantaneous indicators and long-term behavior patterns of the storage system state are integrated, providing rich and multi-dimensional data representation for subsequent online support vector machine model training.

[0098] 2. Online support vector machine model online update part:

[0099] Among them, the online support vector is a variant of the traditional support vector machine, designed specifically for online learning scenarios. The online support vector machine dynamically adjusts the decision boundary according to the current model and new data points. This incremental learning approach enables the online support vector machine to efficiently process large-scale data streams while maintaining low computational complexity and memory usage.

[0100] The online support vector machine model online update part mainly includes: model online update and dynamic case generation. The online support vector machine outputs classification results based on the feature vectors formed by the log processing part, and the dynamic case generation generates executable test cases based on the output classification results combined with dynamic parameterized test templates. Specifically:

[0101] (1) Online support vector machine (i.e., target model) supports incremental training:

[0102] The online support vector machine model can use real-time logs as training sample sets during use.

[0103] To solve the nonlinear classification problem, the online support vector machine model selects a kernel function using a radial basis function kernel (Radial Basis Function Kernel, abbreviated as RBF, which is a kernel function) to handle nonlinear features. The RBF kernel function can map non-linearly separable data to a high-dimensional space, thereby achieving linear classification in that space. The application of the RBF kernel function enhances the model's ability to process complex system logs, enabling it to recognize and classify patterns that are not easily distinguishable in low-dimensional space.

[0104] Unlike traditional models that need to be periodically or completely retrained, the online support vector machine model employs a linear approximation sequential minimal optimization for support vector machine (LASVM) algorithm, which allows the model to update its parameters every time it receives 100 new data. This small-batch parameter update strategy not only greatly reduces the time and computational resource requirements of the online support vector machine model training, but also ensures that the online support vector machine model can quickly adapt to new environments and new challenges of the system without going through a time-consuming global retraining process. That is, the online update of the online support vector machine model weight is realized, and the whole model does not need to be retrained, and the change of system behavior and the new emerging mode are quickly adapted.

[0105] (2) Dynamic use case generation:

[0106] A mapping rule library (i.e., a preset rule library) of state / risk patterns to test strategies is established, and this rule library is dynamically extensible. When the online support vector machine identifies a specific system state pattern or detects an abnormal event, the mapping rule library constructed in advance will match the preset or template-based test strategy according to the identified state / abnormal type, then combine the current system context information (e.g., the identified slow disk ID, high-latency network node, etc., i.e., the running parameters) to supplement the instantiated test case template (i.e., the test template) to form the corresponding test case (i.e., the first test case), and finally execute the dynamically generated test case through the device (i.e., the target device) containing the automation framework, and the test result (i.e., the execution result) is captured and sent back to the storage system, which is used as new training data with labels to train the online support vector machine.

[0107] In summary, the optional embodiment of the present application introduces an automatic case dynamic generation module, realizes the association of automatic testing and storage logs, and in the traditional automatic testing process, real-time collection of storage logs and structured data processing of storage logs using a stream processing mode, incremental training of the introduced online support vector machine model, and dynamic generation of relevant use cases for abnormal analysis and error prompts in the log, as a supplement to the traditional single automatic test, to improve the test range.

[0108] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software on a general hardware platform, and of course, can also be realized by hardware, but in many cases, the former is a better implementation. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product in essence or in the form of a part of the prior art. The computer software product is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk) and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device) to execute the method described in each embodiment of the present application.

[0109] In this embodiment, a test case generation device is also provided, which is used to implement the above embodiments and preferred embodiments, and will not be described again. Although the device described in the following embodiments is preferably realized in software, hardware or a combination of software and hardware is also possible and is contemplated.

[0110] Figure 4 is a structural block diagram of the test case generation device according to the embodiments of the present application, as shown in Figure 4 The device includes:

[0111] The extraction module 42 is configured to perform feature extraction on the first log data generated in real time by the storage system to obtain a first feature vector corresponding to the first log data, wherein the first log data at least includes one of the following: performance indicator data of the storage system at a current time, a storage state, an operation record, and a storage error record.

[0112] The input module 44 is configured to input the first feature vector into a target model, so that the target model classifies the running state of the storage system at the current time based on the first feature vector to generate a classification result, wherein the classification result at least includes one of the following: a normal running state, and an abnormal running state of different abnormal categories.

[0113] The generation module 46 is configured to generate a first test case corresponding to the storage system according to the classification result.

[0114] The test case generation apparatus provided in the present application determines a first feature vector based on the first log data of the storage system, and then inputs the first feature vector into the target model, so that the target model outputs a classification result corresponding to the running state of the storage system based on the first feature vector. The classification result includes: a normal running state, and abnormal running states of different abnormal categories. The first test case corresponding to the storage system is generated according to the classification result. That is, the present application determines the classification result corresponding to the running state of the storage system in combination with the first log data, and then generates the first test case corresponding to the storage system according to different classification results. Through the present application, the problem that the test case in the related art is developed based on the established test plan and cannot generate a targeted test case for testing the storage system can be solved, and then the targeted test case is generated according to the first log data of the storage system to test the storage system.

[0115] In one example embodiment, the extraction module 42 is further configured to: perform statistics on scalar data in the first log data based on a sliding window to obtain a multi-dimensional scalar feature vector corresponding to the first log data, wherein the scalar data includes performance index data of the storage system; encode sequence data in the first log data by using a long short-term memory network to obtain a multi-dimensional sequence feature vector corresponding to the first log data, wherein the sequence data at least includes one of the following: a storage state, an operation record, and a storage error record of the storage system; and perform merging processing on the multi-dimensional scalar feature vector and the multi-dimensional sequence feature vector to generate the first feature vector.

[0116] In one example embodiment, the input module 44 is further configured to: obtain target log data of the storage system in a target time period, and perform feature extraction on the target log data to obtain a second feature vector corresponding to the target log data, wherein the target time period is a time period before the current time; input the second feature vector into a kernel function to re-characterize the second feature vector in a high-dimensional feature space by using the kernel function to obtain a re-characterized second feature vector; input the re-characterized second feature vector into a target algorithm corresponding to the target model, so that the target algorithm calculates a first model parameter corresponding to the target model based on the re-characterized second feature vector; and modify the model parameter corresponding to the target model to the first model parameter to train the target model.

[0117] In an example embodiment, the input module 44 is further configured to acquire a plurality of historical log data of the storage system at a plurality of time nodes within the target time period, and add a time stamp in each historical log data; add a target mark in each historical log data, where the target mark is used to indicate a service flow of each historical log data; associate the plurality of historical log data based on the time stamp and the service flow to generate a full-link view corresponding to the storage system; and determine target log data for recording an abnormal event in the storage system based on the full-link view, where the abnormal event is an event executed by the storage system in an abnormal running state different from an abnormal category.

[0118] In an example embodiment, the generation module 46 is further configured to, in a case where the classification result indicates that the running state of the storage system at the current time is an abnormal running state of any abnormal category, determine a test template having a corresponding relationship with the any abnormal category in a preset rule base, where the preset rule base includes a corresponding relationship between each abnormal category and each test template; acquire a running parameter corresponding to a current running state of the storage system; supplement the test template based on the running parameter; and generate the first test case according to the supplemented test template.

[0119] In an example embodiment, the generation module 46 is further configured to execute the first test case based on a target device, and determine an execution result corresponding to the first test case; generate second log data corresponding to the first test case according to the execution result, and perform feature extraction on the second log data to acquire a third feature vector corresponding to the second log data; and update a model parameter in the target model according to the third feature vector.

[0120] In an example embodiment, the generating module 46 is further configured to, in a case where it is determined that the plurality of second test cases corresponding to the storage system have been generated, execute the plurality of second test cases in one of the following manners: manner one: obtaining a first urgency level corresponding to each second test case, and determining a first execution order of the plurality of second test cases according to the first urgency level; and executing the plurality of second test cases in the first execution order; manner two: constructing a plurality of parallel nodes, and executing the plurality of second test cases in parallel based on the plurality of parallel nodes; manner three: determining whether there are a plurality of third test cases with identical test steps in the plurality of second test cases; performing merging processing on the plurality of third test cases; processing the merged test cases and other test cases except the merged test cases in the plurality of second test cases in parallel; or, determining a second urgency level of the merged test cases and the other test cases; determining a second execution order of the merged test cases and the other test cases based on the second urgency level; and executing the merged test cases and the other test cases in the second execution order.

[0121] The features of the embodiments of the test case generation device can be understood with reference to the related descriptions of the embodiments of the test case generation method, which will not be repeated here.

[0122] The embodiments of the present application also provide an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the embodiments of the test case generation method.

[0123] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the embodiments of the test case generation method when running.

[0124] In an example embodiment, the computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0125] The embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to perform the steps in any of the embodiments of the test case generation method.

[0126] The embodiment of the present application further provides another computer program product comprising a nonvolatile computer readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above test case generation method embodiments.

[0127] Those skilled in the art will further appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or both, and that the implementation decisions are within the skill of an informed technician. The functions of the examples have been described in terms of one or more functionalities to facilitate explanation of the present application. The functions can be implemented in hardware or software depending upon the application and design constraints. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0128] The above has introduced in detail a test case generation method provided by the present application. The principles and implementation manners of the present application are described herein by applying specific examples, and the above embodiment description is only for helping to understand the method of the present application and its core idea. It should be noted that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for generating a test case, characterized in that: include: Performing feature extraction on first log data generated in real time by the storage system to obtain a first feature vector corresponding to the first log data, wherein the first log data includes at least one of the following: performance indicator data, storage status, operation records, and storage error records of the storage system at a current time; Inputting the first feature vector into a target model so that the target model classifies the operating state of the storage system at the current time based on the first feature vector and generates a classification result, wherein the classification result includes at least one of the following: a normal operating state and an abnormal operating state of different abnormal types; A first test case corresponding to the storage system is generated according to the classification result.

2. The test case generation method according to claim 1, characterized in that: Performing feature extraction on first log data generated in real time by the storage system to obtain a first feature vector corresponding to the first log data includes: performing statistics on scalar data in the first log data based on a sliding window to obtain a multidimensional scalar feature vector corresponding to the first log data, wherein the scalar data includes: performance indicator data of the storage system; Encoding sequence data in the first log data using a long short-term memory network to obtain a multidimensional sequence feature vector corresponding to the first log data, wherein the sequence data includes at least one of the following: a storage state, an operation record, and a storage error record of the storage system; The multidimensional scalar feature vector and the multidimensional sequence feature vector are merged to generate the first feature vector.

3. The test case generation method according to claim 1, characterized in that: Before inputting the first feature vector into the target model, the method further includes: Obtain target log data of the storage system within a target time period, and perform feature extraction on the target log data to obtain a second feature vector corresponding to the target log data, wherein the target time period is a time period before the current time; Inputting the second eigenvector into a kernel function to re-characterize the second eigenvector in a high-dimensional feature space through the kernel function to obtain a re-characterized second eigenvector; Inputting the re-characterized second eigenvector into a target algorithm corresponding to the target model, so that the target algorithm calculates first model parameters corresponding to the target model based on the re-characterized second eigenvector; The model parameters corresponding to the target model are modified to the first model parameters to train the target model.

4. The method for generating a test case according to claim 3, wherein: Obtaining target log data of the storage system within a target time period includes: Acquire multiple historical log data of multiple time nodes of the storage system within the target time period, and add a timestamp to each historical log data; Adding a target tag to each of the historical log data, wherein the target tag is used to indicate the business flow of each of the historical log data; Associating the plurality of historical log data based on the timestamp and the service flow to generate a full-link view corresponding to the storage system; Target log data for recording abnormal events in the storage system is determined based on the full-link view, wherein the abnormal events are events executed when the storage system is in an abnormal operating state of different abnormal types.

5. The test case generation method according to claim 1, characterized in that: Generating a first test case corresponding to the storage system according to the classification result includes: If the classification result indicates that the operation state of the storage system at the current time is an abnormal operation state of any abnormal type, determining a test template corresponding to the abnormal type in a preset rule base, wherein the preset rule base includes: a correspondence between each abnormal type and each test template; Obtaining operating parameters corresponding to the current operating state of the storage system; Supplementing the test template based on the operating parameters; The first test case is generated according to the supplemented test template.

6. The test case generation method according to claim 1, characterized in that: After generating a first test case corresponding to the storage system according to the classification result, the method further includes: Executing the first test case based on the target device, and determining an execution result corresponding to the first test case; generating second log data corresponding to the first test case according to the execution result, and performing feature extraction on the second log data to obtain a third feature vector corresponding to the second log data; Model parameters in the target model are updated according to the third eigenvector.

7. The test case generation method according to claim 1, characterized in that: After generating a first test case corresponding to the storage system according to the classification result, the method further includes: When it is determined that a plurality of second test cases corresponding to the storage system have been generated, executing the plurality of second test cases in one of the following ways: Method 1: obtaining a first urgency level corresponding to each second test case, and determining a first execution order of the plurality of second test cases according to the first urgency level; Execute the plurality of second test cases in sequence according to the first execution order; Method 2: construct multiple parallel nodes, and execute the multiple second test cases in parallel based on the multiple parallel nodes; Method three: determining whether there are multiple third test cases with consistent test steps among the multiple second test cases; Merging the multiple third test cases; Processing the merged test case and other test cases in the plurality of second test cases except the merged test case in parallel; or, determining a second urgency level of the merged test case and the other test cases; determining a second execution order of the merged test case and the other test cases based on the second urgency; The merged test case and the other test cases are executed in sequence based on the second execution order.

8. A test case generation device, characterized in that: include: an extraction module, configured to perform feature extraction on first log data generated in real time by the storage system to obtain a first feature vector corresponding to the first log data, wherein the first log data includes at least one of the following: performance indicator data, storage status, operation records, and storage error records of the storage system at a current time; an input module, configured to input the first feature vector into a target model, so that the target model classifies the operating state of the storage system at the current time based on the first feature vector and generates a classification result, wherein the classification result includes at least one of the following: a normal operating state and an abnormal operating state of different abnormal types; A generation module is used to generate a first test case corresponding to the storage system according to the classification result.

9. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the test case generation method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the test case generation method according to any one of claims 1 to 7 are implemented.