Performance test method and device, electronic equipment and storage medium
By storing performance test data in column format and aggregating operations, the problem of performance test data occupies a large amount of computing resources is solved, and more efficient data storage and analysis is achieved.
Patent Information
- Application Number
- CN202411768189.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-06
AI Technical Summary
In performance testing, the generated performance test data occupies a large amount of computing resources, resulting in reduced processing efficiency.
Save performance test data in column format and generate performance test reports through columnar aggregation operations to reduce the need for reading and computing.
Reduces disk space, improves data analysis efficiency, and reduces memory footprint, thereby alleviating the memory pressure when generating performance test reports.
Smart Images

Figure CN119938462A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a performance testing method, device, electronic device and storage medium. Background Art
[0002] When a large number of concurrent user requests are simulated through performance testing tools (such as Locust) to test the performance of a device, a large amount of performance test data is generally generated, which requires more computing resources to process the large amount of performance test data and also reduces processing efficiency. Summary of the invention
[0003] The main purpose of the embodiments of the present application is to propose a performance testing method, device, electronic device and storage medium to reduce the computing resources occupied by performance testing data and improve processing efficiency.
[0004] To achieve the above object, an embodiment of the present application provides a performance testing method, which includes the following steps:
[0005] Storing the performance test data in a columnar format; wherein the performance test data is obtained by performing a performance test on the device under test;
[0006] The performance test data stored in the column format is read to generate a performance test report of the device under test.
[0007] In some embodiments, the reading of the performance test data stored in the columnar format to generate a performance test report of the device under test includes the following steps:
[0008] Reading the performance test data stored in the columnar format, and performing a columnar aggregation operation on the read data to obtain an aggregate value;
[0009] The performance test report of the device under test is generated according to the aggregate value.
[0010] In some embodiments, the reading of the performance test data stored in the columnar format and performing a columnar aggregation operation on the read data to obtain an aggregate value includes the following steps:
[0011] Reading a target column in the performance test data stored in the columnar format;
[0012] The column aggregation operation is performed on the data obtained by reading the target column to obtain the aggregate value.
[0013] In some embodiments, reading the target column in the performance test data stored in the columnar format comprises the following steps:
[0014] Using DuckDB to read the target column in the performance test data stored in the columnar format;
[0015] The step of performing the column aggregation operation on the data obtained by reading the target column to obtain the aggregate value includes the following steps:
[0016] The column aggregation operation is performed on the data obtained by reading the target column using an SQL statement to obtain the aggregate value.
[0017] In some embodiments, generating the performance test report of the device under test according to the aggregate value comprises the following steps:
[0018] Determine the value of each performance indicator according to the aggregate value;
[0019] The performance test report of the device under test is generated using the values of the various performance indicators.
[0020] In some embodiments, storing the performance test data in a columnar format includes at least one of the following steps:
[0021] If the performance test duration of each of the devices under test is less than a set duration threshold, the performance test data is generated as non-column format data, and after the performance test is completed, the non-column format data is converted into the column format data for storage;
[0022] Alternatively, if the performance test duration of each of the devices to be tested is equal to or greater than the set duration threshold, the performance test data is generated as data in the columnar format for storage.
[0023] In some embodiments, the number of the devices to be tested is at least 2, and storing the performance test data in a columnar format includes the following steps:
[0024] The performance test data in different data formats are stored as Parquet format files.
[0025] In some embodiments, before storing the performance test data in a columnar format, the method further includes the following steps:
[0026] Collect the performance test data corresponding to one or at least two of the devices under test.
[0027] In some embodiments, the collecting of the performance test data corresponding to one or at least two of the devices under test comprises the following steps:
[0028] When one or at least two of the devices under test run a preset test script, status data of one or at least two of the devices under test are collected as the performance test data.
[0029] In some embodiments, the collecting of the performance test data corresponding to one or at least two of the devices under test comprises the following steps:
[0030] The performance test data generated by at least two of the devices under test during concurrent testing are collected.
[0031] In some embodiments, the method further comprises the following steps:
[0032] Get query data;
[0033] determining associated data of the query data in the performance test data stored in the columnar format;
[0034] The column where the associated data is located is decompressed to obtain the target data.
[0035] In order to achieve the above-mentioned purpose, a performance testing device is provided in another aspect of an embodiment of the present application, wherein the device comprises:
[0036] A data storage unit, used to store performance test data in a column format; wherein the performance test data is obtained by performing performance testing on the device under test;
[0037] The data reading and report generating unit is used to read the performance test data stored in the column format and then generate a performance test report of the device under test.
[0038] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned performance testing method when executing the computer program.
[0039] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned performance testing method is implemented.
[0040] The embodiments of the present application include at least the following beneficial effects:
[0041] The present application can store performance test data in a columnar format, wherein the performance test data is obtained by performing a performance test on the device to be tested; the performance test data stored in a columnar format is read to generate a performance test report for the device to be tested. The present application stores performance test data in a columnar format, which can greatly reduce the space occupied by the disk, thereby greatly alleviating the storage pressure of the disk; in addition, when querying and analyzing performance test data, only the data of the required columns can be read, thereby greatly reducing the reading operation and improving the efficiency of analyzing data; the small amount of data read can reduce the amount of calculation, which can significantly reduce the memory usage, thereby alleviating the pressure on memory usage when generating performance test reports. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0043] Figure 1 A schematic diagram of a performance testing method provided in an embodiment of the present application;
[0044] Figure 2 An example diagram of performance test data provided in an embodiment of the present application;
[0045] Figure 3 A schematic diagram of the structure of a performance testing device provided in an embodiment of the present application;
[0046] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0048] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0049] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0050] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0051] The embodiments of the present application provide a performance testing method, device, electronic device and storage medium. The technical solution of the present application includes: collecting performance test data corresponding to multiple devices under test; storing the performance test data in a columnar format; performing columnar aggregation operations on the performance test data stored in a columnar format to obtain an aggregate value; and generating a performance test report for each device under test based on the aggregate value. The present application stores performance test data in a columnar format, which can greatly reduce the space occupied by the disk, thereby greatly alleviating the storage pressure of the disk; in addition, when querying and analyzing performance test data, only the data of the required columns can be accessed, thereby greatly reducing the read operation and improving the efficiency of analyzing data; the performance data of the device under test can be directly determined at one time through the aggregate value, which reduces the amount of calculation and can significantly reduce the memory usage, thereby alleviating the pressure on memory usage.
[0052] The embodiment of the present application provides a performance testing method, which relates to the field of computer technology. A performance testing method provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or it can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms and other basic cloud computing services. The server can also be a node server in a blockchain network; the software can be an application that implements a performance testing method, etc., but is not limited to the above forms.
[0053] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0054] Reference Figure 1 The present application embodiment provides a performance testing method, which may include but is not limited to S110 to S120, as follows:
[0055] S110: Storing performance test data in a column format; wherein the performance test data is obtained by performing performance testing on the device under test.
[0056] First, the performance tests that may be involved in the embodiments of the present application are described. By way of example, the performance tests may include concurrency testing, load testing, stress testing, configuration testing, reliability testing, capacity testing, and the like.
[0057] Specifically, the above-mentioned performance tests of each type are described as follows:
[0058] Load testing: Gradually increase the load on the device under test, test the performance changes of the device under test, and determine the maximum load that the device under test can withstand while meeting the performance indicators.
[0059] Stress testing: Gradually increase the stress on the device under test so that some resources of the device under test reach saturation or the edge of collapse, thereby determining the maximum stress that the device under test can withstand.
[0060] Concurrency testing: Simulates concurrent user access and tests whether there are deadlocks or other performance issues when multiple users concurrently access the same application or module.
[0061] Configuration test: Adjust the software and hardware environment and test the impact of various environments on the performance of the device under test to determine the optimal allocation principle.
[0062] Reliability test: Load a certain business pressure on the device under test, run it continuously for a period of time, and test the stability of the device under given conditions.
[0063] Capacity test: Under certain conditions, test the maximum number of users and maximum storage capacity that the device under test can support.
[0064] It is understandable that, during the performance test of the device under test, the status data of the device under test under various test conditions can be used as performance test data.
[0065] Optionally, this embodiment can be applied to any device with data processing capabilities (referred to as a host in the embodiment of this application). Therefore, this embodiment can store performance test data in a columnar format in the database of the host. In a database based on column storage, data can be stored according to columns as logical storage units, and data in a column exists in a continuous storage form in the storage medium. The column storage of the embodiment of this application can be applied to the following example scenarios: online analytical processing systems, such as data warehouses, business intelligence systems, reporting systems, etc.; big data analysis; read-only and data archiving: data is not often modified after being written, and is mainly used for reading and analysis.
[0066] Furthermore, it should be noted that when storing in columnar format, the disk can store only the values from a single column, which is more efficient than storing the values of an entire row. The compressed data will be smaller after columnar storage.
[0067] Exemplarily, the column storage format of this embodiment may include the following example formats:
[0068] Parquet: Supports efficient data compression and encoding, suitable for big data processing in the Hadoop ecosystem.
[0069] Apache ORC: can be used in Hive and other Hadoop ecosystem components to provide efficient storage and query performance.
[0070] ColumnStore: Focuses on fast reading and supports efficient analytical queries.
[0071] Google BigQuery: It uses column-based storage technology to optimize query performance and can be used for large-scale data analysis.
[0072] ClickHouse: Designed for fast data analysis, supports high-concurrency queries, and can be used for real-time data analysis.
[0073] Vertica: Focuses on large-scale data analysis, provides high-performance query and analysis capabilities, and can be used for enterprise-level applications.
[0074] Amazon Redshift: uses columnar storage to optimize performance and is suitable for processing large amounts of data analysis and for data warehouses.
[0075] SAP HANA: An in-memory database that uses columnar storage technology and is suitable for real-time data analysis.
[0076] Column storage can usually use efficient compression algorithms because the same column has the same data type, has more duplicate data, and has a high compression rate. The smaller the disk space occupied by the data, the less time the query engine spends on reading (the reading includes reading the data from the disk into the memory and reading the data from the memory into the CPU). After compressing the data, it generally needs to be decompressed when it is processed again. Therefore, considering the processing time of the entire process from compression to decompression, this embodiment may not use the compression ratio as the only processing indicator, but comprehensively consider the compression ratio and decompression speed to minimize the above processing time.
[0077] In addition, the column-based storage of performance test data in the embodiment of the present application may also include nested data structures and some complex data types, such as arrays and mappings, etc., based on which the compression efficiency can be further improved, that is, the disk space required to store performance test data can be effectively reduced.
[0078] Therefore, thanks to columnar storage, the host can quickly query and aggregate performance test data.
[0079] S120: Read the performance test data stored in the column format to generate a performance test report for the device under test.
[0080] Specifically, the performance test data stored in a columnar format is efficiently compressed. Therefore, this embodiment can read a small amount of data and then decompress it. For example, the host reads the performance test data stored in a columnar format in the database into the memory for decompression, and the required data can be obtained, and then a performance test report for the device under test is generated based on the read data.
[0081] Exemplarily, the performance test report may be in the form of a table, different types of statistical graphics (eg, line graphs, pie charts, etc.), a text summary, or a video report, etc. The specific form may be determined according to actual needs.
[0082] It is understandable that the performance test report can enable users to clearly understand the performance of each device under test, including existing defects and corresponding modification suggestions, and can provide real and effective reference data for improving the device under test.
[0083] The beneficial effects of this embodiment include at least:
[0084] The performance test data of the device under test is stored in a columnar format, the storage method of the performance test data is optimized, the performance test data is efficiently compressed, and the disk space occupancy is greatly reduced. Then, the reading of the efficiently compressed performance test data can reduce the reading operation, and then the performance test report is generated according to the read data, thereby improving the efficiency of generating the performance test report.
[0085] Next, the method provided by this application is further described.
[0086] As an optional implementation, before S110 stores the performance test data in a columnar format, the embodiment of the present application may further include S100:
[0087] S100: Collecting a plurality of performance test data corresponding to one or at least two of the devices under test.
[0088] For example, the host collects performance test data corresponding to one or at least two devices under test, wherein this embodiment can perform a performance test on one device under test separately and then collect the performance test data of the device under test separately, or can perform a performance test on at least two devices under test and then collect the performance test data of each device under test separately.
[0089] It should be noted that each device under test generates its own performance test data. More specifically, each device under test generates its own multiple performance test data, and the host collects the performance test data of the device under test.
[0090] In some optional embodiments, the host may collect performance test data in the following manner: the host establishes a communication connection with the device under test, and dynamically obtains performance test data through the network, that is, the device under test continuously generates performance test data during the performance test process, and then dynamically sends it to the host through the network; or, each device under test first completes the performance test locally, and then sends the collected performance test data to the host through the network, and the host receives the performance test data sent by each device under test. In addition, in this embodiment, the performance test data of each device under test can also be pre-stored in a readable storage medium, and the host collects the performance test data therein by reading the readable storage medium.
[0091] Then, the device under test of the embodiment of the present application is described. The device under test can be an electronic device with data processing capabilities, such as a client, a network terminal, a server terminal, etc., and specific devices can include mobile phones, tablet computers, computers, servers, etc. The host is an electronic device that can process the performance test data generated by the above-mentioned device under test, such as a computer, a server, etc.
[0092] The beneficial effects of this embodiment include at least:
[0093] By collecting performance test data of several devices under test, it is possible to analyze the performance of one device under test individually or the performance of multiple devices under test, which can adapt to various test scenarios.
[0094] In some optional embodiments, step S100 of collecting the performance test data corresponding to one or at least two of the devices under test may include S101:
[0095] S101: When one or at least two of the devices under test run a preset test script, status data of one or at least two of the devices under test are collected as the performance test data.
[0096] Specifically, this embodiment can run a preset test script on one or at least two devices under test. The test script can be written specifically according to the requirements of the performance test so that the devices under test can obtain the required performance test data when running the test script.
[0097] Considering that each device under test may be an electronic device of a different type and with a different operating system, a fixed test script may not be able to run on all types of devices under test at the same time. In order to adapt to various types of devices under test, the test script of this embodiment can also be adaptively configured according to the hardware configuration, operating system, etc. of the device under test, such as adaptively configuring the relevant parameters of the test script so that the test script can successfully run on each device under test. Exemplarily, this embodiment can implement large-scale performance testing based on the test script, and the specific implementation method may include: multiple devices under test run the same test script, and the test script is adaptively configured with relevant parameters to adapt to each device under test of different types or different operating systems, and each device under test runs the test script to generate corresponding performance test data, and the host can collect the performance test data in this embodiment, and then store it in a columnar format, and then read the performance test data stored in the columnar format to generate a performance test report.
[0098] Among them, the performance testing that can be achieved by running the test script on the device under test can include not only concurrency testing, but also load testing, stress testing, configuration testing, reliability testing, capacity testing, and other optional forms of performance testing.
[0099] The beneficial effects of this embodiment may include:
[0100] The device under test runs the preset test script, and the host is responsible for collecting the performance test data generated by the device under test. The device under test performs performance tests locally and autonomously, freeing up the computing resources of the host and achieving performance testing that is not restricted by distance.
[0101] As another implementation, step S100 may include S102:
[0102] S102: Collect the performance test data generated by at least two of the devices under test during concurrent testing.
[0103] Specifically, in order to achieve large-scale performance testing, in this embodiment, concurrent testing can be performed in at least two devices under test, and multiple devices under test simulate multiple users concurrently accessing the same application or module, for example, multiple users access the same server at the same time. Each device under test can be deployed in a distributed manner, and it can be connected to the host through a relatively close communication network (such as various networking forms of the Internet of Things, Bluetooth, WIFI, etc.), or it can be connected to the host through remote communication, and then each device under test performs concurrent testing to generate performance test data, and then dynamically sends the performance test data to the host, or it can be sent to the host after the test is completed, and the host collects the performance test data sent by each device under test, and then stores it in a column format. Optionally, the host can also collect performance test data through channels such as mobile readable storage media and cloud servers. The specific collection method can refer to the detailed description of S100 above.
[0104] The beneficial effects of this embodiment include at least:
[0105] Large-scale concurrent testing can be achieved by concurrently testing at least two devices under test and then collecting corresponding performance test data.
[0106] In some optional embodiments, step S110 stores the performance test data in a columnar format, which may include S111:
[0107] S111: Storing the performance test data in different data formats in the column format.
[0108] It is understandable that since each device under test may be installed with a different operating system or be in different performance test conditions, the data format of the performance test data generated by it may be different. After collecting the performance test data in different data formats, the host can perform necessary data preprocessing, such as data cleaning and standardization, and then store the preprocessed performance test data in a columnar format in the database.
[0109] The beneficial effects of this embodiment include at least:
[0110] It is compatible with performance test data of various data formats generated by different operating systems and different performance test conditions, thereby improving the scope of application and compatibility of this embodiment.
[0111] Considering that the duration of the performance test is a key condition affecting the capacity of the performance test data, therefore, as another implementation, step S110 stores the performance test data in a columnar format, which may include at least one of S112 or S113:
[0112] S112: if the performance test duration of each of the devices under test is less than a set duration threshold, the performance test data is generated as non-column format data, and after the performance test is completed, the non-column format data is converted into the column format data for storage.
[0113] Specifically, this embodiment can determine the set duration threshold, and the duration that does not reach the set duration threshold is regarded as a short-time test. Considering that the amount of short-time test data may not be huge, in order to enable the host to collect performance test data as soon as possible, in the time period from the start of the performance test to the set duration threshold, the host can generate the collected performance test data into non-column format data (such as non-column formats such as csv, txt, etc.), based on the high write efficiency of non-column format data, the host's reading ability is used to quickly receive data and then store it in the database. When the performance test is completed, the non-column format data in the database is converted into column format data to compress the space occupied by the performance test data, which not only improves the efficiency of collecting performance test data, but also reduces the space occupied by the disk.
[0114] S113: If the performance test duration of each of the devices under test is equal to or greater than the set duration threshold, the performance test data is generated into data in the columnar format for storage.
[0115] Specifically, when the performance test duration is equal to or greater than the set duration threshold, the performance test is treated as a long test, and the long test will generate a large amount of performance test data. At the same time, considering that the long test pays more attention to the analysis and conclusion of the final performance test data, therefore, in order to reduce disk space occupied and improve the continuous stability of the performance test, the host of this embodiment can directly generate the collected performance test data into columnar format data, and then store it in the database.
[0116] It can be understood that this embodiment includes three implementation methods: 1) When the performance test duration is less than the set duration threshold, a short test is performed. The performance test data is generated as non-column format data, and after the performance test is completed, the non-column format data is converted into column format data for storage; 2) When the performance test duration is equal to or greater than the set duration threshold, a long test is performed. At this time, the performance test data is directly generated as column format data for storage from the beginning of the test; 3) When the performance test duration is greater than the set duration threshold, a short test is performed first, and then a long test is performed. Before the performance test duration is less than the set duration threshold, a short test is performed, the performance test data is generated as non-column format data, and when the set duration threshold is reached, the non-column format data is converted into column format data for storage; after the set duration threshold is reached, a long test is performed, and the performance test data is generated as column format data for storage.
[0117] The beneficial effects of this embodiment include at least:
[0118] Since the write performance of non-columnar format data is better than that of columnar format data, the short-term performance test host can directly generate non-columnar format data to make full use of the reading capability and collect performance test data faster. When the stress test time is very long, that is, the performance test data becomes big data, the performance requirements for data landing are not very high, and more emphasis is placed on the analysis and conclusions of the final stress test data. Therefore, the host can directly generate the performance test data into columnar format data. The columnar format data is compressed, which is more conducive to space saving and balancing the overall load of the host, and can also improve the continuous stability of the performance test.
[0119] In some other embodiments, the number of the devices to be tested is at least 2, and step S110 stores the performance test data in a columnar format, including step S114:
[0120] S114: storing the performance test data in different data formats as Parquet format files.
[0121] It can be understood that, in this embodiment, the performance test data in various data formats can be stored as Parquet format files, for example, the performance test data in various data formats can be converted into Parquet format files and stored in the database of the host.
[0122] The features of Parquet format files include efficient data storage, fast query performance, high compression ratio, and support for complex data structures.
[0123] Thanks to the characteristics of the Parquet format file, the beneficial effects of this embodiment include at least:
[0124] Efficient column reading: Only the required columns need to be read, which reduces disk I / O operations (read operations) and improves query performance.
[0125] High compression rate: Since the data types and value ranges of the same column are similar, it is easier to perform efficient compression, such as Run Length Encoding and Delta Encoding.
[0126] Vectorized processing: facilitates CPU instruction-level parallelism and vectorized computing, improving processing speed.
[0127] Secondly, Parquet format files have a high compression ratio and efficient encoding method, which further saves storage space, including:
[0128] High compression ratio: Parquet format files use an efficient compression algorithm, which can significantly reduce the storage space occupied by files.
[0129] Multiple encoding methods: Use multiple column encoding methods, such as Run Length Encoding, Delta Encoding, Bit Packing, etc., to further improve the compression ratio.
[0130] In addition, Parquet format files also support complex data structures and efficient data processing, including:
[0131] Support for complex data structures: Parquet format files support complex data types and nested data structures and can be used to process large-scale data sets.
[0132] Multi-platform support: Parquet format files can be used across multiple computing platforms, such as Apache Hadoop, Apache Spark, Apache Hive, etc.
[0133] Schema evolution: Parquet format files support schema evolution, allowing changes to the data schema without destroying existing data.
[0134] In summary, the Parquet format file adopted in the embodiment of the present application provides an efficient and fast solution for big data processing and analysis through its unique columnar storage method, high compression ratio and efficient encoding method, as well as support for complex data structures and multi-platform support.
[0135] In some embodiments, step S120 reads the performance test data stored in the column format to generate a performance test report of the device under test, which may include S121-S122:
[0136] S121: Read the performance test data stored in the columnar format, and perform a columnar aggregation operation on the read data to obtain an aggregate value.
[0137] Exemplarily, this embodiment may utilize an embedded analytical database to read performance test data stored in a columnar format, wherein the embedded analytical database may include DuckDB, EdgeDB, or FoundationDB, etc.
[0138] Specifically, DuckDB is an analytical database similar to SQLite, with a rich SQL dialect and efficient data processing capabilities. DuckDB is particularly suitable for online analytical processing, capable of processing large data sets and executing complex SQL queries. DuckDB has no external dependencies and does not run as an independent process, but runs directly within the address space of the application, which makes it very convenient to transfer data between the application and the database. EdgeDB is a database that combines the ease of use and immediacy of NoSQL, the relational modeling capabilities of a graph database, and the guarantees and consistency of SQL. EdgeDB uses the PostgreSQL core, provides strong field types and a SQL-like query language, and also supports the relational modeling capabilities of a graph database. FoundationDB is a multi-mode database that stores data internally as key-value pairs, but can be organized into relational tables, graphs, documents, and many other data structures. FoundationDB supports ACID transactions, and horizontal expansion and replication functions are both directly available.
[0139] It should be noted that an aggregation operation refers to a calculation performed on a data set, which is used to summarize and count the data. The aggregate value obtained by the aggregation operation can be used in scenarios such as data analysis, report generation, and business intelligence, thereby helping users quickly obtain the overall trends and characteristics of the data.
[0140] Aggregation operation is a data batch processing operation, which can group data (or not group data, that is, each record is treated as a group), and then perform multiple batch processing operations on each group of data, and then return the result, that is, the aggregate value.
[0141] Aggregation operations can generally include: single-action aggregation, aggregation pipeline, and MapReduce (programming model). Single-action aggregation provides simple access to common aggregation processes, and operations aggregate documents from a single collection. Aggregation pipeline is a framework for data aggregation. Based on the concept of data processing pipeline, documents enter a multi-stage pipeline to transform documents into aggregation results. MapReduce operation has two stages: the map stage that processes each document and emits one or more objects for each input document, and the reduce stage that combines the output of map operations.
[0142] Taking MongoDB as an example, single-action aggregate functions include:
[0143] db.collection.estimatedDocumentCount(); db.collection.count(); db.collection.dist inct(). The above functions are used to return the count of documents in a collection or view, the count of documents matching a query, and to find the distinct values of a specified field.
[0144] The aggregation pipeline operation defines a series of stages, passes the processing results of the previous stage to the next stage for further processing, and finally returns a new result set.
[0145] It can be understood that this embodiment can perform column aggregation operations on performance test data stored in a columnar format in the host, and aggregate and calculate several performance indicators at one time. The performance indicators can be set arbitrarily according to actual needs, that is, the aggregation value obtained each time the aggregation operation is performed may be different.
[0146] In some optional embodiments, the host can receive the aggregation operation instruction input by the user through an interactive interface, an input box, a button or a remote communication information, and in response to the aggregation operation instruction, read the performance test data stored in the host database in a columnar format into the host memory, and then perform the aggregation operation in the memory to obtain the corresponding aggregation value. Among them, the performance test data stored in the columnar format is highly compressed, and the occupied disk space and memory are correspondingly less, so the data read into the host memory can occupy less memory resources. At the same time, the aggregation operation can directly calculate each required performance parameter at one time. Compared with calculating the performance parameters one by one in the memory, it can reduce the number of repeated calculations and further reduce the occupied memory resources, which can greatly improve the efficiency of memory utilization.
[0147] It should be noted that the performance test data includes different performance data of multiple devices under test. In the embodiment of the present application, the user can arbitrarily specify part or all of the data of one or more devices under test for aggregation, and then generate corresponding aggregation operation instructions. The host executes the aggregation operation instructions to specifically aggregate the performance test data of the required devices under test, thereby achieving accurate customized aggregation.
[0148] S122: Generate the performance test report of the device under test according to the aggregate value.
[0149] It can be understood that the aggregate value includes various performance test data of the device to be tested. Optionally, in this embodiment, the aggregate value can be analyzed by an analysis tool, and then a performance test report can be generated according to the analysis result.
[0150] In some more specific embodiments, the step of analyzing the aggregated value may include at least one of the following analysis types: demand analysis, data acquisition, data preprocessing, analysis and modeling, model evaluation and optimization, and deployment.
[0151] The steps of each of the above analysis types may include:
[0152] Demand analysis: This can be used as the first step in aggregate value analysis to determine the direction and method of subsequent analysis. By clarifying the goals and requirements of aggregate value analysis, it can provide guidance for subsequent steps.
[0153] Data acquisition: Based on the results of demand analysis, relevant data is collected and extracted from the aggregated value. The extracted data is the basis of data analysis. The more accurate and complete the extracted data is, the more accurate and reliable the results of data analysis will be.
[0154] Data preprocessing: In addition to the required target data, the aggregate value may also contain other unnecessary data. In addition, in order to uniformly calculate and measure data of different data types in the aggregate value, the present embodiment may preprocess the aggregate value, including cleaning, merging, transforming and standardizing the data to improve the quality and consistency of the data and make it suitable for analysis and modeling.
[0155] Analysis and modeling: Discover valuable information in the data and draw conclusions through methods such as comparative analysis, grouping analysis, cross analysis, regression analysis, and models and algorithms such as clustering, classification, and association rules.
[0156] Model evaluation and optimization: Evaluate the established model, use different indicators to assess its performance, and optimize it as needed to improve the accuracy and reliability of the model.
[0157] Deployment: Apply the analysis results and conclusions to actual production systems to realize the practical application value of data analysis.
[0158] After analyzing the aggregated values, the analysis results are obtained, and then the performance test report generated based on the analysis results generally includes:
[0159] Overview: Provide a brief description of the purpose and scope of the test.
[0160] Test execution status: describes the execution status and results of the test cases, including the number of passed cases, the number of failed cases, and the number of unexecuted cases.
[0161] Bugs and Defects: List the bugs and defects found, including a detailed description of the bug, steps to reproduce it, and priority.
[0162] Performance and stability evaluation: Evaluate and analyze performance and stability data.
[0163] Conclusion and Recommendations: Summarize the test results and make recommendations for improvement and optimization.
[0164] It is understandable that the performance test report can enable users to clearly understand the performance of each device under test, including existing defects and corresponding modification suggestions, and can provide real and effective reference data for improving the device under test.
[0165] The beneficial effects of this embodiment include at least:
[0166] The performance test data of each device under test is stored in a columnar format, which optimizes the storage method of the performance test data and greatly reduces the disk space occupied. Then, the analysis method of the performance test data is improved. By aggregating the performance test data, the calculation of each data is reduced, thereby reducing the memory space occupied and improving the efficiency of generating performance test reports.
[0167] Furthermore, step S121 reads the performance test data stored in the columnar format, and performs a columnar aggregation operation on the read data to obtain an aggregate value, which may include:
[0168] S1211: Read a target column in the performance test data stored in the columnar format, and perform the columnar aggregation operation on the data obtained by reading the target column to obtain the aggregate value.
[0169] Specifically, the performance test data stored in a columnar format can directly read the required data, so this embodiment can first determine the column where the required data is located as the target column, and skip other columns that are not target columns, that is, skip unnecessary data, and directly read the required data in the target column, and then perform aggregation operations on the read data to obtain the corresponding aggregate value.
[0170] The beneficial effects of this embodiment may include:
[0171] The characteristics of the column storage format are fully utilized to directly read the column where the performance test data is located when reading it, which improves the reading efficiency.
[0172] In some more specific embodiments, step S1211 reads a target column in the performance test data stored in the columnar format, and performs the columnar aggregation operation on the data obtained by reading the target column to obtain the aggregate value, which may include S12111 to S12112:
[0173] S12111: Using DuckDB to read the target column in the performance test data stored in the columnar format.
[0174] Specifically, DuckDB uses a column storage format, making it faster and more efficient to analyze large amounts of performance test data; DuckDB can run efficiently in memory, supports fast SQL queries, and can be seamlessly integrated with Python, R, and other languages, making it suitable for a variety of data analysis environments. DuckDB provides a SQLite compatibility layer, allowing applications that previously used SQLite to use DuckDB through relinking or library reloading. Efficient query processing: vector processing is performed to maximize the use of CPU cache, providing convenience and efficiency similar to SQLite.
[0175] Therefore, this embodiment can use SQL query based on DuckDB to quickly determine the target column and read the performance data of the target column.
[0176] S12112: Using a SQL statement, perform the column aggregation operation on the data obtained by reading the target column to obtain the aggregate value.
[0177] For example, this embodiment can use the aggregate function in the SQL statement to perform aggregation operations. The aggregate function can perform calculations on a set of performance test data rows and return calculated values. The returned calculated values can be used to summarize and aggregate data to extract meaningful information and conclusions.
[0178] Aggregate functions process performance test data by performing calculations on a given column on all rows, aggregating the results of the calculations into a calculated value, and returning the aggregated calculated value, which is usually displayed in the header of the row or column. Aggregate functions are often used in conjunction with the GROUP BY clause to group specific columns and then perform aggregate calculations on each group.
[0179] Advantages of aggregate functions: Summarize performance test data to obtain meaningful conclusions, identify trends and patterns in performance test data, and simplify the writing of complex queries.
[0180] Specifically, the aggregate functions in SQL statements can include:
[0181] COUNT() - counts the number of rows;
[0182] SUM() - calculates the sum of the values in a given column;
[0183] AVG() - calculates the average of the values in a given column;
[0184] MAX() - returns the maximum value in a given column;
[0185] MIN() - Returns the minimum value in a given column.
[0186] The beneficial effects of this embodiment include at least:
[0187] In the step of obtaining the indicator values of the performance test data, the integrated single SQL statement operation result can be used for one-time calculation, which greatly improves the efficiency of calculating the performance indicator values.
[0188] In some optional embodiments, step S122 generates the performance test report of the device under test according to the aggregate value, which may include S1221:
[0189] S1221: Determine the value of each performance indicator according to the aggregation value, and generate the performance test report of each device under test using the value of each performance indicator.
[0190] Exemplarily, the performance indicators of this embodiment may include the protocol name of the device to be tested, the number of devices to be tested, whether the test is successful, the test time, etc. The aggregate value may include the values of the above-mentioned performance indicators. According to the aggregate value, the average value, maximum value, minimum value and other processed values of the above-mentioned performance indicator values can also be obtained. Then, the host can use data analysis tools, such as data analysis algorithms or models, to analyze the values of each performance indicator and the above-mentioned processed values, so as to generate a performance test report. Among them, the performance test report can be in the form of a table, different types of statistical graphics (such as line charts, pie charts, etc.), text summary or video report, etc., and the detailed description of the above-mentioned step S130 can be referred to for details.
[0191] The beneficial effects of this embodiment include at least:
[0192] Generate a performance test report based on the numerical values of various performance indicators, accurately reflect the key data of the performance test, and provide users with effective reference data.
[0193] Considering that the stored performance test data may need to be queried and analyzed, the embodiment of the present application may further include steps S131 to S133:
[0194] S131: Obtain query data.
[0195] Specifically, the host may obtain query data input by the user, and the query data may be input in the form of voice, text, or image.
[0196] When the query data is input in the form of voice, the host can record the user's voice through a recording device, or directly obtain a pre-recorded voice, or directly capture a certain audio segment in the audio or video as the query data.
[0197] When the query data is input in the form of text or in the form of an image, the input can be made through an input box or an interactive interface that establishes a communication connection with the host.
[0198] S132: Determine associated data of the query data in the performance test data stored in the columnar format.
[0199] Exemplarily, the host may query the database for data related to the query data as associated data. The association may be determined based on the similarity between the query data and the performance test data, for example, taking a number of performance test data ranked high in similarity as associated data, or taking performance test data whose similarity reaches a preset threshold as associated data.
[0200] It should be noted that the performance test data is stored in a columnar format, and the columnar storage format allows unnecessary columns to be skipped, that is, the present embodiment can directly determine the columns related to the query data, thereby improving query efficiency.
[0201] S133: Decompress the column where the associated data is located to obtain target data.
[0202] It should be noted that the performance test data stored in columnar format can store the same type of data in the same column and store it after being compressed using an efficient compression algorithm. Therefore, this embodiment can first determine the column where the associated data is located, which may be one or more columns. The target data can be obtained by simply decompressing the column where the associated data is located, thereby further improving query efficiency.
[0203] The user inputs query data to the host, with the purpose of obtaining data that meets the user's needs through the query data. After decompression, the data that meets the user's needs, namely, the target data, can be obtained from the associated data.
[0204] The beneficial effects of this embodiment include at least:
[0205] By making full use of the columnar storage characteristics of performance test data, queries only need to determine the required columns, and then decompress the required columns to obtain the target data that meets the user's query purpose, greatly improving query efficiency.
[0206] Next, the solution of the embodiment of the present application will be introduced and explained in detail with reference to specific application examples.
[0207] Before describing this embodiment in detail, some related technologies involved in this embodiment are described first, as follows:
[0208] Locust (literally translated as locust, meaning that it can simulate thousands of concurrent users like locusts) is an open source performance testing tool. It is written in Python and is used to simulate concurrent requests from a large number of users to test the performance of various devices and systems. Locust is an easy-to-use distributed load testing tool, mainly used to simulate multiple users operating various devices and systems to test the performance of various devices and systems. The features of Locust include an event-based operation mechanism that can support thousands of concurrent users in a single process without using callbacks, and use lightweight processes to run through gevent. The main functions of Locust include: Distributed testing: supports running load tests in parallel on multiple devices and systems to improve test efficiency. User behavior definition: The behavior of each virtual user is defined through Python scripts, which can simulate complex usage scenarios. Web interface monitoring: Provides a web-based interface to display detailed information and statistical results of the test in real time, which is convenient for real-time monitoring and result analysis.
[0209] Advantages of Locust include:
[0210] Ease of use: Compared with other performance testing tools, Locust is easier to get started and use.
[0211] Flexibility: Using Python to write test plans does not require operation through the UI interface, providing higher flexibility.
[0212] Scalability: It supports distributed testing and has third-party plug-in support, making it easy to expand.
[0213] Locust has many advantages, such as user simulation, distributed testing, and real-time monitoring, but it also has problems that users criticize. For example, it consumes a lot of computing resources when generating performance test reports.
[0214] The high resource consumption when Locust generates performance test reports is mainly manifested in the following two aspects:
[0215] 1. High consumption of hard disk resources;
[0216] 2. High consumption of memory resources.
[0217] High consumption of hard disk resources is mainly reflected in the data storage of distributed files: when the number of distributed stress testing nodes reaches hundreds and there are many collection protocol points, the data set of a single node will become very large, which will eventually lead to the entire stress testing data set becoming very large. For example, if the data of a single node is 2G, 500 stress testing nodes will generate 1T of data, which places very high requirements on data storage.
[0218] High memory resource consumption is mainly reflected in the generation of the final stress test report, which reads all data into the memory, performs calculations in the memory, and involves sliced data when summarizing the calculations, resulting in memory usage of 2 to 3 times the actual data. For example, if 1G of dot data (i.e., performance test data) needs to generate a performance test report, it will consume about 3G of memory. The above problem will seriously occupy memory resources in the case of massive data.
[0219] Therefore, this embodiment, based on open source, changes the storage format of Locust's collected and marked data to achieve the purpose of reducing storage space, and at the same time changes the implementation method of Locust's report generation, greatly reducing the dependence on memory, and achieving the overall reduction of the problem of high resource dependence. That is, based on open source, the improved technical solutions of this embodiment include:
[0220] 1) Optimize the data storage format of open source projects: Change the format of local storage of dot data. Usually, dot data is stored as .txt or .csv data. Both of these formats occupy a large amount of space. After multiple comparisons and attempts, this embodiment finally selects the .Parquet format to store data.
[0221] 2) Reconstruct the way of generating reports for open source projects: For statistics and calculations of big data, the first thing that comes to mind is Pandas, but Pandas operations are very dependent on memory, that is, Pandas is pure memory calculation. Through various trial solutions, this embodiment finally reconstructs the report generation module and selects DuckDB to implement the entire generation logic.
[0222] The specific solutions of this embodiment may include:
[0223] First, the data storage format is explained. Figure 2 , Figure 2 This is an example graph of Locust's performance test data, from Figure 2 It can be seen that the status of each protocol in a unit of time includes the protocol name, quantity, success mark, time consumption and other data. Locust can record Figure 2 The data shown is used to store the performance data of the devices involved in the stress test, and finally summarize and analyze them into a report.
[0224] The existing Locust stores performance test data in csv or txt format. Since the existing Locust saves in csv format by default, the collected and stored data is quite large. When performing stability stress testing and continuous high-concurrency stress testing, the storage space occupied will continue to grow, and the requirements for disk storage space will be very high. To solve the space occupation problem, this embodiment has been tried and verified by many parties, and finally chose to use Parquet format to store data.
[0225] Parquet column storage has the following advantages:
[0226] 1. Column storage: Data is stored in columns rather than rows, which means that queries only need to access data in the required columns, which can greatly reduce I / O operations (read operations) and is particularly suitable for analytical workloads.
[0227] 2. Compression: Parquet format uses compression technology to reduce storage space and only decompresses data in necessary columns when reading, which improves efficiency.
[0228] 3. The columnar storage format allows skipping unnecessary columns, thereby improving query performance.
[0229] 4. Support for complex data structures: Parquet format supports nested data structures and complex types, such as arrays and mappings.
[0230] In terms of design and implementation, the main strategy for converting the Locust stress testing and point data into Parquet format files in this embodiment is as follows:
[0231] 1. Short-term stress testing, that is, when the stress testing task time is less than the set duration threshold, for example, 1 hour, a csv or txt file is generated, and finally, at the end of the stress testing, the csv or txt file is converted into a Parquet format file.
[0232] 2. For long-term stress testing, that is, if the stress testing task time is greater than the set duration threshold, for example, 1 hour, a file in Parquet format is directly generated.
[0233] The main reason for the above strategy is that Locust's writing performance for csv or txt files is better than that for files in Parquet format. For short-term stress testing, directly writing txt files can make full use of the host's io capabilities and land data faster. When the stress testing time is too long, that is, the dotted data becomes big data, the performance requirements for data landing are not very high, and more attention is paid to the analysis and conclusions of the final stress testing data. Therefore, this embodiment can directly write Parquet files, which is more conducive to space saving and balancing the overall load of the host, and can also continuously and stably perform stress testing.
[0234] Optionally, the Locust stress test data after the modification in this embodiment may include: Figure 2 The shown data shows the status of each protocol within a unit of time, including protocol name, quantity, success mark, time consumption and other data, and the above data is stored as a Parquet format file.
[0235] Based on the above reasons, with respect to the stress testing duration, the data landing strategy of Locust after the transformation in this embodiment is more conducive to the efficient and stable operation of Locust.
[0236] Since Locust itself is implemented in Python, in Python, this embodiment can use the pyarrow library to process Parquet files. Pyarrow is a Python library of the Apache Arrow project that provides a rich API for reading and writing Parquet files.
[0237] Python uses the pandas library to read the csv file into a DataFrame (a data structure similar to a two-dimensional array or table), and then uses the pyarrow library to save the DataFrame as a Parquet file.
[0238] In summary, Parquet format column storage is suitable for large-scale data analysis, data warehouse and data lake scenarios, especially for applications that need to quickly query large amounts of data and save storage space. Due to the efficient compression and query performance of the Parquet format, it is often used on big data processing platforms such as Apache Hadoop and Apache Spark.
[0239] In an example, the same amount of performance test data is stored as 421M in csv format and 18M in Parquet format. It can be seen that for the same amount of data, the disk space occupied by the file in csv format is 23 times the disk space occupied by the file in Parquet format. Therefore, after changing the data storage format in this embodiment, more than 20 times of space occupation can be saved, greatly alleviating the pressure on disk storage.
[0240] Then, the optimization of generating the performance test report in the embodiment of the present application is described.
[0241] The existing Locust is implemented in Python as a whole. When generating the performance test report, the Python open source library Pandas is used to process the collected files. That is, Pandas is used to process a bunch of csv files generated by distributed management. That is, the resource consumption of report generation is basically caused by Pandas operations.
[0242] Pandas's reliance on memory is mainly reflected in the following aspects:
[0243] 1. Data loading: Pandas needs to load the data completely into memory before processing, so it is limited by the available memory size. When the amount of data is large, sufficient memory is required to load and process the data.
[0244] 2. Data structure: Pandas' core data structures, DataFrame and Series, are memory-based data structures, so they are stored in RAM and are not suitable for processing ultra-large-scale data.
[0245] 3. Operation: When Pandas performs data operations and transformations, it usually needs to load the data into memory and use the data structure in memory to perform operations.
[0246] Due to the above memory dependency, when processing large-scale data, special attention should be paid to the memory limitations of the system and consideration should be given to using distributed computing or other big data processing frameworks to process data that exceeds the memory range.
[0247] Given that Pandas itself is highly dependent on memory, it is not possible to use Pandas to reduce memory costs. Therefore, after comparative attempts, this embodiment finally chooses to use DuckDB to achieve the generation of the entire report.
[0248] DuckDB is a memory-centric SQL database management system with the following advantages:
[0249] 1. Performance advantage: DuckDB performs well in processing analytical queries. DuckDB uses an optimized column storage and vectorized execution engine, which enables DuckDB to have excellent performance when processing large-scale data.
[0250] 2. Low latency: Since DuckDB is designed as a memory-centric database, DuckDB is usually able to provide low-latency query response time, which can improve efficiency in scenarios that require fast interactive queries.
[0251] 3. Memory usage advantage: DuckDB is optimized for memory usage, especially when processing large data sets. DuckDB provides efficient query processing by intelligently managing memory, allowing large-scale data to be processed with relatively less memory, thus avoiding the performance bottleneck caused by traditional disk access.
[0252] In general, DuckDB's advantages lie in its high performance, low latency, and optimized memory usage for analytical workloads, making it suitable for application scenarios that require fast and efficient processing of large-scale data.
[0253] Exemplarily, the steps of generating a performance test report using DuckDB include:
[0254] 1. DuckDB loads the stored performance test data, that is, specifies a series of Parquet files that DuckDB needs to process;
[0255] 2. Since Parquet files are stored in columnar format, the performance of aggregation operations on a column is extremely high. DuckDB performs columnar aggregation operations on Parquet files in SQL statement style to obtain aggregate values.
[0256] 3. Aggregate values can include the average, maximum, minimum, tp distribution value (probability value of t distribution or value of t statistic) of various performance indicators, etc., which can all be calculated in the form of aggregate SQL statements.
[0257] 4. In specific implementation, there is no need to calculate the specific values of multiple performance indicators separately, but the integrated aggregate value can be used for one-time calculation, which can greatly improve the efficiency of Locust in generating performance test reports in this embodiment.
[0258] The various performance test data generated by Locust in this embodiment are read in DuckDB mode, and the maximum, average, tp and other data of various performance indicators can be obtained by executing the corresponding SQL statements. The performance and resource consumption are better than Pandas, especially in terms of memory consumption. The overall memory usage is less than one-third of that of Pandas, which greatly saves memory requirements.
[0259] The beneficial effects of this embodiment include at least:
[0260] This embodiment can reduce the existing Locust's occupation of computing resources and save the hardware cost of performance testing, which is mainly reflected in the savings in disk space and the memory occupation required for performance test report generation. By optimizing the storage method of performance test data, the disk space occupied by performance test data is greatly reduced. By reconstructing the existing Locust report generation module and rewriting the entire report generation architecture, the memory usage is also halved.
[0261] Reference Figure 3 The embodiment of the present application also provides a performance testing device, which can implement the above-mentioned performance testing method, and the device includes:
[0262] A data storage unit, used to store performance test data in a column format; wherein the performance test data is obtained by performing performance testing on the device under test;
[0263] The data reading and report generating unit is used to read the performance test data stored in the column format and then generate a performance test report of the device under test.
[0264] In some embodiments, the data reading and report generating unit comprises:
[0265] A data reading unit, used for reading the performance test data stored in the columnar format, and performing a columnar aggregation operation on the read data to obtain an aggregate value;
[0266] A report generating unit is used to generate the performance test report of the device under test according to the aggregate value.
[0267] In some embodiments, the data reading unit includes:
[0268] A target column reading unit, configured to read a target column in the performance test data stored in the columnar format;
[0269] The data aggregation unit is used to perform the column aggregation operation on the data obtained by reading the target column to obtain the aggregation value.
[0270] In some embodiments, the target column reading unit comprises:
[0271] a target column reading subunit, configured to read the target column in the performance test data stored in the columnar format by using DuckDB;
[0272] The data aggregation unit comprises:
[0273] The data aggregation subunit is used to perform the column aggregation operation on the data obtained by reading the target column using a SQL statement to obtain the aggregate value.
[0274] In some embodiments, the report generation unit includes:
[0275] The report generating subunit is used to determine the value of each performance indicator according to the aggregate value, and generate the performance test report of each device under test by using the value of each performance indicator.
[0276] In some embodiments, the data storage unit comprises:
[0277] A first data storage subunit is used to generate the performance test data into non-column format data before the performance test duration of each of the devices under test is less than a set duration threshold, and convert the non-column format data into the column format data for storage after the performance test is completed;
[0278] The second data storage subunit is used to generate the performance test data into the columnar format for storage after the performance test duration of each of the devices under test reaches the set duration threshold.
[0279] In some embodiments, the number of the devices under test is at least 2, and the data compatible storage unit includes:
[0280] The third data storage subunit is used to store the performance test data in different data formats as Parquet format files.
[0281] In some embodiments, the apparatus further comprises:
[0282] The data collection unit is used to collect the performance test data corresponding to one or at least two of the devices under test before storing the performance test data in a column format.
[0283] In some embodiments, the data acquisition unit includes:
[0284] The first data collection subunit is used to collect status data of one or at least two of the devices under test as the performance test data when one or at least two of the devices under test run a preset test script.
[0285] In some embodiments, the data acquisition unit includes:
[0286] The second data collection subunit is used to collect the performance test data generated by at least two of the devices under test during concurrent testing.
[0287] In some embodiments, the device further includes a data query unit, wherein the data query unit includes:
[0288] A query data acquisition unit, used to acquire query data;
[0289] an associated data locating unit, configured to determine associated data of the query data in the performance test data stored in the columnar format;
[0290] The associated data decompression unit is used to decompress the column where the associated data is located to obtain target data.
[0291] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0292] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned performance test method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0293] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0294] See also Figure 4 , Figure 4 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0295] The processor 401 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0296] The memory 402 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 402 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 402, and the processor 401 calls and executes a performance testing method of the embodiment of this application;
[0297] Input / output interface 403, used to implement information input and output;
[0298] Communication interface 404, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);
[0299] Bus 405 , which transmits information between various components of the device (e.g., processor 401 , memory 402 , input / output interface 403 , and communication interface 404 );
[0300] The processor 401 , the memory 402 , the input / output interface 403 and the communication interface 404 are connected to each other in communication within the device via the bus 405 .
[0301] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned performance testing method is implemented.
[0302] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0303] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0304] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0305] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0306] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0307] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0308] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0309] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0310] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0311] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0312] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0313] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0314] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A performance testing method, characterized in that: The method comprises the following steps: Storing the performance test data in a columnar format; wherein the performance test data is obtained by performing a performance test on the device under test; The performance test data stored in the column format is read to generate a performance test report of the device under test.
2. A performance testing method according to claim 1, characterized in that: The step of reading the performance test data stored in the columnar format and generating a performance test report of the device under test comprises the following steps: Reading the performance test data stored in the columnar format, and performing a columnar aggregation operation on the read data to obtain an aggregate value; The performance test report of the device under test is generated according to the aggregate value.
3. A performance testing method according to claim 2, characterized in that: The step of reading the performance test data stored in the columnar format and performing a columnar aggregation operation on the read data to obtain an aggregate value includes the following steps: Reading a target column in the performance test data stored in the columnar format; The column aggregation operation is performed on the data obtained by reading the target column to obtain the aggregate value.
4. A performance testing method according to claim 3, characterized in that: The reading of the target column in the performance test data stored in the columnar format comprises the following steps: Using DuckDB to read the target column in the performance test data stored in the columnar format; The step of performing the column aggregation operation on the data obtained by reading the target column to obtain the aggregate value includes the following steps: The column aggregation operation is performed on the data obtained by reading the target column using an SQL statement to obtain the aggregate value.
5. A performance testing method according to claim 2, characterized in that: Generating the performance test report of the device under test according to the aggregate value comprises the following steps: Determine the value of each performance indicator according to the aggregate value; The performance test report of the device under test is generated using the values of the various performance indicators.
6. A performance testing method according to claim 1, characterized in that: The storing of the performance test data in a columnar format comprises at least one of the following steps: If the performance test duration of each of the devices under test is less than a set duration threshold, the performance test data is generated as non-column format data, and after the performance test is completed, the non-column format data is converted into the column format data for storage; Alternatively, if the performance test duration of each of the devices to be tested is equal to or greater than the set duration threshold, the performance test data is generated as data in the columnar format for storage.
7. A performance testing method according to claim 1, characterized in that: The number of the devices to be tested is at least 2, and the storing of the performance test data in a columnar format comprises the following steps: The performance test data in different data formats are stored as Parquet format files.
8. A performance testing method according to claim 1, characterized in that: Before storing the performance test data in a columnar format, the method further includes the following steps: Collecting the performance test data corresponding to one or at least two of the devices under test includes: When one or at least two of the devices under test run a preset test script, status data of one or at least two of the devices under test are collected as the performance test data; or, The performance test data generated by at least two of the devices under test during concurrent testing are collected.
9. A performance testing method according to any one of claims 1 to 8, characterized in that: The method further comprises the following steps: Get query data; determining associated data of the query data in the performance test data stored in the columnar format; The column where the associated data is located is decompressed to obtain the target data.
10. A performance testing device, characterized in that: The device comprises: A data storage unit, used to store performance test data in a column format; wherein the performance test data is obtained by performing performance testing on the device under test; The data reading and report generating unit is used to read the performance test data stored in the column format and then generate a performance test report of the device under test.
11. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements a performance testing method as described in any one of claims 1 to 8 when executing the computer program.
12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a performance testing method according to any one of claims 1 to 8 is implemented.