Data processing
By constructing a log feature matrix and using a type recognition model, the problem of identifying multi-service log data is solved, and the management efficiency of storage devices is improved.
Patent Information
- Application Number
- PCT/IB2025/050214
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2025-01-09
- Publication Date
- 2025-08-07
AI Technical Summary
The prior art is difficult to effectively identify and classify multi-service log data, resulting in inefficient management of storage devices.
By obtaining the multi-source log data of the associated storage device, a log feature matrix is constructed, and its input type identification model is used to obtain the service type information of the sub-log data, and then select the sub-log data of the target service type to form the target log data group.
The type identification of multi-source log data is realized, the type identification efficiency is improved, and data support is provided for subsequent supervision of storage devices.
Smart Images

Figure IB2025050214_07082025_PF_FP_ABST
Abstract
Description
Data processing technology field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular to data processing.
[0002] With the development of the Internet industry, business pressure is also increasing. More and more businesses will share storage devices such as cloud disks and magnetic disks. There are also huge differences between different businesses and different business loads, which brings huge challenges to the performance of storage devices. Due to the wide variety of businesses running on storage devices, it is necessary to distinguish different businesses and different data types when analyzing log data. In related technologies, only the data type of a single business can be identified, and the identification efficiency is low. It is impossible to identify and classify log data for multiple businesses. Therefore, a more effective data processing method is urgently needed to solve the above problems.
[0003] In view of this, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a data processing system, a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in related technologies.
[0004] According to a first aspect of an embodiment of the present specification, a data processing method is provided, comprising: obtaining multi-source log data from an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; constructing a log feature matrix based on the multi-source log data, and inputting the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; and selecting sub-log data corresponding to a target business type from the multi-source log data to form a target log data group according to the business type information corresponding to the sub-log data in the multi-source log data.
[0005] According to a second aspect of an embodiment of the present specification, a data processing device is provided, comprising: an acquisition module configured to acquire multi-source log data from an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; a construction module configured to construct a log feature matrix based on the multi-source log data, and input the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; and a selection module configured to select sub-log data corresponding to a target business type in the multi-source log data to form a target log data group according to the business type information corresponding to the sub-log data in the multi-source log data.
[0006] According to a third aspect of an embodiment of the present specification, a data processing system is provided, comprising a server, a data processing node, and a feature processing node; the server is configured to determine a storage device in response to a data processing request, invoke a block storage service to obtain multi-source log data associated with the storage device, and forward the multi-source log data to the data processing node via an object storage service; the data processing node is configured to construct a log based on the multi-source log data. Feature matrix, and sending the log feature matrix to the feature processing node; the feature processing node is used to input the log feature matrix into the type recognition model, obtain the business type information corresponding to the sub-log data in the multi-source log data and feed it back to the server; the server is also used to select the sub-log data corresponding to the target business type in the multi-source log data to form a target log data group according to the business type information corresponding to the sub-log data in the multi-source log data.
[0007] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned data processing method.
[0008] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned data processing method are implemented.
[0009] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.
[0010] One embodiment of the present specification obtains multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; constructs a log feature matrix based on the multi-source log data, and inputs the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; and selects sub-log data corresponding to a target business type from the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data to form a target log data group for executing a corresponding supervision task for the storage device. This implements type recognition of the multi-source log data of the associated storage device, improving type recognition efficiency. This provides data support for subsequent supervision of the storage device.
[0011] FIG1 is a schematic diagram of a processing process of a data processing method provided by one embodiment of this specification;
[0012] FIG2 is a flow chart of a data processing method provided by one embodiment of this specification;
[0013] FIG3 is a flowchart of a data processing method according to an embodiment of the present disclosure;
[0014] FIG4 is a signal processing diagram of a data processing method provided by one embodiment of this specification;
[0015] FIG5 is a schematic structural diagram of a data processing device provided by one embodiment of this specification;
[0016] FIG6 is a schematic diagram of a data processing system provided by one embodiment of this specification;
[0017] FIG7 is a schematic diagram of a processing process of a data processing system provided by one embodiment of this specification;
[0018] FIG8 is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0019] The following description sets forth numerous specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art may make similar generalizations without departing from the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0020] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "an," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and
[0021] It should be understood that while terms such as "first," "second," and so on may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, "first" could also be referred to as "second," and similarly, "second" could also be referred to as "first," without departing from the scope of one or more embodiments of this specification. Depending on the context, the term "if" as used herein could be interpreted as meaning "when," "when," or "in response to a determination."
[0022] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0023] First, the terms involved in one or more embodiments of this specification are explained.
[0024] Cloud disk: Cloud storage block device, also referred to as cloud disk.
[0025] IO2Matrix: If an IO record lasts N minutes, we get an access sequence for N minutes. Through mathematical methods, statistical methods, and machine learning strategies, we convert the abstract access sequence into a feature matrix, which is the 10-data conversion matrix.
[0026] PB: Protobuf (Protocol Buffers) is a data description language used to describe a lightweight and efficient structured data storage format.
[0027] ICA: Independent Component Analysis, used to find the independent parts that make up the signal, corresponding to high-order statistical analysis.
[0028] Blind Source Separation (BSS) is the process of recovering the independent components of the source signal from the observed signal, without knowing the parameters of the source signal and the transmission channel. This process is also called Independent Component Analysis (ICA).
[0029] IOPS: (Input / Output Operations Per Second) is a measurement method used for performance testing of computer storage devices (such as hard disk drives (HDDs), solid-state drives (SSDs), or storage area networks (SANs)). It can be regarded as the number of reads and writes per second.
[0030] BPS: (Byte per second) refers to the number of bytes transmitted per unit time.
[0031] RAR: (read after read) Read after read, a data reading and writing behavior, refers to two consecutive operations on the same block address, the first operation is a read operation, and the second operation is also a read operation.
[0032] RAW: (read after write) Read after write, a data read and write behavior, refers to two consecutive operations on the same block address, the first operation is a read operation, and the second operation is a write operation.
[0033] WAR: (write after read) Write after read, a data read and write behavior, refers to two consecutive operations on the same block address, the first operation is a write operation, and the second operation is a read operation.
[0034] WAW: (write after write) Write after write, a data read and write behavior, refers to two consecutive operations on the same block address, the first operation is a write operation, and the second operation is also a write operation.
[0035] 10: Input, Output, divided into 10 devices and 10 interfaces.
[0036] Mixed I / O Workload: Mixed I / O workload.
[0037] 10 Record: Log record of 10 disk accesses, including operation type, offset, length, and timestamp. The time is a microsecond timestamp, indicating which blocks were accessed at a certain moment and the type of access. Generally, the action is divided into read / write.
[0038] Object storage: Storing data as objects is a method for describing and processing discrete units, called objects. Each object consists of data (file data), metadata (information describing the data), and a global ID (identifying the file's storage path). Users can access the corresponding object online from anywhere and at any time using the ID.
[0039] FIG1 is a schematic diagram of a processing process of a data processing method provided by an embodiment of the present specification; as shown in FIG1 , a storage device is used to store multi-source log data generated when a target business is running. The target business includes sub-businesses of multiple business types, and the sub-businesses are jointly operated in the form of business combinations. When analyzing the multi-source log data associated with the target business, the multi-source log data associated with the target business running on the storage device is obtained. The multi-source log data includes sub-log data of at least two business types corresponding to the target business. The target business includes different business combinations. Accordingly, the storage device corresponds to a mixed load. A log feature matrix is constructed based on the multi-source log data, and the log feature matrix is input into a type recognition model for type recognition to obtain business type information corresponding to each sub-log data in the multi-source log data. According to the business type information corresponding to the sub-log data in the multi-source log data, the sub-log data corresponding to the target business type is selected from the multi-source log data to form a target log data group for executing the supervision task corresponding to the storage device. The type recognition and data classification of the multi-source log data associated with the target business running on the storage device are realized. Distance, improve the efficiency of type recognition.
[0040] This specification provides a data processing method. This specification also relates to a data processing system, a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, each of which is described in detail in the following embodiments.
[0041] 2 , which shows a flow chart of a data processing method according to an embodiment of the present specification, specifically including the following steps.
[0042] Step 202: Acquire multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types.
[0043] Specifically, a storage device is a storage medium, which can be a block device of different forms, such as a cloud disk, a magnetic disk, a hard disk, and other block devices; the storage device is used to store multi-source log data generated when the target business is running. The target business includes sub-businesses of various business types, and the sub-businesses are jointly operated in a business combination. Accordingly, each sub-business corresponds to a business load; multi-source log data is log data such as I / O data and access data generated during the operation of the target business, which is used to record the operation of the business. Accordingly, the business type is the type of log data.
[0044] Based on this, the target business of the associated storage device is determined, and multi-source log data of the target business is obtained. The multi-source log data includes sub-log data of at least two business types corresponding to the target business. There are multiple business combinations in the running process of the target business.
[0045] In practical applications, the log data corresponding to the target service can be log data representing at least two service characteristics, each corresponding to different loads. When acquiring multi-source log data, log data can be collected uniformly within the storage system. The collected log data can be IO record data corresponding to multiple disks. The collected log data can be divided based on the disks to which the data belongs, obtaining multi-source log data corresponding to each disk (storage device). The log data in this embodiment can be stored on a certain storage medium or on the computer to which the disk is mounted.
[0046] Furthermore, after acquiring multi-source log data associated with the target service running on a storage device, it's impossible to determine the service type corresponding to each sub-log data item because the sub-log data within the multi-source log data is in a pending classification state. Subsequent processing using a type recognition model is required to determine the service type information for each sub-log data item within the multi-source log data item.
[0047] Furthermore, considering that when the storage device is a disk, each disk corresponds to multiple storage blocks. When the log data of each storage block is stored separately, when obtaining the multi-source log data corresponding to the storage device, it is necessary to obtain the log data corresponding to each storage block respectively to form the multi-source log data corresponding to the storage device. The specific implementation is as follows: determine at least two storage blocks associated with the storage device, and determine the sub-businesses running on the at least two storage blocks respectively; obtain the sub-log data associated with the sub-businesses running on the at least two storage blocks respectively; based on the sub-log data associated with the sub-businesses running on the at least two storage blocks, generate the multi-source log data of the associated storage device.
[0048] Specifically, a storage block is a storage structure on a disk. A disk is composed of multiple storage blocks, each of which provides data services. A sub-service is a service that runs on each storage block. Correspondingly, sub-log data is log data generated during the operation of a sub-service.
[0049] Based on this, at least two storage blocks associated with the storage device are determined. Sub-services running on the at least two storage blocks are determined. Log files corresponding to the at least two storage blocks are determined, and sub-log data associated with each sub-service is obtained from the log files corresponding to the at least two storage blocks. Multi-source log data associated with the storage device is assembled based on the obtained sub-log data associated with the sub-services running on the at least two storage blocks.
[0050] For example, the storage device may be a disk, which corresponds to multiple storage blocks. Log data corresponding to the storage device may be stored in a single log file, or each storage block may correspond to a log file storing log data. In the case where each storage block corresponds to a log file storing log data, the log data corresponding to each storage block is obtained separately, and the log data corresponding to each storage block is used to form the multi-source log data corresponding to the disk.
[0051] In summary, based on the acquired sub-log data associated with the sub-services running on at least two storage blocks, multi-source log data associated with the target service is composed, thereby improving the comprehensiveness of data acquisition.
[0052] Furthermore, considering that the log data read from the storage system corresponding to the storage device has a specific data format, it is impossible to directly perform feature extraction on the log data. Therefore, after obtaining the log data to be processed, it is also necessary to convert the format of the log data to be processed. The specific implementation is as follows: obtaining the multi-source log data to be processed in a first data format of the storage device; based on the preset conversion relationship between the first data format and the second data format, converting the multi-source log data to be processed in the first data format into multi-source log data in the second data format.
[0053] Specifically, the first data format refers to the storage format of log data in the log data storage space corresponding to the storage device, that is, the data format; the first data format can be a structured data storage format, such as a PB format; the second data format can be any text format. In order to facilitate the processing of the log data, the log data in the structured data storage format is converted into an analyzable text format.
[0054] Based on this, to-be-processed multi-source log data in a first data format is obtained from an associated storage device. A second data format corresponding to the first data format is determined, and a preset conversion relationship between the first data format and the second data format is determined. Based on the conversion relationship between the first data format and the second data format, the to-be-processed multi-source log data in the first data format is converted into multi-source log data in the second data format.
[0055] Continuing with the above example, when the multi-source log data to be processed is in PB format, the conversion relationship between the PB format and the text format is determined, and the format of the multi-source log data to be processed in PB format is converted into multi-source log data in text format based on the conversion relationship.
[0056] In summary, the to-be-processed multi-source log data in the first data format is converted into the multi-source log data in the second data format, thereby facilitating subsequent processing of the multi-source log data based on the second data format.
[0057] Furthermore, considering that new log data is continuously generated on the storage device during the service period, in order to ensure the timeliness of the data, a time interval can be set to obtain multi-source log data associated with the storage device based on a fixed time interval. The specific implementation is as follows: setting a time interval for the storage device; obtaining multi-source log data associated with the storage device according to the time interval.
[0058] Specifically, the time interval refers to a fixed time interval, which may be a time interval at the minute level, such as a one-minute interval, a ten-minute interval, a thirty-minute interval, or the like.
[0059] Based on this, before acquiring multi-source log data of the associated storage device, a time interval for acquiring multi-source log data is set for the storage device. The multi-source log data of the associated storage device is acquired at the time interval, that is, the multi-source log data of the associated storage device is acquired in a polling manner at the time interval.
[0060] Continuing with the previous example, you can set the multi-source log data acquisition task as a scheduled task. Set the interval to 1 hour, and acquire multi-source log data every hour. Also, acquire new log data from the associated storage devices within that hour.
[0061] In summary, multi-source log data of associated storage devices is obtained according to time intervals to ensure the timeliness of the data. When multi-source log processing is required, the multi-source log data obtained at each time interval can be processed separately.
[0062] Step 204: construct a log feature matrix based on the multi-source log data, and input the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data.
[0063] Specifically, after acquiring the multi-source log data of the associated storage device, a log feature matrix can be constructed based on the multi-source log data, and the log feature matrix can be input into the type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data. The log feature matrix is composed of feature types and feature points, and the log data in the multi-source log data is represented in the form of feature points. The feature type is represented by N, and the feature point is represented by M, forming an M*N log feature matrix. The type recognition model is an unsupervised machine learning model constructed based on the independent component analysis algorithm. The business type information indicates the data type of each sub-log data, and the business type information includes but is not limited to RAR read-after-read, WAW write-after-write, WAR read-after-write, and other types.
[0064] Based on this, after acquiring the multi-source log data from the associated storage devices, a log feature matrix is constructed based on the multi-source log data, representing the multi-source log data as feature vectors. The log feature matrix is then input into a type recognition model for type recognition. This model can obtain the business type information corresponding to each sub-log data item in the multi-source log data, thereby identifying the data type of each sub-log data item in the multi-source log data.
[0065] In practical applications, converting multi-source log data into a log feature matrix can be achieved through an analysis platform. Time series feature extraction is used to extract features from multi-source log data and construct a log feature matrix. This log feature matrix can be constructed using the I02Matrix approach. Given a log data duration of K minutes, an access sequence can be obtained. This abstract access sequence is then converted into a feature matrix using the I02Matrix approach. The specific calculation method is as follows: For a K-minute access sequence, determine the following parameters: IOPS: calculate the number of IOPS per minute; BPS: calculate the total length per minute; WAR: calculate the number of WAR requests for each address in this 10-record segment per minute; RAR: calculate the number of RAR requests for each address in this 10-record segment per minute; WAW: calculate the number of WAW requests for each address in this 10-record segment per minute; RAW: calculate the number of RAW requests for each address in this 10-record segment per minute; sequential: calculate the ratio of sequential address accesses to the total IOPS per minute; RW_ration: calculate the ratio of read operations to the total IOPS per minute; IO_size: calculate the average IO_size per minute, using the length as the average per minute. Combining the seven parameters (WAR, RAR, WAW, RAW, sequential, RW_ration, and IO_size) yields a 7*K matrix. Convert 10 records into a feature matrix.
[0066] Furthermore, considering that the type recognition model is an unsupervised machine learning model, the type recognition model contains multiple units for data processing, and each unit needs to cooperate to complete the processing of the log data. The specific implementation is as follows: the log feature matrix is input into the type recognition model, and the conversion unit in the type recognition model is used to convert the log feature matrix into a log signal; the log signal is input into the type recognition unit in the type recognition model, and the log signal is identified by the coefficient matrix of the type recognition unit. The business type information corresponding to the sub-log data in the multi-source log data is determined according to the identification result, and the type recognition model is output.
[0067] Specifically, the conversion unit is used to convert the log feature matrix into a log signal. The type recognition unit is used to predict the business type information of sub-log data in the multi-source log data based on the log signal. The type recognition unit corresponds to the independent component analysis (ICA) algorithm. The prediction of the business type information of sub-log data in the multi-source log data is regarded as a blind signal separation task. The recognition result is the type recognition model's prediction result for each sub-log data in the multi-source log data.
[0068] Based on this, the log feature matrix is input into the type recognition model and sequentially passes through multiple processing units of the type recognition model, including a conversion unit and a type recognition unit. The conversion unit in the type recognition model converts the log feature matrix into a log signal. The log signal is then input into the type recognition unit in the type recognition model, where it is identified using the coefficient matrix of the type recognition unit to obtain an identification result. Based on the identification result, the business type information corresponding to each sub-log data in the multi-source log data is determined, and a type recognition model is output.
[0069] In practical applications, when determining the business type information corresponding to each sub-log data in multi-source log data based on the independent component analysis algorithm (ICA), the following formulas (1) to (3) can be used: X - AS ⑶
[0070] Where X represents a dataset composed of multi-source log data; m represents the log data in the dataset; n represents the number of where n business types of log data in the data set are represented by n; i represents the log data generated at time i, that is, the time series information of the log data; S represents the matrix to be calculated corresponding to the multi-source log data, that is, the recognition result matrix for the multi-source log data; R represents the data source; and A represents the unknown mixing matrix, that is, the coefficient matrix.
[0071] Continuing with the previous example, since the independent component analysis (ICA) algorithm processes signals, when processing multi-source log data based on a type recognition model, the log feature matrix is transformed into a signal using the hybrid payload separation technique within ICA. The log signal is then identified based on the coefficient matrix to obtain the business type information corresponding to each sub-log data item in the multi-source log data.
[0072] In summary, the conversion unit and the type recognition unit in the type recognition model cooperate to complete the prediction of business type information of multi-source log data, thereby improving the prediction accuracy of the business type information and improving the prediction efficiency.
[0073] Furthermore, considering that the acquired multi-source log data is 10 Record data, feature extraction for the multi-source log data needs to be performed in multiple dimensions. The specific implementation is as follows: extracting a storage performance data sequence corresponding to the storage performance dimension from the multi-source log data, and extracting a storage block data sequence corresponding to the storage block dimension from the multi-source log data; constructing a data sequence corresponding to the multi-source log data based on the storage performance data sequence and the storage block data sequence; and converting the data sequence into a log feature matrix.
[0074] Specifically, storage performance dimensions include but are not limited to IOPS sequence (number of reads and writes per second), BPS sequence (number of bytes transferred per unit time), and latency sequence (the difference between the start and end time of each 10 record, where the average latency of multiple 10s can be calculated within one second). The storage block dimension can also be understood as a block-level dimension. In this embodiment, the block-level dimension can be a sequential stream access sequence, RAR, WAR, WAW access sequence, read / write distribution sequence, and 10 size sequence. Accordingly, the storage performance data sequence includes information such as IOPS and BPS. The storage block data sequence includes information about each access sequence. A sequential stream refers to log accesses to consecutive addresses. For example, if a consecutive 10 requests access address locations 1, 2, 3, 4, 5, 7, and 9, then the logs accessing 1, 2, 3, 4, and 5 are sequential streams, and statistics are calculated based on the proportion of this sequential stream.
[0075] Based on this, we extract storage performance data sequences corresponding to storage performance dimensions from multi-source log data, and extract storage block data sequences corresponding to storage block dimensions from multi-source log data. Based on the storage performance data sequences and storage block data sequences, we construct data sequences corresponding to the multi-source log data. We then convert the data sequences into log feature matrices.
[0076] Continuing with the previous example, when extracting features from multi-source log data, we can extract access sequences at both the statistical dimension (storage performance dimension) and the storage block layer dimension (storage block dimension). We can obtain IOPS sequences, BPS sequences, latency sequences, sequential stream access sequences, RAR, WAR, and WAW access sequences, read / write distribution sequences, and 10-size sequences, thereby constructing data sequences for multi-source log data.
[0077] In summary, feature extraction is performed on multi-source log data in the storage performance dimension and storage block dimension respectively, and the multi-source log data is converted into a log feature matrix. In the process of predicting log data by the type recognition model, the information of the storage performance dimension and the storage block dimension are integrated, thereby improving the business efficiency of the type recognition model output. The degree of match between the type information and the actual business type information of the log data.
[0078] Furthermore, considering that the type recognition model is an unsupervised model, in order to ensure the prediction accuracy of the type recognition model, after the type recognition model is constructed, it is necessary to train the type recognition model based on the log data so that the type recognition model can achieve a prediction accuracy that meets the usage conditions. The specific implementation is as follows: determine a log data set containing sub-log data of at least two business types; train the type recognition model to be trained based on the log data in the log data set until a type recognition model that meets the training stop conditions is obtained.
[0079] Specifically, the log data set includes log data generated by target services running on storage devices in a real-world environment. The log data in the log data set corresponds to at least two service types, where the service type is the data type, including but not limited to RAR read-after-read, WAW write-after-write, and WAR read-after-write. The training stop condition can be when the type recognition model achieves a preset accuracy level when predicting the type of the log data. For example, if the prediction accuracy of the log data type reaches 90% for 100 log data items, the training stop condition can also be when a preset number of training rounds is reached. The training stop condition can also be when the training time is reached.
[0080] Based on this, the log data file corresponding to the storage device is determined. Sub-log data containing at least two business types is obtained from the log data corresponding to the storage device. After data cleansing, a log data set is constructed. The log data in the log data set is input into the type recognition model to be trained. The type recognition model to be trained is trained based on the log data in the log data set until the type recognition model to be trained meets the training stop condition. This completes the training of the type recognition model.
[0081] Continuing with the above example, when training a type recognition model, log data can be obtained from the log data file corresponding to the storage device. After performing data cleaning operations such as removing abnormal data on the obtained log data, a log data set is constructed for training the type recognition model. The type recognition model is trained based on the log data in the log data set until the training type recognition model meets the training stop condition. A trained type recognition model is obtained. The type recognition model can be used to predict business type information in actual log data.
[0082] In summary, the type recognition model to be trained is trained based on log data to obtain a type recognition model that meets the training stop condition, thereby improving the prediction accuracy of the type recognition model.
[0083] Step 206: According to the business type information corresponding to the sub-log data in the multi-source log data, sub-log data corresponding to the target business type are selected from the multi-source log data to form a target log data group.
[0084] Specifically, after constructing the log feature matrix based on the multi-source log data and inputting the log feature matrix into the type recognition model to obtain the business type information corresponding to the sub-log data in the multi-source log data, the sub-log data corresponding to the target business type can be selected from the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data to form a target log data group for executing the supervision task corresponding to the storage device, wherein the target business type is the target type determined based on the business type information of each sub-log data. Any business type information The corresponding types can all be used as target business types; accordingly, the business type information of each log data in the target log data group is the same.
[0085] Based on this, after constructing a log feature matrix based on multi-source log data and inputting the log feature matrix into the type recognition model to obtain the business type information corresponding to the sub-log data in the multi-source log data, the target business type is determined based on the business type information of each sub-log data in the multi-source log data. Based on the business type information corresponding to the sub-log data in the multi-source log data, sub-log data corresponding to the target business type is selected from the multi-source log data, and the sub-log data are combined to form a target log data group. The target log data is used to execute the corresponding supervision task for the storage device.
[0086] In practical applications, any business type can be used as a target business type, and a target log data group corresponding to the target business type can be constructed. The business type information corresponding to the sub-log data in the multi-source log data is used to determine the type of data corresponding to the multi-source log data. Based on the business type information of the sub-log data, the multi-source log data can be divided into log data groups corresponding to each type. This allows each type of sub-log data to be grouped together.
[0087] Furthermore, considering that multi-source log data corresponds to multiple business loads, and each business load corresponds to a different business type, in order to facilitate the supervision of multiple business loads, a type label can be added to each sub-log data based on the business type information corresponding to the sub-log data in the multi-source log data. The specific implementation is as follows: according to the business type information corresponding to the sub-log data in the multi-source log data, a type label is added to the sub-log data in the multi-source log data; according to the type label corresponding to each sub-log data, the multi-source log data is divided into at least one target log data group, and the at least one target log data group is used to execute the supervision task corresponding to the storage device.
[0088] Specifically, a type tag refers to tag information used to indicate the read and write type of sub-log data. This tag information includes, but is not limited to, RAR read-after-read, WAW write-after-write, and WAR read-after-write. A supervision task is used to monitor and manage at least one service load corresponding to a target service, including but not limited to stress testing and traffic playback of the service load.
[0089] Based on this, the business type information corresponding to each sub-log data in the multi-source log data is determined, and type tags are added to the sub-log data in the multi-source log data according to the business type information corresponding to each sub-log data. The multi-source log data is divided into at least one target log data group according to the type tag corresponding to each sub-log data. Each target log data group stores multi-source log data with the same type tag. The at least one target log data group is used to execute a supervision task corresponding to the storage device, thereby enabling monitoring and management of at least one business load corresponding to the storage device.
[0090] Continuing with the above example, if the multi-source log data contains 10 sub-log data, and the service type information of each sub-log data is read-after-read, write-after-write, read-after-write, read-after-read, write-after-write, read-after-write, read-after-read, write-after-write, read-after-write, and write-after-read, read-write tags are added to each sub-log data, and sub-log data with the same read-write tags are stored in a log data group.
[0091] In summary, by dividing the multi-source log data into at least one target log data group according to the type label corresponding to each sub-log data, the sub-log data with the same type label are stored in one data group. In order to facilitate the execution of subsequent regulatory tasks, it also realizes standardized data storage.
[0092] Furthermore, during the operation of the storage device, the storage performance and traffic conditions of the storage device can be detected at any time based on the target log data group. The specific implementation is as follows: receiving a device detection request associated with the supervision task, wherein the supervision task is used to monitor and manage the storage device; determining the target log data included in the target log data group based on the device detection request; and determining the detection information of the storage device based on the target log data as the task execution result of the supervision task.
[0093] Specifically, a device detection request is a computer instruction submitted to a storage device for detecting the storage device. Correspondingly, the detection information is the detection result of the storage device obtained based on the analysis of the target log data, which is used to indicate information such as the storage performance and traffic status of the storage device.
[0094] Based on this, upon receiving a device detection request associated with a supervisory task, the target log data group is determined based on the time information and data type information carried in the device detection request. The target log data is extracted from the target log data group. Based on an analysis of the target log data combined with the time information, the storage device detection information is determined and used as the supervisory task execution result.
[0095] Continuing with the previous example, after receiving a disk test request, the target log data from the disk is retrieved based on the time and data type information carried in the test request. Based on the volume of the target log data, performance stress testing and traffic flow testing can be performed on the disk.
[0096] In summary, by testing a storage device based on a device detection request and obtaining a detection result, storage performance of the storage device can be tested and usage of the storage device can be optimized.
[0097] One embodiment of this specification obtains multi-source log data associated with a target service on an associated storage device, where the multi-source log data includes sub-log data of at least two service types. A log feature matrix is constructed based on the multi-source log data and input into a type recognition model to obtain service type information corresponding to the sub-log data in the multi-source log data. Based on the service type information corresponding to the sub-log data in the multi-source log data, sub-log data corresponding to the target service type is selected from the multi-source log data to form a target log data group for executing a corresponding storage device supervision task. This method achieves type recognition of the multi-source log data associated with the storage device, improving type recognition efficiency and providing data support for subsequent storage device supervision.
[0098] The following, in conjunction with FIG3 , further illustrates the data processing method provided in this specification, using its application in cloud disk read and write data processing as an example. FIG3 illustrates a flowchart of a data processing method provided in one embodiment of this specification, specifically including the following steps.
[0099] Step 302: Collect multi-source log data associated with storage devices in the storage system.
[0100] In practical applications, storage devices can be cloud disks, magnetic disks, and other devices. When collecting data, you can set up a scheduled task to collect multi-source log data corresponding to the storage device at fixed intervals.
[0101] Step 304: Determine at least two storage blocks associated with the storage device, and determine the multi-source log data corresponding to each storage block in the multi-source log data.
[0102] Step 306: Perform format conversion on the multi-source log data corresponding to each storage block.
[0103] The acquired log data is in PB format, which can be parsed into any parseable text format. It should be noted that since a storage block can correspond to multiple business loads, the log data corresponding to each storage block can be treated as multi-source log data for data classification.
[0104] Step 308: For the multi-source log data corresponding to each storage block, extract the storage performance data sequence corresponding to the statistical dimension from the multi-source log data, and extract the storage block data sequence corresponding to the storage block dimension from the multi-source log data.
[0105] Step 310: Construct a data sequence corresponding to the multi-source log data based on the storage performance data sequence and the storage block data sequence.
[0106] Feature extraction is performed on the log data of each storage block, including statistical features and storage block features. Statistical features include IOPS (input and output per second) and BPS (bytes transferred per unit time). Storage block feature extraction includes sequential stream sequences, access sequences such as RAR, WAR, and WAW, read and write distribution sequences, and size sequences. An N*M feature matrix is constructed, representing N features and M feature points.
[0107] Step 312: Input the data sequence into the type recognition model for recognition, and obtain type information corresponding to the log data in the multi-source log data.
[0108] A type recognition model based on ICA (independent component analysis) is constructed, and the type recognition task of multi-source log data is regarded as a blind signal separation task to realize the type recognition of individual log data in multi-source log data.
[0109] In practice, as shown in Figure 4, the storage block system uses the acquired multi-source log data as the raw signal. This raw signal is a mixed signal with multiple payloads. The signal separation system provides a type recognition model based on ICA (Independent Component Analysis) to separate the raw signals.
[0110] Step 314: According to the type information corresponding to the log data in the multi-source log data, log data corresponding to the target business type is selected from the multi-source log data to form a target log data group.
[0111] After determining the type information corresponding to the log data in the multi-source log data, the type information is used as a label for the log data, and the multi-source log data is separated and stored according to the type information.
[0112] Step 316: Supervise the storage device based on the type information corresponding to the log data in the multi-source log data.
[0113] In practical applications, determining the type information corresponding to the log data in the storage device's multi-source log data allows identification and separation of mixed service loads. This also allows the impact of different services on overall performance and the signal proportion at different times to be determined. This type information can be used to perform stress testing and traffic replay on the storage device.
[0114] Corresponding to the above method embodiments, this specification also provides a data processing device embodiment. FIG5 shows a schematic structural diagram of a data processing device provided in one embodiment of this specification. As shown in FIG5 , the device includes an acquisition module 502, a construction module 504, and a selection module 506.
[0115] The acquisition module 502 is configured to acquire multi-source log data associated with a target service running on a storage device, wherein the multi-source log data includes sub-log data of at least two service types corresponding to the target service.
[0116] The construction module 504 is configured to construct a log feature matrix based on the multi-source log data, and input the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data.
[0117] The selection module 506 is configured to select sub-log data corresponding to a target business type from the multi-source log data according to business type information corresponding to the sub-log data in the multi-source log data to form a target log data group.
[0118] In an optional embodiment, the construction module 504 is further configured to: input the log feature matrix into a type recognition model, and use a conversion unit in the type recognition model to convert the log feature matrix into a log signal; input the log signal into a type recognition unit in the type recognition model, identify the log signal using a coefficient matrix of the type recognition unit, determine the business type information corresponding to the sub-log data in the multi-source log data based on the identification result, and output the type recognition model.
[0119] In an optional embodiment, the construction module 504 is further configured to: extract a storage performance data sequence corresponding to a storage performance dimension from the multi-source log data, and extract a storage block data sequence corresponding to a storage block dimension from the multi-source log data; construct a data sequence corresponding to the multi-source log data based on the storage performance data sequence and the storage block data sequence; and convert the data sequence into a log feature matrix.
[0120] In an optional embodiment, the acquisition module 502 is further configured to: determine at least two storage blocks associated with the storage device, and determine sub-services respectively running on the at least two storage blocks; respectively acquire sub-log data associated with the sub-services running on the at least two storage blocks; and generate multi-source log data associated with the storage device based on the sub-log data associated with the sub-services running on the at least two storage blocks.
[0121] In an optional embodiment, the acquisition module 502 is further configured to: acquire the to-be-processed multi-source log data in a first data format from the management storage device; and convert the to-be-processed multi-source log data in the first data format into multi-source log data in a second data format based on a preset conversion relationship between the first data format and the second data format.
[0122] In an optional embodiment, the acquisition module 502 is further configured to: set a time interval for the storage device; and the acquiring of multi-source log data associated with the storage device includes: acquiring the multi-source log data associated with the storage device according to the time interval.
[0123] In an optional embodiment, the selection module 506 is further configured to: According to the business type information corresponding to the neutron log data, a type label is added to the neutron log data in the multi-source log data; according to the type label corresponding to each sub-log data, the multi-source log data is divided into at least one target log data group, and the at least one target log data group is used to execute the supervision task corresponding to the storage device.
[0124] In an optional embodiment, the construction module 504 is further configured to: determine a log data set containing sub-log data of at least two business types; and train the type recognition model to be trained based on the log data in the log data set until a type recognition model that meets the training stop condition is obtained.
[0125] In an optional embodiment, the selection module 506 is further configured to: receive a device detection request associated with the supervision task, wherein the supervision task is used to monitor and manage the storage device; determine the target log data included in the target log data group based on the device detection request; and determine the detection information of the storage device based on the target log data as the task execution result of the supervision task.
[0126] In summary, one embodiment of this specification obtains multi-source log data from an associated storage device, where the multi-source log data includes sub-log data of at least two business types; constructs a log feature matrix based on the multi-source log data, and inputs the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; and, based on the business type information corresponding to the sub-log data in the multi-source log data, selects sub-log data corresponding to a target business type from the multi-source log data to form a target log data group for executing a corresponding storage device supervision task. This achieves type recognition of the multi-source log data associated with the storage device, improving type recognition efficiency and providing data support for subsequent storage device supervision.
[0127] The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the aforementioned data processing method are based on the same concept. For details not described in detail in the technical solution of the data processing device, please refer to the description of the technical solution of the aforementioned data processing method.
[0128] 6 shows a schematic diagram of a data processing system according to an embodiment of this specification. The data processing system 600 includes a server 610, a data processing node 620, and a feature processing node 630.
[0129] The server 610 is configured to determine a storage device in response to a data processing request, invoke a block storage service to obtain multi-source log data associated with the storage device, and forward the multi-source log data to the data processing node 620 via an object storage service.
[0130] The data processing node 620 is configured to construct a log feature matrix based on the multi-source log data, and send the log feature matrix to the feature processing node 630.
[0131] The feature processing node 630 is configured to input the log feature matrix into a type recognition model, obtain business type information corresponding to sub-log data in the multi-source log data, and feed the information back to the server 610.
[0132] The server 610 is further configured to select sub-log data corresponding to a target business type from the multi-source log data according to business type information corresponding to the sub-log data in the multi-source log data to form a target log data group.
[0133] Based on this, the data processing system for performing type recognition on multi-source log data includes a server 610, a data processing node 620, and a feature processing node 630. The data processing node 620 has a feature extraction function, which is used to extract features from multi-source log data; the feature processing node 630 has a feature processing function, which is used to input the log feature matrix into the type recognition model for type recognition.
[0134] In practical applications, when performing type identification on multi-source log data corresponding to storage devices, the process flow can be seen in Figure 7. The server dispatches the block storage service to obtain multi-source log data corresponding to various businesses running on cloud disks, disks, and other devices, and temporarily stores this data in object storage. After obtaining the multi-source log data, the analysis service extracts features from the multi-source log data. The resulting feature sequence is then stored in the database. This feature sequence is processed using a type identification model built using the independent component analysis algorithm to determine the type of each log data item in the multi-source log data, thereby assigning a label to each log data item. This allows for identification and separation of mixed workloads when cloud disks and disks correspond to a variety of business combinations.
[0135] In summary, by integrating blind source separation technology and analyzing multi-source log data, mixed business types can be identified and separated. Based on the identification and separation results, performance stress testing and traffic playback can be performed on cloud disks and disks, enabling oversight of these systems.
[0136] Furthermore, considering that storage devices continuously generate log data while providing services, and that various services are running on the storage devices, different traffic flows may be generated at different times due to the influence of time and service requirements. Therefore, when acquiring multi-source log data, it is also necessary to consider the time factor and acquire data based on preset time intervals. In a specific implementation, the server 610 is further configured to call the block storage service to acquire, based on preset time intervals, to-be-processed multi-source log data in a third data format from the associated storage device; and convert the to-be-processed multi-source log data in the third data format into multi-source log data in the fourth data format based on a preset conversion relationship between the third data format and the fourth data format.
[0137] The third data format is a storage format of multi-source log data in object storage, such as a PB format. Correspondingly, the fourth data format is a text-type data format obtained by performing format conversion on the basis of the third data format. The fourth data format can be set according to actual needs and is not limited in this embodiment.
[0138] Continuing with the previous example and referring to Figure 7, we can set up a scheduled task to poll log data stored on object storage to obtain multi-source log data from disks and cloud drives. The log data obtained from object storage is in petabyte format and needs to be parsed into a parseable text format to facilitate subsequent feature extraction and type identification.
[0139] The server 610 is further configured to input the log feature matrix into a type recognition model, convert the log feature matrix into a log signal using a conversion unit in the type recognition model, input the log signal into a type recognition unit in the type recognition model, identify the log signal using a coefficient matrix of the type recognition unit, and determine business type information corresponding to the sub-log data in the multi-source log data based on the identification result. And output the type recognition model.
[0140] The server 610 is further configured to extract a storage performance data sequence corresponding to a storage performance dimension from the multi-source log data, and extract a storage block data sequence corresponding to a storage block dimension from the multi-source log data; construct a data sequence corresponding to the multi-source log data based on the storage performance data sequence and the storage block data sequence; and convert the data sequence into a log feature matrix.
[0141] The server 610 is further configured to determine at least two storage blocks associated with the storage device, and determine sub-businesses respectively running on the at least two storage blocks; obtain sub-log data associated with the sub-businesses running on the at least two storage blocks; and generate multi-source log data associated with the target business based on the sub-log data associated with the sub-businesses running on the at least two storage blocks.
[0142] The server 610 is further configured to determine a log data set containing sub-log data of at least two business types; and train a type recognition model to be trained based on the log data in the log data set until a type recognition model that meets a training stop condition is obtained.
[0143] The server 610 is further configured to receive a device detection request associated with the supervisory task, wherein the supervisory task is configured to monitor and manage the storage device; determine target log data included in the target log data group based on the device detection request; and determine detection information of the storage device based on the target log data as a task execution result of the supervisory task.
[0144] One embodiment of this specification obtains multi-source log data from associated storage devices, where the multi-source log data includes sub-log data of at least two business types. A log feature matrix is constructed based on the multi-source log data and input into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data. Based on the business type information corresponding to the sub-log data in the multi-source log data, sub-log data corresponding to a target business type is selected from the multi-source log data to form a target log data group for executing a corresponding storage device supervision task. This method achieves type recognition of the multi-source log data associated with the storage device, improving type recognition efficiency and providing data support for subsequent storage device supervision.
[0145] The above is a schematic diagram of a data processing system according to this embodiment. It should be noted that the technical solution of this data processing system and the technical solution of the aforementioned data processing method are based on the same concept. For details not described in detail in the technical solution of the data processing system, please refer to the description of the technical solution of the aforementioned data processing method.
[0146] FIG8 shows a block diagram of a computing device 800 according to one embodiment of this specification. Components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830. A database 850 is used to store data.
[0147] The computing device 800 also includes an access device 840 that enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), and the like. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0148] In one embodiment of the present specification, the aforementioned components of computing device 800 and other components not shown in FIG. 8 may also be connected to one another, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG. 8 is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0149] Computing device 800 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 800 can also be a mobile or stationary server.
[0150] The processor 820 is configured to execute the following computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by the processor.
[0151] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the aforementioned data processing method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned data processing method.
[0152] An embodiment of this specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.
[0153] The above is an illustrative embodiment of a computer-readable storage medium. It should be noted that the technical solution of this storage medium and the technical solution of the aforementioned data processing method share the same concept. For details not described in detail in the technical solution of the storage medium, refer to the description of the technical solution of the aforementioned data processing method.
[0154] An embodiment of this specification further provides a computer program product, including a computer program or instructions, which implements the steps of the above-mentioned data processing method when executed by a processor.
[0155] The above is an illustrative embodiment of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the aforementioned data processing method share the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the aforementioned data processing method.
[0156] While the foregoing description describes certain embodiments of the present disclosure, other embodiments are within the scope of the following claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0157] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0158] It should be noted that, for ease of description, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of this specification are not limited to the order of the actions described, as certain steps may be performed in a different order or simultaneously, depending on the embodiments of this specification. Furthermore, those skilled in the art should also be aware that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily required for the embodiments of this specification.
[0159] In the above embodiments, the description of each embodiment is given with emphasis. For parts not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0160] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of the embodiments described herein. These embodiments are selected and described in detail herein to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
Claims 1. A data processing method, comprising: Acquire multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; construct a log feature matrix based on the multi-source log data, and input the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; and select sub-log data corresponding to a target business type from the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data to form a target log data group.
2. The method according to claim 1, wherein inputting the log feature matrix into a type recognition model to obtain business type information corresponding to sub-log data in the multi-source log data comprises: The log feature matrix is input into a type recognition model, and the log feature matrix is converted into a log signal using a conversion unit in the type recognition model; the log signal is input into a type recognition unit in the type recognition model, and the log signal is identified by a coefficient matrix of the type recognition unit. The business type information corresponding to the sub-log data in the multi-source log data is determined according to the identification result, and the type recognition model is output.
3. The method according to claim 1, wherein constructing a log feature matrix based on the multi-source log data comprises: Extracting a storage performance data sequence corresponding to a storage performance dimension from the multi-source log data, and extracting a storage block data sequence corresponding to a storage block dimension from the multi-source log data; Based on the storage performance data sequence and the storage block data sequence, a data sequence corresponding to the multi-source log data is constructed; and the data sequence is converted into a log feature matrix.
4. The method according to claim 1, wherein obtaining multi-source log data of an associated storage device comprises: Determining at least two storage blocks associated with the storage device, and determining sub-services respectively running on the at least two storage blocks; Sub-log data associated with sub-services running on at least two storage blocks are respectively obtained; and multi-source log data associated with the storage device is generated based on the sub-log data associated with sub-services running on at least two storage blocks.
5. The method according to claim 1, wherein obtaining multi-source log data of an associated storage device comprises: Acquire to-be-processed multi-source log data in a first data format associated with a storage device; Based on a preset conversion relationship between the first data format and the second data format, the to-be-processed multi-source log data in the first data format is converted into multi-source log data in the second data format.
6. The method according to claim 1, before acquiring multi-source log data of associated storage devices, further comprising: Setting a time interval for the storage device; The acquiring of multi-source log data associated with the storage device includes: acquiring the multi-source log data associated with the storage device according to the time interval.
7. The method according to claim 1, wherein the selecting, according to the business type information corresponding to the sub-log data in the multi-source log data, the sub-log data corresponding to the target business type in the multi-source log data to form a target log data group comprises: Adding a type tag to the sub-log data in the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data; The multi-source log data are divided into at least one target log data group according to a type label corresponding to each sub-log data, and the at least one target log data group is used to execute a supervision task corresponding to the storage device.
8. The method according to claim 1, wherein the training of the type recognition model comprises: Determining a log data set including sub-log data of at least two business types; The type recognition model to be trained is trained based on the log data in the log data set until a type recognition model that meets the training stop condition is obtained.
9. The method according to claim 7, further comprising: after selecting sub-log data corresponding to a target business type from the multi-source log data to form a target log data group; Receive a device detection request associated with the supervision task, wherein the supervision task is used to monitor and manage the storage device; determine target log data included in the target log data group based on the device detection request; and determine detection information of the storage device based on the target log data as a task execution result of the supervision task.
10. A data processing system, comprising a server, a data processing node, and a feature processing node; the server, configured to determine a storage device in response to a data processing request, invoke a block storage service to obtain multi-source log data associated with the storage device, and forward the multi-source log data to the data processing node via an object storage service; the data processing node, configured to construct a log feature matrix based on the multi-source log data, and send the log feature matrix to the feature processing node; the feature processing node, configured to input the log feature matrix into a type recognition model, obtain business type information corresponding to sub-log data in the multi-source log data, and feed the result back to the server; the server, further configured to select sub-log data corresponding to a target business type from the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data to form a target log data group.
11. The data processing system according to claim 10, wherein the server is further configured to call the block storage service to obtain the to-be-processed multi-source log data in the third data format of the associated storage device based on a preset time interval; and convert the to-be-processed multi-source log data in the third data format into the to-be-processed multi-source log data based on a preset conversion relationship between the third data format and the fourth data format. The log data is converted into multi-source log data in a fourth data format.
12. A data processing device, comprising: An acquisition module is configured to acquire multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; a construction module is configured to construct a log feature matrix based on the multi-source log data, and input the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; and a selection module is configured to select sub-log data corresponding to a target business type in the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data to form a target log data group.
13. A computing device, comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the data processing method according to any one of claims 1 to 9 are implemented.
14. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the data processing method according to any one of claims 1 to 9.
15. A computer program product, comprising a computer program or instructions, which, when executed by a processor, implements the steps of the data processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Time delay performance detection method, device and equipment for storage system
CN112000543A
Parameter configuration method and system of storage system and related device
CN112130759A
Method and apparatus for tracking performance of storage system
CN113138903A
Service verification method and device
CN113157911A
Method for predicting random contributors
CN115605811A