Data processing method, system and device

By constructing a log feature matrix and using the type identification model to identify the service types of multi-source log data, the problem of inefficient type identification of storage devices under multi-service load is solved, the recognition efficiency is improved and the supervision of storage devices is supported.

CN120429436APending Publication Date: 2025-08-05HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410133398.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the prior art, storage devices cannot effectively identify and classify multi-source log data when facing multi-service loads, resulting in inefficient identification.

Method used

By obtaining the multi-source log data of the associated storage device, a log feature matrix is constructed, and inputting its input type identification model to identify the service type information of each sub-log data, and then selecting the sub-log data group of the target service type.

Benefits of technology

The type identification of multi-source log data is realized, the type identification efficiency is improved, and data support is provided for subsequent supervision of storage devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429436A_ABST
    Figure CN120429436A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, system and device.The data processing method comprises the steps that multi-source log data of associated storage equipment is obtained, and the multi-source log data comprises sub-log data of at least two service types; constructing a log feature matrix based on the multi-source log data, and inputting the log feature matrix into a type recognition model to obtain service type information corresponding to sub-log data in the multi-source log data; and according to the business type information corresponding to the sub-log data in the multi-source log data, selecting the sub-log data corresponding to a target business type from the multi-source log data to form a target log data group. The type identification of the multi-source log data associated with the storage device is realized, and the type identification efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of computer technology, and more particularly to a data processing method, system, and device. Background Art

[0002] With the development of the internet industry, business pressures are also increasing, and more and more businesses will share storage devices such as cloud drives and magnetic disks. Different businesses and different business loads also have huge differences, which poses a huge challenge to the performance of storage devices. Due to the wide variety of businesses running on storage devices, it is necessary to distinguish different businesses and different data types when analyzing log data. The existing technology can only identify the data type of a single business, and the identification efficiency is low. It is impossible to identify and classify log data for multiple businesses. Therefore, a more effective data processing method is urgently needed to solve the above problems. Summary of the Invention

[0003] In view of this, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a data processing system, a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0004] According to a first aspect of an embodiment of this specification, there is provided a data processing method, including:

[0005] Acquire multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types;

[0006] Constructing a log feature matrix based on the multi-source log data, and inputting the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data;

[0007] According to the business type information corresponding to the sub-log data in the multi-source log data, sub-log data corresponding to the target business type are selected from the multi-source log data to form a target log data group.

[0008] According to a second aspect of the embodiments of this specification, a data processing device is provided, including:

[0009] an acquisition module configured to acquire multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types;

[0010] A construction module is configured to construct a log feature matrix based on the multi-source log data, and input the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data;

[0011] The selection module is configured to select sub-log data corresponding to a target business type from the multi-source log data to form a target log data group according to business type information corresponding to the sub-log data in the multi-source log data.

[0012] According to a third aspect of an embodiment of this specification, there is provided a data processing system, including a server, a data processing node, and a feature processing node;

[0013] The server is configured to determine a storage device in response to a data processing request, call a block storage service to obtain multi-source log data associated with the storage device, and forward the multi-source log data to the data processing node via an object storage service;

[0014] The data processing node is configured to construct a log feature matrix based on the multi-source log data and send the log feature matrix to the feature processing node;

[0015] The feature processing node is configured to input the log feature matrix into a type recognition model, obtain business type information corresponding to sub-log data in the multi-source log data, and feed the information back to the server;

[0016] The server is further configured to select sub-log data corresponding to a target business type from the multi-source log data to form a target log data group according to business type information corresponding to the sub-log data in the multi-source log data.

[0017] According to a fourth aspect of the embodiments of this specification, a computing device is provided, including:

[0018] memory and processor;

[0019] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned data processing method are implemented.

[0020] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned data processing method are implemented.

[0021] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program or instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.

[0022] One embodiment of the present specification obtains multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; constructs a log feature matrix based on the multi-source log data, and inputs the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; according to the business type information corresponding to the sub-log data in the multi-source log data, selects sub-log data of the corresponding target business type from the multi-source log data to form a target log data group for executing the corresponding supervision task of the storage device. This realizes the type recognition of the multi-source log data of the associated storage device and improves the efficiency of type recognition. This provides data support for the subsequent supervision of the storage device. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a schematic diagram of a processing process of a data processing method provided by an embodiment of this specification;

[0024] Figure 2 is a flow chart of a data processing method provided by one embodiment of this specification;

[0025] Figure 3 This is a flowchart of a data processing method provided by one embodiment of this specification;

[0026] Figure 4 This is a signal processing diagram of a data processing method provided by one embodiment of this specification;

[0027] Figure 5 This is a schematic diagram of the structure of a data processing device provided by one embodiment of this specification;

[0028] Figure 6 is a schematic diagram of a data processing system provided by one embodiment of this specification;

[0029] Figure 7 This is a schematic diagram of a processing process of a data processing system provided by an embodiment of this specification;

[0030] Figure 8 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION

[0031] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0032] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0033] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0034] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0035] First, the terms involved in one or more embodiments of this specification are explained.

[0036] Cloud disk: cloud storage block device, also referred to as cloud disk.

[0037] IO2Matrix: If an IO record lasts N minutes, we obtain an access sequence for N minutes. Using mathematical methods, statistical methods, and machine learning strategies, we convert the abstract access sequence into a feature matrix, which is the IO data conversion matrix.

[0038] PB: Protobuf (Protocol Buffers) is a data description language used to describe a lightweight and efficient structured data storage format.

[0039] ICA: Independent Component Analysis, used to find the independent parts that make up the signal, corresponding to high-order statistical analysis.

[0040] Blind Source Separation (BSS) is the process of recovering the independent components of a source signal from the observed signal alone, without knowing the source signal or the transmission channel parameters. This process is also known as Independent Component Analysis (ICA).

[0041] IOPS: (Input / Output Operations Per Second) is a measurement method used to test the performance of computer storage devices such as hard disk drives (HDDs), solid-state drives (SSDs), or storage area networks (SANs). It can be regarded as the number of reads and writes per second.

[0042] BPS: (Byte per second) refers to the number of bytes transmitted per unit time.

[0043] RAR: (read after read) read after read, a data reading and writing behavior, refers to two consecutive operations on the same block address, the first operation is a read operation, and the second operation is also a read operation.

[0044] RAW: (read after write) read after write, a data reading and writing behavior, refers to two consecutive operations on the same block address, the first operation is a read operation, and the second operation is a write operation.

[0045] WAR: (write after read) write after read, a data read and write behavior, refers to two consecutive operations on the same block address, the first operation is a write operation, and the second operation is a read operation.

[0046] WAW: (write after write) write after write, a data read and write behavior, refers to two consecutive operations on the same block address, the first operation is a write operation, and the second operation is also a write operation.

[0047] IO: Input, Output, divided into two parts: IO devices and IO interfaces.

[0048] Mixed I / O Workload: Mixed I / O workload.

[0049] IO Record: Log record of disk IO and disk access, including operation type, offset, length, timestamp, etc. The time is a microsecond timestamp, indicating which blocks are accessed at a certain moment and the type of access. In general, the action is divided into read / write.

[0050] Object storage: Storing data as objects is a method for describing and processing discrete units, called objects. Each object consists of data (file data), metadata (information describing the data), and a global ID (identifying the file's storage path). Users can access the corresponding object online from any location and at any time using the ID.

[0051] Figure 1 This is a schematic diagram of a data processing method provided by an embodiment of this specification; Figure 1 As shown, the storage device is used to store the multi-source log data generated when the target business is running. The target business contains sub-businesses of various business types, and the sub-businesses are jointly operated in the form of business combinations. When analyzing the multi-source log data associated with the target business, the multi-source log data associated with the target business running on the storage device is obtained. The multi-source log data contains sub-log data of at least two business types corresponding to the target business. The target business contains different business combinations. Accordingly, the storage device corresponds to a mixed load. A log feature matrix is constructed based on the multi-source log data, and the log feature matrix is input into the type recognition model for type recognition to obtain the business type information corresponding to each sub-log data in the multi-source log data. According to the business type information corresponding to the sub-log data in the multi-source log data, the sub-log data corresponding to the target business type is selected in the multi-source log data to form a target log data group for executing the supervision task corresponding to the storage device. Type recognition and data separation of the multi-source log data associated with the target business running on the storage device are realized, thereby improving the efficiency of type recognition.

[0052] In this specification, a data processing method is provided. This specification also involves a data processing system, a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.

[0053] See also Figure 2 , Figure 2 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0054] Step 202: Acquire multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types.

[0055] Specifically, the storage device is a storage medium, which can be a block device in different forms, such as a cloud disk, a magnetic disk, a hard disk and other block devices; the storage device is used to store the multi-source log data generated when the target business is running. The target business includes sub-businesses of various business types, and the sub-businesses are jointly operated in the form of business combinations. Accordingly, each sub-business corresponds to a business load; the multi-source log data is the IO data, access data, etc. generated during the operation of the target business, which are used to record the operation of the business. Accordingly, the business type is the type of log data.

[0056] Based on this, the target business of the associated storage device is determined, and multi-source log data of the target business is obtained. The multi-source log data includes sub-log data of at least two business types corresponding to the target business. There are multiple business combinations in the operation of the target business.

[0057] In practical applications, the log data corresponding to the target business can be log data of at least two business characteristics, and the business characteristics correspond to different loads. When obtaining multi-source log data, log data collection can be uniformly performed in the storage system. The collected log data can be IO Record data corresponding to multiple disks. The collected log data can be divided according to the different disks to which the data belongs, and multi-source log data corresponding to each disk (storage device) can be obtained. The log data in this embodiment can be log data stored in certain storage media, or it can be log data stored on the computer to which the disk is mounted.

[0058] Furthermore, after acquiring multi-source log data associated with the target business running on the storage device, it is impossible to determine the business type corresponding to each sub-log data because the sub-log data in the multi-source log data is in a pending classification state. Subsequent processing using the type recognition model is required to determine the business type information of each sub-log data in the multi-source log data.

[0059] Furthermore, considering that when the storage device is a disk, each disk corresponds to multiple storage blocks. When the log data of each storage block is stored separately, when obtaining the multi-source log data corresponding to the storage device, it is necessary to obtain the log data corresponding to each storage block respectively to form the multi-source log data corresponding to the storage device. The specific implementation is as follows:

[0060] Determine at least two storage blocks associated with the storage device, and determine the sub-businesses running on the at least two storage blocks respectively; obtain sub-log data associated with the sub-businesses running on the at least two storage blocks respectively; and generate multi-source log data of the associated storage device based on the sub-log data associated with the sub-businesses running on the at least two storage blocks.

[0061] Specifically, a storage block is a storage structure on a disk. A disk is composed of multiple storage blocks, and each storage block provides data services. A sub-business is a business running on each storage block. Correspondingly, sub-log data is log data generated during the operation of a sub-business.

[0062] Based on this, at least two storage blocks associated with the storage device are determined. Sub-services running on the at least two storage blocks are determined. Log files corresponding to the at least two storage blocks are determined, and sub-log data associated with each sub-service is obtained from the log files corresponding to the at least two storage blocks. Multi-source log data associated with the storage device is assembled based on the obtained sub-log data associated with the sub-services running on the at least two storage blocks.

[0063] For example, a storage device can be a disk, which corresponds to multiple storage blocks. The log data corresponding to the storage device can be stored in a single log file, or each storage block can correspond to a log file storing log data. In the case where each storage block corresponds to a log file storing log data, the log data corresponding to each storage block is obtained separately, and the log data corresponding to each storage block is used to form the multi-source log data corresponding to the disk.

[0064] In summary, based on the acquired sub-log data associated with the sub-services running on at least two storage blocks, multi-source log data associated with the target service is composed, thereby improving the comprehensiveness of data acquisition.

[0065] Furthermore, considering that the log data read from the storage system corresponding to the storage device has a specific data format, it is impossible to directly perform feature extraction on the log data. Therefore, after obtaining the log data to be processed, it is necessary to perform format conversion on the log data to be processed. The specific implementation is as follows:

[0066] Obtaining to-be-processed multi-source log data in a first data format from a storage device; and converting the to-be-processed multi-source log data in the first data format into multi-source log data in a second data format based on a preset conversion relationship between the first data format and the second data format.

[0067] Specifically, the first data format refers to the storage format of log data in the log data storage space corresponding to the storage device, that is, the data format; the first data format can be a structured data storage format, such as a PB format; the second data format can be any text format. In order to facilitate the processing of log data, the log data in the structured data storage format is converted into an analyzable text format.

[0068] Based on this, to-be-processed multi-source log data in a first data format is obtained from an associated storage device. A second data format corresponding to the first data format is determined, as well as a preset conversion relationship between the first data format and the second data format. Based on the conversion relationship between the first data format and the second data format, the to-be-processed multi-source log data in the first data format is converted into multi-source log data in the second data format.

[0069] Continuing with the above example, when the multi-source log data to be processed is in PB format, the conversion relationship between the PB format and the text format is determined, and the format of the multi-source log data to be processed in PB format is converted based on the conversion relationship into multi-source log data in text format.

[0070] In summary, the to-be-processed multi-source log data in the first data format is converted into the multi-source log data in the second data format, thereby facilitating subsequent processing of the multi-source log data based on the second data format.

[0071] Furthermore, considering that new log data is continuously generated on the storage device during the service period, in order to ensure the timeliness of the data, a time interval can be set to obtain multi-source log data of the associated storage device based on a fixed time interval. The specific implementation is as follows:

[0072] A time interval is set for the storage device; and multi-source log data associated with the storage device is acquired according to the time interval.

[0073] Specifically, the time interval refers to a fixed time interval, which may be a time interval at the minute level, such as a one-minute interval, a ten-minute interval, a thirty-minute interval, or the like.

[0074] Based on this, before obtaining the multi-source log data of the associated storage device, a time interval for obtaining the multi-source log data is set for the storage device. The multi-source log data of the associated storage device is obtained at the time interval, that is, the multi-source log data of the associated storage device is obtained in a polling manner at the time interval.

[0075] Continuing with the previous example, you can set the multi-source log data acquisition task as a scheduled task. If the interval is set to 1 hour, the multi-source log data will be acquired every hour. New log data for the associated storage devices within that hour will also be acquired.

[0076] In summary, multi-source log data of associated storage devices is obtained according to time intervals to ensure the timeliness of the data. When there is a need for multi-source log processing, the multi-source log data obtained at each time interval can be processed separately.

[0077] Step 204: constructing a log feature matrix based on the multi-source log data, and inputting the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data.

[0078] Specifically, after obtaining the multi-source log data of the associated storage device as mentioned above, a log feature matrix can be constructed based on the multi-source log data, and the log feature matrix can be input into the type recognition model to obtain the business type information corresponding to the sub-log data in the multi-source log data, wherein the log feature matrix is composed of the type of feature and feature points, and the log data in the multi-source log data is represented in the form of feature points; the type of feature is represented by N, and the feature point is represented by M, forming an M*N log feature matrix; the type recognition model is an unsupervised machine learning model constructed based on the independent component analysis algorithm; the business type information represents the data type of each sub-log data, and the business type information includes but is not limited to RAR read-after-read, WAW write-after-write, WAR read-after-write and other types.

[0079] Based on this, after acquiring the multi-source log data from the associated storage device, a log feature matrix is constructed based on the multi-source log data, representing the multi-source log data in the form of a feature vector. The log feature matrix is input into the type recognition model for type recognition, which can obtain the business type information corresponding to each sub-log data in the multi-source log data, thereby achieving the purpose of identifying the data type of each sub-log data in the multi-source log data.

[0080] In practical applications, converting multi-source log data into a log feature matrix can be achieved through an analysis platform. Time series feature extraction is used to extract features from multi-source log data and construct the log feature matrix. The IO2Matrix approach can be used to construct the log feature matrix. Given a log data duration of K minutes, an access sequence can be obtained. This abstract access sequence can then be converted into a feature matrix using the IO2Matrix approach. The specific calculation method is as follows: For a K-minute access sequence, determine the following parameters: IOPS (calculate the number of I / O operations per minute); BPS (calculate the total length per minute); WAR (calculate the number of WARs per address in this I / O record per minute); RAR (calculate the number of RARs per address in this I / O record per minute); WAW (calculate the number of WAWs per address in this I / O record per minute); RAW (calculate the number of RAWs per address in this I / O record per minute); sequential (calculate the ratio of sequential address accesses to the total I / O per minute); RW_ration (calculate the ratio of read operations to the total I / O per minute); and IO size (calculate the average I / O size per minute, using the length as the average per minute). Combining the seven parameters (WAR, RAR, WAW, RAW, sequential, RW_ration, and I / O size) yields a 7*K matrix. This converts I / O records into a feature matrix.

[0081] Furthermore, considering that the type recognition model is an unsupervised machine learning model, it contains multiple units for data processing. Each unit needs to cooperate to complete the processing of log data. The specific implementation is as follows:

[0082] The log feature matrix is input into a type recognition model, and the conversion unit in the type recognition model is used to convert the log feature matrix into a log signal; the log signal is input into a type recognition unit in the type recognition model, and the log signal is identified by the coefficient matrix of the type recognition unit. The business type information corresponding to the sub-log data in the multi-source log data is determined according to the identification result, and the type recognition model is output.

[0083] Specifically, the conversion unit is used to convert the log feature matrix into a log signal; the type identification unit is used to predict the business type information of the sub-log data in the multi-source log data based on the log signal; the type identification unit corresponds to the independent component analysis algorithm (ICA); the prediction of the business type information of the sub-log data in the multi-source log data is used as a blind signal separation task; the recognition result is the type prediction result of the type identification model for each sub-log data in the multi-source log data.

[0084] Based on this, the log feature matrix is input into the type recognition model and sequentially passes through multiple processing units such as the conversion unit and the type recognition unit of the type recognition model. The conversion unit in the type recognition model is used to convert the log feature matrix into a log signal. The log signal is input into the type recognition unit in the type recognition model, and the log signal is recognized by the coefficient matrix of the type recognition unit to obtain a recognition result. Based on the recognition result, the business type information corresponding to each sub-log data in the multi-source log data is determined, and the type recognition model is output.

[0085] In practical applications, when determining the business type information corresponding to each sub-log data in multi-source log data based on the independent component analysis algorithm (ICA), the following formulas (1) to (3) can be used:

[0086]

[0087] S=(s1,s2,...,s n ),S∈R n (2)

[0088] X=AS (3)

[0089] Where X represents a dataset consisting of multi-source log data; m represents the log data in the dataset; n represents the n business types of log data in the dataset; i represents the log data generated at time i, that is, the time series information of the log data; S represents the matrix to be calculated corresponding to the multi-source log data, that is, the recognition result matrix for the multi-source log data; R represents the data source; A represents the unknown mixing matrix, that is, the coefficient matrix.

[0090] Continuing with the previous example, since the independent component analysis (ICA) algorithm processes signals, when processing multi-source log data based on a type recognition model, the log feature matrix is transformed into a signal using the hybrid load separation technology within ICA. The log signal is then identified based on the coefficient matrix to obtain the business type information corresponding to each sub-log data in the multi-source log data.

[0091] In summary, the conversion unit and the type identification unit in the type identification model cooperate to complete the prediction of business type information of multi-source log data, thereby improving the prediction accuracy of business type information and improving prediction efficiency.

[0092] Furthermore, considering that the multi-source log data obtained is IO Record data, feature extraction for multi-source log data needs to be performed in multiple dimensions. The specific implementation is as follows:

[0093] A storage performance data sequence corresponding to a storage performance dimension is extracted from the multi-source log data, and a storage block data sequence corresponding to a storage block dimension is extracted from the multi-source log data; a data sequence corresponding to the multi-source log data is constructed based on the storage performance data sequence and the storage block data sequence; and the data sequence is converted into a log feature matrix.

[0094] Specifically, storage performance dimensions include but are not limited to IOPS sequence (number of reads and writes per second), BPS sequence (number of bytes transferred per unit time), and Latency sequence (the difference between the start and end time of each IO record, and the average latency of multiple IOs can be calculated within one second); the storage block dimension can also be understood as a block-level dimension. In this embodiment, the block-level dimension can be a sequential stream access sequence, RAR, WAR, WAW access sequence, read and write distribution sequence, and IO size sequence; accordingly, the storage performance data sequence includes information such as IOPS and BPS; the storage block data sequence includes information on each access sequence, where a sequential stream refers to the access of a log to a continuous address; for example, a continuous IO request accesses address locations such as 1, 2, 3, 4, 5, 7, and 9, then the accesses to the five logs 1, 2, 3, 4, and 5 are sequential streams, and the statistics are about the proportion of this sequential stream.

[0095] Based on this, storage performance data sequences corresponding to storage performance dimensions are extracted from multi-source log data, and storage block data sequences corresponding to storage block dimensions are extracted from multi-source log data. Based on the storage performance data sequences and storage block data sequences, data sequences corresponding to multi-source log data are constructed. The data sequences are converted into log feature matrices.

[0096] Continuing with the previous example, when extracting features from multi-source log data, we can extract access sequences in both the statistical dimension (storage performance dimension) and the storage block layer dimension (storage block dimension). We can obtain IOPS sequences, BPS sequences, latency sequences, sequential stream access sequences, RAR, WAR, and WAW access sequences, read / write distribution sequences, and IO size sequences, thereby constructing data sequences for the multi-source log data.

[0097] In summary, feature extraction is performed on multi-source log data in the storage performance dimension and storage block dimension respectively, and the multi-source log data is converted into a log feature matrix. Therefore, in the process of log data prediction by the type recognition model, the information of the two dimensions of storage performance dimension and storage block dimension is integrated, thereby improving the matching degree between the business type information output by the type recognition model and the actual business type information of the log data.

[0098] Furthermore, considering that the type recognition model is an unsupervised model, in order to ensure the prediction accuracy of the type recognition model, after the type recognition model is built, it is necessary to train the type recognition model based on log data to ensure that the type recognition model reaches a prediction accuracy that meets the usage conditions. The specific implementation is as follows:

[0099] A log data set including sub-log data of at least two business types is determined; and a type recognition model to be trained is trained based on the log data in the log data set until a type recognition model that meets a training stop condition is obtained.

[0100] Specifically, the log data set includes log data generated by target business operations on storage devices in a real-world environment. The log data in the log data set corresponds to at least two business types, namely data I / O types, including but not limited to RAR read-after-read, WAW write-after-write, and WAR read-after-write. The training stop condition can be when the type recognition model achieves a preset accuracy level when predicting the type of the log data. For example, when predicting 100 log data items, the accuracy of the predicted log data type reaches 90%. The training stop condition can also be when the number of training rounds reaches a preset number. The training stop condition can also be when the training time is set.

[0101] Based on this, the log data file corresponding to the storage device is determined. Sub-log data containing at least two business types is obtained from the log data corresponding to the storage device. After data cleaning, a log data set is constructed. The log data in the log data set is input into the type recognition model to be trained. The type recognition model to be trained is trained based on the log data in the log data set until the type recognition model to be trained meets the training stop condition, thereby obtaining a trained type recognition model.

[0102] Continuing with the above example, when training a type recognition model, log data can be obtained from the log data file corresponding to the storage device. After performing data cleaning operations such as removing abnormal data on the obtained log data, a log data set is constructed for training the type recognition model. The type recognition model is trained based on the log data in the log data set until the training type recognition model meets the training stop condition, resulting in a trained type recognition model. The type recognition model can then be used to predict business type information in actual log data.

[0103] In summary, the type recognition model to be trained is trained based on log data to obtain a type recognition model that meets the training stop condition, thereby improving the prediction accuracy of the type recognition model.

[0104] Step 206: According to the business type information corresponding to the sub-log data in the multi-source log data, sub-log data corresponding to the target business type are selected from the multi-source log data to form a target log data group.

[0105] Specifically, after constructing the log feature matrix based on the multi-source log data and inputting the log feature matrix into the type recognition model to obtain the business type information corresponding to the sub-log data in the multi-source log data, the sub-log data corresponding to the target business type can be selected from the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data to form a target log data group for executing the corresponding supervision task of the storage device, wherein the target business type is the target type determined based on the business type information of each sub-log data, and the type corresponding to any business type information can be used as the target business type; accordingly, the business type information of each log data in the target log data group is the same.

[0106] Based on this, after constructing a log feature matrix based on multi-source log data and inputting the log feature matrix into the type recognition model to obtain the business type information corresponding to the sub-log data in the multi-source log data, the target business type is determined based on the business type information of each sub-log data in the multi-source log data. Based on the business type information corresponding to the sub-log data in the multi-source log data, sub-log data corresponding to the target business type is selected from the multi-source log data, and the sub-log data are used to form a target log data group. The target log data is used to perform the supervision task corresponding to the storage device.

[0107] In practical applications, any business type can be used as the target business type, and a target log data group corresponding to the target business type can be constructed. The business type information corresponding to the sub-log data in the multi-source log data is used to determine the type of data type corresponding to the multi-source log data. Based on the business type information of the sub-log data, the multi-source log data can be divided into log data groups corresponding to each type. This allows each type of sub-log data to be grouped together.

[0108] Furthermore, considering that multi-source log data corresponds to multiple business loads, each business load corresponds to a different business type. In order to facilitate the supervision of multiple business loads, a type label can be added to each sub-log data based on the business type information corresponding to the sub-log data in the multi-source log data. The specific implementation is as follows:

[0109] According to the business type information corresponding to the sub-log data in the multi-source log data, type labels are added to the sub-log data in the multi-source log data; according to the type label corresponding to each sub-log data, the multi-source log data is divided into at least one target log data group, and the at least one target log data group is used to execute the supervision task corresponding to the storage device.

[0110] Specifically, the type tag refers to the tag information used to indicate the read and write type of sub-log data, and is used to indicate the read and write type of each sub-log data, including but not limited to RAR read-after-read, WAW write-after-write, WAR read-after-write, and other types of tag information; the supervision task is used to monitor and manage at least one business load corresponding to the target business, including but not limited to stress testing and traffic playback of the business load.

[0111] Based on this, the business type information corresponding to each sub-log data in the multi-source log data is determined, and type tags are added to the sub-log data in the multi-source log data according to the business type information corresponding to each sub-log data. The multi-source log data is divided into at least one target log data group according to the type tag corresponding to each sub-log data, and each target log data group stores multi-source log data with the same type tag. The at least one target log data group is used to perform a supervisory task corresponding to the storage device, thereby enabling monitoring and management of at least one business load corresponding to the storage device.

[0112] Continuing with the above example, if the multi-source log data contains 10 sub-log data, and the service type information of each sub-log data is read-after-read, write-after-write, read-after-write, read-after-read, write-after-write, read-after-write, read-after-read, write-after-write, read-after-write, and write-after-read, read and write tags are added to each sub-log data, and sub-log data with the same read and write tags are stored in a log data group.

[0113] To sum up, by dividing multi-source log data into at least one target log data group according to the type label corresponding to each sub-log data, the sub-log data with the same type label can be stored in one data group to facilitate the execution of subsequent supervision tasks, and standardized data storage is also achieved.

[0114] Furthermore, during the operation of the storage device, the storage performance and traffic conditions of the storage device can be tested at any time based on the target log data group. The specific implementation is as follows:

[0115] Receive a device detection request associated with the supervision task, wherein the supervision task is used to monitor and manage the storage device; determine the target log data contained in the target log data group based on the device detection request; determine the detection information of the storage device based on the target log data as the task execution result of the supervision task.

[0116] Specifically, the device detection request is a computer instruction submitted to the storage device for detecting the storage device; correspondingly, the detection information is the detection result of the storage device obtained based on the analysis of the target log data, which is used to represent the storage performance, traffic status and other information of the storage device.

[0117] Based on this, upon receiving a device detection request associated with a supervisory task, the target log data group is determined based on the time information and data type information included in the device detection request. The target log data is extracted from the target log data group. Based on an analysis of the target log data combined with the time information, the storage device detection information is determined and used as the supervisory task execution result.

[0118] Continuing with the previous example, after receiving a disk test request, the target log data from the disk is retrieved based on the time and data type information included in the test request. Based on the volume of the target log data, performance stress testing and traffic flow testing can be performed on the disk.

[0119] In summary, the storage device is tested based on the device test request to obtain the test result, thereby realizing the storage performance test of the storage device and optimizing the use of the storage device.

[0120] One embodiment of this specification obtains multi-source log data associated with a target business on an associated storage device, wherein the multi-source log data contains sub-log data of at least two business types; constructs a log feature matrix based on the multi-source log data, and inputs the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; according to the business type information corresponding to the sub-log data in the multi-source log data, selects sub-log data corresponding to the target business type in the multi-source log data to form a target log data group for executing the corresponding supervision task of the storage device. This realizes the type recognition of the multi-source log data of the associated storage device and improves the efficiency of type recognition. This provides data support for the subsequent supervision of the storage device.

[0121] The following combined Figure 3 , taking the application of the data processing method provided in this specification in the cloud disk reading and writing data processing as an example, the data processing method is further explained. Figure 3 A flowchart of a data processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.

[0122] Step 302: Collect multi-source log data associated with storage devices in the storage system.

[0123] In actual applications, storage devices can be cloud drives, magnetic disks, etc. When collecting data, you can set up a scheduled task to collect multi-source log data corresponding to the storage device at fixed intervals.

[0124] Step 304: Determine at least two storage blocks associated with the storage device, and determine the multi-source log data corresponding to each storage block in the multi-source log data.

[0125] Step 306: Perform format conversion on the multi-source log data corresponding to each storage block.

[0126] The log data obtained is in petabyte format, which can be parsed into any parseable text format. It should be noted that since a storage block can correspond to multiple business loads, the log data corresponding to each storage block can be used as multi-source log data for data classification.

[0127] Step 308: For the multi-source log data corresponding to each storage block, extract the storage performance data sequence corresponding to the statistical dimension from the multi-source log data, and extract the storage block data sequence corresponding to the storage block dimension from the multi-source log data.

[0128] Step 310: Construct a data sequence corresponding to the multi-source log data based on the storage performance data sequence and the storage block data sequence.

[0129] Feature extraction is performed on the log data of each storage block, including statistical features and storage block features. Statistical features include IOPS (Input and Output Per Second) and BPS (Bytes Transferred per Unit Time). Storage block feature extraction includes sequential flow sequences, access sequences such as RAR, WAR, and WAW, read and write distribution sequences, and I / O size sequences. An N*M feature matrix is constructed, representing N features and M feature points.

[0130] Step 312: Input the data sequence into the type recognition model for recognition, and obtain type information corresponding to the log data in the multi-source log data.

[0131] A type recognition model based on ICA (independent component analysis) is constructed, and the type recognition task of multi-source log data is regarded as a blind signal separation task to realize the type recognition of individual log data in multi-source log data.

[0132] In practical applications, such as Figure 4 As shown, the storage system uses the acquired multi-source log data as the original signal. In this case, the acquired original signal is a mixed signal of multiple loads. The signal separation system provides a type recognition model based on ICA (Independent Component Analysis) to achieve separation of the original signal.

[0133] Step 314: According to the type information corresponding to the log data in the multi-source log data, log data corresponding to the target business type is selected from the multi-source log data to form a target log data group.

[0134] After determining the type information corresponding to the log data in the multi-source log data, the type information is used as a label for the log data, and the multi-source log data is separated and stored according to the type information.

[0135] Step 316: Supervise the storage device based on the type information corresponding to the log data in the multi-source log data.

[0136] In practical applications, determining the type information corresponding to the log data in the storage device's multi-source log data allows identification and separation of mixed service loads. This also allows the impact of different services on overall performance and the signal proportion at different times to be determined. This type information can be used to perform stress testing and traffic replay on the storage device.

[0137] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 5 FIG1 shows a schematic diagram of the structure of a data processing device provided by an embodiment of this specification. Figure 5 As shown, the device includes:

[0138] An acquisition module 502 is configured to acquire multi-source log data associated with a target business running on a storage device, wherein the multi-source log data includes sub-log data of at least two business types corresponding to the target business;

[0139] A construction module 504 is configured to construct a log feature matrix based on the multi-source log data, and input the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data;

[0140] The selection module 506 is configured to select sub-log data corresponding to a target business type from the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data to form a target log data group.

[0141] In an optional embodiment, the building module 504 is further configured to:

[0142] Inputting the log feature matrix into a type recognition model, and using a conversion unit in the type recognition model to convert the log feature matrix into a log signal;

[0143] The log signal is input into the type recognition unit in the type recognition model, the log signal is identified by the coefficient matrix of the type recognition unit, the business type information corresponding to the sub-log data in the multi-source log data is determined according to the recognition result, and the type recognition model is output.

[0144] In an optional embodiment, the building module 504 is further configured to:

[0145] Extracting a storage performance data sequence corresponding to a storage performance dimension from the multi-source log data, and extracting a storage block data sequence corresponding to a storage block dimension from the multi-source log data;

[0146] Constructing a data sequence corresponding to the multi-source log data based on the storage performance data sequence and the storage block data sequence;

[0147] Convert the data series into a log-feature matrix.

[0148] In an optional embodiment, the acquisition module 502 is further configured to:

[0149] Determining at least two storage blocks associated with the storage device, and determining sub-services respectively running on the at least two storage blocks;

[0150] Obtain sub-log data associated with sub-businesses running on at least two storage blocks respectively;

[0151] Based on the sub-log data associated with the sub-services running on at least two storage blocks, multi-source log data associated with the storage device is generated.

[0152] In an optional embodiment, the acquisition module 502 is further configured to:

[0153] Acquire to-be-processed multi-source log data in a first data format of a management storage device;

[0154] Based on a preset conversion relationship between the first data format and the second data format, the to-be-processed multi-source log data in the first data format is converted into multi-source log data in the second data format.

[0155] In an optional embodiment, the acquisition module 502 is further configured to:

[0156] Setting a time interval for the storage device;

[0157] The acquiring of multi-source log data of associated storage devices includes:

[0158] Multi-source log data associated with the storage device is acquired according to the time interval.

[0159] In an optional embodiment, the selection module 506 is further configured to:

[0160] Adding a type tag to the sub-log data in the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data;

[0161] The multi-source log data are divided into at least one target log data group according to a type label corresponding to each sub-log data. The at least one target log data group is used to execute a supervision task corresponding to the storage device.

[0162] In an optional embodiment, the building module 504 is further configured to:

[0163] Determining a log data set including sub-log data of at least two business types;

[0164] The type recognition model to be trained is trained based on the log data in the log data set until a type recognition model that meets the training stop condition is obtained.

[0165] In an optional embodiment, the selection module 506 is further configured to:

[0166] receiving a device detection request associated with the supervisory task, wherein the supervisory task is used to monitor and manage the storage device;

[0167] determining target log data included in the target log data group based on the device detection request;

[0168] The detection information of the storage device is determined based on the target log data as a task execution result of the supervision task.

[0169] In summary, one embodiment of the present specification obtains multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; constructs a log feature matrix based on the multi-source log data, and inputs the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; and selects sub-log data corresponding to a target business type from the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data to form a target log data group for executing a corresponding supervision task for the storage device. This achieves type recognition of the multi-source log data of the associated storage device and improves type recognition efficiency. This provides data support for subsequent supervision of the storage device.

[0170] The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solution of the data processing device and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the data processing device, please refer to the description of the technical solution of the above-mentioned data processing method.

[0171] Figure 6Schematic diagram of a data processing system according to an embodiment of the present disclosure. The data processing system 600 includes a server 610, a data processing node 620, and a feature processing node 630;

[0172] The server 610 is configured to determine a storage device in response to a data processing request, invoke a block storage service to obtain multi-source log data associated with the storage device, and forward the multi-source log data to the data processing node 620 via an object storage service;

[0173] The data processing node 620 is configured to construct a log feature matrix based on the multi-source log data and send the log feature matrix to the feature processing node 630;

[0174] The feature processing node 630 is configured to input the log feature matrix into a type recognition model, obtain business type information corresponding to sub-log data in the multi-source log data, and feed the information back to the server 610;

[0175] The server 610 is further configured to select sub-log data corresponding to a target business type from the multi-source log data according to business type information corresponding to the sub-log data in the multi-source log data to form a target log data group.

[0176] Based on this, the data processing system for type identification of multi-source log data includes a server 610, a data processing node 620 and a feature processing node 630; the data processing node 620 has a feature extraction function, which is used to extract features from multi-source log data; the feature processing node 630 has a feature processing function, which is used to input the log feature matrix into the type identification model for type identification.

[0177] In actual applications, when performing type identification on multi-source log data corresponding to storage devices, the processing flow can refer to Figure 7 . The server dispatches the block storage service to obtain multi-source log data corresponding to various businesses running on cloud disks, magnetic disks and other devices, and temporarily stores the multi-source log data through object storage. Obtain multi-source log data, extract features from the multi-source log data through the analysis service, obtain the feature sequence, and store it in the database. The feature sequence is processed based on the type recognition model constructed based on the independent component analysis algorithm, and then the type of each log data in the multi-source log data is determined, so as to achieve the purpose of determining the label for each log data in the multi-source log data. It can realize the identification and separation of mixed loads when cloud disks and magnetic disks correspond to multiple business combinations.

[0178] In summary, by integrating blind source separation technology and analyzing multi-source log data, mixed business types can be identified and separated. Based on the identification and separation results, performance stress testing and traffic playback can be performed on cloud disks and disks, achieving supervision of cloud disks and disks.

[0179] Furthermore, considering that the storage device will continuously generate log data in the process of providing services, and since multiple businesses are running on the storage device, different traffic will be generated at different times due to the influence of time and business needs. Therefore, when obtaining multi-source log data, it is also necessary to consider the time factor and obtain data based on a preset time interval. In specific implementation, the server 610 is also used to call the block storage service to obtain the third data format of the associated storage device based on a preset time interval to be processed multi-source log data; based on the conversion relationship between the preset third data format and the fourth data format, the third data format to be processed multi-source log data is converted into multi-source log data in the fourth data format.

[0180] Among them, the third data format is the storage format of multi-source log data in object storage, such as PB format; correspondingly, the fourth data format is the text type data format obtained after format conversion based on the third data format. The fourth data format can be set according to actual needs, and this embodiment does not impose any restrictions on this.

[0181] Using the above example, combined with Figure 7 By setting up a scheduled task to poll log data stored in object storage, you can retrieve multi-source log data from disks and cloud drives. Log data retrieved from object storage is in petabyte format and needs to be parsed into a parseable text format to facilitate subsequent feature extraction and type identification.

[0182] The server 610 is further configured to input the log feature matrix into a type recognition model, and utilize a conversion unit in the type recognition model to convert the log feature matrix into a log signal; input the log signal into a type recognition unit in the type recognition model, identify the log signal using a coefficient matrix of the type recognition unit, determine the business type information corresponding to the sub-log data in the multi-source log data based on the identification result, and output the type recognition model.

[0183] The server 610 is further configured to extract a storage performance data sequence corresponding to a storage performance dimension from the multi-source log data, and extract a storage block data sequence corresponding to a storage block dimension from the multi-source log data; construct a data sequence corresponding to the multi-source log data based on the storage performance data sequence and the storage block data sequence; and convert the data sequence into a log feature matrix.

[0184] The server 610 is further used to determine at least two storage blocks associated with the storage device, and determine the sub-businesses running on the at least two storage blocks respectively; obtain the sub-log data associated with the sub-businesses running on the at least two storage blocks respectively; and generate multi-source log data associated with the target business based on the sub-log data associated with the sub-businesses running on the at least two storage blocks.

[0185] The server 610 is further configured to determine a log data set containing sub-log data of at least two business types; and train a type recognition model to be trained based on the log data in the log data set until a type recognition model that meets a training stop condition is obtained.

[0186] The server 610 is also used to receive a device detection request associated with the supervision task, wherein the supervision task is used to monitor and manage the storage device; determine the target log data included in the target log data group based on the device detection request; and determine the detection information of the storage device based on the target log data as the task execution result of the supervision task.

[0187] One embodiment of the present specification obtains multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; constructs a log feature matrix based on the multi-source log data, and inputs the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; according to the business type information corresponding to the sub-log data in the multi-source log data, selects sub-log data of the corresponding target business type from the multi-source log data to form a target log data group for executing the corresponding supervision task of the storage device. This realizes the type recognition of the multi-source log data of the associated storage device and improves the efficiency of type recognition. This provides data support for the subsequent supervision of the storage device.

[0188] The above is a schematic diagram of a data processing system according to this embodiment. It should be noted that the technical solution of the data processing system and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the data processing system, please refer to the description of the technical solution of the above-mentioned data processing method.

[0189] Figure 8 8 shows a block diagram of a computing device 800 according to one embodiment of the present disclosure. Components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.

[0190] The computing device 800 also includes an access device 840 that enables the computing device 800 to communicate via one or more networks 860. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0191] In one embodiment of the present specification, the above components of the computing device 800 and Figure 8 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 8 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0192] Computing device 800 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 800 may also be a mobile or stationary server.

[0193] The processor 820 is configured to execute the following computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by the processor.

[0194] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned data processing method are of the same concept. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical scheme of the above-mentioned data processing method.

[0195] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.

[0196] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical scheme of the above-mentioned data processing method.

[0197] An embodiment of the present specification further provides a computer program product, including a computer program or instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.

[0198] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-mentioned data processing method.

[0199] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0200] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0201] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0202] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0203] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A data processing method, comprising: Acquire multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; Constructing a log feature matrix based on the multi-source log data, and inputting the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; According to the business type information corresponding to the sub-log data in the multi-source log data, sub-log data corresponding to the target business type are selected from the multi-source log data to form a target log data group.

2. The method according to claim 1, wherein inputting the log feature matrix into a type recognition model to obtain business type information corresponding to sub-log data in the multi-source log data comprises: Inputting the log feature matrix into a type recognition model, and converting the log feature matrix into a log signal using a conversion unit in the type recognition model; The log signal is input into the type recognition unit in the type recognition model, the log signal is identified by the coefficient matrix of the type recognition unit, the business type information corresponding to the sub-log data in the multi-source log data is determined according to the recognition result, and the type recognition model is output.

3. The method according to claim 1, wherein constructing a log feature matrix based on the multi-source log data comprises: Extracting a storage performance data sequence corresponding to a storage performance dimension from the multi-source log data, and extracting a storage block data sequence corresponding to a storage block dimension from the multi-source log data; Constructing a data sequence corresponding to the multi-source log data based on the storage performance data sequence and the storage block data sequence; Convert the data series into a log-feature matrix.

4. The method according to claim 1, wherein obtaining multi-source log data of associated storage devices comprises: Determining at least two storage blocks associated with the storage device, and determining sub-services respectively running on the at least two storage blocks; Obtain sub-log data associated with sub-businesses running on at least two storage blocks respectively; Based on the sub-log data associated with the sub-services running on at least two storage blocks, multi-source log data associated with the storage device is generated.

5. The method according to claim 1, wherein obtaining multi-source log data of associated storage devices comprises: Acquire to-be-processed multi-source log data in a first data format associated with a storage device; Based on a preset conversion relationship between the first data format and the second data format, the to-be-processed multi-source log data in the first data format is converted into multi-source log data in the second data format.

6. The method according to claim 1, before acquiring multi-source log data of associated storage devices, further comprising: Setting a time interval for the storage device; The acquiring of multi-source log data of associated storage devices includes: Multi-source log data associated with the storage device is acquired according to the time interval.

7. The method according to claim 1, wherein the selecting sub-log data corresponding to a target business type from the multi-source log data to form a target log data group according to business type information corresponding to the sub-log data in the multi-source log data comprises: Adding a type tag to the sub-log data in the multi-source log data according to the business type information corresponding to the sub-log data in the multi-source log data; The multi-source log data are divided into at least one target log data group according to a type label corresponding to each sub-log data. The at least one target log data group is used to execute a supervision task corresponding to the storage device.

8. The method according to claim 1, wherein the training of the type recognition model comprises: Determining a log data set including sub-log data of at least two business types; The type recognition model to be trained is trained based on the log data in the log data set until a type recognition model that meets the training stop condition is obtained.

9. The method according to claim 1, further comprising: after selecting sub-log data corresponding to a target business type from the multi-source log data to form a target log data group; receiving a device detection request associated with the supervisory task, wherein the supervisory task is used to monitor and manage the storage device; determining target log data included in the target log data group based on the device detection request; The detection information of the storage device is determined based on the target log data as a task execution result of the supervision task.

10. A data processing system comprising a server, a data processing node and a feature processing node; The server is configured to determine a storage device in response to a data processing request, call a block storage service to obtain multi-source log data associated with the storage device, and forward the multi-source log data to the data processing node via an object storage service; The data processing node is configured to construct a log feature matrix based on the multi-source log data and send the log feature matrix to the feature processing node; The feature processing node is configured to input the log feature matrix into a type recognition model, obtain business type information corresponding to sub-log data in the multi-source log data, and feed the information back to the server; The server is further configured to select sub-log data corresponding to a target business type from the multi-source log data to form a target log data group according to business type information corresponding to the sub-log data in the multi-source log data.

11. The data processing system according to claim 10, wherein the server is further configured to call a block storage service to obtain, based on a preset time interval, to-be-processed multi-source log data in a third data format from an associated storage device; and convert the to-be-processed multi-source log data in the third data format into multi-source log data in the fourth data format based on a preset conversion relationship between the third data format and a fourth data format.

12. A data processing device comprising: an acquisition module configured to acquire multi-source log data of an associated storage device, wherein the multi-source log data includes sub-log data of at least two business types; A construction module is configured to construct a log feature matrix based on the multi-source log data, and input the log feature matrix into a type recognition model to obtain business type information corresponding to the sub-log data in the multi-source log data; The selection module is configured to select sub-log data corresponding to a target business type from the multi-source log data to form a target log data group according to business type information corresponding to the sub-log data in the multi-source log data.

13. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the data processing method according to any one of claims 1 to 9 are implemented.

14. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the data processing method according to any one of claims 1 to 9.

15. A computer program product, comprising a computer program or instructions, which implement the steps of the data processing method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Patent Citations

  • Multi-source heterogeneous log collection method and device

    CN112256651A

  • Log label labeling method, device and equipment and readable storage medium

    CN114756521A

  • Multi-source log data processing method for electric power information system

    CN116756617A

  • Analysis apparatus, analysis method and recording medium for recording analysis program

    US20090248863A1