Model training method and device, detection method and device, related equipment, storage medium and computer program product
By generating adversarial network training to detect storage node abnormalities, using self-attention mechanism and distributed processing, the accuracy and efficiency of storage node abnormality detection in the prior art are solved, and efficient and accurate abnormality detection is achieved.
Patent Information
- Application Number
- CN202510037731.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to accurately and efficiently detect whether there are abnormalities in storage nodes. Especially in large-scale data storage scenarios, it is difficult to detect abnormal storage nodes in a timely and accurate manner.
The generative adversarial network is used for training, and the first sub-model generates sample data. The second sub-model distinguishes between real data and generated data. The second sub-model has a self-attention mechanism. The service-related information of the storage node is obtained and preprocessed through distributed processing. The generation adversarial network is used to detect whether there are abnormalities in the storage node.
It realizes high-accuracy and efficient storage node abnormality detection, generates training sample data independently generated by adversarial networks, improves model performance, and the self-attention mechanism takes into account global information and processes complex data characteristics, improving detection efficiency and accuracy.
Smart Images

Figure CN120011809A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence (AI), and in particular to a model training method, a detection method, an apparatus, related equipment, a storage medium and a computer program product. Background Art
[0002] In the related art, storage nodes can provide data storage services (such as object storage services), and the maintenance of storage nodes usually includes monitoring the behavior patterns of storage nodes (which can also be understood as detecting whether there are abnormalities in storage nodes). When an abnormal storage node (which can also be understood as an abnormal node) is found, the corresponding alarm or processing mechanism can be triggered to ensure the reliability and stability of the services provided by the storage node.
[0003] However, there is currently no effective solution for accurately and efficiently detecting whether storage nodes have abnormalities. Summary of the invention
[0004] To solve related technical problems, the embodiments of the present application provide a model training method, a detection method, an apparatus, a terminal, an electronic device, a storage medium and a computer program product.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The embodiment of the present application provides a model training method, which is applied to a storage system including multiple storage nodes, wherein the storage nodes are at least used to store data, and the method includes:
[0007] Obtain service-related information of multiple normal storage nodes;
[0008] Structuring the service-related information of the plurality of normal storage nodes obtained to obtain structured service-related information;
[0009] Preprocessing the structured service-related information to obtain processed service-related information;
[0010] The generative adversarial network is trained using the processed service-related information, and the generative adversarial network includes a first sub-model and a second sub-model. During the training process, the first sub-model is used to generate sample data, and the second sub-model is used to distinguish between real data and generated data. The second sub-model has a self-attention mechanism; the trained second sub-model is used to detect whether there is an abnormality in the storage node.
[0011] In the above solution, the structured processing of the service-related information of the plurality of normal storage nodes obtained includes:
[0012] A distributed processing method is adopted to perform structured processing on the service-related information of multiple normal storage nodes.
[0013] In the above solution, the preprocessing of structured service-related information includes:
[0014] A distributed processing method is used to pre-process structured service-related information.
[0015] In the above solution, the loss function corresponding to the first sub-model is associated with the Minkowski distance.
[0016] In the above scheme, the second sub-model adopts a multi-layer perceptron (MLP) architecture, and the multiple layers of the MLP architecture include one or more feature processing modules, and the feature processing module is associated with the self-attention mechanism.
[0017] The present application also provides a detection method, including:
[0018] Acquire service-related information of a first storage node, where the first storage node is at least used to store data;
[0019] Structuring the service-related information of the first storage node to obtain structured service-related information;
[0020] Preprocessing the structured service-related information to obtain processed service-related information;
[0021] The processed service-related information is used in combination with a second sub-model to detect whether the first storage node has an abnormality, and the second sub-model is trained using any of the above-mentioned model training methods.
[0022] In the above solution, the first storage node belongs to a storage system, and the storage system includes multiple storage nodes;
[0023] In the process of detecting whether a storage node is abnormal, at least one of the following is satisfied:
[0024] Service-related information of multiple storage nodes is structured using distributed processing;
[0025] The structured service-related information of multiple storage nodes is pre-processed in a distributed processing manner;
[0026] Multiple storage nodes are processed in a distributed manner using the second sub-model to detect whether there are abnormalities in the storage nodes.
[0027] The embodiment of the present application further provides a model training device, which is arranged in a storage system including a plurality of storage nodes, wherein the storage nodes are at least used to store data, including:
[0028] A first acquisition unit, used to acquire service related information of multiple normal storage nodes;
[0029] A first processing unit is used to perform structured processing on the service-related information of the plurality of normal storage nodes to obtain structured service-related information;
[0030] A second processing unit is used to pre-process the structured service related information to obtain processed service related information;
[0031] A training unit is used to train a generative adversarial network using processed service-related information, wherein the generative adversarial network includes a first sub-model and a second sub-model. During the training process, the first sub-model is used to generate sample data, and the second sub-model is used to distinguish between real data and generated data, and the second sub-model has a self-attention mechanism; the trained second sub-model is used to detect whether there is an abnormality in the storage node.
[0032] The present application also provides a detection device, including:
[0033] A second acquisition unit is used to acquire service related information of a first storage node, where the first storage node is at least used to store data;
[0034] A third processing unit is used to perform structured processing on the service related information of the first storage node to obtain structured service related information;
[0035] A fourth processing unit, configured to pre-process the structured service-related information to obtain processed service-related information;
[0036] A detection unit is used to use the processed service-related information in combination with a second sub-model to detect whether there is an abnormality in the first storage node, and the second sub-model is trained using any of the above-mentioned model training methods.
[0037] The embodiment of the present application further provides a first electronic device, comprising: a first processor and a first memory for storing a computer program that can be run on the processor,
[0038] Wherein, the first processor is used to execute the steps of any of the above-mentioned model training methods when running the computer program.
[0039] The embodiment of the present application further provides a second electronic device, comprising: a second processor and a second memory for storing a computer program that can be run on the processor,
[0040] Wherein, the second processor is used to execute the steps of any of the above detection methods when running the computer program.
[0041] An embodiment of the present application also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned model training methods are implemented, or the steps of any of the above-mentioned detection methods are implemented.
[0042] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any of the above-mentioned model training methods, or implements the steps of any of the above-mentioned detection methods.
[0043] The model training method, detection method, device, related equipment, storage medium and computer program product provided by the embodiment of the present application, the storage system obtains service-related information of multiple normal storage nodes; the storage system includes multiple storage nodes, and the storage nodes are at least used to store data; the service-related information of the obtained multiple normal storage nodes is structured to obtain structured service-related information; the structured service-related information is preprocessed to obtain processed service-related information; the processed service-related information is used to train the generative adversarial network, the generative adversarial network includes a first sub-model and a second sub-model, during the training process, the first sub-model is used to generate sample data, the second sub-model is used to distinguish between real data and generated data, and the second sub-model has a self-attention mechanism; the trained second sub-model is used to detect whether the storage node is abnormal. At the same time, the electronic device obtains the service-related information of the first storage node, and the first storage node is at least used to store data; the service-related information of the first storage node is structured to obtain structured service-related information; the structured service-related information is preprocessed to obtain processed service-related information; the processed service-related information is combined with the second sub-model to detect whether the first storage node is abnormal, and the second sub-model is trained using any of the above model training methods. The solution provided in the embodiment of the present application detects anomalies of storage nodes through a generative adversarial network. On the one hand, the generative adversarial network can autonomously generate sample data required for training, which can ensure the scale of sample data used for training and improve the performance of the model obtained through training. On the other hand, the training efficiency of the generative adversarial network is high and the model performance obtained through adversarial training is better. At the same time, a self-attention mechanism is introduced into the generative adversarial network, and the performance of the model obtained through training can be further improved by utilizing the characteristics of the self-attention mechanism that considers global information and can process complex data. In this way, when the trained model is used to detect whether there are anomalies in the storage node, high-accuracy and efficient detection can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1A flowchart of a model training method according to an embodiment of the present application;
[0045] Figure 2 A schematic diagram of a flow chart of feature processing performed by a feature processing module in an embodiment of the present application;
[0046] Figure 3 A flowchart of another model training method according to an embodiment of the present application;
[0047] Figure 4 A schematic diagram of the structure of an abnormal monitoring system for object storage service node maintenance based on deep learning is used as an example of this application;
[0048] Figure 5 A flowchart of an abnormality monitoring method for object storage service node maintenance based on deep learning is provided as an example of the application of this application;
[0049] Figure 6 This is a schematic diagram of the structure of a model training device according to an embodiment of the present application;
[0050] Figure 7 This is a schematic diagram of the structure of a detection device according to an embodiment of the present application;
[0051] Figure 8 This is a schematic diagram of the structure of a first electronic device according to an embodiment of the present application;
[0052] Fig. 9 This is a schematic diagram of the structure of a second electronic device according to an embodiment of the present application;
[0053] Fig.10 This is a schematic diagram of the detection system structure of the embodiment of the present application. DETAILED DESCRIPTION
[0054] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0055] Nodes that provide object storage services (also referred to as object storage service nodes, hereinafter referred to as storage nodes) may have abnormalities when providing storage services, such as abnormal data access patterns, node performance issues, security threats, etc. In related technologies, methods for detecting abnormalities in storage nodes generally include the following:
[0056] 1) Anomaly detection based on log analysis: Storage nodes can generate log information when providing storage services. The log information includes data access patterns, access events and other information. Based on this, when performing anomaly detection, the log information generated by the storage node can be analyzed to detect (also understood as identifying) whether the log information contains abnormal patterns and / or abnormal events, so as to determine (also understood as judging) the health status of the storage node. Specifically, a rule engine or pattern matching algorithm can be used to detect abnormal patterns and / or abnormal events contained in the log information, and when abnormal patterns and / or abnormal events are detected in the log information, a corresponding alarm or processing mechanism is triggered;
[0057] 2) Anomaly detection based on monitoring indicators: When performing anomaly detection, various performance indicators of storage nodes can be monitored to determine whether the storage nodes are abnormal; the performance indicators of storage nodes may include: network bandwidth, disk usage, central processing unit (CPU) utilization, etc. Specifically, the values of real-time monitored performance indicators (also called indicator values) can be analyzed using technical means such as threshold detection or statistical analysis. When the value of a performance indicator exceeds a preset threshold or an abnormal change occurs, a corresponding alarm or processing mechanism can be triggered;
[0058] 3) Collaborative anomaly detection based on distributed architecture: When there are multiple storage nodes and the multiple storage nodes can communicate with each other (that is, they can pass messages), a collaborative mechanism can be established between the multiple storage nodes to achieve state synchronization through shared information and collaborative work between storage nodes, so as to timely discover and handle abnormal situations.
[0059] However, in actual applications, the effect of anomaly detection using the above method may not be good, that is, it is difficult to accurately and timely detect anomalies of storage nodes.
[0060] For example, when anomaly detection is performed based on log analysis, if the amount of data inside the storage node is large (which can also be understood as in the scenario of large-scale data storage), since the storage node storing large-scale data may be frequently accessed, the storage node will generate a large amount of log information, which will increase the amount of calculation and time required to analyze the log information, making it difficult to detect anomalies in a timely manner; or, when anomaly detection is performed based on monitoring indicators, the performance indicators obtained from monitoring may be irrelevant to the anomaly. In this case, it is impossible to determine the anomaly by analyzing the performance indicators; or, when anomaly detection is performed collaboratively based on a distributed architecture, if the scale of stored data is large and / or the communication distance between multiple storage nodes is far, the communication resources and time resources between multiple storage nodes may be difficult to meet the requirements for state synchronization, resulting in the inability to detect anomalies in a timely manner.
[0061] In order to improve the accuracy and efficiency of anomaly detection, a scheme for anomaly detection based on machine learning has been proposed in the related art. In the scheme for anomaly detection based on machine learning, a model for detecting storage node anomalies based on a machine learning algorithm (hereinafter referred to as an anomaly detection model) is first constructed, and the anomaly detection model is trained using the historical data of the storage node to obtain a trained anomaly detection model; then, the trained anomaly detection model can be used to detect anomalies in the storage node, or to predict anomalies that may occur in the storage node. Specifically, the anomaly detection model can be trained using supervised learning, unsupervised learning, or semi-supervised learning, so that the trained anomaly detection model can learn the normal behavior pattern of the storage node; then, the trained anomaly detection model can be used to analyze the real-time data related to the storage node, and by determining whether the real-time data conforms to the normal behavior pattern, it is determined whether an anomaly occurs in the storage node. Specifically, the following steps may be included:
[0062] Step 1: Data collection: Collect (also known as acquisition) various monitoring indicator data of storage nodes, such as CPU utilization, memory usage, network traffic, etc.; then execute step 2;
[0063] Step 2: Feature engineering: Extract features from the collected monitoring indicator data to obtain features that can describe the behavior patterns of storage nodes, such as one or more of statistical features, spectrum features, and time series features (one or more can also be understood as at least one);
[0064] Step 3: Data preprocessing: Perform one or more preprocessing operations such as cleaning, normalization, and dimension reduction on the extracted features to obtain preprocessed data to improve the training effect of the model;
[0065] Step 4: Model training: Establish an anomaly detection model based on a machine learning algorithm, and use the preprocessed data in step 3 to train the anomaly detection model to obtain a trained anomaly detection model; wherein the machine learning algorithm may include a support vector machine (SVM) or a random forest, etc.;
[0066] Step 5: Anomaly detection: Use the trained anomaly detection model to predict the real-time data of the storage node to obtain the prediction result, and then determine whether the storage node has an abnormal situation based on the prediction result; specifically, if the prediction result exceeds the preset threshold, it can be considered that the storage node has an abnormal situation (it can also be understood as detecting an abnormal situation or determining that the behavior pattern of the storage node is abnormal);
[0067] Step 6: Alarm and processing: When an abnormality is detected in a storage node, the corresponding alarm mechanism is triggered and / or appropriate processing measures are taken. The processing measures may include notifying the administrator, starting the backup node, or reallocating the task, etc.
[0068] However, the scheme of anomaly detection based on machine learning may be difficult to meet actual needs. For example, the historical data used to train the anomaly detection model may be outdated (i.e., it cannot reflect the actual situation of the storage node) and / or limited in number, which will lead to poor performance (i.e., low accuracy) of the anomaly detection model trained with these historical data; for another example, the anomaly detection model in the related technology may not be able to identify abnormal situations, making it difficult to accurately and timely find abnormal storage nodes; for another example, when the scale of data stored in the storage node is large (which can also be understood as in the scenario of massive data), it may be difficult to meet the computing power and time resource consumption required for data collection, preprocessing and other operations.
[0069] Based on this, in various embodiments of the present application, anomalies of storage nodes are detected by generative adversarial networks. On the one hand, the generative adversarial networks can autonomously generate sample data required for training, which can ensure the scale of sample data used for training and improve the performance of the model obtained by training. On the other hand, the training efficiency of the generative adversarial networks is high and the model performance obtained by adversarial training is better. At the same time, the self-attention mechanism is introduced into the generative adversarial network, and the performance of the model obtained by training can be further improved by utilizing the characteristics of the self-attention mechanism that considers global information and can process complex data. In this way, when the trained model is used to detect whether there are anomalies in the storage nodes, high-accuracy and efficient detection can be achieved.
[0070] The present application embodiment provides a model training method, which is applied to a storage system including multiple storage nodes, wherein the storage nodes are used to store data at least. Figure 1As shown, the method includes:
[0071] Step 101: Obtain service-related information of multiple normal storage nodes;
[0072] Step 102: Structuring the service-related information of the plurality of normal storage nodes obtained to obtain structured service-related information;
[0073] Step 103: pre-process the structured service-related information to obtain processed service-related information;
[0074] Step 104: Use the processed service-related information to train a generative adversarial network, where the generative adversarial network includes a first sub-model and a second sub-model. During the training process, the first sub-model is used to generate sample data, and the second sub-model is used to distinguish between real data and generated data. The second sub-model has a self-attention mechanism; the trained second sub-model is used to detect whether there is an abnormality in the storage node.
[0075] Here, in actual application, the storage system may include an object storage system, and the storage node may include an object storage node; the object storage system may specifically provide object storage services, and the object storage system includes multiple object storage nodes, and the object storage node may be understood as an entity (or a basic component unit) of the object storage system, and the object storage node may be specifically used to store, retrieve, and manage (can also be understood as processing, which may include backup, monitoring, etc.) object data, etc.
[0076] The storage node may also be called a storage server, a storage bucket, etc. The embodiment of the present application does not limit the name of the storage node.
[0077] When providing storage services for large-scale data, the storage system can store data in a dispersed manner on multiple storage nodes, i.e., perform distributed storage. In the case of distributed storage, the storage system can detect each storage node in parallel (i.e., detect multiple storage nodes at the same time) to detect whether each storage node has an abnormality, which can fully utilize the computing resources of the storage system and improve detection efficiency.
[0078] Before detecting whether a storage node is abnormal, the storage system can obtain service-related information from a storage node in a normal state, and use the obtained service-related information to train a model for detecting abnormal storage nodes, that is, a model for detecting whether a storage node is abnormal. The specific implementation of the service-related information can be set according to actual needs, such as including the amount of storage data of the storage node, the temperature of the processor (such as the central processing unit (CPU), the processor load and other information, and the service-related information can also be called node service information. The embodiment of the present application does not limit the name and specific implementation of the service-related information.
[0079] In actual application, the storage system may predetermine that multiple storage nodes among the storage nodes are in a normal state, that is, predetermine multiple normal storage nodes. Here, the manner in which the storage system determines multiple normal storage nodes may be understood according to relevant technologies, such as determining multiple normal storage nodes based on historical detection results, etc., and the embodiments of the present application do not limit this.
[0080] After determining multiple normal storage nodes, in step 101, the storage system can obtain service-related information of the multiple normal storage nodes. Specifically, in the storage system, the service-related information corresponding to each normal storage node can be obtained through the first sub-processor corresponding to the normal storage node. That is to say, the service-related information of multiple normal storage nodes can be obtained in parallel by using the multiple first sub-processors in the storage system, which can also be understood as obtaining the service-related information of multiple normal storage nodes in a distributed processing manner. By processing the data of each normal storage node in a parallel processing manner, the computing resources of each first sub-processor in the storage system can be fully utilized, the processing efficiency can be improved, and the processing time can be saved. Among them, the storage system may include multiple sub-processors, and the multiple sub-processors include the multiple first sub-processors.
[0081] After obtaining service-related information of multiple normal storage nodes, in step 102, multiple first sub-processors in the storage system may perform structured processing on the service-related information obtained by themselves to obtain structured service-related information. That is, in one embodiment, the specific implementation of step 102 may include:
[0082] A distributed processing method is adopted to perform structured processing on the service-related information of multiple normal storage nodes.
[0083] Here, in actual application, the structured processing of service-related information may include: using service-related information to generate structured service-related information (also understood as structured data) according to a preset structure, that is, data that conforms to the same preset structure, so that the storage system can manage and analyze service-related information, which helps to improve the efficiency and accuracy of subsequent information processing. Among them, the structured service-related information may include storage node statistics (also referred to as storage bucket statistics) and data distribution information (such as object distribution data); the storage node statistics may include information such as the amount of stored data (also understood as the number of objects), the amount of storage space occupied, etc.; the data distribution information may include information such as data naming rules, the frequency of use of common data naming prefixes, and the frequency of use of common data naming suffixes.
[0084] After the structured processing, in step 103, the multiple second sub-processors in the storage system can respectively obtain the structured service-related information corresponding to each normal storage node, and pre-process the structured service-related information; wherein the second sub-processor may include the first sub-processor, or may include other sub-processors in the storage system except the first sub-processor. That is, in one embodiment, the specific implementation of step 103 may include:
[0085] A distributed processing method is used to pre-process structured service-related information.
[0086] Here, in actual application, preprocessing the structured service-related information may include: performing data encoding operations and / or normalization operations on the structured service-related information to obtain information suitable for subsequent model training (i.e., the processed service-related information).
[0087] Specifically, the data encoding operation for the structured service-related information includes: converting the non-numeric information (also understood as non-numeric features) in the structured service-related information into numerical representation (also understood as conversion into numerical information), such as one-hot encoding the name of the storage node, label encoding the naming rules, etc. The numerical information obtained through the data encoding operation can be directly used in the subsequent model training process, which can reduce the complexity and training time of model training.
[0088] Normalizing structured service-related information includes: scaling numerical information, such as scaling the corresponding value of the information to a range of 0-1, or using a standardization method to make the values of each numerical information have a similar range (also understood as a scale). By performing normalization operations, the dimensional differences between different information can be eliminated. When the model is trained using the information after the normalization operation, faster model convergence speed, higher model accuracy and stability can be obtained.
[0089] In actual applications, since there are few cases where storage nodes are abnormal, it is usually difficult to obtain enough abnormal sample data (i.e., data related to abnormal storage nodes) for training the model used to distinguish whether the storage nodes are abnormal (i.e., the model that distinguishes normal storage nodes from abnormal storage nodes). In this case, if the model is still trained using normal sample data (i.e., data related to normal storage nodes) and abnormal sample data, it may lead to class imbalance and it is difficult to ensure the accuracy of the trained model.
[0090] Based on this, when training the model, a small number of abnormal sample data can be discarded, and the sample data generated by the generator model in the generative adversarial network (i.e., generated data) can be used as abnormal sample data (also understood as abnormal data). In this way, the generated sample data and normal sample data can be used to train the discriminator model, so that the discriminator model forms a feature pattern mainly around the normal sample data, and the trained discriminator model can be used to accurately identify data that is not different from the normal sample data, which has a better exclusion effect. Therefore, when the abnormal sample data is input into the discriminator model, it can be accurately identified, so as to determine that the storage node corresponding to the sample data is abnormal, and the class imbalance problem can be solved.
[0091] Specifically, after obtaining the processed service-related information, in step 104, the storage system can use the processed service-related information to perform generative adversarial network training to obtain a model that can accurately and efficiently detect whether there are abnormalities in the storage node. The generative adversarial network includes a first sub-model for generating sample data and a second sub-model for distinguishing between real data and generated data. The first sub-model can also be understood as a generator model, and the second sub-model can also be understood as a discriminator model. The embodiments of the present application do not limit the names of the first sub-model and the second sub-model.
[0092] In actual application, the processed service-related information can be input into the first sub-model as real data (which can also be understood as real samples or real sample data), and the first sub-model is used to extract the features of the real data, and the extracted features are used to learn the distribution of the real data; then, based on the extracted features and the learned data distribution, generated data close to the real data is generated. Then, the generated data and the real data (i.e., the processed service-related information) can be input into the second sub-model, and the second sub-model can be used to determine which data belongs to the generated data and which data belongs to the real data to obtain a judgment result. Then, the error between the judgment result and the actual result (i.e., the real result) can be used to determine the performance of the second sub-model. Specifically, when the error is greater than a preset threshold, it can be considered that the performance of the second sub-model cannot meet the demand. At this time, the parameters of the first sub-model can be adjusted to achieve the use of the first sub-model to generate generated data that is closer to the real data; and by adjusting the parameters of the second sub-model, the second sub-model can be used to more sensitively (i.e., more accurately) distinguish between real data and generated data. The above-mentioned generative adversarial training is re-performed using the first sub-model and the second sub-model after adjusting the parameters (i.e., using the first sub-model to generate sample data, and using the second sub-model to distinguish between real data and generated data). A second sub-model with a smaller error can be obtained, that is, the accuracy and robustness of the second sub-model are improved (i.e., the second sub-model is improved); when the error is less than or equal to the preset threshold, it can be considered that the performance of the second sub-model can meet the demand. At this time, the second sub-model can be used as a model for detecting whether there is an abnormality in the storage node (i.e., the trained second sub-model).
[0093] The first sub-model and the second sub-model are described below respectively.
[0094] For the first sub-model: the first sub-model can adopt a long short-term memory network (LSTM) model. The LSTM model can be used to accurately and efficiently extract the features of time series data. Under the premise that the processed service-related information is associated with time (that is, it belongs to time series data), the first sub-model adopts the LSTM model to achieve better feature extraction effect.
[0095] In order to generate data closer to real data, in one embodiment, the loss function of the first sub-model may adopt a loss function associated with the Minkowski distance. In this case, the training objective of the first sub-model may be expressed as formula (1):
[0096] min(E X~pt[||C(R(X))-1||2+||C(R(X)-1||]) (1)
[0097] Wherein, X represents the input of the first sub-model (i.e., real data, which can also be understood as training data), R(X) represents the output of the first sub-model (i.e., generated data), C(R(X)) represents the output of the second sub-model, i.e., the distinction result of the second sub-model on the data generated by the first sub-model, ||||2 represents the Euclidean norm (which can also be understood as the Euclidean distance), |||| represents the Manhattan norm (which can also be understood as the Manhattan distance), and ||C(R(X))-1||2 represents the reconstruction loss (which can also be expressed as L R ), ||C(R(X))-1|| is part of the loss function of the generative adversarial network, which can also be expressed as pt represents the data distribution pattern of a normal storage node, and E represents the expected value, that is, the calculation result of the expected value calculation of pt based on the input of the first sub-model.
[0098] At the same time, since the real data input into the first sub-model (i.e., the processed service-related information) is only associated with normal storage nodes, the data generated by the first sub-model (i.e., the output of the first sub-model) is also close to the data corresponding to normal storage nodes. After training the second sub-model with the data generated by the first sub-model, a second sub-model (i.e., the trained second sub-model) that can distinguish between real data and generated data can be obtained. Since the data corresponding to the abnormal storage node is different from the data corresponding to the normal storage node, i.e., it is not close to the real data, the trained second sub-model can accurately identify the data corresponding to the abnormal storage node as generated data. In this way, the identification result (i.e., the output of the second sub-model) can be used to characterize the input data as a storage node for generated data, and it can be determined as a storage node with an abnormality (i.e., an abnormal storage node).
[0099] For the second sub-model: the second sub-model can adopt an MLP architecture, which includes multiple layers, including an input layer, one or more hidden layers, and an output layer; at the same time, the MLP architecture has high flexibility, scalability, and robustness. In view of the fact that there may be a large amount of data overlap in the service-related information of the storage node, the model using the MLP architecture can efficiently, flexibly, and accurately distinguish between generated data and real data.
[0100] In actual application, the input of the second sub-model may include data that needs to be distinguished (which can also be understood as data to be distinguished), and the output of the second sub-model may include a single value, which represents the probability that the input data belongs to the generated data, and can also be understood as the probability that the storage node corresponding to the input data has an abnormality. The training objective of the second sub-model can be expressed as formula (2):
[0101] max(E X~pz (λ||R(X)||+||C(R(X))||2+||X-1||2)) (2)
[0102] Among them, pz represents a set of subsequences extracted from the random space, λ represents a random parameter, and λ can be specifically set to 1.85.
[0103] In order to further improve the generalization ability and robustness of the second sub-model, a self-attention mechanism is introduced into the second sub-model, that is, the second sub-model has a self-attention mechanism. After the introduction of the self-attention mechanism, the second sub-model can be used to ignore irrelevant information and focus on key information (that is, have attention), thereby increasing the accuracy and efficiency of distinguishing real data from generated data.
[0104] Specifically, when the second sub-model adopts the MLP architecture, in one embodiment, one or more feature processing modules may be included between multiple layers of the MLP architecture, and the feature processing modules are associated with the self-attention mechanism. The above steps can also be understood as setting one or more self-attention mechanism modules between multiple layers included in the MLP architecture.
[0105] Exemplarily, the process of processing features by the feature processing module may include: Figure 2As shown, assuming that the second sub-model adopts the MLP architecture, and the feature processing module is located between the Ath layer and the A+1th layer in the MLP architecture, at this time, the output sequence of the Ath layer can be used as the input sequence of the feature processing module. After the feature processing module obtains the input sequence, it performs 1*1 convolution on each data in the input sequence (which can also be understood as each element) to obtain the corresponding linear mapping f of the query vector, the linear mapping g of the key vector, and the linear mapping h of the value vector; then, the linear mapping f of the query vector and the linear mapping g of the key vector can be calculated for similarity to obtain a first weight, and the first weight is normalized to obtain a normalized first weight; finally, the normalized first weight and the linear mapping h of the value vector are weightedly summed to obtain an attention value (which can be expressed as attention in English and can also be understood as an output feature), and the attention value is used as the input of the A+1 layer. Among them, the query vector represents the mapping of the current position (i.e., the location where the data is located); the key vector represents the mapping of other positions, which is used to measure the correlation between other positions and the current position; the value vector represents the specific information of different positions in the input sequence; similarity functions such as dot product, concatenation, and perceptron can be used for similarity calculation, and softmax function can be used for normalization processing, which is not limited in the embodiment of the present application; the initial value of the first weight can be set to 0. In the initial case, the output of the feature module is the same as the input, and it can also be understood as only relying on the elements of the input sequence itself (i.e., local information). As the training of the second sub-model proceeds, the first weight of the feature module can be updated, such as increasing the non-local weight, so as to learn the correlation between different positions, capture the long-distance dependency in the sequence, and enrich the features obtained in the A+1 layer in the MLP architecture, thereby improving the performance of the second sub-model (such as discrimination accuracy).
[0106] From the above description, it can be seen that using the processed service-related information to train the generative adversarial network can generate sample data that is closer to the real data by continuously optimizing the first sub-model, and continuously optimizing the second sub-model to achieve a more accurate effect of distinguishing the real data from the generated data. At this time, the overall training goal of the generative adversarial network can be expressed as formula (3)
[0107]
[0108] Among them, G represents the first sub-model, D represents the second sub-model, and ν(D,G) represents the overall loss function of the generative adversarial network.
[0109] After the second sub-model is trained, the second sub-model can be deployed in the storage system to detect whether there is an abnormality in the storage node using the deployed second sub-model. Specifically, a second sub-model can be deployed for each storage node (which can also be understood as a distributed deployment of the second sub-model) to achieve parallel detection of each storage node and improve detection efficiency.
[0110] Based on this, the embodiment of the present application further provides a detection method, which is applied to the third sub-processor of the storage system, such as Figure 3 As shown, the method includes:
[0111] Step 301: Obtain service-related information of a first storage node, where the first storage node is at least used to store data;
[0112] Step 302: Structuring the service-related information of the first storage node to obtain structured service-related information;
[0113] Step 303: pre-process the structured service-related information to obtain processed service-related information;
[0114] Step 304: Utilize the processed service-related information in combination with the second sub-model to detect whether there is an abnormality in the first storage node, where the second sub-model is trained using the above-mentioned model training method.
[0115] Here, in actual application, the storage system may include multiple third sub-processors, and the third sub-processors may include the first sub-processor and / or the second sub-processor. Each of the third sub-processors may correspond to a storage node, and is used to execute the above detection method for the storage node.
[0116] Specifically, in one embodiment, the first storage node belongs to a storage system, and the storage system includes a plurality of storage nodes;
[0117] In the process of detecting whether a storage node is abnormal, at least one of the following is satisfied:
[0118] Service-related information of multiple storage nodes is structured using distributed processing;
[0119] The structured service-related information of multiple storage nodes is pre-processed in a distributed processing manner;
[0120] Multiple storage nodes are processed in a distributed manner using the second sub-model to detect whether there are abnormalities in the storage nodes.
[0121] In actual application, the specific implementation of steps 301 to 303 can be understood by referring to the above steps 101 to 103, and the embodiments of the present application are not limited to this.
[0122] The method provided by the embodiment of the present application is that the storage system obtains service-related information of multiple normal storage nodes; the storage system includes multiple storage nodes, and the storage nodes are at least used to store data; the service-related information of the obtained multiple normal storage nodes is structured to obtain structured service-related information; the structured service-related information is preprocessed to obtain processed service-related information; the processed service-related information is used to train the generative adversarial network, and the generative adversarial network includes a first sub-model and a second sub-model. During the training process, the first sub-model is used to generate sample data, and the second sub-model is used to distinguish between real data and generated data, and the second sub-model has a self-attention mechanism; the trained second sub-model is used to detect whether the storage node is abnormal. At the same time, the electronic device obtains the service-related information of the first storage node, and the first storage node is at least used to store data; the service-related information of the first storage node is structured to obtain structured service-related information; the structured service-related information is preprocessed to obtain processed service-related information; the processed service-related information is combined with the second sub-model to detect whether the first storage node is abnormal, and the second sub-model is trained using any of the above-mentioned model training methods. The solution provided in the embodiment of the present application detects anomalies of storage nodes through a generative adversarial network. On the one hand, the generative adversarial network can autonomously generate sample data required for training, which can ensure the scale of sample data used for training and improve the performance of the model obtained through training. On the other hand, the training efficiency of the generative adversarial network is high and the model performance obtained through adversarial training is better. At the same time, a self-attention mechanism is introduced into the generative adversarial network, and the performance of the model obtained through training can be further improved by utilizing the characteristics of the self-attention mechanism that considers global information and can process complex data. In this way, when the trained model is used to detect whether there are anomalies in the storage node, high-accuracy and efficient detection can be achieved.
[0123] The present application is described in further detail below in conjunction with application examples.
[0124] The application example of this application proposes an abnormal monitoring system for object storage service node maintenance based on deep learning (i.e., the above storage system), such as Figure 4As shown, it includes an object storage service module, a structured data generation module, a data preprocessing module, a generative adversarial network module, and an anomaly monitoring module. Among them, the object storage service module includes N data nodes (i.e., the above-mentioned storage nodes), and the data nodes are used to store object data (i.e., provide object storage services); the structured data generation module includes N sub-structured modules; the data preprocessing module includes N sub-preprocessing modules; the anomaly monitoring module includes N sub-monitoring modules, and N is an integer greater than 1; at the same time, each data node corresponds to a sub-structured module, a sub-preprocessing module, and a sub-detection module, and the sub-structured module, sub-preprocessing module, and sub-detection module corresponding to each data node can use the same processor or different processors. The object storage service module, structured data generation module, data preprocessing module, generative adversarial network module, and anomaly monitoring module can be deployed centrally or distributed as multiple devices.
[0125] Based on the above system, the application example of this application also proposes an abnormal monitoring method for object storage service node maintenance based on deep learning. By applying deep learning technology to object storage and monitoring service node maintenance, the normal operation of the object storage service can be ensured. Figure 5 As shown, the following steps are included:
[0126] Step 501: Each sub-structuring module obtains the node service information of the corresponding data node (i.e. the above service related information), and generates structured data (i.e. the above structured service related information, which can also be understood as structured training data) using the obtained node service information, and then sends the structured data to the corresponding sub-preprocessing module;
[0127] Step 502: After receiving the structured data, the sub-preprocessing module preprocesses the structured data to obtain processed data (i.e. the processed service-related information);
[0128] Here, in actual application, the preprocessing may include performing data encoding operations and / or normalization operations.
[0129] Step 503: The generative adversarial network module obtains the data processed in all the sub-preprocessing modules, and uses all the acquired data to train the generative adversarial network to obtain a trained discriminator model;
[0130] The generative adversarial network includes a generator model (i.e., the first sub-model mentioned above) and a discriminator model (i.e., the second sub-model mentioned above); the generator adopts an LSTM architecture to generate sample data similar to normal node data, and the training objective of the generator can be expressed as formula (1); the discriminator adopts an MLP architecture based on a self-attention mechanism (also referred to as an aMLP architecture) to distinguish generated data from real data, and the training objective of the discriminator can be expressed as formula (2).
[0131] In actual application, the discriminator is used to distinguish generated data from real data to obtain a distinction result. According to the distinction result, it can be determined which data is real data. The generative adversarial network module can further use the determined real data to train the generator.
[0132] In actual application, steps 501 to 503 may also be referred to as a training process. After obtaining the trained discriminator model, steps 504 to 505 may be performed to detect each data node, and steps 504 to 505 may also be referred to as a detection process.
[0133] Step 504: deploying the trained discriminator model in each sub-detection module;
[0134] Step 505: Each sub-structured module obtains the node service information of the corresponding data node, and uses the obtained node service information to generate structured data (which can also be understood as structured monitoring data), and then sends the structured data to the corresponding sub-preprocessing module; after receiving the structured data, the sub-preprocessing module preprocesses the structured data to obtain processed data, and sends the processed data to the sub-detection module; the sub-detection module uses the deployed trained discriminator model and the received processed data to detect whether the data node has an abnormality and obtain a detection result.
[0135] Here, in actual application, the detection result may include a probability value of the existence of an abnormality in the data node. When the probability value is greater than a preset threshold, the data node can be considered to be an abnormal node (i.e., there is an abnormality), which can trigger an alarm and / or related response processing (which can also be understood as an early warning).
[0136] The solution provided in the application example of this application applies the generative adversarial network to the anomaly monitoring of the object storage service node. Specifically, LSTM and MLP are used as the generator and discriminator models of the generative adversarial network respectively, and the self-attention mechanism is added to the MLP, which can better improve the discriminator's recognition ability (which can also be understood as having an efficient algorithm design) and strengthen the abnormal data monitoring ability; at the same time, the model is optimized in combination with the characteristics of the object storage service node (such as the throughput of the data stored in the storage service node, CPU usage rate and other information) to make it more suitable for anomaly detection needs and improve detection effect and efficiency. At the same time, the use of distributed processing for model training and detection can improve detection efficiency, and has strong scalability and can adapt to large-scale data processing scenarios. In other words, it can better cope with the anomaly monitoring needs of massive data in the object storage service and provide a more efficient and reliable anomaly detection solution.
[0137] Specific advantages include:
[0138] Efficient algorithm design: In response to the needs of large-scale data processing, we designed efficient algorithms and data processing processes to improve processing speed and efficiency. Specifically, by proposing a loss function associated with the Minkowski distance, we increased the training rate and optimized the time complexity of the algorithm, which can process large-scale data faster and reduce processing time and resource consumption;
[0139] High scalability: In view of the ever-increasing data size in the object storage service, a system architecture and algorithm model with good scalability are designed. Through reasonable distributed storage (i.e., data storage through multiple data nodes), the space complexity is optimized to avoid excessive load on a single storage node; at the same time, through distributed computing (i.e., structured processing, preprocessing, and detection for each data node separately) strategy, the system's processing capacity can be seamlessly expanded according to actual needs (such as adding new data nodes, etc.), better adapting to the growing data size and maintaining high performance;
[0140] Highly customized: The model is optimized and adapted according to the characteristics of the object storage service node, providing customized solutions for the needs of this field. By fully understanding the behavior patterns and abnormal situations of the object storage service node, a more accurate and reliable anomaly detection model can be designed to improve the accuracy and robustness of anomaly detection;
[0141] Low cost: By applying deep learning models to object storage service node monitoring, the monitoring cost is greatly reduced compared to manual monitoring. The automated anomaly monitoring function can analyze node status and performance indicators in real time and provide accurate anomaly detection and early warning. Monitoring personnel can make secondary judgments and processing based on the monitoring results provided by the model, which improves monitoring efficiency and reduces human resource requirements. Once this automated monitoring method is trained and deployed, there is almost no additional cost, saving time and human resources;
[0142] Identify subtle faults: By extracting feature changes, identify subtle faults, and solve the problem that is difficult to detect by human monitoring. In object storage services, there are some subtle faults that may not seem significant individually, but may cause the entire system to crash as they accumulate. By applying deep learning models, changes in node features can be analyzed, these fault signs can be identified in a timely manner, and early warnings and repairs can be carried out.
[0143] In order to implement the model training method of the embodiment of the present application, the embodiment of the present application also provides a model training device, which is arranged in a storage system including multiple storage nodes, and the storage nodes are at least used to store data, such as Figure 6 As shown, the device comprises:
[0144] A first acquisition unit 601 is used to acquire service related information of multiple normal storage nodes;
[0145] A first processing unit 602 is configured to perform structured processing on the service related information of the plurality of normal storage nodes to obtain structured service related information;
[0146] The second processing unit 603 is used to pre-process the structured service related information to obtain processed service related information;
[0147] The training unit 604 is used to train a generative adversarial network using processed service-related information, wherein the generative adversarial network includes a first sub-model and a second sub-model. During the training process, the first sub-model is used to generate sample data, and the second sub-model is used to distinguish between real data and generated data, and the second sub-model has a self-attention mechanism; the trained second sub-model is used to detect whether there is an abnormality in the storage node.
[0148] In one embodiment, the first processing unit 602 is specifically configured to:
[0149] A distributed processing method is adopted to perform structured processing on the service-related information of multiple normal storage nodes.
[0150] In one embodiment, the second processing unit 603 is specifically configured to:
[0151] A distributed processing method is used to pre-process structured service-related information.
[0152] In actual application, the first acquisition unit 601 can be implemented by a processor in the model training device in combination with a communication interface, and the first processing unit 602, the second processing unit 603, and the training unit 604 can be implemented by a processor in the model training device.
[0153] It should be noted that: the model training device provided in the above embodiment only uses the division of the above program units as an example when performing model training. In actual applications, the above processing can be assigned to different program units as needed, that is, the internal structure of the device is divided into different program units to complete all or part of the processing described above. In addition, the model training device provided in the above embodiment and the model training method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0154] In order to implement the detection method of the embodiment of the present application, the embodiment of the present application also provides a detection device, which is arranged on an electronic device, such as Figure 7 As shown, the device comprises:
[0155] A second acquisition unit 701 is used to acquire service related information of a first storage node, where the first storage node is at least used to store data;
[0156] The third processing unit 702 is used to perform structured processing on the service related information of the first storage node to obtain structured service related information;
[0157] The fourth processing unit 703 is used to pre-process the structured service-related information to obtain processed service-related information;
[0158] The detection unit 704 is used to use the processed service-related information in combination with the second sub-model to detect whether there is an abnormality in the first storage node, and the second sub-model is trained using any of the above-mentioned model training methods.
[0159] In actual application, the second acquisition unit 701 can be implemented by a processor in the detection device in combination with a communication interface, and the third processing unit 702, the fourth processing unit 703, and the detection unit 704 can be implemented by a processor in the detection device.
[0160] It should be noted that: when the detection device provided in the above embodiment performs detection, only the division of the above program units is used as an example. In actual applications, the above processing can be assigned to different program units as needed, that is, the internal structure of the device is divided into different program units to complete all or part of the processing described above. In addition, the detection device provided in the above embodiment and the detection method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0161] Based on the hardware implementation of the above program modules, and in order to implement the model training method of the embodiment of the present application, the embodiment of the present application also provides a first electronic device, which is applied to a storage system including multiple storage nodes, and the storage nodes are at least used to store data, such as Figure 8 As shown, the first electronic device 800 includes:
[0162] The first communication interface 801 is capable of exchanging information with other devices;
[0163] A first processor 802 is connected to the first communication interface 801 to implement information interaction with other devices, and is used to execute the model training method provided by one or more of the above technical solutions when running a computer program;
[0164] A first memory 803 , in which the computer program is stored.
[0165] Specifically, the first communication interface 801 is used to:
[0166] Obtain service-related information of multiple normal storage nodes;
[0167] The first processor 802 is configured to:
[0168] The service-related information of multiple normal storage nodes obtained is structured to obtain structured service-related information; the structured service-related information is preprocessed to obtain processed service-related information; and the processed service-related information is used to train a generative adversarial network, wherein the generative adversarial network includes a first sub-model and a second sub-model. During the training process, the first sub-model is used to generate sample data, and the second sub-model is used to distinguish between real data and generated data, and the second sub-model has a self-attention mechanism; the trained second sub-model is used to detect whether there is an abnormality in the storage node.
[0169] In one embodiment, the first processor 802 is specifically configured to:
[0170] A distributed processing method is adopted to perform structured processing on the service-related information of multiple normal storage nodes.
[0171] In one embodiment, the first processor 802 is specifically configured to:
[0172] A distributed processing method is used to pre-process structured service-related information.
[0173] It should be noted that the specific processing process of the first processor 802 and the first communication interface 801 can be understood by referring to the above-mentioned model training method.
[0174] Of course, in actual application, the various components in the first electronic device 800 are coupled together through the bus system 804. It can be understood that the bus system 804 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 804 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 8 Various buses are labeled as bus system 804 .
[0175] The first memory 803 in the embodiment of the present application is used to store various types of data to support the operation of the first electronic device 800. Examples of such data include: any computer program used to operate on the first electronic device 800.
[0176] The method disclosed in the above embodiment of the present application can be applied to the first processor 802, or implemented by the first processor 802. The first processor 802 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware or software instructions in the first processor 802. The above-mentioned first processor 802 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The first processor 802 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in the first memory 803, and the first processor 802 reads the information in the first memory 803 and completes the steps of the above method in combination with its hardware.
[0177] In an exemplary embodiment, the first electronic device 800 can be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned method.
[0178] Based on the hardware implementation of the above program modules, and in order to implement the detection method of the embodiment of the present application, the embodiment of the present application also provides a second electronic device, such as Fig. 9 As shown, the second electronic device 900 includes:
[0179] The second communication interface 901 is capable of exchanging information with other devices;
[0180] A second processor 902, connected to the second communication interface 901 to implement information interaction with other devices, and configured to execute the detection method provided by one or more of the above technical solutions when running a computer program;
[0181] A second memory 903 , in which the computer program is stored.
[0182] Specifically, the second communication interface 901 is used to:
[0183] Acquire service-related information of a first storage node, where the first storage node is at least used to store data;
[0184] The second processor 902 is configured to:
[0185] Structural processing is performed on the service-related information of the first storage node to obtain structured service-related information; the structured service-related information is preprocessed to obtain processed service-related information; the processed service-related information is combined with a second sub-model to detect whether there is an abnormality in the first storage node, and the second sub-model is trained using any of the above-mentioned model training methods.
[0186] It should be noted that the specific processing process of the second processor 902 and the second communication interface 901 can be understood by referring to the above detection method.
[0187] Of course, in actual application, the various components in the second electronic device 900 are coupled together through the bus system 904. It can be understood that the bus system 904 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 904 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig. 9 Various buses are labeled as bus system 904.
[0188] The second memory 903 in the embodiment of the present application is used to store various types of data to support the operation of the second electronic device 900. Examples of such data include: any computer program used to operate on the second electronic device 900.
[0189] The method disclosed in the above embodiment of the present application can be applied to the second processor 902, or implemented by the second processor 902. The second processor 902 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the second processor 902 or an instruction in the form of software. The above-mentioned second processor 902 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The second processor 902 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in the second memory 903, and the second processor 902 reads the information in the second memory 903 and completes the steps of the above method in combination with its hardware.
[0190] In an exemplary embodiment, the second electronic device 900 may be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general processors, controllers, MCUs, Microprocessors, or other electronic components to perform the aforementioned methods.
[0191] It can be understood that the memory (first memory 803, second memory 903) of the embodiment of the present application can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a ferromagnetic random access memory, a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct RAM bus random access memory (DRRAM).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.
[0192] In an exemplary embodiment, the embodiment of the present application further provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a first memory 803 storing a computer program, the computer program can be executed by the first processor 802 of the first electronic device 800 to complete the steps of the aforementioned model training method, and for another example, including a second memory 903 storing a computer program, the computer program can be executed by the second processor 902 of the second electronic device 900 to complete the steps of the aforementioned detection method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.
[0193] In an exemplary embodiment, the embodiment of the present application also provides a computer program product, including a computer program, which can be executed by a first processor 802 of a first electronic device 800 to complete the steps of the aforementioned model training method, or the computer program can be executed by a second processor 902 of a second electronic device 900 to complete the steps of the aforementioned detection method.
[0194] In order to implement the method provided in the embodiment of the present application, the embodiment of the present application also provides a detection system, such as Fig.10 As shown, the system includes: a first electronic device 1001 and a second electronic device 1002.
[0195] Here, it should be noted that the specific processing procedures of the first electronic device 1001 and the second electronic device 1002 have been described in detail above and will not be repeated here.
[0196] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0197] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0198] The above description is only a preferred embodiment of the present application and is not intended to limit the protection scope of the present application.
Claims
1. A model training method, characterized in that: Applied to a storage system comprising a plurality of storage nodes, wherein the storage nodes are at least used to store data, the method comprises: Obtain service-related information of multiple normal storage nodes; Structuring the service-related information of the plurality of normal storage nodes obtained to obtain structured service-related information; Preprocessing the structured service-related information to obtain processed service-related information; The generative adversarial network is trained using the processed service-related information, and the generative adversarial network includes a first sub-model and a second sub-model. During the training process, the first sub-model is used to generate sample data, and the second sub-model is used to distinguish between real data and generated data. The second sub-model has a self-attention mechanism; the trained second sub-model is used to detect whether there is an abnormality in the storage node.
2. The method according to claim 1, characterized in that The structural processing of the service-related information of the plurality of normal storage nodes obtained includes: A distributed processing method is adopted to perform structured processing on the service-related information of multiple normal storage nodes.
3. The method according to claim 1, characterized in that The preprocessing of structured service related information includes: A distributed processing method is used to pre-process structured service-related information.
4. The method according to any one of claims 1 to 3, characterized in that: The loss function corresponding to the first sub-model is associated with the Minkowski distance.
5. The method according to any one of claims 1 to 3, characterized in that: The second sub-model adopts a multi-layer perceptron MLP architecture, and the multiple layers of the MLP architecture include one or more feature processing modules, and the feature processing modules are associated with a self-attention mechanism.
6. A detection method, characterized in that: include: Acquire service-related information of a first storage node, where the first storage node is at least used to store data; Structuring the service-related information of the first storage node to obtain structured service-related information; Preprocessing the structured service-related information to obtain processed service-related information; The processed service-related information is used in combination with a second sub-model to detect whether the first storage node has an abnormality, and the second sub-model is trained using the method described in any one of claims 1 to 5.
7. The method according to claim 6, characterized in that The first storage node belongs to a storage system, and the storage system includes multiple storage nodes; In the process of detecting whether a storage node is abnormal, at least one of the following is satisfied: Service-related information of multiple storage nodes is structured using distributed processing; The structured service-related information of multiple storage nodes is pre-processed in a distributed processing manner; Multiple storage nodes are processed in a distributed manner using the second sub-model to detect whether there are abnormalities in the storage nodes.
8. A model training device, characterized in that: A storage system is provided including a plurality of storage nodes, wherein the storage nodes are used at least to store data, including: A first acquisition unit, used to acquire service related information of multiple normal storage nodes; A first processing unit is used to perform structured processing on the service-related information of the plurality of normal storage nodes to obtain structured service-related information; A second processing unit is used to pre-process the structured service related information to obtain processed service related information; A training unit is used to train a generative adversarial network using processed service-related information, wherein the generative adversarial network includes a first sub-model and a second sub-model. During the training process, the first sub-model is used to generate sample data, and the second sub-model is used to distinguish between real data and generated data, and the second sub-model has a self-attention mechanism; the trained second sub-model is used to detect whether there is an abnormality in the storage node.
9. A detection device, characterized in that: include: A second acquisition unit is used to acquire service related information of a first storage node, where the first storage node is at least used to store data; A third processing unit is used to perform structured processing on the service related information of the first storage node to obtain structured service related information; A fourth processing unit, configured to pre-process the structured service-related information to obtain processed service-related information; A detection unit is used to use the processed service-related information in combination with a second sub-model to detect whether there is an abnormality in the first storage node, and the second sub-model is trained using the method described in any one of claims 1 to 5.
10. A first electronic device, characterized in that: include: a first processor and a first memory for storing a computer program executable on the processor, Wherein, when the first processor is used to run the computer program, the steps of the method described in any one of claims 1 to 5 are executed.
11. A second electronic device, characterized in that: include: a second processor and a second memory for storing a computer program executable on the processor, Wherein, when the second processor is used to run the computer program, the steps of the method described in any one of claims 6 to 7 are executed.
12. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 5, or implements the steps of the method according to any one of claims 6 to 7.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 5, or implements the steps of the method according to any one of claims 6 to 7.