Abnormal storage unit positioning method and distributed storage system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ALIBABA CLOUD COMPUTING CO LTD
- Filing Date
- 2025-02-07
- Publication Date
- 2026-08-07
AI Technical Summary
然而,异常的存储单元会导致其他服务节点故障,甚至引发连锁反应,造成更大范围的服务节点故障,影响了系统的稳定性
[0011]本说明书一个实施例提供的异常存储单元定位方法,应用于分布式存储系统中的控制节点,分布式存储系统包括多个服务节点和控制节点,包括:获取目标存储单元的风险因素,其中,目标存储单元为目标服务节点中的存储单元,目标服务节点为多个服务节点中的故障服务节点,风险因素为导致目标存储单元出现异常的条件;根据风险因素,对目标存储单元进行离群点检测,获得检测结果,其中,检测结果用于指示目标存储单元中的异常存储单元。基于风险因素进行离群点检测,并利用检测结果定位异常存储单元,无需将异常存储单元调度至其他服务节点即可实现异常存储单元精准定位,避免了潜在的连锁故障风险,增强了系统的可靠性和稳定性。
Smart Images

Figure CN122526862A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of computer technology, and in particular to methods for locating abnormal storage units and distributed storage systems. Background Technology
[0002] In large-scale distributed storage systems, when a storage unit malfunctions, it causes the corresponding service node to fail. Since a service node typically handles read and write requests for multiple storage units simultaneously, the malfunction of a single storage unit affects other storage units within the same service node. Therefore, to ensure high availability and stability of the system, quickly and accurately locating the malfunctioning storage unit is crucial.
[0003] Currently, storage units from a failed service node can typically be relocated to other service nodes via a control node, and the malfunctioning storage unit can be located based on the node status of the other service nodes. However, malfunctioning storage units can cause other service nodes to fail, and even trigger a chain reaction, resulting in a wider range of service node failures and affecting system stability. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a method for locating abnormal storage units. One or more embodiments of this specification also relate to an abnormal storage unit location device, a distributed storage system, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, an abnormal storage unit location method is provided, applied to a control node in a distributed storage system, the distributed storage system including multiple service nodes and a control node, including: Obtain the risk factors of the target storage unit, where the target storage unit is the storage unit in the target service node, the target service node is the faulty service node among multiple service nodes, and the risk factor is the condition that causes the target storage unit to malfunction. Based on risk factors, outlier detection is performed on the target storage unit to obtain detection results, which are used to indicate abnormal storage units in the target storage unit.
[0006] According to a second aspect of the embodiments of this specification, an abnormal storage unit location device is provided, applied to a control node in a distributed storage system. The distributed storage system includes multiple service nodes and a control node, comprising: The acquisition module is configured to acquire risk factors for the target storage unit, where the target storage unit is a storage unit in the target service node, the target service node is a faulty service node among multiple service nodes, and the risk factor is the condition that causes the target storage unit to malfunction. The detection module is configured to perform outlier detection on the target storage unit based on risk factors and obtain detection results, wherein the detection results are used to indicate abnormal storage units in the target storage unit.
[0007] According to a third aspect of the embodiments of this specification, a distributed storage system is provided, deployed on a cloud-side device, the distributed storage system including a control node and multiple service nodes; The control node is used to acquire risk factors of the target storage unit, where the target storage unit is a storage unit in the target service node, the target service node is a faulty service node among multiple service nodes, and the risk factor is the condition that causes the target storage unit to be abnormal; based on the risk factors, outlier detection is performed on the target storage unit to obtain the detection results, where the detection results are used to indicate abnormal storage units in the target storage unit.
[0008] According to a fourth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the abnormal storage unit location method proposed in the first aspect above.
[0009] According to a fifth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the abnormal storage cell location method proposed in the first aspect above.
[0010] According to a sixth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the abnormal memory cell location method proposed in the first aspect above.
[0011] This specification provides an embodiment of an abnormal storage unit location method, applied to a control node in a distributed storage system. The distributed storage system includes multiple service nodes and a control node. The method includes: acquiring risk factors for a target storage unit, where the target storage unit is a storage unit within a target service node, the target service node is a faulty service node among the multiple service nodes, and the risk factors are conditions that cause the target storage unit to malfunction; and performing outlier detection on the target storage unit based on the risk factors to obtain detection results, where the detection results are used to indicate abnormal storage units within the target storage unit. By performing outlier detection based on risk factors and locating abnormal storage units using the detection results, accurate location of abnormal storage units can be achieved without scheduling them to other service nodes, avoiding potential cascading failure risks and enhancing the reliability and stability of the system. Attached Figure Description
[0012] Figure 1 This is an architecture diagram of a distributed storage system; Figure 2 This is a schematic diagram of the processing flow of a control unit in a distributed storage system; Figure 3 This is an architecture diagram of another distributed storage system; Figure 4 This is a schematic diagram of the processing flow of the control unit in another distributed storage system; Figure 5 This is an architecture diagram of a distributed storage system provided in one embodiment of this specification; Figure 6 This is a flowchart of an abnormal storage unit location method provided in one embodiment of this specification; Figure 7 This is a schematic diagram of an abnormal storage unit locating device provided in one embodiment of this specification; Figure 8 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0016] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0017] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0018] Distributed storage systems are data storage architectures that distribute data across multiple independent devices, rather than centralizing it in a single storage node. These systems offer higher reliability, availability, and scalability, making them suitable for handling large-scale data storage needs. Distributed storage systems include distributed block storage systems.
[0019] Block storage system: refers to a low-latency, persistent, and highly reliable block-level random access storage system. A block storage system can divide data into fixed-size blocks and perform read and write operations on these blocks as units.
[0020] The control node (BM, Block Master) is one of the key components of the block storage system backend, responsible for managing and coordinating the core functions of resource allocation, data distribution, and fault recovery of the entire storage cluster.
[0021] Block Server (BS): One of the key components of the block storage system backend, responsible for data read and write operations and actual storage tasks. Each BS typically corresponds to one or more physical storage devices and provides a low-latency, high-throughput data access interface for use by front-end compute nodes or virtual machines.
[0022] Cluster: refers to a group of independent computers interconnected by a high-speed computer network, which form a group and are managed as a single system.
[0023] Virtual machine block devices (Device): These are storage resources allocated to virtual machines that mimic the behavior of physical disks. These block devices can be local hard drives, network storage, or persistent storage volumes provided by cloud service providers. They offer virtual machines a direct disk access experience similar to that on a physical server, allowing the operating system to mount and manage these virtual disks as if they were local disks.
[0024] Outliers are data points in a dataset that are significantly different from other observations. These points typically exhibit a clear deviation from the rest of the dataset.
[0025] Explosion radius: refers to the range of storage cells that are ultimately affected by an abnormal storage cell.
[0026] Abnormal cloud disks: These are cloud storage disks in a cloud storage system whose behavior deviates from normal expectations due to various reasons. These abnormalities may be caused by a variety of factors such as hardware failure, software errors, configuration problems, or external attacks.
[0027] In large-scale distributed storage clusters, data is typically divided into multiple blocks and distributed across different storage servers within the cluster. Each storage server serves multiple storage units. If a storage unit malfunctions, it triggers a service process error, causing the corresponding service node to fail. All storage units served on that service node are affected. Simultaneously, the storage unit's scheduling role (control node) triggers rescheduling to restore the affected service. Specifically, the control node reschedules all storage units on the failed service node to other normally functioning service nodes to quickly restore read and write services. If the malfunctioning storage unit is rescheduled to another service node, it will cause that receiving service node to fail again, and so on. Ultimately, this cycle leads to repeated failures across all service nodes in the entire storage service cluster, eventually rendering the entire cluster unusable. Therefore, to mitigate the widespread contagion caused by malfunctioning storage units, it is crucial to quickly and accurately locate malfunctioning storage units and prevent them from being rescheduled to other service nodes, thereby controlling the blast radius of the failure.
[0028] Currently, there are two main methods for locating abnormal storage units. The first method involves comparing the storage unit's scheduling history with the service node's failure history. The second method involves adding a sandbox node to the existing system as a screening target for abnormal storage units. These two methods will be explained in detail below.
[0029] See Figure 1 , Figure 1 An architecture diagram of a distributed storage system is shown. The distributed storage system includes a control node and four service nodes, namely Service Node 1, Service Node 2, Service Node 3, and Service Node 4, which correspond to different servers in the cluster. Service Node 1 includes four storage units: Storage Unit 1, Storage Unit 2, Storage Unit 3, and Storage Unit 4. Service Node 2 includes three storage units: Storage Unit 7, Storage Unit 8, and Storage Unit 9. Service Node 3 includes four storage units: Storage Unit 100, Storage Unit 110, Storage Unit 99, and Storage Unit 200. Service Node 4 includes four storage units: Storage Unit 777, Storage Unit 666, Storage Unit 888, and Storage Unit 123.
[0030] See Figure 2 , Figure 2 This diagram illustrates the processing flow of a control unit in a distributed storage system. Figure 1 In the distributed storage system shown, suppose storage unit 200 in service node 3 malfunctions, causing service node 3 to fail. At this time, the control node will reschedule all storage units in service node 3 to other service nodes. For example, the control node might reschedule storage unit 100 to service node 2, storage unit 110 to service node 1, storage unit 99 to service node 2, and storage unit 200 to service node 4. Because the control node is unaware that the failure of service node 3 is caused by the malfunction of storage unit 200, rescheduling storage unit 200 to service node 4 will also cause service node 4 to fail due to the malfunctioning storage unit 200. If the control unit then performs further scheduling, it will cause all service nodes in the cluster to fail. The control unit can record the scheduling trajectory of each storage unit and the failure history of the service nodes, and compare the trajectories. If the scheduling trajectory of a storage unit perfectly matches the failure history of a service node, then this storage unit is identified as an malfunctioning storage unit. However, the above approach typically requires scheduling the faulty storage unit three times. This means that the faulty storage unit will affect all storage units on at least three service nodes before the specific faulty storage unit can be located. If multiple storage units are faulty, the impact on the service nodes will be even greater.
[0031] See Figure 3 , Figure 3 This diagram illustrates the architecture of another distributed storage system. The system includes a control node, three service nodes, and five sandbox nodes. The three service nodes correspond to different servers in the cluster, and the five sandbox nodes are started on a separate server within the cluster. Each sandbox node is an empty, streamlined service node that provides normal read and write services. These sandbox nodes will not start if the three service nodes are running normally. The three service nodes are Service Node 1, Service Node 2, and Service Node 3. Service Node 1 contains four storage units: Storage Unit 1, Storage Unit 2, Storage Unit 3, and Storage Unit 4. Service Node 2 contains three storage units: Storage Unit 7, Storage Unit 8, and Storage Unit 9. Service Node 3 contains four storage units: Storage Unit 100, Storage Unit 110, Storage Unit 99, and Storage Unit 200. The five sandbox nodes are Sandbox Node 10001, Sandbox Node 10002, Sandbox Node 10003, Sandbox Node 10006, and Sandbox Node 12000.
[0032] See Figure 4 , Figure 4 This diagram illustrates the processing flow of the control unit in another distributed storage system. Figure 3In the distributed storage system shown, assuming storage unit 200 in service node 3 malfunctions, service node 3 will fail. At this time, the control node will schedule all storage units in service node 3 to sandbox nodes. When scheduling to sandbox nodes, the control unit needs to ensure that one sandbox node corresponds to only one storage unit. For example, the control node schedules storage unit 100 to sandbox node 10001, storage unit 110 to sandbox node 10002, storage unit 99 to sandbox node 10003, and storage unit 200 to sandbox node 10006. Simultaneously, the control unit needs to record the mapping relationship between storage units and sandbox nodes. When a user's read / write request is issued again, sandbox node 10006 will crash again. Sandbox node 10006 will then register with the control node again. The control unit, receiving a registration request from an existing sandbox node, can confirm that the sandbox node has crashed and restarted. At this point, the control node will identify storage unit 200 on sandbox node 10006 as an abnormal storage unit and will not reschedule storage unit 200 to other service nodes. Throughout this process, only the initial stage of scheduling storage unit 200 to sandbox node 10006 affects read / write services. After the observation period (e.g., a sandbox node provides normal service for 20 seconds without crashing again), these storage units on sandbox node 10006 can be reassigned to other service nodes. However, this solution requires adding sandbox nodes to the existing system as abnormal storage unit screening targets, wasting resources and increasing operational complexity.
[0033] To address the problems of the aforementioned solutions, this specification proposes an abnormal storage unit location scheme. This scheme obtains risk factors for the target storage unit, where the target storage unit is a storage unit within a target service node, the target service node is a faulty service node among multiple service nodes, and the risk factors are the conditions that cause the target storage unit to malfunction. Based on the risk factors, outlier detection is performed on the target storage unit to obtain detection results, which are used to indicate abnormal storage units within the target storage unit. By detecting outliers based on risk factors and using the detection results to locate abnormal storage units, this scheme achieves rapid and accurate location of abnormal storage units without scheduling them to other service nodes. This avoids potential cascading failure risks, reduces the cluster's catastrophic radius, eliminates the need for additional sandbox nodes, enhances system reliability and stability, and reduces operational complexity.
[0034] This specification provides a method for locating abnormal storage units. It also relates to an abnormal storage unit location device, a distributed storage system, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0035] See Figure 5 , Figure 5 This specification illustrates an architecture diagram of a distributed storage system provided in one embodiment of the specification. The distributed storage system is deployed on a cloud-side device and includes a control node 502 and multiple service nodes 504. Control node 502 is used to obtain risk factors of the target storage unit, wherein the target storage unit is a storage unit in the target service node, the target service node is a faulty service node among multiple service nodes 504, and the risk factor is the condition that causes the target storage unit to be abnormal; based on the risk factors, outlier detection is performed on the target storage unit to obtain detection results, wherein the detection results are used to indicate abnormal storage units in the target storage unit.
[0036] The solution implemented in this specification detects outliers based on risk factors and uses the detection results to locate abnormal storage units. This eliminates the need to schedule abnormal storage units to other service nodes, thereby avoiding potential cascading failures and enhancing the reliability and stability of the system.
[0037] See Figure 6 , Figure 6 A flowchart illustrating an embodiment of an abnormal storage unit location method provided in this specification is shown. This method is applied to the control node of a distributed storage system, which includes multiple service nodes and a control node. The method specifically includes the following steps: Step 602: Obtain the risk factors of the target storage unit, where the target storage unit is the storage unit in the target service node, the target service node is the faulty service node among multiple service nodes, and the risk factors are the conditions that cause the target storage unit to malfunction.
[0038] It should be noted that the target service node refers to a service node in a distributed storage system that has failed due to the abnormality of its included storage units. The target storage unit refers to a specific storage entity located within the target service node. Target storage units include virtual block devices, cloud disks, etc. Since the target service node has failed due to the abnormality of its storage units, it includes at least one abnormal storage unit. The purpose of the abnormal storage unit location method proposed in the embodiments of this specification is to locate the specific abnormal storage unit within the target service node.
[0039] Risk factors refer to various conditions that may cause storage units to malfunction. These risk factors may lead to performance degradation, data loss, or service interruption. Risk factors include, but are not limited to, security risk factors and operational risk factors. Security risk factors refer to conditions that cause storage units to malfunction due to attacks or data breaches. Security risk factors directly affect the security of storage units. Security risk factors include changes in access permissions, access intrusion, expired security patches, etc. Operational risk factors refer to potential risk factors related to the operational behavior, performance, and configuration status of storage units. Operational risk factors affect the internal operation and performance of storage units. Normally, normally functioning storage units will not malfunction. Malfunctioning storage units usually experience behavioral changes, performance anomalies, or configuration changes during operation; therefore, operational risk factors of storage units can be determined based on operational change information.
[0040] In practical applications, there are various ways to obtain the risk factors of the target storage unit, and the specific method should be selected according to the actual situation. This specification does not impose any limitations on these methods in the embodiments. In one possible implementation of this specification, the risk factors of the target storage unit can be read from other data acquisition devices or databases. In another possible implementation of this specification, pre-configured risk factors can be determined as the risk factors of the target storage unit.
[0041] In one optional embodiment of this specification, the risk factors include operational risk factors; the risk factors for obtaining the target storage unit described above may include the following steps: Monitor the operation of the target storage unit and, based on the operational change information obtained from the monitoring, determine the operational risk factors of the target storage unit.
[0042] It should be noted that operational change information refers to information related to the operational behavior, performance, and configuration status of a storage unit. Operational change information may affect the normal operation and health of the storage unit. Operational change information includes at least one of the following: creation behavior, migration behavior, read / write change information, configuration change information, capacity change information, memory change information, processor exception information, scheduling exception information, and traffic change information. Creation behavior indicates whether a storage unit was created within a preset time. Since newly created storage units have not undergone sufficient testing or optimization, creation behavior may cause the storage unit to malfunction due to unknown problems or configuration errors. Migration behavior indicates whether a storage unit was migrated from one location to another within a preset time. Migration operations may cause the storage unit to malfunction due to data consistency issues. Read / write change information indicates whether there have been significant changes in read / write requests for the storage unit. Significant changes in read / write requests may cause the storage unit to malfunction due to significant changes in read / write operation frequency, throughput, etc. Configuration change information indicates whether the performance parameters, access permissions, or other configuration attributes of the storage unit have been changed within a preset time. Configuration changes may introduce new problems and risks, leading to storage unit malfunctions. Capacity change information indicates whether the available space capacity of the storage unit is lower than a preset capacity threshold. A capacity below the preset threshold may cause the storage unit to malfunction due to performance degradation and inability to write new data. Memory change information indicates whether the memory resource usage of the storage unit is higher than a first preset usage threshold or lower than a second preset usage threshold, i.e., abnormally high or abnormally low. Abnormal memory resource usage may cause the storage unit to malfunction due to processing delays and task backlogs. Processor anomaly information indicates whether the core processor (CPU, Central Processing Unit) associated with the storage unit exhibits abnormal behavior. Processor anomalies may cause the storage unit to malfunction due to performance degradation and task failures. Scheduling anomaly information indicates whether the scheduling mode of tasks or processes on the storage unit exhibits abnormal behavior within a preset time. Scheduling anomalies may cause the storage unit to malfunction due to task failures and service interruptions. Traffic change information indicates whether the amount of data transmitted by the storage unit has changed abnormally. Abnormal traffic change may cause the storage unit to malfunction due to communication interruptions and network congestion.
[0043] For example, assuming the operational change information of the target storage unit is "migrated from one location to another within one hour", then the risk factor of the target storage unit is "recently migrated".
[0044] The solution implemented in this specification uses multi-dimensional operational change information of storage units to characterize the risk factors of storage units, enabling the subsequent screening of storage units with significantly different behavior from normal storage units as abnormal storage units through an outlier detection mechanism.
[0045] In one optional embodiment of this specification, before obtaining the risk factors of the target storage unit, the target service node among multiple service nodes can be determined. That is, before obtaining the risk factors of the target storage unit, the following steps may also be included: Perform status monitoring on service nodes to obtain node status; If a node is in a faulty state, the target service node is determined.
[0046] It should be noted that status monitoring refers to the process by which the control node continuously monitors the operational status of the service node. Methods for monitoring the status of service nodes include, but are not limited to, heartbeat-based monitoring and application programming interface (API)-based monitoring; the specific method is chosen based on the actual situation, and this specification does not impose any limitations on this approach. Node status refers to the result obtained from monitoring the status of the service node. Node status is a set of information describing the current operational status of the service node, such as fault status and non-fault status. The node status can be used to assess whether a service node is currently experiencing a fault. If the node status is faulty, it indicates that the service node contains at least one abnormal storage unit; therefore, this service node is the target service node.
[0047] Using the scheme of the embodiments in this specification, the control node monitors the status of the service node. Once the node status is faulty, the corresponding service node is identified as the target service node, so that the abnormal storage unit of the target service node can be located in a timely manner and the service of other normal storage units in the target service node can be restored in a timely manner.
[0048] Step 604: Based on risk factors, perform outlier detection on the target storage unit and obtain the detection results, wherein the detection results are used to indicate abnormal storage units in the target storage unit.
[0049] It should be noted that outlier detection is used to identify anomalous storage units within a target storage unit where risk factors deviate from other storage units. Depending on the risk factors, outlier detection methods for the target storage unit include, but are not limited to, detection methods based on Principal Component Analysis (PCA), detection methods based on local outliers, and detection methods based on the isolated forest model. The specific method chosen depends on the actual situation, and the embodiments in this specification do not impose any limitations on this. PCA is a data dimensionality reduction technique that primarily projects high-dimensional data into a low-dimensional space, and during projection, seeks to maximize the variance of each data point's coordinates to the low-dimensional coordinates to preserve as much original data information as possible during the dimensionality reduction process. Typically, multiple dimensions of features can be defined for an object (e.g., multi-dimensional operational change information in the embodiments of this specification). When analyzing a large number of object feature values, it is desirable to compress the multi-dimensional features of each object into a small number of features that can represent the main information. PCA analyzes the features of all objects in each dimension. For example, if there are m objects and each object has n features, and if the first feature of each object is x1=1 and the second feature is x2=2, then these two features can be represented by x2=2*x1. Therefore, the two features x1 and x2 can be simplified into one feature. So, when the numerical representation of each corresponding feature of all analyzed objects is basically consistent, PCA can reduce the multi-dimensional coordinate system to a one-dimensional coordinate system. However, if there are outliers in the analyzed objects, and these outliers differ significantly from other objects in some features, PCA will fail to achieve good dimensionality reduction. This is because a small number of dimensions cannot represent multiple dimensions. For example, in the above example, assuming there are a total of 1000 objects, the first 900 objects have the first feature x1=1 and the second feature x2=2, but the x1 and x2 values of the last 100 objects are completely random and have no correlation. In this case, x2=2*x1 cannot be used for dimensionality reduction. For the last 100 objects, the dimensionality reduction effect of PCA will be directly reduced. These objects are called outliers in PCA dimensionality reduction. Therefore, the core of PCA outlier detection lies in utilizing the dimensionality reduction characteristics of the data and evaluating the fitness of the samples to the overall data structure through the reconstruction process, thereby identifying those outliers that exhibit abnormal behavior.
[0050] Furthermore, based on risk factors, outlier detection is performed on the target storage units. After obtaining the detection results, abnormal storage units can be directly located from the target storage units. An abnormal storage unit refers to a target storage unit confirmed to exhibit abnormal behavior after outlier detection, such as an abnormal cloud disk. An abnormal storage unit may have already failed or exhibit behavioral patterns that could lead to future failures. The detection results can indicate whether the data reconstruction error corresponding to the target storage unit meets the outlier determination criteria. If the outlier determination criteria are met, the target storage unit is determined to be an abnormal storage unit; if the outlier determination criteria are not met, the target storage unit is determined not to be an abnormal storage unit. For example, if the detection results indicate that the data reconstruction error corresponding to a certain target storage unit is significantly higher than that of other target storage units, then that target storage unit is considered an outlier, i.e., an abnormal storage unit.
[0051] The solution implemented in this specification detects outliers based on risk factors and uses the detection results to locate abnormal storage units. This eliminates the need to schedule abnormal storage units to other service nodes, thereby avoiding potential cascading failures and enhancing the reliability and stability of the system.
[0052] In practical applications, when performing outlier detection on target storage units based on risk factors, the outlier detection process can be initiated immediately after the target service node is identified, or it can be initiated when a batch of target service nodes appear (e.g., more than 3 service nodes fail within 5 minutes). Outlier detection is performed on the risk factors corresponding to all target service nodes, prioritizing the containment of abnormal storage units within a fixed range to prevent the spread of the fault range.
[0053] In one optional embodiment of this specification, outlier detection based on principal component analysis is used as an example to illustrate the outlier detection process. That is, the outlier detection of the target storage unit based on risk factors and the acquisition of detection results may include the following steps: Based on risk factors, determine the data to be tested for the target storage unit; The data to be tested is subjected to dimensionality reduction processing to obtain dimensionality-reduced test data; The data to be detected is reconstructed based on the dimensionality reduction detection data to obtain the reconstructed data, and the data reconstruction error is determined based on the reconstructed data and the data to be detected. The detection result of the target storage unit is determined based on the data reconstruction error.
[0054] It should be noted that the data to be detected refers to a set of numerical values describing the target storage unit based on risk factors. Dimensionality reduction is the process of converting high-dimensional data into a low-dimensional representation, aiming to reduce data complexity while preserving as much important information as possible from the original data. Reconstruction is the process of attempting to reconstruct a new dataset that approximates the original data to be detected based on the dimensionality-reduced detection data. The reconstruction process attempts to capture and reproduce the essential structure of the input data. Data reconstruction error refers to the similarity or magnitude of the error between the reconstructed data and the data to be detected. The detection result refers to the result determined based on the data reconstruction error, used to indicate which target storage units exhibit abnormal behavior.
[0055] For example, suppose there are 11 target storage units, each corresponding to 5 dimensions of operational change information. The first operational change information is: whether it was created within 5 minutes; the second operational change information is: whether it is a migrated storage unit; the third operational change information is: whether there was a significant change in read / write request size within 5 minutes; the fourth operational change information is: whether there was a configuration change within 5 minutes; and the fifth operational change information is: whether there was a significant change in traffic within 5 minutes. If the first target storage unit has no operational change information for all five dimensions, then the risk factor for the first target storage unit is determined to be empty, and the data to be tested for the first target storage unit is [0,0,0,0,0]. If the second target storage unit has a yes operational change information for the first dimension, and no operational change information for the other four dimensions, then the risk factor for the second target storage unit is determined to be "created within 5 minutes," and the data to be tested for the second target storage unit is [1,0,0,0,0]. If the second operational change information for the third target storage unit is "Yes," and the operational change information for the other four dimensions is "No," then the risk factor for the second target storage unit is determined to be "Migrated," and the data to be tested for the third target storage unit is [0,1,0,0,0]. If the operational change information for the fourth target storage unit for all five dimensions is "No," then the risk factor for the fourth target storage unit is determined to be empty, and the data to be tested for the fourth target storage unit is [0,0,0,0,0]. If the operational change information for the fifth target storage unit for all five dimensions is "No," then the risk factor for the fifth target storage unit is determined to be empty, and the data to be tested for the fifth target storage unit is [0,0,0,0,0]. If the operational change information for the sixth target storage unit for all five dimensions is "No," then the risk factor for the sixth target storage unit is determined to be empty, and the data to be tested for the sixth target storage unit is [0,0,0,0,0]. If the seventh target storage unit has no operational change information for any of the five dimensions mentioned above, then the risk factor for the seventh target storage unit is determined to be empty, and the data to be tested for the seventh target storage unit is [0,0,0,0,0]. If the eighth target storage unit has no operational change information for any of the five dimensions mentioned above, then the risk factor for the eighth target storage unit is determined to be empty, and the data to be tested for the eighth target storage unit is [0,0,0,0,0]. If the fifth operational change information for the ninth target storage unit is yes, and the operational change information for the other four dimensions is no, then the risk factor for the ninth target storage unit is determined to be "a huge change in traffic within 5 minutes", and the data to be tested for the ninth target storage unit is [0,0,0,0,1]. If the tenth target storage unit has no operational change information for any of the five dimensions mentioned above, then the risk factor for the tenth target storage unit is determined to be empty, and the data to be tested for the tenth target storage unit is [0,0,0,0,0].If the operational change information for the eleventh target storage unit is negative for all five dimensions mentioned above, then the risk factors for the eleventh target storage unit are determined to be empty, and the data to be tested for the eleventh target storage unit is [0,0,0,0,0].
[0056] The above example demonstrates how to determine the data to be tested in a target storage unit based on risk factors: from sklearn.decomposition import PCA # Import the PCA class from the sklearn library from sklearn.preprocessing import StandardScaler import matplotlib.pyplot as plt # Generate sample data, 11 storage units, each storage unit has 5 dimension definitions. # First column: Was it created within the last 5 minutes? # Second column: Whether it is a migrated storage unit # Third column: Has there been a significant change in read / write request size within 5 minutes? # Fourth column: Are there any configuration changes within 5 minutes? # Fifth column: Was there a significant change in traffic within 5 minutes? data=[[0,0,0,0,0],[1,0,0,0,0],[0,1,0,0,0],[0,0,0,0,0],[0,0,0,0,0],[0, 0,0,0,0],[0,0,0,0,0],[0,0,0,0,0],[0,0,0,0,1],[0,0,0,0,0],[0,0,0,0,0]] After determining the target storage unit's data to be tested based on risk factors, PCA can be used to reduce the original 5-dimensional data to 2-dimensionality, obtaining dimensionality-reduced detection data. Dimensionality-reduced detection data makes it easier to identify patterns and structures within the data to be tested. After obtaining the dimensionality-reduced detection data, the data to be tested can be reconstructed based on it, resulting in new reconstructed data. Since the reconstructed data differs from the original data to be tested, the data reconstruction error can be determined based on the reconstructed data and the original data to be tested. A larger data reconstruction error indicates that the corresponding target storage unit is less consistent with the main pattern of the data, and therefore more likely to be an outlier. When determining the detection results of target storage units based on the data reconstruction error, outlier identification criteria can be set (e.g., a data reconstruction error exceeding 70% is considered an outlier) to determine which target storage units are considered abnormal. Since the index of the data to be tested starts from 0, based on the output, the target storage units corresponding to the 2nd (index 1) data [1,0,0,0,0], the 3rd (index 2) data [0,1,0,0,0], and the 9th (index 8) data [0,0,0,0,1] are determined to be abnormal storage units. The above process can be implemented using the following program: # PCA Dimensionality Reduction pca = PCA(n_components=2) data_pca = pca.fit_transform(data_scaled) # Determine data reconstruction error data_reconstructed = pca.inverse_transform(data_pca) reconstruction_error = np.linalg.norm(data_scaled - data_reconstructed, axis=1) # Outlier Detection threshold = np.percentile(reconstruction_error, 70) outliers = np.where(reconstruction_error>threshold) print(data) print(outliers) Output result: (array([1,2,8])) During PCA dimensionality reduction, a PCA object is created and specified to be reduced to 2 dimensions (n_components=2). The `fit_transform()` method is used to fit and apply the PCA transformation to the data to be detected, `data_scaled`, to obtain the dimensionality-reduced detection data, `data_pca`. When determining the data reconstruction error, the `inverse_transform()` method transforms the dimensionality-reduced detection data back to the original dimension, resulting in the reconstructed data, `data_reconstructed`. Based on the reconstructed data and the data to be detected, the Euclidean norm (`np.linalg.norm`) is used to determine the data reconstruction error, and the error is calculated along the direction of each sample (i.e., row) (`axis=1`), resulting in an array representing the data reconstruction error for each sample. For outlier detection, data points with a reconstruction error exceeding 70% are considered outliers. The `np.where()` function is used to find the indices of data points with reconstruction errors greater than a set threshold; these data points are considered outliers. The output shows the indices of the identified outliers. `(array([1,2,8]))` indicates that data points with indices 1, 2, and 8 are considered outliers.
[0057] The solution implemented in this specification uses the PCA outlier detection mechanism, which only requires adding processing logic to the control unit. It does not require adding new components to the system architecture to locate abnormal storage units, thus avoiding potential cascading failure risks, reducing the cluster's explosion radius, and lowering the complexity of system operation and maintenance.
[0058] In practical applications, methods for measuring data reconstruction error include, but are not limited to, Mean Squared Error (MSE) and Root Mean Squared Error (RMSE). There are various ways to determine the detection result of the target storage unit based on the data reconstruction error; the specific method chosen depends on the actual situation, and this specification does not impose any limitations on this approach. In one possible implementation of this specification, the detection result of the target storage unit can be determined based on outlier identification conditions and the data reconstruction error. In another possible implementation of this specification, the Isolation Forest algorithm can be used to directly identify outliers from the data reconstruction error, thereby determining the detection result of the target storage unit.
[0059] In one optional embodiment of this specification, determining the data reconstruction error based on the reconstructed data and the data to be detected may include the following steps: Based on the reconstructed data and the data to be detected, a data reconstruction view is generated; Based on the data reconstruction view, determine the data reconstruction error.
[0060] It's important to note that generating a data reconstruction view based on reconstructed data and the data to be tested refers to converting the reconstructed data and the data to be tested into graphs, charts, or other visual representations, thereby obtaining a data reconstruction view that allows for a more intuitive understanding and analysis of the reconstructed data and the data to be tested. A data reconstruction view is a visual representation that graphically displays the relationship between the reconstructed data and the data to be tested. Through a data reconstruction view, the differences between the reconstructed data and the data to be tested can be intuitively identified. Data reconstruction views include, but are not limited to, scatter plots and parallel coordinate plots.
[0061] By applying the solutions in the embodiments of this specification, the reconstructed data and the data to be detected are displayed through graphical tools, thereby enabling the rapid and accurate determination of data reconstruction errors.
[0062] In one optional embodiment of this specification, the detection result of determining the target storage unit based on the data reconstruction error may include the following steps: Obtain the conditions for identifying outliers; Based on the outlier determination conditions and data reconstruction errors, the detection results of the target storage unit are determined.
[0063] It should be noted that outlier identification criteria refer to the specific standards or rules used to identify outliers from the data to be detected. Outlier identification criteria include, but are not limited to, data reconstruction errors exceeding an error threshold (e.g., 70%) being considered outliers, or N data points with large reconstruction errors being considered outliers. The specific criteria should be selected based on the actual situation, and this specification does not impose any limitations on them in the embodiments. The error threshold is used to determine the accuracy of the filtered abnormal storage units. If you want to locate as many abnormal storage units as possible, you can set the error threshold smaller; if you don't want to mistakenly identify normal storage units as abnormal storage units due to locating too many abnormal storage units, you can set the error threshold larger.
[0064] By applying the schemes in the embodiments of this specification, the detection results of the target storage unit can be determined efficiently and accurately by using outlier point determination conditions and data reconstruction errors.
[0065] In one optional embodiment of this specification, before performing dimensionality reduction processing on the data to be detected to obtain dimensionality-reduced detection data, the following steps may also be included: The data to be tested is normalized to obtain updated data to be tested. The data to be tested is subjected to dimensionality reduction processing to obtain dimensionality-reduced detection data, including: The updated data to be detected is subjected to dimensionality reduction processing to obtain dimensionality-reduced detection data.
[0066] It should be noted that normalization can be understood as standardization. Normalization refers to the process of transforming the data to be detected to a fixed range. Through normalization, all features of the data to be detected can have similar scales, preventing certain features from dominating the outlier detection process due to their large original scale. Normalization methods include, but are not limited to, min-max scaling and Z-score standardization; the specific method should be selected according to the actual situation, and the embodiments in this specification do not impose any limitations on this. In practical applications, the following is an example of a program that performs normalization on the data to be detected to obtain updated data: # Data normalization processing scaler = StandardScaler() data_scaled = scaler.fit_transform(data) `StandardScaler` is a class in the scikit-learn library used for normalizing datasets. Calling `StandardScaler()` creates a new scaler object. This object internally stores parameters such as the mean and standard deviation of the data for subsequent data transformation and inverse transformation. `fit_transform` combines the `fit` and `transform` steps. `fit` calculates the mean and standard deviation of each feature of the input data and stores them in the scaler object. `transform` standardizes the data using the previously calculated mean and standard deviation, that is, subtracting the mean from the value of each feature and dividing by its standard deviation.
[0067] By applying the scheme of the embodiments in this specification, the updated data to be detected is obtained by normalizing the data to be detected, and the updated data to be detected is then subjected to dimensionality reduction processing to obtain dimensionality-reduced detection data. This ensures that all features of the data to be detected are on the same scale, making the outlier detection process more accurate.
[0068] In one optional embodiment of this specification, after performing outlier detection on the target storage unit based on risk factors and obtaining the detection results, the following steps may be further included: From the target service nodes, identify the storage units to be scheduled, excluding the abnormal storage units; The storage units to be scheduled are scheduled to be stored on auxiliary service nodes, where the auxiliary service nodes are service nodes other than the target service node.
[0069] It should be noted that the storage unit to be scheduled refers to the target storage unit within the target service node, excluding the identified faulty storage units. Therefore, the storage unit to be scheduled is healthy and functioning normally. The auxiliary service node refers to a non-faulty service node other than the target service node among the multiple service nodes in the cluster. When the target service node fails, the auxiliary service node can take over the storage unit to be scheduled, ensuring uninterrupted service for that storage unit.
[0070] For example, suppose the cluster includes two service nodes. The first service node is the target service node that failed, and the second service node is the auxiliary service node. If the target service node includes three target storage units, and based on the detection results, the first target storage unit is determined to be an abnormal storage unit, then the second and third target storage units can be scheduled to the second service node for storage.
[0071] By applying the solutions in the embodiments of this specification, the storage unit to be scheduled can be scheduled to an auxiliary service node for storage, thereby restoring the service of the storage unit to be scheduled and ensuring uninterrupted service.
[0072] Corresponding to the above embodiments of the abnormal storage unit location method, this specification also provides embodiments of the abnormal storage unit location device. Figure 7 A schematic diagram of an abnormal storage unit locating device according to one embodiment of this specification is shown. Figure 7 As shown, this device is used in the control node of a distributed storage system. The distributed storage system includes multiple service nodes and a control node. The device includes: The acquisition module 702 is configured to acquire risk factors of the target storage unit, wherein the target storage unit is a storage unit in the target service node, the target service node is a faulty service node among multiple service nodes, and the risk factor is the condition that causes the target storage unit to malfunction. The detection module 704 is configured to perform outlier detection on the target storage unit based on risk factors and obtain detection results, wherein the detection results are used to indicate abnormal storage units in the target storage unit.
[0073] Optionally, the detection module 704 is further configured to: determine the data to be detected for the target storage unit based on risk factors; perform dimensionality reduction processing on the data to be detected to obtain dimensionality-reduced detection data; reconstruct the data to be detected based on the dimensionality-reduced detection data to obtain reconstructed data; determine the data reconstruction error based on the reconstructed data and the data to be detected; and determine the detection result of the target storage unit based on the data reconstruction error.
[0074] Optionally, the detection module 704 is further configured to generate a data reconstruction view based on the reconstructed data and the data to be detected; and to determine the data reconstruction error based on the data reconstruction view.
[0075] Optionally, the detection module 704 is further configured to acquire outlier determination conditions and determine the detection result of the target storage unit based on the outlier determination conditions and the data reconstruction error.
[0076] Optionally, the device further includes: a processing module configured to normalize the data to be detected to obtain updated data to be detected; and a detection module 704 further configured to perform dimensionality reduction processing on the updated data to be detected to obtain dimensionality-reduced detection data.
[0077] Optionally, the risk factors include operational risk factors; the acquisition module 702 is further configured to monitor the operation of the target storage unit and determine the operational risk factors of the target storage unit based on the monitored operational change information.
[0078] Optionally, the device further includes: a monitoring module configured to monitor the status of the service node and obtain the node status; and to determine the service node as the target service node if the node status is faulty.
[0079] Optionally, the device further includes: a determining module configured to determine, from the target service node, a storage unit to be scheduled other than the abnormal storage unit; and to schedule the storage unit to be scheduled to an auxiliary service node for storage, wherein the auxiliary service node is a service node other than the target service node.
[0080] The solution implemented in this specification detects outliers based on risk factors and uses the detection results to locate abnormal storage units. This eliminates the need to schedule abnormal storage units to other service nodes, thereby avoiding potential cascading failures and enhancing the reliability and stability of the system.
[0081] The above is a schematic scheme of an abnormal storage cell location device according to this embodiment. It should be noted that the technical solution of this abnormal storage cell location device and the technical solution of the abnormal storage cell location method described above belong to the same concept. For details not described in detail in the technical solution of the abnormal storage cell location device, please refer to the description of the technical solution of the abnormal storage cell location method described above.
[0082] Figure 8 This specification illustrates a structural block diagram of a computing device according to one embodiment. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.
[0083] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., Network Interface Card (NIC)), such as an IEEE 802.11 Wireless Local Area Networks (WLAN) interface, a Wi-MAX (World Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0084] In one embodiment of this specification, the above-described components of the computing device 800 and Figure 8 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 8 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0085] The computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 800 can also be a mobile or stationary server.
[0086] The processor 820 is used to execute computer programs / instructions, which, when executed by the processor, implement the steps of the above-described abnormal memory location method.
[0087] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-described abnormal storage unit location method belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-described abnormal storage unit location method.
[0088] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described abnormal storage unit location method.
[0089] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the abnormal storage cell location method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the abnormal storage cell location method described above.
[0090] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described abnormal memory cell location method.
[0091] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described abnormal storage unit location method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described abnormal storage unit location method.
[0092] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0093] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0094] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0095] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0096] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A method for locating abnormal storage units, applied to a control node in a distributed storage system, the distributed storage system comprising multiple service nodes and the control node, the method comprising: Obtain risk factors for the target storage unit, wherein the target storage unit is a storage unit in the target service node, the target service node is a faulty service node among the plurality of service nodes, and the risk factor is a condition that causes the target storage unit to malfunction; Based on the risk factors, outlier detection is performed on the target storage unit to obtain detection results, wherein the detection results are used to indicate abnormal storage units in the target storage unit.
2. The method according to claim 1, wherein the step of performing outlier detection on the target storage unit based on the risk factors to obtain detection results includes: The target storage unit is described based on the risk factors to obtain the data to be tested for the target storage unit; The data to be detected is subjected to dimensionality reduction processing to obtain dimensionality-reduced detection data; The data to be detected is reconstructed based on the dimensionality reduction detection data to obtain reconstructed data, and the data reconstruction error is determined based on the reconstructed data and the data to be detected. The detection result of the target storage unit is determined based on the data reconstruction error.
3. The method according to claim 2, wherein determining the data reconstruction error based on the reconstructed data and the data to be detected includes: Based on the reconstructed data and the data to be detected, a data reconstruction view is generated; Based on the data reconstruction view, the data reconstruction error is determined.
4. The method according to claim 2, wherein determining the detection result of the target storage unit based on the data reconstruction error includes: Obtain the conditions for identifying outliers; The detection result of the target storage unit is determined based on the outlier determination conditions and the data reconstruction error.
5. The method according to claim 2, further comprising, before performing dimensionality reduction processing on the data to be detected to obtain dimensionality-reduced detection data: The data to be detected is normalized to obtain updated data to be detected; The step of performing dimensionality reduction processing on the data to be detected to obtain dimensionality-reduced detection data includes: The updated data to be detected is subjected to dimensionality reduction processing to obtain dimensionality-reduced detection data.
6. The method according to claim 1, wherein the risk factors include operational risk factors; The risk factors for acquiring the target storage unit include: The operation of the target storage unit is monitored, and based on the operational change information obtained from the monitoring, the operational risk factors of the target storage unit are determined.
7. The method according to any one of claims 1 to 6, further comprising, before obtaining the risk factors of the target storage unit: The status of the service nodes is monitored to obtain the node status; If the node is in a fault state, the service node is determined to be the target service node.
8. The method according to any one of claims 1 to 6, wherein after performing outlier detection on the target storage unit based on the risk factors and obtaining the detection result, it further includes: From the target service nodes, identify the storage units to be scheduled, excluding the abnormal storage units; The storage unit to be scheduled is scheduled to an auxiliary service node for storage, wherein the auxiliary service node is a service node other than the target service node.
9. A distributed storage system deployed on a cloud-side device, the distributed storage system comprising a control node and multiple service nodes; The control node is used to acquire risk factors for the target storage unit, wherein... The target storage unit is a storage unit in the target service node, the target service node is a faulty service node among the plurality of service nodes, and the risk factor is a condition that causes the target storage unit to malfunction. Based on the risk factors, outlier detection is performed on the target storage unit to obtain detection results, wherein the detection results are used to indicate abnormal storage units in the target storage unit.
10. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 8.
11. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 8.
12. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 8.