Slow disk detection method and device and first storage and calculation integrated hard disk

By utilizing the internal communication of the in-memory computing hard drive, storage latency can be quickly obtained for slow disk detection, solving the problem of low disk performance detection efficiency and achieving efficient and accurate slow disk detection.

CN120909858APending Publication Date: 2025-11-07SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510666357.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing technologies, disk performance testing is inefficient, especially when testing is performed remotely, which may reduce efficiency.

Method used

It adopts the architecture of a storage-computing hard drive, which enables it to have computing capabilities. Through direct communication between the computing unit and the storage unit inside the storage-computing hard drive, it can quickly obtain storage latency for slow disk detection.

Benefits of technology

It improves the efficiency and accuracy of slow disk detection, enabling rapid detection of whether a disk is slow within the hard drive itself, and adapts to dynamic monitoring of load changes, thus improving the timeliness and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909858A_ABST
    Figure CN120909858A_ABST
Patent Text Reader

Abstract

The invention provides a slow disk detection method and device and a first storage and calculation integrated hard disk. The method is applied to a calculation unit in a first storage and calculation integrated hard disk, the first storage and calculation integrated hard disk further comprises a first storage unit, and the method comprises the steps that a first moment when the calculation unit responds to a data operation request from a client side and sends an operation instruction to the first storage unit is obtained; obtaining a second moment when the calculation unit receives an operation result (a result indicating the execution of the operation instruction for reading or writing) sent by the first storage unit; determining a first storage delay corresponding to the first storage unit based on the first moment and the second moment; and based on the first storage time delay, determining a slow disk detection result of the first storage and calculation integrated hard disk. A storage and calculation integrated hard disk architecture is adopted, so that a calculation unit in the storage and calculation integrated hard disk is directly communicated with a storage unit in the storage and calculation integrated hard disk, the storage and calculation time delay is quickly obtained, slow disk detection is performed by using the storage time delay, and the efficiency of slow disk detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a slow disk detection method, device and first storage-computing integrated hard disk. BACKGROUND

[0002] The stability of disk performance directly affects the quality of overall service, and slow disk detection needs to be performed. The disk itself is a passive response device and does not have any detection capability.

[0003] Currently, disk data can be uniformly collected by a slow disk detection tool and reported to a remote end. The remote end can determine that the disk is a slow disk in the case that the read-write delay is high, or the proportion of high read-write delay in the continuous sampling period exceeds a certain number.

[0004] However, slow disk detection by the remote end can reduce the efficiency of slow disk detection. SUMMARY

[0005] Embodiments of the present application provide a slow disk detection method, device and first storage-computing integrated hard disk. The storage-computing integrated hard disk adopts a storage-computing integrated architecture, so that the storage-computing integrated hard disk has computing capability. Subsequently, the computing unit inside the storage-computing integrated hard disk and the storage unit inside the storage-computing integrated hard disk directly communicate, quickly obtain storage-computing delay, and perform slow disk detection using the storage delay, thereby improving the efficiency of slow disk detection.

[0006] In a first aspect, embodiments of the present application provide a slow disk detection method applied to a computing unit, the computing unit being located in a first storage-computing integrated hard disk, the first storage-computing integrated hard disk further comprising a first storage unit, and the method comprising:

[0007] obtaining a first time point, wherein the first time point is used to indicate a time point at which the computing unit sends an operation instruction to the first storage unit in response to a data operation request, and the data operation request comes from a client; obtaining a second time point, wherein the second time point is a time point at which the computing unit receives an operation result sent by the first storage unit, and the operation result is used to indicate a result of reading or writing performed by the first storage unit on the operation instruction; determining a first storage delay corresponding to the first storage unit based on the first time point and the second time point, wherein the first storage delay is used to indicate a time length required by the first storage unit for reading or writing; and determining a slow disk detection result of the first storage-computing integrated hard disk based on the first storage delay, wherein the slow disk detection result is used to indicate whether the first storage-computing integrated hard disk is a slow disk.

[0008] In this scheme, the storage-computing integrated hard disk adopts a storage-computing integrated architecture, so that the storage-computing integrated hard disk has computing capability. Subsequently, the computing unit inside the storage-computing integrated hard disk and the storage unit inside the storage-computing integrated hard disk directly communicate, quickly obtain storage-computing delay, and perform slow disk detection using the storage delay, thereby improving the efficiency of slow disk detection.

[0009] In a possible implementation, the method further includes: obtaining a third time at which the data operation request is received; determining a request response time delay based on the third time and the second time; wherein the request response time delay is used to indicate a time length consumed for completing the data operation request; and determining whether the slow network exists based on the first storage time delay and the request response time delay.

[0010] In this solution, based on the all-in-one storage and computing hard disk architecture, the computing unit inside the all-in-one storage and computing hard disk can directly communicate with the storage unit inside the all-in-one storage and computing hard disk, the request response time delay is collected, and whether the slow network exists is further analyzed based on the request response time delay and the storage time delay, so as to realize the double detection of the slow disk and the slow network.

[0011] In one example of the implementation, determining whether the slow network exists based on the first storage time delay and the request response time delay includes: determining a network time delay based on the first storage time delay and the request response time delay; and determining whether the slow network exists based on the network time delay.

[0012] In a possible implementation, the method further includes:

[0013] obtaining load data of the first storage unit; and determining the slow disk detection result of the first all-in-one storage and computing hard disk based on the first storage time delay, including: determining the slow disk detection result of the first all-in-one storage and computing hard disk based on the first storage time delay and the load data of the first storage unit.

[0014] In this solution, the storage time delay is comprehensively analyzed in combination with the load data, so that the slow disk is dynamically monitored by adapting the change of the load, and the accuracy of the slow disk detection is improved.

[0015] In one example of the implementation, determining the slow disk detection result of the first all-in-one storage and computing hard disk based on the first storage time delay and the load data of the first storage unit includes: determining a storage time delay threshold based on the load data; and determining the slow disk detection result of the first all-in-one storage and computing hard disk based on the first storage time delay and the storage time delay threshold.

[0016] In this solution, the storage time delay threshold is corrected according to the load data to obtain the latest storage time delay threshold, so that the storage time delay threshold is dynamically updated by adapting the change of the load, and the accuracy of the slow disk detection is improved.

[0017] In a possible implementation, before determining the slow disk detection result of the first all-in-one storage and computing hard disk based on the first storage time delay, the method further includes: determining that data processing performed by the computing unit in response to the data processing request is normal, an operating system deployed by the computing unit is normal, and the first storage unit is in a healthy state.

[0018] In the scheme, on the basis of adopting the storage-computing integrated hard disk architecture, the computing unit inside the storage-computing integrated hard disk can perform self data processing, self operating system and health condition checking and analysis of the storage unit. In the case that the data processing is normal, the self operating system is normal and the health condition of the storage unit is normal, it is determined that the environment outside the storage unit is normal. Subsequently, the storage latency is used for slow disk detection, thereby improving the accuracy of slow disk detection.

[0019] In a possible implementation, the method further includes:

[0020] acquiring a second storage latency sent by a second storage-computing integrated hard disk; wherein the first storage-computing integrated hard disk and the second storage-computing integrated hard disk are located in a same distributed system, the second storage-computing integrated hard disk is a normal disk, the second storage-computing integrated hard disk includes a second storage unit, and the second storage latency is used to indicate a time length required by the second storage unit for reading and / or writing;

[0021] determining a slow disk detection result of the first storage-computing integrated hard disk based on the first storage latency, including: comparing the first storage latency and the request response latency to determine the slow disk detection result of the first storage-computing integrated hard disk.

[0022] In the scheme, the storage-computing integrated hard disks communicate with each other, so that the storage-computing integrated hard disks can acquire the storage latency of other storage-computing integrated hard disks. By comparing the storage latencies of different storage-computing integrated hard disks, the accuracy and efficiency of slow disk detection are improved.

[0023] In a second aspect, an embodiment of the present application provides a slow disk detection device. The slow disk detection device includes a plurality of modules, each module being configured to perform each step of the slow disk detection method provided in the first aspect of the present application. The division of the modules is not limited herein. For the specific functions performed by each module of the slow disk detection device and the beneficial effects achieved, reference can be made to the functions of each step of the slow disk detection method provided in the first aspect of the present application, which will not be described herein again.

[0024] Exemplarily, the slow disk detection device is applied to a computing unit, the computing unit being located in a first storage-computing integrated hard disk, the first storage-computing integrated hard disk further including a first storage unit, and the device includes:

[0025] an acquisition module, configured to acquire a first time point, wherein the first time point is used to indicate a time point at which the computing unit sends an operation instruction to the first storage unit in response to a data operation request, the data operation request being from a client; and acquire a second time point, wherein the second time point is a time point at which the computing unit receives an operation result sent by the first storage unit, the operation result being used to indicate a result of the first storage unit performing the operation instruction for reading or writing;

[0026] The time delay determination module is configured to determine a first storage time delay corresponding to the first storage unit based on the first time and the second time, where the first storage time delay is used to indicate a time length required for the first storage unit to perform reading or writing;

[0027] The detection module is configured to determine a slow disk detection result of the first storage-compute integrated hard disk based on the first storage time delay, where the slow disk detection result is used to indicate whether the first storage-compute integrated hard disk is a slow disk.

[0028] In a possible implementation, the acquisition module is further configured to acquire a third time at which the data operation request is received;

[0029] The time delay determination module is further configured to determine a request response time delay based on the third time and the second time, where the request response time delay is used to indicate a time length required for completing the data operation request;

[0030] The detection module is further configured to determine whether there is a slow network based on the first storage time delay and the request response time delay.

[0031] In a possible implementation, the detection module is further configured to determine a network time delay based on the first storage time delay and the request response time delay, and determine whether there is a slow network based on the network time delay.

[0032] In a possible implementation, the acquisition module is further configured to acquire load data of the first storage unit;

[0033] The detection module is configured to determine a slow disk detection result of the first storage-compute integrated hard disk based on the first storage time delay and the load data of the first storage unit.

[0034] In a possible implementation, the detection module is configured to determine a storage time delay threshold based on the load data, and determine a slow disk detection result of the first storage-compute integrated hard disk based on the first storage time delay and the storage time delay threshold.

[0035] In a possible implementation, the detection module is configured to determine a slow disk detection result of the first storage-compute integrated hard disk based on the first storage time delay in a case where the data processing performed by the compute unit in response to the data processing request is normal, an operating system deployed by the compute unit is normal, and the first storage unit is in a healthy state.

[0036] In a possible implementation, the acquisition module is configured to acquire a second storage time delay sent by a second storage-compute integrated hard disk, where the first storage-compute integrated hard disk and the second storage-compute integrated hard disk are located in a same distributed system, the second storage-compute integrated hard disk is a normal disk, the second storage-compute integrated hard disk includes a second storage unit, and the second storage time delay is used to indicate a time length required for the second storage unit to perform reading and / or writing;

[0037] The detection module is configured to compare the first storage delay and the request response delay, and determine a slow disk detection result of the first storage and computing integrated hard disk.

[0038] In a third aspect, the embodiments of the present application provide a first storage and computing integrated hard disk, which comprises a computing unit and a first storage unit, wherein

[0039] The computing unit is configured to obtain a first time point, wherein the first time point is used to indicate a time point at which the computing unit sends an operation instruction to the first storage unit in response to a data operation request, and the data operation request is from a client; and obtain a second time point, wherein the second time point is a time point at which the computing unit receives an operation result sent by the first storage unit, and the operation result is used to indicate a result of reading or writing performed by the first storage unit according to the operation instruction.

[0040] The computing unit is configured to determine a first storage delay corresponding to the first storage unit based on the first time point and the second time point, wherein the first storage delay is used to indicate a time length required by the first storage unit for reading or writing.

[0041] The computing unit is configured to determine a slow disk detection result of the first storage and computing integrated hard disk based on the first storage delay, wherein the slow disk detection result is used to indicate whether the first storage and computing integrated hard disk is a slow disk.

[0042] In a possible implementation, the computing unit is configured to obtain a third time point at which the data operation request is received; determine a request response delay based on the third time point and the second time point, wherein the request response delay is used to indicate a time length required for completing the data operation request; and determine whether there is a slow network based on the first storage delay and the request response delay.

[0043] In a possible implementation, the computing unit is configured to determine a network delay based on the first storage delay and the request response delay; and determine whether there is a slow network based on the network delay.

[0044] In a possible implementation, the computing unit is configured to obtain load data of the first storage unit; and determine the slow disk detection result of the first storage and computing integrated hard disk based on the first storage delay and the load data of the first storage unit.

[0045] In a possible implementation, the computing unit is configured to determine a storage delay threshold value based on the load data; and determine the slow disk detection result of the first storage and computing integrated hard disk based on the first storage delay and the storage delay threshold value.

[0046] In a possible implementation, the computing unit is configured to determine, based on the first storage latency, a slow disk detection result of the first storage-computing hard disk in a case where the computing unit is determined to be normal in data processing, an operating system deployed by the computing unit is determined to be normal, and the first storage unit is determined to be in a healthy state.

[0047] In a possible implementation, the computing unit is configured to obtain a second storage latency sent by a second storage-computing hard disk; the first storage-computing hard disk and the second storage-computing hard disk are located in a same distributed system, the second storage-computing hard disk is a normal disk, the second storage-computing hard disk comprises a second storage unit, and the second storage latency is used to indicate a time length required for the second storage unit to perform reading and / or writing; the first storage latency and the request response latency are compared to determine a slow disk detection result of the first storage-computing hard disk.

[0048] In a fourth aspect, an embodiment of the present application provides a first storage-computing hard disk, comprising a computing unit and a first storage unit; the computing unit is configured to execute an instruction stored in the first storage unit, so that the first storage-computing hard disk executes the method provided in the first aspect.

[0049] In a fifth aspect, an embodiment of the present application provides a storage-computing hard disk cluster, the storage-computing hard disk cluster comprising at least one storage-computing hard disk, each storage-computing hard disk comprising a computing unit and a storage unit; the computing unit of the at least one storage-computing hard disk is configured to execute an instruction stored in the storage unit of the at least one storage-computing hard disk, so that the storage-computing hard disk cluster executes the method provided in the first aspect.

[0050] In a sixth aspect, an embodiment of the present application provides a computer storage unit, the computer storage unit storing an instruction, when the instruction is executed on a computing device cluster or a storage-computing hard disk cluster, the computing device cluster or the storage-computing hard disk cluster executes the method provided in the first aspect.

[0051] In a seventh aspect, an embodiment of the present application provides a computer program product comprising an instruction, when the instruction is executed on a computing device cluster or a storage-computing hard disk cluster, the computing device cluster or the storage-computing hard disk cluster executes the method provided in the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 FIG. 1 is a system architecture diagram of a distributed system provided by an embodiment of the present application;

[0053] Figure 2 FIG. 2 is a system architecture diagram of another distributed system provided by an embodiment of the present application;

[0054] Figure 3a FIG. 3 is a flow diagram of a first slow disk detection scenario provided by an embodiment of the present application;

[0055] Figure 3b is a flowchart of a first slow disk detection method provided by an embodiment of the present application;

[0056] Figure 4a is a schematic diagram of a first storage and computing integrated disk responding to a data operation request scenario provided by an embodiment of the present application;

[0057] Figure 4b is a schematic diagram of a second storage and computing integrated disk responding to a data operation request scenario provided by an embodiment of the present application;

[0058] Figure 4c is a schematic diagram of a third storage and computing integrated disk responding to a data operation request scenario provided by an embodiment of the present application;

[0059] Figure 5 is a flowchart of a second slow disk detection scenario provided by an embodiment of the present application;

[0060] Figure 6a is a flowchart of a third slow disk detection scenario provided by an embodiment of the present application;

[0061] Figure 6b is a flowchart of a second slow disk detection method provided by an embodiment of the present application;

[0062] Figure 7a is a flowchart of a fourth slow disk detection scenario provided by an embodiment of the present application;

[0063] Figure 7b is a flowchart of a third slow disk detection method provided by an embodiment of the present application;

[0064] Figure 8 is a flowchart of a fifth slow disk detection scenario provided by an embodiment of the present application;

[0065] Figure 9 is a structural schematic diagram of a slow disk detection device provided by an embodiment of the present application;

[0066] Figure 10 is a structural schematic diagram of a computing device provided by an embodiment of the present application;

[0067] Figure 11 is a structural schematic diagram of a computing device cluster provided by an embodiment of the present application;

[0068] Figure 12 is a schematic diagram of a computing device in a computer cluster connected through a network provided by an embodiment of the present application. DETAILED DESCRIPTION

[0069] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below with reference to the drawings.

[0070] In the description of the embodiments of the present application, the words such as "exemplary", "for example", or "for instance" are used to mean serving as an example, instance or illustration. Any embodiment or design solution described as "exemplary", "for example" or "for instance" in the embodiments of the present application should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. In fact, the words such as "exemplary", "for example" or "for instance" are used to present related concepts in a specific manner.

[0071] In the description of the embodiments of the present application, the term "and / or" is merely used to describe an association relationship of associated objects, and indicates that there can be three relationships, for example, A and / or B can represent three cases of existence of A alone, existence of B alone and existence of A and B simultaneously. In addition, unless otherwise specified, the term "multiple" means two or more. For example, multiple systems mean two or more systems, and multiple terminals mean two or more terminals.

[0072] In addition, the terms "first", "second" are used only for descriptive purposes, and should not be construed as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include one or more features. The terms "include", "contain", "have" and their variants mean "include but are not limited to", unless otherwise specifically emphasized.

[0073] In the following, some terms in the embodiments are explained. It should be noted that these explanations are for the convenience of understanding by those skilled in the art, and do not constitute a limitation on the scope of protection required by the present application.

[0074] The so-called "Computing in Memory" refers to changing the traditional computing-centered architecture to a data-centered architecture, directly utilizing the memory for data processing, thereby fusing data storage and computing in the same chip, greatly improving the parallelism and energy efficiency of computing, and being particularly suitable for the field of deep learning neural networks, such as wearable devices, mobile devices, smart home and other scenarios.

[0075] Self-Monitoring, Analysis, and Reporting Technology (SMART): A technology used to monitor and report the health status of storage devices, such as solid-state drives (SSDs). SMART technology collects and analyzes operational data from the hard drive to provide detailed information about the health of the hard drive, helping users and administrators to identify potential problems and take action in a timely manner. SMART is a built-in health monitoring system for hard drives that records the operating status of the hard drive every day, such as the number of bad sectors, power-on time, temperature, etc. SMART can only tell you the current health status of the hard drive, not the exact time when the hard drive will fail. For example, SMART parameters are like a "physical examination report" for the hard drive. If the report shows "increased number of bad sectors" or "high temperature", it means that the hard drive has started to "cough", and it may need to back up data or replace the hard drive.

[0076] Slow disk: Refers to the phenomenon that in a distributed storage system, due to the decline in hard disk performance or excessive load, the read-write delay increases, affecting the performance of the entire storage system. This phenomenon is similar to the effect of the wooden barrel, the overall performance of the system is determined by the slowest hard disk, not the fastest hard disk.

[0077] iostat command: A tool for monitoring system I / O device load, mainly used to monitor CPU usage and disk I / O performance. Through iostat, users can understand the performance of the system's I / O subsystem, diagnose I / O bottlenecks, and optimize system performance. The output of iostat can be divided into two parts: CPU usage statistics and disk I / O statistics. CPU usage statistics: Displays the percentage of CPU in different states such as user state, system state, idle state, etc. Disk I / O statistics: Displays the I / O operation of each device, including read / write times, data transfer volume, etc.

[0078] Computational load: Refers to the workload of a computer system processing various tasks and running programs. This includes CPU usage, memory consumption, disk I / O, and network traffic, etc. When the computational load is too high, the performance of the computer system may decrease, the response time may be delayed, and even problems such as crashes or freezes may occur.

[0079] Network load: Refers to the workload and data transmission pressure borne by network communication devices. This includes network bandwidth usage, network traffic size, request processing capacity, etc. When the network load is too high, the network speed may slow down, the delay may increase, and even problems such as network congestion and packet loss may occur.

[0080] Input / Output Operations Per Second (IOPS): refers to the number of I / O requests that the system can handle per unit of time, generally in units of I / O requests per second, and the I / O request is usually a read or write data operation request.

[0081] Embodiments of the present application provide a slow disk detection method, which is applied to a storage-computing integrated hard disk.

[0082] On the one hand, the method offloads slow disk detection logic to the storage-computing integrated hard disk, which can detect whether it is a slow disk and report the slow disk detection result to an upper layer such as a business warning platform, thereby improving the timeliness of slow disk detection.

[0083] On the other hand, in the scenario of unilateral reading of the storage-computing integrated hard disk, unilateral reading refers to that the client sends a write instruction to the storage-computing integrated hard disk, then the storage-computing integrated hard disk requests data to be written from the client, then the client sends the data to be written, and the storage-computing integrated hard disk writes the data to be written in the storage unit; the method can directly detect the time delay of the storage unit in storing data, and at least based on the time delay, slow disk detection is performed.

[0084] Here, only a brief description of the method is given, and detailed contents of the method are described below.

[0085] Figure 1 An architecture example diagram of a distributed storage system to which embodiments of the present application are applied is shown. As shown in the figure, Figure 1 The distributed storage system includes a plurality of clients 120 and a plurality of storage-computing integrated hard disks 110. Each client 120 can communicate with any storage-computing integrated hard disk 110. The storage-computing integrated hard disk 110 can include a computing unit and a storage unit. The computing unit can be a dedicated computing unit such as an NPU, an FPGA, an ASIC, or an AI acceleration core. The computing unit supports real-time I / O monitoring and data analysis, supports slow disk detection logic, and the slow disk detection logic can be customized and adjusted. The storage unit can be a non-volatile memory such as a magnetic medium or a flash chip. The storage unit and the computing unit interact with each other. The storage unit is used to provide storage resources, for example, to store data.

[0086] The client 120 is used to determine a data operation request. The data operation request can be generated by processing a data access request from the outside, or can be generated by processing a data access request generated inside the distributed storage system. Illustratively, the processing can include metadata management, duplicate data deletion, data compression, data verification, virtualization storage space, and address conversion.

[0087] The computing unit in the storage-computing integrated hard disk 110 is configured to receive a data operation request of the client 120, and operate the storage unit in the storage-computing integrated hard disk 110 based on the data operation request. For example, the operation on the storage unit can include reading data in the storage unit, writing data into the storage unit, or deleting data in the storage unit.

[0088] In addition, the computing unit in the storage-computing integrated hard disk 110 supports real-time I / O monitoring and data analysis, and supports slow disk detection logic, which can be customized and adjusted.

[0089] For example, as shown in FIG. 1, the storage-computing integrated disk 110 can include a computing unit, a storage unit, a collection unit, a communication unit, and a control unit. Figure 1

[0090] The computing unit can be a special computing unit, which can be a customized computing logic or a special accelerator, such as an NPU, an FPGA, an ASIC, or an AI acceleration core. The collection unit can be a chip register. The collection unit is configured to collect multi-dimensional index data required for slow disk detection in real time, such as the time delay of the storage unit, the load of the storage unit, the health data of the storage unit, and system logs. The collection unit can also collect other related information of the storage-computing integrated disk.

[0091] The communication unit can be a network card. The communication unit is configured to communicate with the upper-layer client and other storage-computing integrated hard disks. In addition, the communication unit is also configured to report the slow disk detection result to the upper-layer service warning platform, and receive the slow disk detection logic.

[0092] The storage unit can be a magnetic medium, a flash chip, or other non-volatile memory. The storage unit and the computing unit interact with each other.

[0093] The control unit is configured to control the state of each unit in the storage-computing integrated disk, such as the partial isolation of the storage unit, or various operations such as data recovery.

[0094] ​It should be noted that in the storage-computing integrated design, in-memory computing or near-memory computing can be used; wherein the in-memory computing is to directly embed the computing logic into the storage unit (new type of non-volatile memory), and the data is calculated in the storage unit to realize the direct interaction at the physical layer. The near-memory computing is to place the computing unit close to the storage unit, and shorten the communication path through the high-speed interconnection. Therefore, the computing unit (such as customized computing logic or special accelerator) can be physically close to or embedded in the storage unit (such as NAND flash memory), so that the computing unit directly processes data near the storage unit, bypasses the traditional bus transmission (such as PCIe or memory bus), and realizes the direct communication at the logical level. In some possible scenarios, the computing unit can be a special controller, and the direct access to the storage unit such as a flash memory chip is realized through the optimized controller.

[0095] In actual application, a user accesses the distributed storage system through the client 120. In some possible scenarios, the user accesses the distributed system through a terminal device, and in other possible scenarios, the distributed storage system is deployed in the cloud, the cloud includes a cloud management platform, and the user accesses the client 120 by accessing the cloud management platform through the terminal device to access the distributed storage system. The terminal device can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The exemplary embodiments of the terminal device involved in the embodiments of the present application include, but are not limited to, electronic devices loaded with iOS, android, Windows, Harmony OS or other operating systems. The embodiments of the present application do not make specific limitation on the type of electronic device. The cloud management platform can be a separate electronic device or a software module integrated on an electronic device, and the embodiments of the present application do not limit the specific deployment manner and deployment position of the cloud management platform.

[0096] In a possible scenario, the client 120 and the storage-computing integrated hard disk 110 are deployed separately. In this scenario, the distributed storage system 100 includes a computing node cluster and a storage node cluster. As shown in FIG. 1, the computing node cluster includes one or more computing nodes 130, and each computing node 130 can communicate with each other. The computing node 130 deploys the client 120, and the client 120 can be a virtual machine or a virtual container or the like. The computing node 130 is a computing device, such as a server, a desktop computer or a controller of a storage array. Figure 2 As shown in FIG. 2, the computing node 130 at least includes a processor 131, a memory 132 and a network card 133. Figure 2 As shown in FIG. 2, the computing node 130 at least includes a processor 131, a memory 132 and a network card 133.

[0097] The processor 131 and the memory 132 are configured to provide computing resources. Specifically, the processor 131 is a central processing unit (CPU) configured to process data access requests from outside the application computing node 130 or requests generated inside the computing node 130. For example, when the processor 131 receives a write data request sent by a user, the processor 131 temporarily stores the data in the write data request in the memory 132. When the total amount of data in the memory 132 reaches a certain threshold, the processor 131 sends the data stored in the memory 132 to the storage-computing integrated hard disk 110 for persistent storage. In addition, the processor 131 is configured to perform data computation or processing, such as metadata management, data deduplication, data compression, data verification, virtual storage space, address translation, and the like. Figure 2 Only one processor 131 is shown in the figure, but in actual applications, the number of processors 131 is often multiple, and one processor 131 has one or more CPU cores. The number of processors 131 and the number of CPU cores are not limited in this embodiment.

[0098] The memory 132 refers to an internal memory that exchanges data directly with the processor, which can read and write data at any time and at a very fast speed, and serves as temporary data storage for the operating system or other programs that are running. The memory includes at least two types of memories, for example, the memory can be a random access memory or a read-only memory (Read Only Memory, ROM). For example, the random access memory is a dynamic random access memory (Dynamic Random Access Memory, DRAM) or a storage class memory (Storage Class Memory, SCM). The DRAM is a semiconductor memory, which, like most random access memories (Random Access Memory, RAM), is a type of volatile memory device. The SCM is a hybrid storage technology that combines the characteristics of traditional storage devices and memories. The storage class memory can provide faster read and write speeds than hard drives, but slower computing speeds than DRAM, and is more cost-effective than DRAM. However, the DRAM and SCM are only exemplary in this embodiment, and the memory can also include other random access memories, such as static random access memories (Static Random Access Memory, SRAM), etc. For read-only memories, for example, they can be programmable read-only memories (Programmable Read Only Memory, PROM), erasable programmable read-only memories (Erasable Programmable Read Only Memory, EPROM), etc. In addition, the memory 113 can also be a dual in-line memory module or a dual in-line memory module (Dual In-line Memory Module, DIMM), i.e., a module composed of dynamic random access memory (DRAM), and can also be a solid state drive (SSD). In practical applications, multiple memories 132 and different types of memories 132 can be configured in the computing node 130. This embodiment does not limit the number and type of the memory 132. In addition, the memory 132 can be configured to have a power retention function. The power retention function refers to that when the system is powered off and then powered on again, the data stored in the memory 132 will not be lost. The memory with the power retention function is called a non-volatile memory.

[0099] The network card 133 is used to communicate with the storage node 140. For example, when the total amount of data in the memory 132 reaches a certain threshold, the computing node 130 can send a request to the storage node 140 through the network card 133 to persistently store the data. In addition, the computing node 130 can also include a bus for communication between the components inside the computing node 130. In terms of function, sinceFigure 2 The main function of the computing node 130 in the system is to compute services, and when storing data, it can use remote memory to achieve persistent storage, so it has less local memory than a regular server, thereby achieving cost and space savings. However, this does not mean that the computing node 130 cannot have local memory. In actual implementation, the computing node 130 can also be built-in with a small amount of hard disk or externally connected with a small amount of hard disk.

[0100] Any one of the computing nodes 130 can access any one of the storage nodes 140 in the storage node cluster through the network. The storage node cluster includes a plurality of storage nodes 140. One storage node 140 includes one or more storage-compute integrated hard disks 110. The storage-compute integrated hard disk 110 is configured to communicate with the computing node 130. The storage-compute integrated hard disk 110 is configured to store data. The storage-compute integrated hard disk 110 is configured to write data into a storage unit in the storage-compute integrated hard disk 110, or read or delete data from the storage unit in the storage-compute integrated hard disk 110 according to a read / write data request sent by the computing node 130.

[0101] The network can be a wired network or a wireless network. For example, the wired network can be a cable network, a fiber network, a digital data network (DDN), etc., and the wireless network can be a telecommunication network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a ZigBee network, a Global System for Mobile Communications (GSM) network, etc., or any combination thereof. It can be understood that the network can use any known network communication protocol to realize communication between different client layers and the gateway. The network communication protocol can be various wired or wireless communication protocols, such as an Ethernet, a universal serial bus (USB), a firewire, a global system for mobile communications (GSM), a general packet radio service (GPRS), a code division multiple access (CDMA), a wideband code division multiple access (WCDMA), a time-division code division multiple access (TD-SCDMA), a long term evolution (LTE), a new radio (NR), etc.

[0102] It should be understood that the distributed storage system to which the embodiments of the present application are applicable can also be applicable to other distributed storage systems Figure 1 and Figure 2 The embodiments of the present application are not limited to the distributed storage system shown. For example, the computing node 130 and the storage-computing integrated hard disk 110 can also be integrated.

[0103] Next, in combination with the system provided above, a slow disk detection method provided by the embodiments of the present application is described in detail.

[0104] Figure 3a A scenario diagram of the slow disk detection method provided by the embodiment of the application is shown. As shown in the figure, Figure 3a the slow disk detection method provided by the embodiment of the application can include: 1. The client 120 issues a data operation request to the computing unit 1; 2. The computing unit 1 sends an operation instruction to the storage unit 1 in response to the data operation request; 3. The storage unit 1 sends an operation result to the computing unit 1, and the operation result is used to indicate the result of reading or writing by the storage unit 1 in response to the operation instruction; 4. The computing unit 1 determines the storage delay corresponding to the storage unit 1 based on the time when the computing unit 1 sends the operation instruction to the storage unit 1 and the time when the computing unit 1 receives the operation result sent by the storage unit 1; 5. The computing unit 1 determines whether the storage-computing integrated hard disk 1 is a slow disk based on the storage delay corresponding to the storage unit 1, thereby improving the detection efficiency of the slow disk. Here, only exemplary description is made, and the details can be referred to the description of Figure 3b and Figure 3b . Wherein, the storage-computing integrated hard disk 1 can be any storage-computing integrated hard disk 130 in the distributed system.

[0105] Figure 3b is a flowchart of the slow disk detection method provided by the embodiment of the application. The embodiment can be applied to the computing unit 1 in the storage-computing integrated hard disk 1.

[0106] As shown in the figure, Figure 3b the slow disk detection method provided by the embodiment of the application at least includes the following steps:

[0107] Step 301, the client 120 sends a data operation request to the computing unit 1.

[0108] In this embodiment, the data operation request is used to request data operation. Optionally, the data operation request can be a read request, which can include an operation type of reading and a case of reading out data such as data identifier (used to uniquely identify data); optionally, the data operation request can be a write request; exemplary, the write request can include a write type and a data size to be written; exemplary, the write request can include a write type and data to be written; wherein, the write type can be append write or overwrite.

[0109] Step 302, the computing unit 1 sends an operation instruction to the storage unit 1 in response to the data operation request.

[0110] In this embodiment, optionally, the data operation request can be a read request, which can include an operation type of reading and a case of reading out data such as data identifier (used to uniquely identify data); correspondingly, the operation instruction is used to instruct the storage unit 1 to read out the data to be read out.

[0111] Optionally, the data operation request can be a write request; for example, the write request can include a write type and a data size to be written; correspondingly, the computing unit 1 can send a read request to the client 120, and the client 120 sends the data to be written to the storage-computing integrated hard disk 1 in response to the read request; correspondingly, the operation instruction is used to instruct the storage unit 1 to store the data to be written; for example, the write request can include a write type and data to be written; correspondingly, the operation instruction is used to instruct the storage unit 1 to store the data to be written.

[0112] Step 303, the storage unit 1 performs reading or writing in response to the operation instruction to obtain an operation result; wherein the operation result is used to indicate the result of the storage unit 1 performing the operation instruction to read or write.

[0113] In this embodiment, the storage unit 1 writes or reads data according to the operation instruction to obtain an operation result. In the case of reading data, the operation result can be the read data; in the case of writing data, the operation result can be the written data.

[0114] In some possible scenarios, such as Figure 4a As shown in the figure, 1. The client 120 sends a data operation request to the storage-computing integrated hard disk 110: read data; 2. The computing unit responds to the data operation request and sends an operation instruction to the storage unit: read data; 3. The storage unit sends an operation result to the computing unit: read data.

[0115] In some possible scenarios, such as Figure 4b As shown in the figure, 1. The client 120 sends a data operation request to the storage-computing integrated hard disk 110: write type and data to be written; 2. The computing unit responds to the data operation request and sends an operation instruction to the storage unit: data to be written; 3. The storage unit sends an operation result to the computing unit: data has been written.

[0116] In some possible scenarios, such as Figure 4c As shown in the figure, 1. The client 120 sends a data operation request to the storage-computing integrated hard disk 110: write type; 2. The computing unit responds to the data operation request and sends a read request to the client 120; 3. The client 120 sends data to be written to the storage-computing integrated hard disk 110; 4. The computing unit sends an operation instruction to the storage unit: data to be written; 5. The storage unit sends an operation result to the computing unit: data has been written.

[0117] Step 304, the storage unit 1 sends an operation result to the computing unit 1.

[0118] Step 305, the computing unit 1 acquires a first time and a second time; wherein the first time is used to indicate the time when the computing unit 1 sends an operation instruction to the storage unit 1 in response to a data operation request, and the second time delay is used to indicate the time when the computing unit 1 receives the operation result sent by the storage unit 1.

[0119] In this embodiment, the computing unit 1 can record the first time when the operation instruction is sent in the process of executing step 302; subsequently, in the case of receiving the operation result, the second time when the operation instruction is received is recorded.

[0120] Step 306, the computing unit 1 determines the first storage time delay corresponding to the storage unit 1 based on the first time and the second time; wherein the first storage time delay is used to indicate the time required for the storage unit 1 to read or write.

[0121] In this embodiment, the computing unit 1 can take the time length between the second time and the first time as the first storage time delay.

[0122] Step 307, the computing unit 1 determines the slow disk detection result of the storage-computing integrated hard disk 1 based on the first storage time delay corresponding to the storage unit 1; wherein the slow disk detection result is used to indicate whether the storage-computing integrated hard disk 1 is a slow disk.

[0123] In some optional implementations of this embodiment, the computing unit 1 can collect the first storage time delay in the manner of the foregoing steps 301 to 306 to obtain N (a positive integer greater than or equal to 1) first storage time delays, and the first storage time delay corresponds to a collection time, which can be the time when the computing unit 1 obtains the first storage time delay; a storage time delay threshold value is acquired; then, based on the N first storage time delays and the storage time delay threshold value, the slow disk detection result is determined, for example, in the case that any first storage time delay is greater than or equal to the storage time delay threshold value, or a plurality of consecutive first storage time delays are all greater than or equal to the storage time delay threshold value, the storage-computing integrated hard disk 1 is determined to be a slow disk; for another example, a preset proportion of the N first storage time delays is greater than or equal to the storage time delay threshold value, the storage-computing integrated hard disk 1 is determined to be a slow disk.

[0124] Optionally, the obtaining, by the computing unit 1, of the storage latency threshold value can include: obtaining load data of the first storage unit; determining the storage latency threshold value based on the load data of the first storage unit. The load data can include IOPS, read-write bandwidth, and I / O size, the I / O size refers to the amount of data involved in each input / output operation, that is, the amount of data requested by the application at a time, the amount of data can include the size and quantity of data blocks, and the data block is the smallest fixed unit of reading and writing of the storage-computing integrated hard disk 1 (such as a 4KB block of a file system or an 8KB page of a database); the read-write bandwidth refers to the amount of data that the storage unit 1 can transmit per unit time, which is usually measured by the number of bits transmitted per second (bps) or the number of bytes (Bps), the higher the read-write bandwidth, the faster the data transmission speed; the greater the load of the storage unit 1 indicated by the load data, such as the greater the I / O size, IOPS, and read-write bandwidth, the greater the storage latency threshold value, for example, the storage-computing integrated hard disk 1 can obtain a previously used storage latency threshold value, correct the storage latency threshold value according to the load data, and obtain the latest storage latency threshold value, so that the storage latency threshold value is dynamically updated to adapt to changes in the load, thereby improving the accuracy of slow disk detection.

[0125] In this embodiment, the storage-computing integrated hard disk adopts a storage-computing integrated architecture, so that the storage-computing integrated hard disk has computing capability, and subsequently, the computing unit inside the storage-computing integrated hard disk and the storage unit inside the storage-computing integrated hard disk directly communicate with each other, quickly obtain the storage-computing latency, and perform slow disk detection using the storage latency, thereby improving the efficiency of slow disk detection.

[0126] In this embodiment, for the storage-computing integrated hard disk 1, in order to ensure the accuracy of slow disk analysis, the computing unit 1 and the storage unit 1 can be determined to be normal before slow disk detection, for example, the computing unit 1 determines that the data processing performed by the computing unit 1 in response to a data processing request is normal, the operating system deployed by the computing unit 1 is normal, and the storage unit 1 is in a healthy state.

[0127] For example, as shown in Figure 5 Based on steps 1 to 4 shown in Figure 3a , the steps further include: 5. The computing unit 1 determines that the data processing performed by the computing unit 1 in response to a data processing request is normal, the operating system deployed by the computing unit 1 is normal, and the storage unit 1 is in a healthy state; 6. The computing unit 1 determines whether the storage-computing integrated hard disk 1 is a slow disk based on the storage latency of the storage unit 1, thereby improving the detection efficiency of the slow disk.

[0128] Optionally, the computing unit 1 can obtain system logs; the system logs are used to indicate the running status of the operating system deployed by the computing unit 1, and then determine whether the operating system deployed by the computing unit 1 is normal based on the system logs.

[0129] Optionally, the computing unit 1 can acquire health data of the storage unit 1, such as data collected by SMART technology, and the health data is used to indicate the health condition of the storage unit 1, and then whether the storage unit 1 is healthy can be determined based on the health data of the storage unit 1. For example, the health data can include, but is not limited to, the temperature of the first storage unit, the number of bad sectors, the number of startup retries, the number of recalibration retries, the power-on time, the startup time, and the number of start-stop.

[0130] Optionally, the computing unit 1 can acquire a data processing delay, which is used to indicate the duration of data processing performed by the computing unit 1 in response to a data processing request, and then whether the data processing performed by the computing unit 1 in response to the data processing request is normal or not can be determined based on the data processing delay. For example, in the case where the data processing delay is greater than or equal to a preset data processing delay threshold, it is determined that the data processing performed by the computing unit 1 in response to the data processing request is not normal, otherwise it is normal. In an optional example, the computing unit 1 can acquire the data processing delay threshold, which can include: acquiring load data of the first storage unit; determining the data processing delay threshold based on the load data of the first storage unit. The greater the load of the storage unit 1 indicated by the load data, such as the I / O size, IOPS and read-write bandwidth, the greater the data processing delay threshold. For example, the storage-computing integrated hard disk 1 can acquire the data processing delay threshold used before, correct the data processing delay threshold according to the load data, and obtain the latest data processing delay threshold, so that the data processing delay threshold is dynamically updated to adapt to the change of the load, and the accuracy of the slow disk detection is improved.

[0131] In this scheme, on the basis of the storage-computing integrated hard disk architecture, the computing unit inside the hard disk can perform its own data processing, its own operating system and the health condition of the storage unit. In the case where the data processing is normal, the operating system is normal and the health condition of the storage unit is normal, it is determined that the environment outside the storage unit is normal, and then the slow disk detection is performed by using the storage delay, so that the accuracy of the slow disk detection is improved.

[0132] In this embodiment, the storage-computing integrated hard disk 1 can not only detect whether it is a slow disk, but also detect whether there is a slow network between the client 120 and the storage-computing integrated hard disk 1.

[0133] Figure 6a A scene schematic diagram of the slow disk detection method provided by the embodiment of the application is shown. As shown in Figure 6a The slow disk detection method provided by the embodiment of the application is used in Figure 3aOn the basis of the above, the method can further comprise: 6. The computing unit 1 determines a request response time delay based on a time when the computing unit 1 receives the data operation request and a time when the computing unit 1 receives the operation result sent by the storage unit 1; and 7. The computing unit 1 determines whether there is a slow network based on the request response time delay and the storage time delay corresponding to the storage unit 1, thereby improving the detection efficiency of the slow network. Here, only exemplary descriptions are given, and details can be found in the description of Figure 6b and the description of Figure 6b .

[0134] Figure 6b A flowchart of another slow disk detection method provided by an embodiment of the application is shown.

[0135] As shown in Figure 6b , in the embodiment of the application, on the basis of steps 301 to 307, at least the following steps can be further included:

[0136] Step 601: The computing unit 1 acquires a third time when the data operation request is received.

[0137] Step 602: The computing unit 1 determines a request response time delay based on the third time and the second time; wherein the request response time delay is used to indicate the time length taken to complete the data operation request.

[0138] In the embodiment, the computing unit 1 can take the time length between the third time and the second time as the request response time delay.

[0139] Step 603: The computing unit 1 determines whether there is a slow network based on the first storage time delay and the request response time delay.

[0140] In the embodiment, the computing unit 1 can determine a network time delay based on the first storage time delay and the request response time delay, such as the difference between the first storage time delay and the request response time delay; then, the computing unit 1 acquires a network time delay threshold, and subsequently determines whether there is a slow network based on the network time delay and the network time delay threshold.

[0141] Optionally, the computing unit can collect network time delays to obtain M (a positive integer greater than or equal to 1) network time delays, the network time delays correspond to collection times, and the collection times can be the times when the computing unit obtains the network time delays; acquire a network time delay threshold; then, determine a slow disk detection result based on the M network time delays and the network time delay threshold, such as determining that the storage-computing integrated hard disk is a slow disk in the case that any network time delay is greater than or equal to the network time delay threshold, or in the case that a plurality of continuous network time delays are all greater than or equal to the network time delay threshold; or such as determining that the storage-computing integrated hard disk is a slow disk in the case that a preset proportion of the M network time delays are greater than or equal to the network time delay threshold.

[0142] In the scheme, on the basis of adopting the storage-computing integrated hard disk architecture, the computing unit inside the storage-computing integrated hard disk can directly communicate with the storage unit inside the storage-computing integrated hard disk, collect the request response time delay, and based on the request response time delay and the storage time delay, further analyze whether there is a slow network, so as to realize the double detection of slow disk and slow network.

[0143] Another slow disk detection method provided by the embodiment of the application is provided.

[0144] Figure 7a A scene schematic diagram of another slow disk detection method provided by the embodiment of the application is shown. Figure 7a As shown in the figure, the slow disk detection method provided by the embodiment of the application can include: 1. The client 120 issues a data operation request to the computing unit 1; 2. The computing unit 1 responds to the data operation request and sends an operation instruction to the storage unit 1; 3. The storage unit 1 sends an operation result to the computing unit 1, and the operation result is used to indicate the result of reading or writing of the storage unit 1 in response to the operation instruction; 4. The computing unit 1 determines the storage time delay corresponding to the storage unit 1 based on the time when the computing unit 1 sends the operation instruction to the storage unit 1 and the time when the computing unit 1 receives the operation result sent by the storage unit 1; 5. The computing unit 2 in the storage-computing integrated hard disk 2 sends the storage time delay corresponding to the storage unit 2 to the storage-computing integrated hard disk 1; 6. The computing unit 1 determines whether the storage-computing integrated hard disk 1 is a slow disk based on the storage time delay corresponding to the storage unit 1 and the storage time delay corresponding to the storage unit 2, thereby improving the detection efficiency of the slow disk. Here, only exemplary description is made, and the detailed content can be referred to the description of Figure 7b and Figure 7b ; wherein the storage-computing integrated hard disk 2 is any storage-computing integrated hard disk 110 other than the storage-computing integrated hard disk 1 in the distributed storage system.

[0145] Wherein the storage-computing integrated hard disk 2 can be any storage-computing integrated hard disk other than the storage-computing integrated hard disk 1 in the distributed storage system, and the determination method of the storage time delay is referred to the determination method of the storage time delay of the storage-computing integrated hard disk 110.

[0146] The application will take the storage-computing integrated hard disk 2 as the storage-computing integrated hard disk 2, and the storage-computing integrated hard disk 2 includes the computing unit 2 (i.e. the computing unit 2) and the storage unit 2 (i.e. the storage unit 2) as an example to describe the scheme of the application. The storage-computing integrated hard disk 2 can be any storage-computing integrated hard disk 130 in the distributed system.

[0147] Figure 7b A flowchart of another slow disk detection method provided by the embodiment of the application is shown.

[0148] As Figure 7bAs shown, on the basis of steps 301 to 305, the embodiment of the present application includes at least the following steps:

[0149] Step 701, the storage-computing integrated hard disk 2 sends the second storage time delay to the computing unit 1; wherein the storage-computing integrated hard disk 2 includes the storage unit 2, and the second storage time delay is used to indicate the time length required for the storage unit 2 to store data for reading and / or writing.

[0150] In this embodiment, the storage-computing integrated hard disk 2 can include the computing unit 2, and the computing unit 2 acquires the second storage time delay. The acquisition manner can refer to the content of the acquisition of the first storage time delay by the computing unit 1, and will not be described in detail.

[0151] It is worth noting that the storage-computing integrated hard disk 2 sends the second storage time delay to the computing unit 1 in the case that it is determined that the storage-computing integrated hard disk 2 is not a slow disk.

[0152] For example, the manner in which the storage-computing integrated hard disk 2 determines whether it is a slow disk can refer to the description of step 307, and will not be described in detail. In addition, the storage-computing integrated hard disk 2 can determine whether it is a slow disk based on the second storage time delay in the case that it is determined that the data processing performed by the computing unit 2 in response to the data processing request is normal, the operating system deployed by the computing unit 2 is normal, and the storage unit 2 is in a healthy state, and send the second storage time delay to the computing unit 1 in the case that it is determined that the storage-computing integrated hard disk 2 is a normal disk.

[0153] Optionally, the computing unit 2 can acquire a system log; wherein the system log is used to indicate the running condition of the operating system deployed by the computing unit 2, and then determine whether the operating system deployed by the computing unit 2 is normal based on the system log.

[0154] Optionally, the computing unit 2 can acquire the health data of the storage unit 2, such as the data collected by the SMART technology, and the health data is used to indicate the health condition of the storage unit 2, and then determine whether the storage unit 2 is healthy based on the health data of the storage unit 2. For example, the health data can include but is not limited to the temperature of the second storage unit, the number of bad sectors, the number of startup retries, the number of recalibration retries, the power-on time, the startup time, and the number of start-stop times.

[0155] Optionally, the computing unit 2 can obtain a data processing time delay, the data processing time delay being used to indicate a time length of data processing performed by the computing unit 2 in response to a data processing request; based on the data processing time delay, it is determined whether the data processing performed by the computing unit 2 in response to the data processing request is normal; for example, in a case where the data processing time delay is greater than or equal to a preset data processing time delay threshold, it is determined that the data processing performed by the computing unit 2 in response to the data processing request is not normal, otherwise it is normal; optionally, in an example, the computing unit 2 obtaining the data processing time delay threshold can include: obtaining load data of the second storage unit; based on the load data of the second storage unit, determining the data processing time delay threshold; the greater the load of the storage unit 2 indicated by the load data, such as the greater the I / O size, IOPS and read-write bandwidth, the greater the data processing time delay threshold; for example, the storage-computing integrated hard disk 2 can obtain a previously used data processing time delay threshold, correct the data processing time delay threshold according to the load data, and obtain the latest data processing time delay threshold, so that the data processing time delay threshold is dynamically updated to adapt to the change of the load, thereby improving the accuracy of the slow disk detection.

[0156] Step 702, the storage-computing integrated hard disk 1 determines a slow disk detection result of the storage-computing integrated hard disk 1 based on the first storage time delay and the second storage time delay.

[0157] In some optional implementations of the present embodiment, the storage-computing integrated hard disk 1 compares the difference between the first storage time delay and the second storage time delay, and in a case where the difference is small, such as less than or equal to a preset threshold, it is determined that the storage-computing integrated hard disk 1 is not a slow disk, otherwise it is determined that the storage-computing integrated hard disk 1 is a slow disk.

[0158] In some optional implementations of the present embodiment, the second storage integrated disk can also send load data of the storage unit 2; then, based on the load data of the storage unit 1 and the load data of the storage unit 2, the storage-computing integrated hard disk 1 can determine whether the loads of the storage unit 1 and the storage unit 2 are similar, such as calculating a difference 1 between the I / O size of the storage unit 1 and the I / O size of the storage unit 2, a difference 2 between the IOPS of the storage unit 1 and the IOPS of the storage unit 2, and / or a difference 3 between the read-write bandwidth of the storage unit 1 and the storage unit 2, in a case where the difference 1, the difference 2 and / or the difference 3 is small, such as less than or equal to a preset threshold, it can be considered that the loads of the storage unit 1 and the storage unit 2 are similar; subsequently, the difference between the first storage time delay and the second storage time delay is compared, and in a case where the difference is small, such as less than or equal to a preset threshold, it is determined that the storage-computing integrated hard disk 1 is not a slow disk, otherwise it is determined that the storage-computing integrated hard disk 1 is a slow disk.

[0159] In some optional implementations of the embodiment, the second storage integrated disk can also send a second storage latency threshold, and the storage-computing integrated hard disk 1 compares a difference 1 between the first storage latency and the second storage latency, and a difference 2 between the first storage latency threshold and the second storage latency threshold. In a case where the difference 1 and the difference 2 are both small, for example, less than or equal to a preset threshold, it is determined that the storage-computing integrated hard disk 1 is not a slow disk, otherwise, it is determined that the storage-computing integrated hard disk 1 is a slow disk.

[0160] In the embodiment, the storage-computing integrated hard disks communicate with each other, so that the storage-computing integrated hard disks can obtain the storage latencies of other storage-computing integrated hard disks. By comparing the storage latencies of different storage-computing integrated hard disks, the accuracy and efficiency of slow disk detection are improved.

[0161] As shown in FIG. 1, Figure 8 The embodiment provides a specific scenario of a slow disk detection method.

[0162] For the storage-computing integrated hard disk 2, the specific steps can include the following steps:

[0163] 21. The client 120 issues a data operation request to the computing unit 1; 22. The computing unit 1 sends an operation instruction to the storage unit in response to the data operation request; 23. The storage unit 1 sends an operation result to the computing unit 1; 24. The computing unit 1 determines a storage latency 2 corresponding to the storage unit 1 based on a time when the computing unit 1 sends the operation instruction to the storage unit 1, and a time when the computing unit 1 receives the operation result sent by the storage unit 1; 25. The computing unit 1 determines that the data processing performed by the computing unit 1 in response to the data processing request is normal, the operating system deployed by the computing unit 1 is normal, and the storage unit 1 is in a healthy state, and sends the storage latency 2 to the storage-computing integrated hard disk 1.

[0164] For the storage-computing integrated hard disk 1, the specific steps can include the following steps:

[0165] 11. The client 120 issues a data operation request to the computing unit 1; 12. The computing unit 1 sends an operation instruction to the storage unit in response to the data operation request; 13. The storage unit 1 sends an operation result to the computing unit 1; 14. The computing unit 1 determines a storage latency 1 corresponding to the storage unit 1 based on a time when the computing unit 1 sends the operation instruction to the storage unit 1, and a time when the computing unit 1 receives the operation result sent by the storage unit 1; 15. The computing unit 1 determines that the data processing performed by the computing unit 1 in response to the data processing request is normal, the operating system deployed by the computing unit 1 is normal, and the storage unit 1 is in a healthy state; 16. The computing unit 1 determines whether the storage-computing integrated hard disk 1 is a slow disk based on the storage latency 1 and the storage latency 2.

[0166] It should be noted that the above-mentioned storage and computing integrated hard disk 1, computing unit 1, storage unit 1, storage and computing integrated hard disk 2, computing unit 2, storage unit 2 are only as an exemplary naming manner, and in some possible scenarios, they can also be called first storage and computing integrated hard disk, first computing unit, first storage unit, second storage and computing integrated hard disk, second computing unit, second storage unit.

[0167] The application also provides a slow disk detection device, which is applied to the storage and computing integrated hard disk 110, as shown in the figure, comprising: Figure 9

[0168] The acquisition module is configured to acquire a first time point, wherein the first time point is used to indicate a time point at which the computing unit sends an operation instruction to the first storage unit in response to a data operation request, and the data operation request comes from a client; and acquire a second time point, wherein the second time point is a time point at which the computing unit receives an operation result sent by the first storage unit, and the operation result is used to indicate a result of reading or writing performed by the first storage unit on the operation instruction.

[0169] The time delay determination module is configured to determine a first storage time delay corresponding to the first storage unit based on the first time point and the second time point, wherein the first storage time delay is used to indicate a time length required by the first storage unit for reading or writing.

[0170] The detection module is configured to determine a slow disk detection result of the first storage and computing integrated hard disk based on the first storage time delay, wherein the slow disk detection result is used to indicate whether the first storage and computing integrated hard disk is a slow disk.

[0171] The acquisition module, the time delay determination module and the detection module can be implemented by software or by hardware. For example, the implementation of the acquisition module is introduced as follows. Similarly, the implementation of the time delay determination module and the detection module can refer to the implementation of the acquisition module.

[0172] ​As an example of a software functional unit, the obtaining module can include code running on a computing instance. The computing instance can include at least one of a storage-computing integrated hard disk, a virtual machine, and a container. Further, the computing instance can be one or more. For example, the obtaining module can include code running on multiple storage-computing integrated hard disks / virtual machines / containers. It should be noted that the multiple storage-computing integrated hard disks / virtual machines / containers used to run the code can be distributed in the same region, or can be distributed in different regions. Further, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ), or can be distributed in different AZs, each AZ including one data center or multiple data centers in close geographical proximity. Generally, one region can include multiple AZs.

[0173] Similarly, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC), or can be distributed in multiple VPCs. Generally, one VPC is set in one region, and communication between two VPCs in the same region and between VPCs in different regions needs to be set in each VPC to set a communication gateway, and interconnection between VPCs is realized through the communication gateway.

[0174] As an example of a hardware functional unit, the obtaining module can include at least one storage-computing integrated hard disk. Alternatively, the obtaining module can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD can be implemented by a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0175] The plurality of storage-computing integrated hard disks included in the acquisition module can be distributed in the same region or in different regions. The plurality of storage-computing integrated hard disks included in the acquisition module can be distributed in the same AZ or in different AZs. Similarly, the plurality of storage-computing integrated hard disks included in the acquisition module can be distributed in the same VPC or in multiple VPCs. The plurality of storage-computing integrated hard disks can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0176] It should be noted that in other embodiments, the acquisition module can be used to perform any step in the slow disk detection method, such as any step in the methods shown in Figure 3b , Figure 6b and Figure 7b The acquisition module can be used to perform any step in the slow disk detection method, such as any step in the methods shown in Figure 3b , Figure 6b and Figure 7b The delay determination module can be used to perform any step in the slow disk detection method, such as any step in the methods shown in Figure 3b , Figure 6b and Figure 7b The detection module can be used to perform any step in the slow disk detection method, such as any step in the methods shown in Figure 3b , Figure 6b and Figure 7b The steps implemented by the acquisition module, the delay determination module, and the detection module can be specified as needed, and the acquisition module, the delay determination module, and the detection module respectively implement different steps in the slow disk detection method to realize the overall function of the slow disk detection device.

[0177] For example, the acquisition module and the delay determination module can be deployed in the collection unit of the storage-computing integrated hard disk 110 shown in Figure 1 The detection module can be deployed as detection logic in the calculation unit of the storage-computing integrated hard disk 110 shown in Figure 1 .

[0178] The present application also provides a computing device 1000. As shown in Figure 10 The computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. The processor 1004, the memory 1006, and the communication interface 1008 communicate through the bus 1002. The computing device 1000 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1000.

[0179] The bus 1002 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, or the like. For ease of representation, Figure 10 only one line is used, but it does not mean that there is only one bus or only one type of bus. The bus 1002 can include a path for transmitting information between various components (for example, the memory 1006, the processor 1004, the communication interface 1008) of the computing device 1000.

[0180] The processor 1004 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), or the like.

[0181] The memory 1006 can include a volatile memory (for example, a random access memory (RAM)), and the processor 1004 can further include a non-volatile memory (for example, a read-only memory (ROM), a flash memory, a mechanical hard disk drive (HDD), or a solid state drive (SSD)).

[0182] The memory 1006 stores executable program code, and the processor 1004 executes the executable program code to respectively implement the functions of the foregoing acquisition module, and the time delay determination module and the detection module, thereby implementing the slow disk detection method, such as the method shown in Figure 3b 、 Figure 6b and Figure 7b . That is, the memory 1006 stores instructions for executing the slow disk detection method, such as the method shown in Figure 3b 、 Figure 6b and Figure 7b .

[0183] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card, a transceiver, or the like, to implement communication between the computing device 1000 and other devices or communication networks.

[0184] The embodiments of the present application also provide a computing device cluster. The computing device cluster comprises at least one computing device. The computing device can be a server, for example, a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.

[0185] As shown in Figure 11 , the computing device cluster comprises at least one computing device 1000. The same instructions for performing the slow disk detection method can be stored in the memory 1006 of one or more computing devices 1000 in the computing device cluster.

[0186] In some possible implementations, partial instructions for performing the slow disk detection method, such as the partial instructions of the methods shown in Figure 3b 、 Figure 6b and Figure 7b , can also be stored in the memory 1006 of one or more computing devices 1000 in the computing device cluster, respectively. In other words, the combination of one or more computing devices 1000 can collectively execute the instructions for performing the slow disk detection method, such as the instructions of the methods shown in Figure 3b 、 Figure 6b and Figure 7b .

[0187] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster can store different instructions, respectively, for performing partial functions of the slow disk detection apparatus. That is, the instructions stored in the memory 1006 in different computing devices 1000 can implement the functions of the obtaining module, and one or more of the time delay determination module and the detection module.

[0188] In some possible implementations, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc. Figure 12 A possible implementation is shown. As shown in Figure 12 , two computing devices 1000A and 1000B are connected through a network. Specifically, the communication interface in each computing device is connected to the network. In this type of possible implementation, the memory 1006 in the computing device 1000A stores instructions for performing the functions of the obtaining module. At the same time, the memory 1006 in the computing device 1000B stores instructions for performing the functions of the time delay determination module and the detection module.

[0189] Figure 12The connection mode between the illustrated computing device clusters can be that a large amount of storage delay needs to be collected for the slow disk detection method provided in the present application, and therefore the functions implemented by the detection module are considered to be executed by the computing device 1000B.

[0190] It should be understood that Figure 12 The functions of the computing device 1000A illustrated in the above embodiment can also be completed by multiple computing devices 1000. Similarly, the functions of the computing device 1000B can also be completed by multiple computing devices 1000.

[0191] It should be noted that the above computing device 1000 is only an example. In another possible implementation, the computing device 1000 can be a storage-computing integrated hard disk 110, at which time the processor 1004 can be a computing unit in the storage-computing integrated hard disk 110, the memory 1006 can be a storage unit in the storage-computing integrated hard disk 110, and the communication interface 1008 can be a communication unit in the storage-computing integrated hard disk 110.

[0192] The present application also provides a computer program product containing instructions. The computer program product can be a software or program product containing instructions, which can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes the at least one computing device or at least one storage-computing integrated hard disk to execute a slow disk detection method, such as the method shown in Figure 3b , Figure 6b and Figure 7b instructions of the method shown in

[0193] The present application also provides a computer readable storage medium. The computer readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk) and the like. The computer readable storage medium includes instructions that instruct at least one computing device or at least one storage-computing integrated hard disk to execute a slow disk detection method, such as the method shown in Figure 3b , Figure 6b and Figure 7b instructions of the method shown in

[0194] In the above embodiments, the description of each embodiment has its own focus, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0195] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0196] The above generally describes the basic principles of the present application in conjunction with specific embodiments. It should be noted that the advantages, benefits and effects mentioned in the present application are only examples and are not intended to limit the various embodiments of the present application. It should not be considered that these advantages, benefits and effects are necessarily required for each embodiment of the present application. In addition, the specific details of the above disclosure are only for the purpose of illustration and understanding, and are not intended to limit the present disclosure to the specific details described above.

[0197] The block diagrams of the devices, apparatuses, equipment, systems involved in the present disclosure are only illustrative examples and are not intended to require or imply the connection, arrangement, configuration shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words, which mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.

[0198] It should also be noted that in the devices, apparatuses and methods of the present disclosure, each component or step can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions of the present disclosure.

[0199] The above description has been given for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain modifications, alterations, changes, additions and sub-combinations thereof.

[0200] It should be understood that the various numerical numbers involved in the embodiments of the present application are only for the purpose of differentiation for convenience of description, and are not intended to limit the scope of the embodiments of the present application.

Claims

1. A slow disk detection method, characterized by, The method is applied to a computing unit in a first storage-computing integrated hard disk, and the first storage-computing integrated hard disk further comprises a first storage unit, and the method comprises the following steps: obtaining a first time point; wherein the first time point is used to indicate a time point at which the computing unit sends an operation instruction to the first storage unit in response to a data operation request, and the data operation request is from a client; obtaining a second time point; wherein the second time point is a time point at which the computing unit receives an operation result sent by the first storage unit, and the operation result is used to indicate a result of reading or writing performed by the first storage unit on the operation instruction; determining a first storage time delay corresponding to the first storage unit based on the first time point and the second time point; wherein the first storage time delay is used to indicate a time length required by the first storage unit for reading or writing; determining a slow disk detection result of the first storage-computing integrated hard disk based on the first storage time delay; wherein the slow disk detection result is used to indicate whether the first storage-computing integrated hard disk is a slow disk.

2. The method of claim 1, wherein, The method further comprises the following steps: obtaining a third time point at which the data operation request is received; determining a request response time delay based on the third time point and the second time point; wherein the request response time delay is used to indicate a time length consumed for completing the data operation request; determining whether there is a slow network based on the first storage time delay and the request response time delay.

3. The method of claim 2, wherein, The step of determining whether there is a slow network based on the first storage time delay and the request response time delay comprises the following steps: determining a network time delay based on the first storage time delay and the request response time delay; determining whether there is a slow network based on the network time delay.

4. The method according to any one of claims 1 to 3, characterized in that, The method further comprises the following steps: obtaining load data of the first storage unit; The step of determining the slow disk detection result of the first storage-computing integrated hard disk based on the first storage time delay comprises the following steps: determining the slow disk detection result of the first storage-computing integrated hard disk based on the first storage time delay and the load data of the first storage unit.

5. The method according to any one of claims 1 to 4, characterized in that, The step of determining the slow disk detection result of the first storage-computing integrated hard disk based on the first storage time delay and the load data of the first storage unit comprises the following steps: determining a storage time delay threshold based on the load data; determining the slow disk detection result of the first storage-computing integrated hard disk based on the first storage time delay and the storage time delay threshold.

6. The method according to any one of claims 1 to 5, characterized in that, Before the step of determining the slow disk detection result of the first storage-computing integrated hard disk based on the first storage time delay, the method further comprises the following step: determining that data processing performed by the computing unit in response to the data processing request is normal, that an operating system deployed by the computing unit is normal, and that the first storage unit is in a healthy state.

7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises the following steps: obtaining a second storage time delay sent by a second storage-computing integrated hard disk; wherein the first storage-computing integrated hard disk and the second storage-computing integrated hard disk are located in a same distributed system, the second storage-computing integrated hard disk is a normal disk, the second storage-computing integrated hard disk comprises a second storage unit, and the second storage time delay is used to indicate a time length required by the second storage unit for reading and / or writing; The slow disk detection result of the first storage and calculation integrated hard disk is determined based on the first storage time delay. The slow disk detection result of the first storage and calculation integrated hard disk is determined by comparing the first storage time delay and the request response time delay.

8. A slow disc detection apparatus, characterized by, The application is applied to a computing unit, and the computing unit is located in a first storage and calculation integrated hard disk. The first storage and calculation integrated hard disk further comprises a first storage unit. The device comprises: The first time point is used to indicate a time point at which the computing unit sends an operation instruction to the first storage unit in response to a data operation request from a client. The second time point is a time point at which the computing unit receives an operation result sent by the first storage unit. The operation result is used to indicate a result of reading or writing performed by the first storage unit on the operation instruction. The first storage time delay corresponding to the first storage unit is determined based on the first time point and the second time point. The first storage time delay is used to indicate a time length required by the first storage unit for reading or writing. The slow disk detection result of the first storage and calculation integrated hard disk is determined based on the first storage time delay. The slow disk detection result is used to indicate whether the first storage and calculation integrated hard disk is a slow disk.

9. A first storage-compute integrated hard disk, characterized in that, The first storage and calculation integrated hard disk comprises a computing unit and a first storage unit. The first time point is used to indicate a time point at which the computing unit sends an operation instruction to the first storage unit in response to a data operation request from a client. The second time point is a time point at which the computing unit receives an operation result sent by the first storage unit. The operation result is used to indicate a result of reading or writing performed by the first storage unit on the operation instruction. The first storage time delay corresponding to the first storage unit is determined based on the first time point and the second time point. The first storage time delay is used to indicate a time length required by the first storage unit for reading or writing. The slow disk detection result of the first storage and calculation integrated hard disk is determined based on the first storage time delay. The slow disk detection result is used to indicate whether the first storage and calculation integrated hard disk is a slow disk.

10. A storage-computing integrated hard disk cluster, characterized in that, The at least one storage and calculation integrated hard disk comprises a computing unit and a storage unit. The computing unit of the at least one storage and calculation integrated hard disk is configured to execute instructions stored in the storage unit of the at least one storage and calculation integrated hard disk, so that the computing unit of the at least one storage and calculation integrated hard disk executes the method in any one of claims 1 to 7.

11. A computer program product comprising instructions, characterized in that, When the instructions are executed by a storage and calculation integrated hard disk cluster, the storage and calculation integrated hard disk cluster executes the method in any one of claims 1 to 7.

12. A computer-readable storage medium, characterized in that, The computer program instructions are executed by a storage and calculation integrated hard disk cluster, so that the storage and calculation integrated hard disk cluster executes the method in any one of claims 1 to 7.