Capacity Monitoring Method, Related Device and Storage Medium of Distributed File System
By setting alarm thresholds and performing directory-level scanning in a distributed file system, an exception directory list is generated, and multi-threaded compression and migration expansion technology is used to solve massive file management problems and achieve efficient capacity monitoring and management.
Patent Information
- Application Number
- CN202210799470.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-07-08
AI Technical Summary
The existing technology lacks efficient methods to monitor and manage massive files in distributed file systems, resulting in difficulty in managing data warehouse capacity.
By monitoring the capacity of network attached storage volumes at preset times, setting secondary and primary alarm thresholds, performing directory-level scanning, generating an exception directory list, and using multi-threaded compression and migration expansion technology to deal with exception directories.
It realizes efficient monitoring and management of distributed file system capacity, reduces business interruption time, improves capacity utilization efficiency, and reduces the need for manual intervention.
Smart Images

Figure CN115129683B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method for monitoring the capacity of a distributed file system, related devices, and a storage medium. Background Art
[0002] With the continuous development of bank informatization, the interaction between various business systems has become increasingly frequent. Therefore, each commercial bank has established its own data warehouse for data sharing among internal systems.
[0003] Traditional data warehouses generally use file interaction for data transfer. That is, the source system uses file transfer to push data to the NAS storage used in the buffer layer of the data warehouse. After data quality verification, the data is stored in the data warehouse for downstream application systems to use.
[0004] However, with the continuous development of banking business, especially the increase in the number of systems and the continuous growth of data volume, the number, size, and data types of source files received by the data warehouse are increasing. Currently, there is no efficient monitoring method for these massive files. Summary of the Invention
[0005] In view of this, this application provides a method for monitoring the capacity of a distributed file system, related devices, and a storage medium, which is used to efficiently monitor the capacity of the distributed file system.
[0006] The first aspect of this application provides a method for monitoring the capacity of a distributed file system, including:
[0007] Monitoring the capacity of the network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume;
[0008] Judging whether the current capacity of the network attached storage volume exceeds the secondary alarm threshold;
[0009] If it is judged that the current capacity of the network attached storage volume exceeds the secondary alarm threshold, then judge whether the current capacity of the network attached storage volume exceeds the primary alarm threshold; wherein, the primary alarm threshold is greater than the secondary alarm threshold;
[0010] If it is judged that the current capacity of the network attached storage volume does not exceed the primary alarm threshold, scan each directory space in the network attached storage volume to obtain a first scan result, and return to execute the step of monitoring the capacity of the network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume;
[0011] If it is determined that the current capacity of the network - attached storage volume exceeds the main warning threshold, scan each directory space in the network - attached storage volume to obtain a second scan result;
[0012] Determine an abnormal directory list according to the first scan result and the second scan result.
[0013] Optionally, the determining an abnormal directory list according to the first scan result and the second scan result includes:
[0014] For the current size of each directory in the second scan result, determine whether the current size of the directory exceeds the space size threshold of the directory; wherein, the space size threshold of the directory is the product of the mean value statistically fixed within a quarter and a preset percentage;
[0015] If it is determined that the current size of the directory exceeds the space size threshold of the directory, determine whether the current date is a special business date of the business component;
[0016] If it is determined that the current date is not a special business date of the business component, determine that the directory is an abnormal directory and store it in the abnormal directory list.
[0017] Optionally, the capacity monitoring method of the distributed file system further includes:
[0018] If it is determined that the current size of the directory does not exceed the space size threshold of the directory or it is determined that the current date is a special business date of the business component, determine the size growth percentage of the directory when obtaining the second scan result compared to when obtaining the first scan result;
[0019] Determine whether the size growth percentage is greater than the growth percentage of the network - attached storage volume;
[0020] If it is determined that the size growth percentage is greater than the growth percentage of the network - attached storage volume, determine that the directory is an abnormal directory and store it in the abnormal directory list.
[0021] Optionally, after determining the abnormal directory list according to the first scan result and the second scan result, it further includes:
[0022] Obtain the number of CPU cores of the compression - dedicated server, and determine the compression parameters of the multi - thread compression program according to the number of CPU cores of the compression - dedicated server; wherein, the compression parameters at least include the number of threads and the compression ratio of the multi - thread compression program;
[0023] Call the multi-threaded compression program to sequentially obtain the abnormal directories in the abnormal directory list according to the sizes of the abnormal directories in the abnormal directory list, and perform emergency compression on the abnormal directories;
[0024] After completing the emergency compression of each abnormal directory, determine whether the current capacity of the network attached storage volume exceeds the main warning threshold;
[0025] If it is determined that the current capacity of the network attached storage volume exceeds the main warning threshold, continue to execute the step of calling the multi-threaded compression program to sequentially obtain the abnormal directories in the abnormal directory list according to the sizes of the abnormal directories in the abnormal directory list, and perform emergency compression on the abnormal directories;
[0026] If it is determined that the current capacity of the network attached storage volume does not exceed the main warning threshold, the multi-threaded compression program stops obtaining the abnormal directories in the abnormal directory list, and stops this compression after the emergency compression of the obtained abnormal directories is completed.
[0027] Optionally, before performing the emergency compression on the abnormal directories, it further includes:
[0028] Determine whether the abnormal directory has been subjected to emergency compression on the same day;
[0029] If it is determined that the abnormal directory has been subjected to emergency compression on the same day, no further emergency compression is performed on the abnormal directory.
[0030] Optionally, after scanning each directory space in the network attached storage volume to obtain a second scan result if it is determined that the current capacity of the network attached storage volume exceeds the main warning threshold, it further includes:
[0031] Obtain the sizes of each business component directory stored in the network attached storage volume;
[0032] For each business component, calculate the difference between the current usage rate of the network attached storage file system and the space ratio where the business component is located;
[0033] When the difference between the difference value and the secondary warning threshold is the smallest and lower than the secondary warning threshold, select the business component for migration and expansion.
[0034] Optionally, the selection of the business component for migration and expansion includes:
[0035] Select new network attached storage volumes with different performances according to the existing labels of the business component directories;
[0036] Create the same directory structure on the new network - attached storage volume as on the network - attached storage volume, and link the component pointers on the directory network - attached storage to the new network - attached storage volume;
[0037] Create a new date directory link on the new network - attached storage volume that links back to the network - attached storage volume.
[0038] The second aspect of the present application provides a capacity monitoring device for a distributed file system, including:
[0039] A first monitoring unit, configured to monitor the capacity of the network - attached storage volume at preset intervals to obtain the current capacity of the network - attached storage volume;
[0040] A first judgment unit, configured to judge whether the current capacity of the network - attached storage volume exceeds a secondary alarm threshold;
[0041] A second judgment unit, configured to, if the first judgment unit determines that the current capacity of the network - attached storage volume exceeds the secondary alarm threshold, judge whether the current capacity of the network - attached storage volume exceeds a primary alarm threshold; wherein, the primary alarm threshold is greater than the secondary alarm threshold;
[0042] A first scanning unit, configured to, if the second judgment unit determines that the current capacity of the network - attached storage volume does not exceed the primary alarm threshold, scan each directory space in the network - attached storage volume to obtain a first scanning result, and activate the first monitoring unit to execute monitoring the capacity of the network - attached storage volume at preset intervals to obtain the current capacity of the network - attached storage volume;
[0043] A second scanning unit, configured to, if the second judgment unit determines that the current capacity of the network - attached storage volume exceeds the primary alarm threshold, scan each directory space in the network - attached storage volume to obtain a second scanning result;
[0044] A determination unit, configured to determine a list of abnormal directories according to the first scanning result and the second scanning result.
[0045] Optionally, the determination unit includes:
[0046] A third judgment unit, configured to, for the current size of each directory in the second scanning result, judge whether the current size of the directory exceeds the space size threshold of the directory; wherein, the space size threshold of the directory is the product of the mean value fixed - statistically within a quarter and a preset percentage;
[0047] A fourth judgment unit, configured to, if the third judgment unit determines that the current size of the directory exceeds the space size threshold of the directory, judge whether the current date is a special business date of the business component;
[0048] A first determination subunit, configured to, if the fourth judgment unit determines that the current date is not a special business date of the business component, determine that the directory is an abnormal directory and store it in the abnormal directory list.
[0049] Optionally, the capacity monitoring device of the distributed file system further includes:
[0050] A second determination subunit, configured to, if the third judgment unit determines that the current size of the directory does not exceed the space size threshold of the directory or the fourth judgment unit determines that the current date is a special business date of the business component, determine the size growth percentage of the directory when obtaining the second scan result compared to when obtaining the first scan result;
[0051] A fifth judgment unit, configured to judge whether the size growth percentage is greater than the growth percentage of the network attached storage volume;
[0052] A third determination subunit, configured to, if the fifth judgment unit determines that the size growth percentage is greater than the growth percentage of the network attached storage volume, determine that the directory is an abnormal directory and store it in the abnormal directory list.
[0053] Optionally, the capacity monitoring device of the distributed file system further includes:
[0054] A first acquisition unit, configured to acquire the number of CPU cores of the compression dedicated server and determine the compression parameters of the multi-threaded compression program according to the number of CPU cores of the compression dedicated server; wherein, the compression parameters at least include the number of threads and the compression ratio of the multi-threaded compression program;
[0055] An invocation unit, configured to invoke the multi-threaded compression program to sequentially acquire the abnormal directories in the abnormal directory list according to the sizes of the abnormal directories in the abnormal directory list, and perform emergency compression on the abnormal directories;
[0056] The second judgment unit is further configured to, after completing the emergency compression of each abnormal directory, judge whether the current capacity of the network attached storage volume exceeds the main alarm threshold;
[0057] An activation unit, configured to, if the second determination unit determines that the current capacity of the network attached storage volume exceeds the main warning threshold, continue to activate the invocation unit to execute the invocation of the multi-threaded compression program to sequentially obtain the abnormal directories in the abnormal directory list according to the sizes of the abnormal directories in the abnormal directory list, and perform emergency compression on the abnormal directories;
[0058] A first stop unit, configured to, if the second determination unit determines that the current capacity of the network attached storage volume does not exceed the main warning threshold, the multi-threaded compression program stops obtaining the abnormal directories in the abnormal directory list, and stops this compression after the emergency compression of the obtained abnormal directories is completed.
[0059] Optionally, the capacity monitoring device of the distributed file system further includes:
[0060] A sixth determination unit, configured to determine whether the abnormal directory has been subjected to emergency compression on the same day;
[0061] A second stop unit, configured to, if the sixth determination unit determines that the abnormal directory has been subjected to emergency compression on the same day, no longer perform emergency compression on the abnormal directory.
[0062] Optionally, the capacity monitoring device of the distributed file system further includes:
[0063] A second acquisition unit, configured to acquire the size of each service component directory stored in the network attached storage volume;
[0064] A calculation unit, configured to, for each service component, calculate the difference between the current usage rate of the network attached storage file system and the space ratio where the service component is located;
[0065] A selection unit, configured to, when the difference between the difference value and the secondary warning threshold is the smallest and lower than the secondary warning threshold, select the service component for migration and expansion.
[0066] Optionally, the selection unit includes:
[0067] A selection unit, configured to select new network attached storage volumes with different performances according to the existing labels of the service component directories;
[0068] An establishment unit, configured to establish the same directory structure as that on the network attached storage volume on the new network attached storage volume, and link the component links on the directory network attached storage to the new network attached storage volume;
[0069] A linking unit, configured to create a date directory link back to the network attached storage volume on the new network attached storage volume.
[0070] A third aspect of the present application provides an electronic device, including:
[0071] One or more processors;
[0072] A storage device storing one or more programs thereon;
[0073] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the capacity monitoring method of the distributed file system according to any one of the first aspect.
[0074] A fourth aspect of the present application provides a computer storage medium storing a computer program thereon, wherein when the computer program is executed by a processor, the capacity monitoring method of the distributed file system according to any one of the first aspect is implemented.
[0075] As can be seen from the above solutions, the present application provides a capacity monitoring method, related device and storage medium for a distributed file system. The capacity monitoring method of the distributed file system includes: monitoring the capacity of a network-attached storage volume at preset intervals to obtain the current capacity of the network-attached storage volume; determining whether the current capacity of the network-attached storage volume exceeds a secondary alarm threshold; if it is determined that the current capacity of the network-attached storage volume exceeds the secondary alarm threshold, determining whether the current capacity of the network-attached storage volume exceeds a primary alarm threshold; wherein the primary alarm threshold is greater than the secondary alarm threshold; if it is determined that the current capacity of the network-attached storage volume does not exceed the primary alarm threshold, scanning each directory space in the network-attached storage volume to obtain a first scan result, and returning to execute the step of monitoring the capacity of the network-attached storage volume at preset intervals to obtain the current capacity of the network-attached storage volume; if it is determined that the current capacity of the network-attached storage volume exceeds the primary alarm threshold, scanning each directory space in the network-attached storage volume to obtain a second scan result; determining an abnormal directory list according to the first scan result and the second scan result. Thus, efficient monitoring of the capacity of the distributed file system is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0077] Figure 1 It is a specific flowchart of a capacity monitoring method for a distributed file system provided by an embodiment of the present application;
[0078] Figure 2 The specific flowchart of a migration and expansion method provided by another embodiment of the present application;
[0079] Figure 3 The specific flowchart of a migration and expansion method provided by another embodiment of the present application;
[0080] Figure 4 The specific flowchart of a method for determining an abnormal directory list provided by another embodiment of the present application;
[0081] Figure 5 The specific flowchart of a compression method provided by another embodiment of the present application;
[0082] Figure 6 The schematic diagram of a capacity monitoring device for a distributed file system provided by another embodiment of the present application;
[0083] Figure 7 The schematic diagram of an electronic device for implementing a capacity monitoring method of a distributed file system provided by another embodiment of the present application. Specific embodiments
[0084] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0085] It should be noted that the concepts such as "first" and "second" mentioned in the present application are only used to distinguish different devices, modules or units, and are not used to limit the order or mutual dependence relationship of the functions performed by these devices, modules or units. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0086] The embodiments of the present application provide a capacity monitoring method for a distributed file system, as Figure 1 shown, specifically including the following steps:
[0087] S101. Monitor the capacity of the network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume.
[0088] Among them, Network Attached Storage (NAS) is a storage technology that integrates distributed and independent data into a large, centralized data center based on network protocols to facilitate access to different hosts and application servers. The preset time (such as 15 minutes) is set and changed in advance by technicians or relevant authorized staff, and is not limited here.
[0089] A Distributed File System (DFS) means that the physical storage resources managed by the file system are not necessarily directly connected to the local node, but are connected to the node (which can be simply understood as a computer) through a computer network; or it is a complete hierarchical file system formed by combining several different logical disk partitions or volume labels. DFS provides a logically tree-shaped file system structure for resources located anywhere on the network, making it more convenient for users to access shared files distributed on the network. The role of a separate DFS shared folder is an access point relative to other shared folders on the network.
[0090] It should be noted that the narrow sense of file system capacity monitoring refers to the capacity monitoring and INODE monitoring of the file system. The broad sense of capacity monitoring not only includes the above two monitoring indicators, but also includes a series of monitoring indicators such as trend monitoring, hot data monitoring, and data operation index monitoring. These indicators can well reflect the capacity and performance usage of the file system when aggregated together.
[0091] It should be noted that based on the current capacity of all NAS volumes, the usage rate of the current network attached storage file system space and the growth percentage threshold of the current network attached storage file system space can be deduced.
[0092] S102. Determine whether the current capacity of the network attached storage volume exceeds the secondary alarm threshold.
[0093] Among them, the secondary alarm threshold is set and changed in advance by technicians or relevant authorized staff, and is not limited here.
[0094] Specifically, if it is determined that the current capacity of the network attached storage volume exceeds the secondary alarm threshold, then step S103 is executed; if it is determined that the current capacity of the network attached storage volume does not exceed the secondary alarm threshold, it means that there is no abnormality in the current network attached storage file system and no operation is required.
[0095] In the specific implementation process of the present application, when it is determined that the current capacity of the network - attached storage volume exceeds the secondary warning threshold, a secondary warning can also be generated to remind the staff that the current capacity of the NAS volume has exceeded the secondary warning threshold.
[0096] S103. Determine whether the current capacity of the network - attached storage volume exceeds the primary warning threshold.
[0097] Among them, the primary warning threshold is set and changed in advance by technicians or relevant authorized staff, which is not limited here. And the primary warning threshold is greater than the secondary warning threshold.
[0098] Specifically, if it is determined that the current capacity of the network - attached storage volume does not exceed the primary warning threshold, step S104 is executed; if it is determined that the current capacity of the network - attached storage volume exceeds the primary warning threshold, step S105 is executed.
[0099] In the specific implementation process of the present application, when it is determined that the current capacity of the network - attached storage volume exceeds the primary warning threshold, a primary warning can also be generated to remind the staff that the current capacity of the NAS volume has exceeded the primary warning threshold.
[0100] S104. Scan each directory space in the network - attached storage volume to obtain a first scan result.
[0101] Specifically, scan each directory space in the network - attached storage volume to obtain a first scan result, and record the first scan result in the database.
[0102] S105. Scan each directory space in the network - attached storage volume to obtain a second scan result.
[0103] Specifically, scan each directory space in the network - attached storage volume to obtain a second scan result, and record the second scan result in the database.
[0104] Optionally, in another embodiment of the present application, after obtaining the second scan result, to achieve rapid expansion, two NAS volume storage resource pools with high performance and low performance for expansion are prepared in advance. When the capacity upper limit triggers the primary warning, the expansion program is automatically triggered, and according to the preset algorithm logic, automatic migration and expansion are performed, reducing the service interruption time and improving the service continuity. As Figure 2 shown, it includes:
[0105] S201. Obtain the size of each business component directory stored in the network - attached storage volume.
[0106] S202. For each business component, calculate the difference between the usage rate of the current network - attached storage file system and the proportion of the space where the business component is located.
[0107] S203. When the difference between the difference and the secondary alarm threshold is the smallest and lower than the secondary alarm threshold, select a service component for migration and expansion.
[0108] Optionally, in another embodiment of the present application, an implementation manner of step S203 is as Figure 3 shown and includes:
[0109] S301. Select new network attached storage volumes with different performances according to the existing labels in the service component directory.
[0110] S302. Establish the same directory structure on the new network attached storage volume as that on the network attached storage volume, and link the components on the directory network attached storage to the new network attached storage volume.
[0111] Among them, the link in this step is a symbolic link: also known as a soft link, which is a special type of file that contains a reference pointing to another file or directory in the form of an absolute path or a relative path. The operation of a symbolic link is transparent: a program that reads and writes to a symbolic link file will behave as if it is operating directly on the target file. Some programs that need to specially handle symbolic links (such as backup programs) may recognize and operate on them directly. A symbolic link file only contains a text string, which is interpreted by the operating system as a path pointing to another file or directory. It is an independent file, and its existence does not depend on the target file. If a symbolic link is deleted, the target file it points to is not affected. If the target file is moved, renamed, or deleted, any symbolic link pointing to it still exists, but they will point to a non-existent file.
[0112] S303. Create a new date directory link on the new network attached storage volume and link it back to the network attached storage volume.
[0113] The present application provides an adaptive algorithm for expansion through the soft link method, avoiding data migration and storage capacity reaching the limit. The expansion solution is unaware of the application, reducing the modification of the application. At the same time, it reduces the service interruption time and improves the service continuity.
[0114] Since the entire directory structure is concatenated through link files and there is no substantial change, the application system using the data can be completely unaware, and the entire data migration can be completed within a data retention period.
[0115] S106. Determine the abnormal directory list according to the first scan result and the second scan result.
[0116] It can be seen that compared with the single file system alarm, the present application adds a directory-level alarm for NAS, which can predict and quickly locate abnormal directories in NAS in advance.
[0117] Optionally, in another embodiment of the present application, an implementation manner of step S106 is as follows Figure 4 shown, including:
[0118] S401. For the current size of each directory in the second scan result, determine whether the current size of the directory exceeds the space size threshold of the directory.
[0119] Among them, the space size threshold of the directory is the product of the mean value statistically fixed within a quarter and a preset percentage. The preset percentage (such as 120%) is set and changed in advance by technicians or relevant authorized staff, and is not limited here.
[0120] The mean value refers to the sum of all data in a set of data divided by the number of data, which is a measure representing the central tendency of a set of data and is an index reflecting the central tendency of the data.
[0121] Specifically, if it is determined that the current size of the directory exceeds the space size threshold of the directory, then step S402 is executed; if it is determined that the current size of the directory does not exceed the space size threshold of the directory, then step S404 is executed.
[0122] S402. Determine whether the current date is a special business date of the business component.
[0123] Among them, the special business date (such as a holiday) is set and changed in advance by technicians or relevant authorized staff, and is not limited here.
[0124] Specifically, if it is determined that the current date is not a special business date of the business component, then step S403 is executed; if it is determined that the current date is a special business date of the business component, then step S404 is executed.
[0125] S403. Determine that the directory is an abnormal directory and store it in the abnormal directory list.
[0126] S404. Determine the percentage increase in the size of the directory when obtaining the second scan result compared to when obtaining the first scan result.
[0127] S405. Determine whether the percentage increase in size is greater than the percentage increase in the network attached storage volume.
[0128] Specifically, if it is determined that the percentage increase in size is greater than the percentage increase in the network attached storage volume, then step S403 is executed.
[0129] The data on the NAS is stored according to the business date. Through actual production tests, the compression ratio of structured text files can reach 1:7. Therefore, when the NAS space has a major warning, compressing the files of one day can reduce the space utilization rate of the target NAS by about 14%, greatly alleviating the NAS space usage situation. Moreover, when the NAS storage is designed, the write priority is set relatively high. Performing file compression processing is actually a read-write operation, and the NAS's response time to write operations is higher than that of delete operations. Therefore, for large files, using compression to reduce the NAS space utilization rate is much faster than deletion. Therefore, in another embodiment of the present application, after obtaining the list of abnormal directories, the abnormal directories in the list of abnormal directories can also be compressed to quickly and efficiently solve the problem of capacity emergency. The specific method is as Figure 5 shown, including the following steps:
[0130] S501. Obtain the number of CPU cores of the compression dedicated server, and determine the compression parameters of the multi-threaded compression program according to the number of CPU cores of the compression dedicated server.
[0131] Among them, the compression parameters at least include the number of threads of the multi-threaded compression program and the compression ratio.
[0132] It should be noted that, in order to relieve the space pressure as soon as possible, in the specific implementation process of the present application, compression can be achieved by means of, but not limited to, the pigz tool to implement multi-threaded compression.
[0133] S502. Call the multi-threaded compression program to sequentially obtain the abnormal directories in the list of abnormal directories according to the sizes of the abnormal directories in the list of abnormal directories, and perform emergency compression on the abnormal directories.
[0134] Optionally, in another embodiment of the present application, before performing emergency compression on the abnormal directories, it further includes:
[0135] Judge whether the abnormal directory has been subjected to emergency compression on the current day.
[0136] Specifically, if it is judged that the abnormal directory has been subjected to emergency compression on the current day, then no further emergency compression will be performed on the abnormal directory.
[0137] S503. After each abnormal directory is emergently compressed, judge whether the current capacity of the network attached storage volume exceeds the major warning threshold.
[0138] Specifically, if it is judged that the current capacity of the network attached storage volume exceeds the major warning threshold, then continue to execute the step of calling the multi-threaded compression program to sequentially obtain the abnormal directories in the list of abnormal directories according to the sizes of the abnormal directories in the list of abnormal directories, and perform emergency compression on the abnormal directories; if it is judged that the current capacity of the network attached storage volume does not exceed the major warning threshold, then execute step S504.
[0139] S504. The multi-threaded compression program stops obtaining the exception directories in the exception directory list, and after the obtained exception directories are urgently compressed, this compression is stopped.
[0140] For example: If the number of CPU cores of the compression dedicated server is 4 cores, then the number of threads of the multi-threaded compression program is 6. According to the sizes of the exception directories in the exception directory list, 4 exception directories in the exception directory list are obtained in sequence, namely exception directory 1 (the largest exception directory in the exception directory list), exception directory 2 (the second largest exception directory in the exception directory list), exception directory 3 (the third largest exception directory in the exception directory list), and exception directory 4 (the fourth largest exception directory in the exception directory list), and urgent compression is performed. After the urgent compression of exception directory 1 or exception directory 2 or exception directory 3 or exception directory 4 is completed, it is judged whether the current capacity of the network attached storage volume exceeds the main warning threshold. If it is judged that the current capacity of the network attached storage volume exceeds the main warning threshold, then among the sizes of the exception directories in the exception directory list, the next exception directory is selected (in this example, it is exception directory 5, which is the fifth largest exception directory in the exception directory list), and the urgent compression step is performed on exception directory 5; if it is judged that the current capacity of the network attached storage volume does not exceed the main warning threshold, then the multi-threaded compression program stops obtaining the exception directories in the exception directory list (that is, does not obtain exception directory 5), and after the obtained exception directories are urgently compressed (that is, the compression of exception directory 1, exception directory 2, exception directory 3, and exception directory 4 is completed), this compression is stopped.
[0141] It can be seen that the compression disposal algorithm of this application supports optional compression ratios and optional numbers of threads. When a capacity warning occurs, it can automatically and quickly reduce the storage space and reduce manual intervention.
[0142] As shown in Table 1, it is the script list used in the embodiment of this application.
[0143] Check script name Function Comment dirspace_check.py Single directory space scan Monitor the capacity of a single directory filesys_check.py Single NAS space scan Monitor the capacity of a single NAS exception_check.py Analysis and judgment of exception directories Analysis and judgment of exception directories compress_pigz.py Multi-threaded compression program Perform multi-threaded compression compress_register.py Compression result registration program Monitor the capacity of a single NAS dir_choose.py Migration directory selection program Select the directory to be migrated and the NAS to be used dir_change.py Expansion migration Expansion migration
[0144] Table 1
[0145] As can be seen from the above solution, the present application provides a method for monitoring the capacity of a distributed file system: monitoring the capacity of a network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume; determining whether the current capacity of the network attached storage volume exceeds a secondary warning threshold; if it is determined that the current capacity of the network attached storage volume exceeds the secondary warning threshold, determining whether the current capacity of the network attached storage volume exceeds a primary warning threshold; wherein, the primary warning threshold is greater than the secondary warning threshold; if it is determined that the current capacity of the network attached storage volume does not exceed the primary warning threshold, scanning each directory space in the network attached storage volume to obtain a first scan result, and returning to execute the step of monitoring the capacity of the network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume; if it is determined that the current capacity of the network attached storage volume exceeds the primary warning threshold, scanning each directory space in the network attached storage volume to obtain a second scan result; determining an abnormal directory list according to the first scan result and the second scan result. Thus, the capacity of the distributed file system can be efficiently monitored.
[0146] Optionally, in another embodiment of the present application, an implementation manner of a capacity monitoring device for a distributed file system is as Figure 6 shown, including:
[0147] A first monitoring unit 601, configured to monitor the capacity of a network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume.
[0148] A first judgment unit 602, configured to determine whether the current capacity of the network attached storage volume exceeds a secondary warning threshold.
[0149] A second judgment unit 603, configured to, if the first judgment unit 602 determines that the current capacity of the network attached storage volume exceeds the secondary warning threshold, determine whether the current capacity of the network attached storage volume exceeds a primary warning threshold.
[0150] Wherein, the primary warning threshold is greater than the secondary warning threshold.
[0151] A first scanning unit 604, configured to, if the second judgment unit 603 determines that the current capacity of the network attached storage volume does not exceed the primary warning threshold, scan each directory space in the network attached storage volume to obtain a first scan result, and activate the first monitoring unit 601 to execute monitoring the capacity of the network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume.
[0152] A second scanning unit 605, configured to, if the second judgment unit 603 determines that the current capacity of the network attached storage volume exceeds the primary warning threshold, scan each directory space in the network attached storage volume to obtain a second scan result.
[0153] Optionally, in another embodiment of the present application, an implementation manner of the capacity monitoring device of the distributed file system further includes:
[0154] A sixth determination unit, configured to determine whether the abnormal directory has been urgently compressed on the current day;
[0155] A second stop unit, configured to, if the sixth determination unit determines that the abnormal directory has been urgently compressed on the current day, no longer perform urgent compression on the abnormal directory.
[0156] Optionally, the capacity monitoring device of the distributed file system further includes:
[0157] A second acquisition unit, configured to acquire the size of each service component directory stored in the network attached storage volume.
[0158] A calculation unit, configured to, for each service component, calculate the difference between the current usage rate of the network attached storage file system and the space ratio where the service component is located.
[0159] A selection unit, configured to, when the difference between the difference value and the secondary alarm threshold is the smallest and lower than the secondary alarm threshold, select a service component for migration and expansion.
[0160] For the specific working process of the units disclosed in the above embodiments of the present application, reference may be made to the corresponding method embodiment content, as Figure 2 shown, which will not be elaborated here.
[0161] Optionally, in another embodiment of the present application, an implementation manner of the selection unit specifically includes:
[0162] A selection unit, configured to select new network attached storage volumes with different performances according to the existing labels of the service component directories.
[0163] A creation unit, configured to create the same directory structure as that on the network attached storage volume on the new network attached storage volume, and link the component on the directory network attached storage to the new network attached storage volume.
[0164] A linking unit, configured to create a date directory link on the new network attached storage volume and link it back to the network attached storage volume.
[0165] For the specific working process of the units disclosed in the above embodiments of the present application, reference may be made to the corresponding method embodiment content, as Figure 3 shown, which will not be elaborated here.
[0166] A determination unit 606, configured to determine an abnormal directory list according to the first scan result and the second scan result.
[0167] For the specific working process of the unit disclosed in the foregoing embodiments of the present application, reference may be made to the corresponding method embodiment content, such as Figure 1 as shown, which will not be elaborated here.
[0168] Optionally, in another embodiment of the present application, an implementation manner of the determining unit 606 specifically includes:
[0169] A third judgment unit, configured to judge whether the current size of each directory in the second scan result exceeds the space size threshold of the directory.
[0170] Wherein, the space size threshold of the directory is the product of the mean value fixedly counted within a quarter and a preset percentage.
[0171] A fourth judgment unit, configured to judge whether the current date is a special business date of the business component if the third judgment unit judges that the current size of the directory exceeds the space size threshold of the directory.
[0172] A first determination subunit, configured to determine that the directory is an abnormal directory and store it in the abnormal directory list if the fourth judgment unit judges that the current date is not a special business date of the business component.
[0173] For the specific working process of the unit disclosed in the foregoing embodiments of the present application, reference may be made to the corresponding method embodiment content, such as Figure 4 as shown, which will not be elaborated here.
[0174] Optionally, in another embodiment of the present application, an implementation manner of the capacity monitoring device of the distributed file system further includes:
[0175] A second determination subunit, configured to determine the size growth percentage of the directory when obtaining the second scan result compared to when obtaining the first scan result if the third judgment unit judges that the current size of the directory does not exceed the space size threshold of the directory or the fourth judgment unit judges that the current date is a special business date of the business component.
[0176] A fifth judgment unit, configured to judge whether the size growth percentage is greater than the growth percentage of the network attached storage volume.
[0177] A third determination subunit, configured to determine that the directory is an abnormal directory and store it in the abnormal directory list if the fifth judgment unit judges that the size growth percentage is greater than the growth percentage of the network attached storage volume.
[0178] For the specific working process of the unit disclosed in the foregoing embodiments of the present application, reference may be made to the corresponding method embodiment content, such as Figure 4 as shown, which will not be elaborated here.
[0179] Optionally, in another embodiment of the present application, an implementation manner of the capacity monitoring device of the distributed file system further includes:
[0180] A first acquisition unit, configured to acquire the number of CPU cores of the compression dedicated server, and determine the compression parameters of the multi-threaded compression program according to the number of CPU cores of the compression dedicated server.
[0181] Wherein, the compression parameters at least include the number of threads and the compression ratio of the multi-threaded compression program.
[0182] A calling unit, configured to call the multi-threaded compression program to sequentially acquire the abnormal directories in the abnormal directory list according to the sizes of the abnormal directories in the abnormal directory list, and perform emergency compression on the abnormal directories.
[0183] A second judgment unit, further configured to, after each abnormal directory is emergently compressed, judge whether the current capacity of the network attached storage volume exceeds the main alarm threshold.
[0184] An activation unit, configured to, if the second judgment unit determines that the current capacity of the network attached storage volume exceeds the main alarm threshold, continue to activate the calling unit to execute the multi-threaded compression program to sequentially acquire the abnormal directories in the abnormal directory list according to the sizes of the abnormal directories in the abnormal directory list, and perform emergency compression on the abnormal directories.
[0185] A first stop unit, configured to, if the second judgment unit determines that the current capacity of the network attached storage volume does not exceed the main alarm threshold, stop the multi-threaded compression program from acquiring the abnormal directories in the abnormal directory list, and stop this compression after the emergency compression of the acquired abnormal directories is completed.
[0186] For the specific working process of the unit disclosed in the above embodiment of the present application, reference may be made to the corresponding method embodiment content, as Figure 5 shown, which will not be elaborated here.
[0187] As can be seen from the above solution, the present application provides a capacity monitoring device for a distributed file system: The first monitoring unit 601 monitors the capacity of the network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume; the first judgment unit 602 judges whether the current capacity of the network attached storage volume exceeds the secondary alarm threshold; if the first judgment unit 602 judges that the current capacity of the network attached storage volume exceeds the secondary alarm threshold, the second judgment unit 603 judges whether the current capacity of the network attached storage volume exceeds the primary alarm threshold; wherein, the primary alarm threshold is greater than the secondary alarm threshold; if the second judgment unit 603 judges that the current capacity of the network attached storage volume does not exceed the primary alarm threshold, the first scanning unit 604 scans each directory space in the network attached storage volume to obtain a first scanning result, and activates the first monitoring unit 601 to execute monitoring of the capacity of the network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume; if the second judgment unit 603 judges that the current capacity of the network attached storage volume exceeds the primary alarm threshold, the second scanning unit 605 scans each directory space in the network attached storage volume to obtain a second scanning result; the determination unit 606 determines an abnormal directory list according to the first scanning result and the second scanning result. Thus, efficient monitoring of the capacity of the distributed file system is achieved.
[0188] Another embodiment of the present application provides an electronic device, as Figure 7 shown, including:
[0189] One or more processors 701.
[0190] A storage device 702, on which one or more programs are stored.
[0191] When the one or more programs are executed by the one or more processors 701, the one or more processors 701 are caused to implement the capacity monitoring method of the distributed file system as described in any one of the above embodiments.
[0192] Another embodiment of the present application provides a computer storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the capacity monitoring method of the distributed file system as described in any one of the above embodiments is implemented.
[0193] In the above embodiments disclosed in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the part of the module, program segment, or code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0194] In addition, the various functional modules in the various embodiments of the present disclosure may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part. If the functions are implemented in the form of software functional modules and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a live broadcast device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0195] Those skilled in the art can implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for monitoring the capacity of a distributed file system, characterized in that, Including: Monitoring the capacity of the network attached storage volume at preset time intervals to obtain the current capacity of the network attached storage volume; Determining whether the current capacity of the network attached storage volume exceeds a secondary warning threshold; If it is determined that the current capacity of the network attached storage volume exceeds the secondary warning threshold, determining whether the current capacity of the network attached storage volume exceeds a primary warning threshold; wherein, the primary warning threshold is greater than the secondary warning threshold; If it is determined that the current capacity of the network attached storage volume does not exceed the primary warning threshold, scanning each directory space in the network attached storage volume to obtain a first scan result, and returning to execute the step of monitoring the capacity of the network attached storage volume at preset time intervals to obtain the current capacity of the network attached storage volume; If it is determined that the current capacity of the network attached storage volume exceeds the primary warning threshold, scanning each directory space in the network attached storage volume to obtain a second scan result; According to the first scan result and the second scan result, determining a list of abnormal directories, including: for the current size of each directory in the second scan result, determining whether the current size of the directory exceeds the space size threshold of the directory; wherein, the space size threshold of the directory is the product of the mean value fixed and statistically counted within a quarter and a preset percentage; if it is determined that the current size of the directory exceeds the space size threshold of the directory, determining whether the current date is a special business date of the business component; if it is determined that the current date is not a special business date of the business component, determining the directory as an abnormal directory and storing it in the list of abnormal directories.
2. The capacity monitoring method according to claim 1, wherein, Also including: If it is determined that the current size of the directory does not exceed the space size threshold of the directory or it is determined that the current date is a special business date of the business component, determining the percentage increase in the size of the directory when obtaining the second scan result compared to when obtaining the first scan result; Determining whether the percentage increase in size is greater than the percentage increase in the network attached storage volume; If it is determined that the percentage increase in size is greater than the percentage increase in the network attached storage volume, determining the directory as an abnormal directory and storing it in the list of abnormal directories.
3. The capacity monitoring method according to claim 1, wherein After determining the list of abnormal directories according to the first scan result and the second scan result, further including: Obtaining the number of CPU cores of the compression dedicated server and determining the compression parameters of the multi-threaded compression program according to the number of CPU cores of the compression dedicated server; wherein, the compression parameters at least include the number of threads and the compression ratio of the multi-threaded compression program; Invoking the multi-threaded compression program to sequentially obtain the abnormal directories in the list of abnormal directories according to the sizes of the abnormal directories in the list of abnormal directories, and performing emergency compression on the abnormal directories; After completing the emergency compression of each abnormal directory, determining whether the current capacity of the network attached storage volume exceeds the primary warning threshold; If it is determined that the current capacity of the network attached storage volume exceeds the main warning threshold, continue to execute the step of calling the multi-threaded compression program to sequentially obtain the abnormal directories in the abnormal directory list according to the sizes of the abnormal directories in the abnormal directory list, and perform emergency compression on the abnormal directories; If it is determined that the current capacity of the network attached storage volume does not exceed the main warning threshold, the multi-threaded compression program stops obtaining the abnormal directories in the abnormal directory list, and stops this compression after the emergency compression of the obtained abnormal directories is completed.
4. The capacity monitoring method according to claim 3, wherein Before performing the emergency compression on the abnormal directories, it further includes: Judging whether the abnormal directories have been emergently compressed on the current day; If it is determined that the abnormal directories have been emergently compressed on the current day, no emergency compression is performed on the abnormal directories.
5. The capacity monitoring method according to claim 1, characterized in that After the step of scanning each directory space in the network attached storage volume to obtain a second scan result if it is determined that the current capacity of the network attached storage volume exceeds the main warning threshold, it further includes: Obtaining the sizes of each business component directory stored in the network attached storage volume; For each business component, calculating the difference between the current usage rate of the network attached storage file system and the space ratio where the business component is located; When the difference between the difference value and the secondary warning threshold is the smallest and lower than the secondary warning threshold, select the business component for migration and expansion.
6. The capacity monitoring method according to claim 5, characterized in that, The step of selecting the business component for migration and expansion includes: Selecting new network attached storage volumes with different performances according to the existing labels of the business component directories; Establishing the same directory structure on the new network attached storage volume as that on the network attached storage volume, and pointing the component link on the directory network attached storage to the new network attached storage volume; Creating a new date directory link on the new network attached storage volume and linking it back to the network attached storage volume.
7. A capacity monitoring device for a distributed file system, characterized in that, It includes: A first monitoring unit, configured to monitor the capacity status of the network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume; A first judgment unit, configured to judge whether the current capacity of the network attached storage volume exceeds the secondary warning threshold; A second judgment unit, configured to judge whether the current capacity of the network attached storage volume exceeds the main warning threshold if the first judgment unit determines that the current capacity of the network attached storage volume exceeds the secondary warning threshold; wherein, the main warning threshold is greater than the secondary warning threshold; A first scanning unit, configured to scan each directory space in the network attached storage volume to obtain a first scan result if the second judgment unit determines that the current capacity of the network attached storage volume does not exceed the main warning threshold, and activate the first monitoring unit to execute the monitoring of the capacity status of the network attached storage volume at preset intervals to obtain the current capacity of the network attached storage volume; A second scanning unit, configured to scan each directory space in the network attached storage volume to obtain a second scan result if the second judgment unit determines that the current capacity of the network attached storage volume exceeds the main warning threshold; A determination unit, configured to determine a list of abnormal directories according to the first scan result and the second scan result; The determination unit includes: a third judgment unit, a fourth judgment unit, and a first determination subunit; The third judgment unit is configured to, for the current size of each directory in the second scan result, judge whether the current size of the directory exceeds the space size threshold of the directory; wherein, the space size threshold of the directory is based on the product of the average value statistically fixed within a quarter and a preset percentage; The fourth judgment unit is configured to, if the third judgment unit determines that the current size of the directory exceeds the space size threshold of the directory, judge whether the current date is a special business date of the business component; The first determination subunit is configured to, if the fourth judgment unit determines that the current date is not a special business date of the business component, determine that the directory is an abnormal directory and store it in the list of abnormal directories.
8. An electronic device, characterized in that, Including: One or more processors; A storage device, on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the capacity monitoring method of the distributed file system according to any one of claims 1 to 6.
9. A computer storage medium, characterized in that, On which a computer program is stored, wherein when the computer program is executed by a processor, the capacity monitoring method of the distributed file system according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Cloud storage quota management method based on Linux kernel monitoring
CN105468989A