Index cleaning method, device and medium thereof

By adopting a balanced and unbalanced index cleaning mode in the ES cluster, index cleaning is carried out based on the disk size and usage of the ES node, which solves the problem that the index cleaning logic in the big data cluster cannot effectively utilize large disk space, and achieves more efficient disk resource utilization.

CN115421666BActive Publication Date: 2025-05-09HANGZHOU DBAPPSECURITY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211150123.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-05-09
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

In a big data cluster environment, the current index cleaning logic cannot effectively utilize large-capacity disk space, resulting in large-capacity disk space being underutilized.

Method used

By obtaining the disk size and usage of all ES nodes, we can determine whether the disk size meets the preset threshold and proportional value, and use balanced cleaning mode or unbalanced cleaning mode for index cleaning. The balanced cleanup mode cleans up when the disk usage of any ES node exceeds the limit, while the non-balanced cleanup mode cleans up only when the disk usage of all ES nodes exceeds the limit.

Benefits of technology

Effectively utilize large-capacity disk space, avoid resource waste caused by automatic index cleaning, and improve the utilization rate of disk storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115421666B_ABST
    Figure CN115421666B_ABST
Patent Text Reader

Abstract

The present application discloses an index cleaning method, device and medium thereof, which relate to the field of big data technology and are used for cleaning the index of an ES cluster. Aiming at the problem that the current index cleaning logic cannot make good use of large-capacity disk space when disks of different sizes are mixed, an index cleaning method is provided. By judging whether the disk sizes of each ES node satisfy the condition that the difference in size does not exceed a preset threshold value and the difference in percentage does not exceed a preset ratio value, the relative relationship between the sizes of the disks is judged; when the disk sizes are relatively slightly different, when the disk usage rate of any ES node exceeds a preset upper limit value, index cleaning is performed on each disk; and when the disk sizes are relatively significantly different, index cleaning is performed on each disk only when the disk usage rates of all ES nodes exceed the upper limit value, so as to ensure that the space of the large-capacity disk is fully utilized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data technology, and in particular to an index cleaning method, device and medium thereof. Background Art

[0002] As a common log analysis tool, ES (Elasticsearch, a data analysis engine) cluster is responsible for collecting system logs, application logs, security logs and other tasks. It generates a large number of indexes every day to facilitate users to retrieve the above logs. However, these indexes will take up disk space, and expired data also needs to be deleted to free up disk space. The current common data cleaning logic is that when the disk usage reaches the upper limit, the system starts the automatic index cleaning function until the disk usage is less than the upper limit or all indexes are deleted.

[0003] However, in a big data cluster environment, it is usually deployed in the form of multiple application container engines (docker) or virtual machines. The disk size of each ES node is not consistent. In the scenario where disks of different sizes are mixed, the current data cleaning logic will cause large-capacity disks to not be fully utilized. Large-capacity disks may still have a lot of space, but the total usage rate has reached the preset upper limit, and automatic index cleaning has to be performed.

[0004] Therefore, technicians in this field are in urgent need of an index cleaning method to solve the problem that the current index cleaning logic cannot make good use of large-capacity disk space when disks of different sizes are mixed. Summary of the invention

[0005] The purpose of the present application is to provide an index cleaning method, device and medium thereof to solve the problem that the current index cleaning logic cannot make good use of large-capacity disk space when disks of different sizes are mixed.

[0006] In order to solve the above technical problems, the present application provides an index cleaning method, comprising:

[0007] Get the disk size and disk usage of all ES nodes;

[0008] Determine whether the disk sizes of each ES node meet the requirement that the difference does not exceed the preset threshold and the difference percentage does not exceed the preset ratio value;

[0009] If all conditions are met, determine whether the disk usage of any ES node exceeds the preset upper limit. If so, perform index cleanup until the disk usage of all ES nodes is lower than the upper limit.

[0010] If not all of them are satisfied, determine whether the disk usage of all ES nodes exceeds the upper limit. If so, perform index cleanup until the disk usage of any ES node is lower than the upper limit.

[0011] Preferably, after determining whether the disk sizes of the ES nodes satisfy the requirement that the difference in size does not exceed a preset threshold and the difference in percentage does not exceed a preset ratio, the method further includes:

[0012] If all conditions are met, it is determined whether the disk usage of any ES node exceeds the preset usage warning value. If so, a warning message is returned;

[0013] If not all conditions are met, it is determined whether the disk usage of all ES nodes exceeds the usage warning value. If so, a warning message is returned;

[0014] Among them, the usage warning value is less than the usage upper limit value.

[0015] Preferably, index cleaning includes:

[0016] Clean up different types of indexes in turn according to the index type and preset cleanup priority.

[0017] Preferably, index cleaning includes:

[0018] Indexes of the same type are cleaned up in the order in which they were created.

[0019] Preferably, index cleaning includes:

[0020] Determine whether each index is expired based on the preset expiration configuration information, and clean up the expired indexes first;

[0021] If the disk usage of the ES node still exceeds the upper limit after the expired index is cleaned up, the non-expired index will be cleaned up.

[0022] Preferably, cleaning up unexpired indexes includes:

[0023] According to the preset protection conditions, only the indexes that do not meet the protection conditions are cleaned up; different index types correspond to different protection conditions.

[0024] Preferably, the expired configuration information is factory index closed configuration information.

[0025] In order to solve the above technical problems, the present application also provides an index cleaning device, comprising:

[0026] The disk size acquisition module is used to obtain the disk size and disk usage of all ES nodes;

[0027] The cleaning mode judgment module is used to judge whether the disk sizes of each ES node meet the requirements that the difference does not exceed the preset threshold and the difference percentage does not exceed the preset ratio value. If all the requirements are met, the balanced cleaning module is triggered; if not, the unbalanced cleaning module is triggered;

[0028] The balanced cleanup module is used to determine whether the disk usage of any ES node exceeds the preset usage upper limit. If so, index cleanup is performed until the disk usage of all ES nodes is lower than the usage upper limit.

[0029] The unbalanced cleanup module is used to determine whether the disk usage of all ES nodes exceeds the upper limit. If so, index cleanup is performed until the disk usage of any ES node is lower than the upper limit.

[0030] Preferably, it also includes:

[0031] The early warning module is used to determine whether the disk size of each ES node satisfies the conditions that the difference in size does not exceed the preset threshold and the difference in percentage does not exceed the preset ratio value. If all conditions are met, it is used to determine whether the disk usage of any ES node exceeds the preset usage warning value. If so, a warning message is returned; if not all conditions are met, it is used to determine whether the disk usage of all ES nodes exceeds the usage warning value. If so, a warning message is returned; wherein the usage warning value is less than the usage upper limit value.

[0032] In order to solve the above technical problems, the present application also provides an index cleaning device, comprising:

[0033] Memory for storing computer programs;

[0034] A processor is used to implement the steps of the above-mentioned index cleaning method when executing a computer program.

[0035] In order to solve the above technical problems, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the index cleaning method as described above are implemented.

[0036] The present application provides an index cleaning method, which determines the relative relationship between the sizes of the disks by determining whether the disk sizes of the ES nodes meet the requirements that the difference does not exceed the preset threshold value and the difference percentage does not exceed the preset ratio value; the balanced cleaning mode is adopted when the disk sizes are relatively small, and when the disk usage rate of any ES node exceeds the preset usage upper limit, the index of each disk is cleaned; and for the case where the disk sizes are relatively large, the unbalanced cleaning mode is adopted, and the index of each disk is cleaned only when the disk usage rate of all ES nodes exceeds the usage upper limit. The above method makes a judgment on the difference in disk size for the case of mixed installation of large and small disks, and adopts different index cleaning modes. When the disk sizes are not much different, if the usage rate of any disk exceeds the limit, it means that the usage rate of other disks is also likely to exceed the limit. At this time, the index of each disk is cleaned to free up disk space; and when the disk sizes are relatively large, the index is cleaned only when the usage rate of all disks exceeds the limit to ensure that the space of large-capacity disks is fully utilized. This method can make better use of disk resources and better meet the needs of actual use of ES clusters.

[0037] The index cleaning device and computer-readable storage medium provided in the present application correspond to the above method and have the same effect as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0039] Figure 1 A flowchart of an index cleaning method provided by the present invention;

[0040] Figure 2 A flowchart of another index cleaning method provided by the present invention;

[0041] Figure 3 A structural diagram of an index cleaning device provided by the present invention;

[0042] Figure 4 A structural diagram of another index cleaning device provided by the present invention. DETAILED DESCRIPTION

[0043] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0044] The core of this application is to provide an index cleaning method, device and medium thereof.

[0045] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0046] Currently, the index cleanup logic of the ES cluster is based on a pre-set cleanup threshold. When there is an ES node whose disk usage exceeds the cleanup threshold, all ES nodes are cleaned up to free up disk space.

[0047] However, in today's big data ES cluster application scenarios, the disk sizes of each ES node deployed in multiple docker or virtual machines are not necessarily the same, and there are even cases where the capacity of individual disks varies greatly. When a small disk is almost full, a large disk may still have a lot of unused space. In other words, the disk with a small capacity will meet the disk usage exceeding the preset cleanup threshold early, while the large disk is still a certain distance away from the cleanup threshold. At this time, cleaning up the large disk that has not reached the cleanup threshold is equivalent to wasting the remaining unused disk space, which is not conducive to the reasonable and full use of disk storage resources.

[0048] Based on the above problems, this embodiment provides an index cleaning method, such as Figure 1 As shown, including:

[0049] S11: Get the disk size and disk usage of all ES nodes.

[0050] S12: Determine whether the disk size of each ES node satisfies that the difference does not exceed a preset threshold and the difference percentage does not exceed a preset ratio value. If so, go to step S13; if not, go to step S14.

[0051] It should be noted that each ES node needs to compare the disk capacity with all other ES nodes. ES nodes that have been compared do not need to be compared repeatedly. That is, after ES node A is compared with ES node B, when ES node B compares with other ES nodes, it does not need to be compared with ES node A repeatedly to improve comparison efficiency.

[0052] As for the difference percentage, it can be based on the capacity of the ES node currently used as the comparison subject, or it can be based on the capacity of the ES node with smaller capacity among the two ES nodes being compared. Considering that when determining the difference percentage for two identical ES nodes, the difference percentage determined based on the ES node with smaller capacity as the benchmark is larger and more likely to exceed the preset ratio value, that is, if the difference percentage exceeds the preset ratio value, then the difference percentage determined based on the ES node with smaller capacity as the benchmark must exceed the preset ratio value. Therefore, in order to reduce the number of calculations, the capacity of the ES node with smaller capacity is often used as the benchmark value for determining the difference percentage.

[0053] For example, the disk capacity of ES node A is 10T, and the disk capacity of ES node B is 12T, then the difference in disk capacity between the two is 2T, and the percentage of difference is 20% based on 10T.

[0054] S13: Determine whether the disk usage of any ES node exceeds the preset upper limit. If so, perform index cleaning until the disk usage of all ES nodes is lower than the upper limit.

[0055] S14: Determine whether the disk usage of all ES nodes exceeds the upper limit. If so, perform index cleaning until the disk usage of any ES node is lower than the upper limit.

[0056] Through step S12, it can be determined whether the disk capacity of each ES node in the current ES cluster is very different. A comprehensive judgment is made from two aspects. On the one hand, the difference in disk capacity is compared with a preset threshold to determine whether the difference is too large. On the other hand, the difference in disk capacity ratio, that is, the difference in relative capacity of the two nodes is compared with a preset ratio value to determine whether the difference is too large.

[0057] When any condition of exceeding the preset threshold or the preset ratio is met, it is considered that the disk capacity in the ES cluster differs too much. At this time, go to step S14 for cleaning, which is called unbalanced cleaning mode. In this mode, it indicates that the disk capacities to be cleaned differ greatly, and the disk usage is uneven. When the usage of a small-capacity disk exceeds the limit, the large disk may not exceed the limit. Therefore, index cleaning is not performed until the disk usage of all ES nodes exceeds the limit, until any ES node does not exceed the limit.

[0058] When none of the above conditions are met, it is considered that the disk capacity in the ES cluster is slightly different. At this time, go to step S13 for cleaning, which is called balanced cleaning mode. This mode indicates that the disk capacity to be cleaned is slightly different, and the disk usage is balanced. When any disk exceeds the limit, other disks may also exceed the limit. Therefore, when the disk usage of any ES node exceeds the limit, all ES nodes are indexed and cleaned until the disk usage of all ES nodes is within the limit.

[0059] It is easy to understand that the above-mentioned preset threshold, preset ratio value and upper limit value are all adjustable, and can be freely configured by the user according to actual needs before performing the above-mentioned index cleaning method. For example, in a possible implementation, the preset threshold is 10T, the preset ratio value is 20%, and the upper limit value is 70%.

[0060] The present application provides an index cleaning method, by obtaining the disk capacity and disk usage of all ES nodes, first judging whether there is a disk pair with a large difference by the disk capacity, when there is a disk pair with a large difference, it means that when the small disk usage exceeds the limit, the large disk may not exceed the limit. At this time, in order to maximize the use of the large disk space, index cleaning is performed only when the disk usage of all ES nodes exceeds the limit, and when the disk usage of any ES node does not exceed the limit, the cleaning is stopped to ensure the full utilization of disk storage resources. When there is no disk pair with a large difference as mentioned above, it means that when any disk exceeds the limit, other disks may have exceeded the limit or will exceed the limit. At this time, in order to ensure the normal operation of the storage system, index cleaning is performed until the disk usage of all ES nodes does not exceed the limit. The above method not only ensures that the disk space of the ES cluster can be cleaned up in time, but also can maximize the use of the storage resources of large-capacity disks. Compared with the current index cleaning strategy, it is more in line with the needs of actual use and has a higher utilization rate of storage resources. At the same time, the above method for judging the difference between the disks is to judge from two angles, namely, the difference and the percentage, through a preset threshold and a preset ratio value, so as to ensure the accuracy of the judgment to the greatest extent.

[0061] In addition, this embodiment also provides a preferred solution, such as Figure 2 As shown, after the above step S12: determining whether the disk sizes of each ES node satisfy the condition that the difference in size does not exceed a preset threshold and the difference in percentage does not exceed a preset ratio value, it also includes:

[0062] If all conditions are met, go to step S15; if not all conditions are met, go to step S16.

[0063] S15: Determine whether the disk usage of any ES node exceeds the preset usage warning value. If so, return warning information.

[0064] S16: Determine whether the disk usage of all ES nodes exceeds the usage warning value. If so, return warning information.

[0065] Among them, the usage warning value is less than the usage upper limit value.

[0066] It is easy to understand that the preconditions of step S15 and step S13 are the same, both of which are aimed at the case where the disk sizes are not much different, that is, the early warning and index cleaning in the balanced cleaning mode are completed. The preconditions of step S16 and step S14 are the same, both of which are aimed at the case where the disk sizes are greatly different, that is, the early warning and index cleaning in the unbalanced cleaning mode are completed.

[0067] It should be noted that there is no order restriction between step S15 and step S13, and between step S16 and step S14. They have different judgment conditions, so the execution order is not restricted in this embodiment. However, considering that the use warning value is less than the use upper limit value, the warning is usually triggered before the index cleaning logic is triggered. Therefore, the judgment of the warning is generally performed first, that is, Figure 2 As shown, step S15 is before step S13, and step S16 is before step S14.

[0068] As for the warning information, in a possible implementation, the warning information includes: a system resource overlimit prompt and the current maximum disk usage of each ES node.

[0069] It should also be noted that this embodiment does not impose too many restrictions on the implementation form of returning warning information, and can be determined based on the actual ES cluster deployment, such as using indicator lights, buzzers, and display screens for sound and light alarms, or by sending push, email, and other information to user terminals held by operation and maintenance personnel, or to devices such as management platforms that are monitored for a long time for warnings.

[0070] A preferred solution provided by this embodiment sets a warning value for use in the corresponding cleanup mode, so that before triggering the index cleanup, an early warning of the disk usage of each ES node is first performed, and the operation and maintenance personnel are promptly prompted to investigate the reasons for the excessive disk usage of the ES node and handle the exception. On the one hand, the speed of resolving ES node storage resource failures is improved, and on the other hand, the operation and maintenance personnel can grasp the disk usage of the current ES node in a timely and detailed manner. In addition, the early warning solution provided by this embodiment is also based on the distinction of cleanup modes according to the disk size of each ES node in the above embodiment, so as to carry out targeted warnings, and can also solve the problem that the storage space of large-capacity disks cannot be fully utilized when the disk capacity of ES nodes varies greatly.

[0071] It is easy to know that in the actual index cleaning process, there are many different index types. Different index types have different importance and different storage cycles. Common index types include: flow, log, event, statistic, alarm, etc.

[0072] Therefore, this embodiment provides a preferred implementation scheme. In the above embodiment, index cleaning is specifically performed as follows:

[0073] Clean up different types of indexes in turn according to the index type and preset cleanup priority.

[0074] In a possible implementation, the cleaning priorities of index types are as follows from low to high: flow, log, event, statistic, alarm. The higher the priority, the more important the index of this type is, and the later it is cleaned. Cleaning different types of indexes in sequence means cleaning up the low-priority indexes first, and then cleaning up the higher-level index types. For example, after cleaning up the log-type index, clean up the event-type index.

[0075] Furthermore, for the same type of indexes, this embodiment also provides a preferred solution for cleaning, including:

[0076] Indexes of the same type are cleaned up in the order in which they were created.

[0077] That is, when cleaning a certain type of index, clean it up in the order of creation time.

[0078] A preferred solution provided by the present embodiment achieves targeted cleaning of indexes by dividing the cleaning priority according to the index type. Generally speaking, a more important index type is defined as a higher priority, that is, the index type that is cleaned later is cleaned, thereby avoiding important indexes from being cleaned by mistake and ensuring the retention of important indexes to the greatest extent. For indexes of the same type, they are cleaned according to the order of creation time. The earlier the index is created, the more likely it is to expire, so it is relatively less important and is cleaned first. The preferred solution provided by the present embodiment can achieve orderly index cleaning, taking into account the importance of each index to the greatest extent, giving priority to cleaning indexes with lower importance and earlier creation time, and providing a more preferred index cleaning logic, thereby better realizing the utilization of ES node storage resources.

[0079] In the actual deployment of ES cluster, as a log analysis tool, it is necessary to collect system logs, application logs, security logs, etc., which will generate a large number of indexes every day. Similarly, these indexes will expire. At present, the factory index shutdown configuration information will be set, as shown in Table 1 below, and the shutdown time will be set for different index types. When the index creation time of this type exceeds the corresponding shutdown time in the factory index shutdown configuration information, the index is closed and can no longer be used, which is equivalent to expiration.

[0080] Table 1 Factory Index Close Configuration Table

[0081] Index Type Expiration time log 180 days flow 30 days event 6 months statistic 6 months alarm 1 year

[0082] Therefore, this embodiment further provides a preferred implementation scheme based on the above embodiment, and the above index cleaning includes:

[0083] Determine whether each index is expired based on the preset expiration configuration information, and clean up the expired indexes first;

[0084] If the disk usage of the ES node still exceeds the upper limit after the expired index is cleaned up, the non-expired index will be cleaned up.

[0085] It should be noted that the above-mentioned expired configuration information can be set by the operation and maintenance personnel according to actual needs, or the configuration information can be directly closed through the above-mentioned factory index and used as expired configuration information without additional configuration. This embodiment does not limit this.

[0086] It is easy to know that the clearing of expired indexes and non-expired indexes can also be carried out in sequence according to the priority clearing mode set in the above embodiment, according to different index types and preset priorities. When the expired index is cleared, the disk usage of the current ES node is judged. When there are still ES nodes whose disk usage exceeds the upper limit, the non-expired index is cleared, otherwise the index clearing process is completed.

[0087] Similarly, for the judgment of whether to clean up the non-expired index, different judgment logics can be executed according to the above different cleaning modes, including:

[0088] When the index cleanup mode is balanced cleanup mode, if the disk usage of any ES node exceeds the upper limit, the non-expired index will be cleaned up.

[0089] When the index cleanup mode is unbalanced, the non-expired indexes are cleaned up when the disk usage of all ES nodes exceeds the upper limit.

[0090] Furthermore, on the basis of cleaning different types of indexes according to priority, when cleaning non-expired indexes, this embodiment re-performs the above steps after completing the cleaning of each index type. If the above conditions for cleaning non-expired indexes are not met, the cleaning is stopped to avoid excessive cleaning of non-expired indexes.

[0091] In addition, since the non-expired indexes are still in use, although the method of cleaning up according to the priority provided in the above embodiment can guarantee the retention of important indexes to the greatest extent, some types of indexes may be completely cleaned up, while some types of indexes are not cleaned up at all, resulting in the loss of some index types. Based on this problem, this embodiment also provides a preferred implementation scheme, and the above-mentioned cleaning up of non-expired indexes specifically includes:

[0092] According to the preset protection conditions, only the indexes that do not meet the protection conditions are cleaned up.

[0093] Different index types correspond to different protection conditions, which may be the number of protected items or the protection time.

[0094] For example, the protection condition set for the index type log is the number of protections, and the number of protections is 30. In this way, when cleaning the log type index, when there are 30 indexes left, no further cleaning will be performed (according to the above embodiment, the cleaning order can be based on the order of creation time), but the index type with the next priority will be cleaned.

[0095] Similarly, the protection condition set for the index type event is the protection days, which is two days. That is, according to the creation time of each event type index, only the event type index that does not meet the protection condition is cleaned up.

[0096] It should be noted that the above protection days can be based on the index cycle. For example, one day is an index cycle, and 0:00 every day is the start time of a cycle. Then the protection condition is two days, that is, the indexes created in the index cycle of today and yesterday are protected. Under this protection logic, for example, even if the current time is 13:00, the indexes created after 13:00 the day before yesterday are not protected. On the contrary, the protection days can be strictly based on time, that is, all indexes created within two days before the current time are protected. For example, if the current time is 13:00, all indexes created after 13:00 the day before yesterday are protected.

[0097] Different protection conditions can be set for different index types. The protection conditions are not limited to the above-mentioned number of days and number of entries. Operation and maintenance personnel can choose to set appropriate protection conditions according to actual needs.

[0098] The preferred solution provided by this embodiment is to judge whether each index is expired according to the index creation time and the preset expiration configuration information, thereby dividing the entire index cleaning process into two stages: expired index cleaning and non-expired index cleaning. Expired index cleaning is performed first. When the conditions for index cleaning are still met after the expired index cleaning is completed, the non-expired index is cleaned to protect the cleaning that is still in use from being deleted by mistake. At the same time, in order to prevent some index types from being completely cleaned and some index types from being not cleaned due to the priority cleaning of index types according to the above embodiment, this embodiment sets a protection condition for each index type, and only cleans the indexes that do not meet the protection condition, and protects the indexes that meet the protection condition, so that each index type can not be completely cleaned, and the occurrence of the problem of mistaken cleaning can be avoided to the greatest extent. In addition, by using the factory index closing configuration information that is usually set as the above-mentioned expiration configuration information, it can well meet the judgment of whether the index is expired, and no additional configuration is required, which reduces the workload required by the operation and maintenance personnel and better realizes the automation of index cleaning.

[0099] In the above embodiment, an index cleaning method is described in detail, and the present application also provides an embodiment corresponding to an index cleaning device. It should be noted that the present application describes the embodiment of the device part from two perspectives, one is based on the functional module perspective, and the other is based on the hardware perspective.

[0100] Based on the functional module perspective, such as Figure 3 As shown, this embodiment provides an index cleaning device, including:

[0101] The disk size acquisition module 21 is used to obtain the disk size and disk usage of all ES nodes;

[0102] The cleaning mode judgment module 22 is used to judge whether the disk sizes of each ES node meet the requirements that the difference does not exceed a preset threshold value and the difference percentage does not exceed a preset ratio value. If both are satisfied, the balanced cleaning module is triggered; if not, the unbalanced cleaning module is triggered;

[0103] The balance cleaning module 23 is used to determine whether the disk usage of any ES node exceeds the preset usage upper limit. If so, index cleaning is performed until the disk usage of all ES nodes is lower than the usage upper limit.

[0104] The unbalanced cleaning module 24 is used to determine whether the disk usage of all ES nodes exceeds the upper limit of usage. If so, index cleaning is performed until the disk usage of any ES node is lower than the upper limit of usage.

[0105] Preferably, it also includes:

[0106] The early warning module is used to determine whether the disk size of each ES node satisfies the conditions that the difference in size does not exceed the preset threshold and the difference in percentage does not exceed the preset ratio value. If all conditions are met, it is used to determine whether the disk usage of any ES node exceeds the preset usage warning value. If so, a warning message is returned; if not all conditions are met, it is used to determine whether the disk usage of all ES nodes exceeds the usage warning value. If so, a warning message is returned; wherein the usage warning value is less than the usage upper limit value.

[0107] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, which will not be repeated here.

[0108] The index cleaning device provided in this embodiment obtains the disk capacity and disk usage of all ES nodes through the disk size acquisition module, first determines whether there is a disk pair with a large difference according to the disk capacity through the cleaning mode judgment module. When there is a disk pair with a large difference, it means that when the small disk usage exceeds the limit, the large disk may not exceed the limit. At this time, in order to maximize the use of the large disk space, the non-balanced cleaning module is used to perform index cleaning only when the disk usage of all ES nodes exceeds the limit, and the cleaning is stopped when the disk usage of any ES node does not exceed the limit, so as to ensure the full utilization of disk storage resources. When there is no disk pair with a large difference, it means that when any disk exceeds the limit, other disks may have exceeded the limit or will exceed the limit. At this time, in order to ensure the normal operation of the storage system, the balanced cleaning module is used to perform index cleaning until the disk usage of all ES nodes does not exceed the limit. This device can not only ensure that the disk space of the ES cluster can be cleaned up in time, but also take into account the maximum use of the storage resources of large-capacity disks. Compared with the current index cleaning strategy, it is more in line with the needs of actual applications and has a higher utilization rate of storage resources. In addition, the above method for judging the difference between the disks is to judge from two perspectives, namely, the difference and the percentage, by using a preset threshold and a preset ratio value, so as to ensure the accuracy of the judgment to the greatest extent.

[0109] Figure 4 A structural diagram of an index cleaning device provided in another embodiment of the present application is shown in FIG. Figure 4 As shown, an index cleaning device includes: a memory 30 for storing a computer program;

[0110] The processor 31 is used to implement the steps of an index cleaning method in the above embodiment when executing a computer program.

[0111] The index cleaning device provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop computer, or a desktop computer.

[0112] Among them, the processor 31 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 31 can be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 31 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 31 may be integrated with a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 31 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.

[0113] The memory 30 may include one or more computer-readable storage media, which may be non-transitory. The memory 30 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 30 is at least used to store the following computer program 301, wherein, after the computer program is loaded and executed by the processor 31, it can implement the relevant steps of an index cleaning method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 30 may also include an operating system 302 and data 303, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 302 may include Windows, Unix, Linux, etc. Data 303 may include but is not limited to an index cleaning method, etc.

[0114] In some embodiments, an index cleaning device may further include a display screen 32 , an input / output interface 33 , a communication interface 34 , a power supply 35 , and a communication bus 36 .

[0115] Those skilled in the art will understand that Figure 4 The structure shown in does not constitute a limitation of an index cleaning device, and may include more or fewer components than shown in the figure.

[0116] An index cleaning device provided in an embodiment of the present application includes a memory and a processor. When the processor executes a program stored in the memory, it can implement the following method: an index cleaning method.

[0117] The index cleaning device provided in this embodiment executes a computer program stored in a memory through a processor to determine whether there are disk pairs with relatively large differences by obtaining the disk capacity and disk usage of all ES nodes. When there are disk pairs with relatively large differences, index cleaning is performed only when the disk usage of all ES nodes exceeds the limit, and cleaning is stopped when the disk usage of any ES node does not exceed the limit, thereby ensuring full utilization of disk storage resources. When there are no disk pairs with relatively large differences, index cleaning is performed when any ES node exceeds the limit until the disk usage of all ES nodes does not exceed the limit. While ensuring that the disk space of the ES cluster can be cleaned up in a timely manner, the storage resources of large-capacity disks can be utilized to the maximum extent, which is more in line with the needs of actual use and has a higher utilization rate of storage resources. At the same time, the above-mentioned method for judging the difference in the size of the disk is to judge from the two perspectives of difference and percentage through a preset threshold and a preset ratio value, thereby ensuring the accuracy of the judgment to the maximum extent.

[0118] Finally, the present application also provides an embodiment corresponding to a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps recorded in the above method embodiment are implemented.

[0119] It is understandable that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0120] A computer-readable storage medium provided by this embodiment, when the computer program stored therein is executed, can be implemented by obtaining the disk capacity and disk usage of all ES nodes to determine whether there are disk pairs with relatively large differences. When there are disk pairs with relatively large differences, index cleaning is performed only when the disk usage of all ES nodes exceeds the limit, and cleaning is stopped when the disk usage of any ES node does not exceed the limit, so as to ensure full utilization of disk storage resources. When there is no such disk pair with relatively large differences, index cleaning is performed when any ES node exceeds the limit until the disk usage of all ES nodes does not exceed the limit. While ensuring that the disk space of the ES cluster can be cleaned up in a timely manner, the storage resources of large-capacity disks can be utilized to the greatest extent, which is more in line with the needs of actual use and has a higher utilization rate of storage resources. At the same time, the above-mentioned method for judging the difference in the size of the disk is to judge from the two perspectives of difference and percentage through a preset threshold and a preset ratio value, so as to ensure the accuracy of the judgment to the greatest extent.

[0121] The above is a detailed introduction to an index cleaning method, device and medium provided by the present application. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referenced to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

[0122] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

Claims

1. An index cleaning method, characterized in that: include: Get the disk size and disk usage of all ES nodes; Determine whether the disk sizes of the ES nodes meet the requirement that the difference in size does not exceed a preset threshold and the difference in percentage does not exceed a preset ratio value; If all of them are satisfied, determine whether the disk usage of any of the ES nodes exceeds the preset usage upper limit. If so, perform index cleaning until the disk usage of all the ES nodes is lower than the usage upper limit. If not all of them are satisfied, determine whether the disk usage of all the ES nodes exceeds the usage upper limit. If so, perform index cleaning until the disk usage of any of the ES nodes is lower than the usage upper limit.

2. The index cleaning method according to claim 1, characterized in that: After determining whether the disk sizes of the ES nodes satisfy the condition that the difference in size does not exceed a preset threshold and the difference in percentage does not exceed a preset ratio value, the method further includes: If all conditions are met, it is determined whether the disk usage of any ES node exceeds the preset usage warning value. If so, a warning message is returned; If not all of them are satisfied, then determine whether the disk usage of all the ES nodes exceeds the usage warning value, and if so, return the warning information; Wherein, the usage warning value is less than the usage upper limit value.

3. The index cleaning method according to claim 1, characterized in that: The index cleaning includes: Clean up different types of indexes in turn according to the index type and preset cleanup priority.

4. The index cleaning method according to claim 3, characterized in that: The index cleaning includes: Indexes of the same type are cleaned up in the order in which they were created.

5. The index cleaning method according to any one of claims 1 to 4, characterized in that: The index cleaning includes: Determine whether each index is expired according to the preset expiration configuration information, and first clean up the expired index; If the disk usage of the ES node still exceeds the upper limit after the expired index is cleaned up, the non-expired index is cleaned up.

6. The index cleaning method according to claim 5, characterized in that: The cleaning of the unexpired indexes includes: According to the preset protection conditions, only the indexes that do not meet the protection conditions are cleaned up; wherein different index types correspond to different protection conditions.

7. The index cleaning method according to claim 5, characterized in that: The expired configuration information is factory index closed configuration information.

8. An index cleaning device, characterized in that: include: The disk size acquisition module is used to obtain the disk size and disk usage of all ES nodes; A cleaning mode judgment module is used to judge whether the disk sizes of the ES nodes meet the following conditions: the difference in size does not exceed a preset threshold, and the difference in percentage does not exceed a preset ratio value. If both conditions are met, a balanced cleaning module is triggered; if not, an unbalanced cleaning module is triggered; The balancing cleanup module is used to determine whether the disk usage of any of the ES nodes exceeds a preset usage upper limit. If so, index cleanup is performed until the disk usage of all the ES nodes is lower than the usage upper limit. The unbalanced cleaning module is used to determine whether the disk usage of all the ES nodes exceeds the upper limit of usage. If so, index cleaning is performed until the disk usage of any ES node is lower than the upper limit of usage.

9. An index cleaning device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the index cleaning method as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the index cleaning method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Dynamic balancing processing method and device for disk spaces and disk system

    CN102799395A

  • Storage space monitoring method and device, electronic terminal and storage medium

    CN109656885A