A method, device, equipment, and storage medium for automatically deleting a message queue cache

By obtaining the configuration and distribution information of topics in the Kafka cluster, generating a delete queue, and analyzing the message consumption status in real time, the problem of untimely data clearing in the Kafka cluster is solved, which improves resource utilization and reduces the pressure on storage resources.

CN113886106BActive Publication Date: 2025-07-22JINAN INSPUR DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111231230.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-07-22
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

After the data processing is completed, the Kafka cluster cannot clear the excess data in real time, resulting in high storage resource consumption and affecting the operation and maintenance problem of insufficient storage resources in the cluster.

Method used

By obtaining the configuration and distribution information of topics in the cluster, generating a deletion queue, and analyzing the consumption status of topic messages in real time, and issuing processing tasks in turn to realize real-time deletion of topic cached data and reducing storage resource pressure.

Benefits of technology

Real-time deletion of topic cached data is realized, improving cluster resource utilization rate, and reducing operation and maintenance problems of insufficient storage resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113886106B_ABST
    Figure CN113886106B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, equipment, and storage medium for automatically deleting message queue caches. It obtains the configuration and distribution information of topics in the cluster, distributes topic deletion requests according to the topic status, and generates a deletion queue; after receiving a topic deletion request, it parses the actual data blocks within the node, obtains the actual list to be deleted, and adds the list to the deletion queue; it sequentially obtains the deletion queue and deletes data blocks according to the list. By setting up a resident deletion process, it can parse the topic message consumption status in real time, issue processing tasks to each node, and through the parsing of data status within the node, it realizes the real-time deletion of topic cache data, reduces the pressure on cluster storage resources, improves the utilization rate of cluster resources, and reduces the operation and maintenance problems caused by insufficient cluster storage resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cluster data clearing mechanisms, and in particular to a method, device, equipment, and storage medium for automatically deleting message queue caches. Background Art

[0002] Kafka is a high-throughput distributed publish-subscribe message system that provides scalable, high-throughput, low-latency, and highly reliable message distribution services, and is widely used in big data fields such as log collection, monitoring data aggregation, streaming data processing, online and offline analysis, etc. Kafka components include Topic, Producer, broker, and Zookeeper; among them, Topic: Messages are classified according to Topic; Producer: The message sender; Consumer: The message receiver; broker: Each kafka instance (server); Zookeeper: Depends on the cluster to save meta information. Zookeeper is an independent component, and Kafka only uses it as a metadata storage backend.

[0003] As an important distributed message queue system in the big data ecosystem, Kafka components usually undertake a large number of data stream tasks, and the amount of data stored by itself is large. However, after the data is processed in the business system, it is usually no longer used, and the disk space will not be automatically released. The data clearing mechanism of Kafka itself only relies on time aging for batch deletion, and in the case of a large amount of data, it cannot clear redundant data in real time according to the message consumption status, resulting in a high consumption of cluster storage resources, and easily leading to the operation and maintenance problem of insufficient storage resources in the Kafka cluster, which in turn affects the upper-layer business system. Summary of the Invention

[0004] In view of the problems of the cluster data clearing mechanism, the present invention provides a method, device, equipment, and storage medium for automatically deleting message queue caches.

[0005] The technical solution of the present invention is as follows:

[0006] In the first aspect, the technical solution of the present invention provides a method for automatically deleting message queue caches, including the following steps:

[0007] Obtain the configuration and distribution information of topics in the cluster, distribute topic deletion requests according to the topic status, and generate a deletion queue;

[0008] After receiving the topic deletion request, parse the actual data blocks within the node, obtain the actual list to be deleted, and add the list to the deletion queue;

[0009] Obtain the deletion queue in sequence, and delete the data blocks according to the list.

[0010] By setting up a resident deletion process, the consumption status of topic messages is parsed in real time, and processing tasks are distributed to each node. Through the parsing of data status within the node, real-time deletion of topic cache data is achieved, reducing the pressure on cluster storage resources, improving the utilization rate of cluster resources, and reducing the operation and maintenance problems caused by insufficient cluster storage resources.

[0011] Further, after the steps of sequentially obtaining the deletion queue and deleting data blocks according to the list, the following steps are also included:

[0012] After the deletion is completed, update the topic status and synchronize the topic status with other nodes.

[0013] If all nodes have completed the deletion request, remove the deletion mark of the topic, release the restriction on the deletion of the topic, and allow the topic to join the deletion queue again.

[0014] Further, the steps of obtaining the configuration and distribution information of topics in the cluster, distributing topic deletion requests according to the topic status, and generating a deletion queue include:

[0015] Obtain the topic list in the cluster;

[0016] Obtain the number of replicas, replica distribution, and message consumption progress value of each topic in the topic list; among them, the replica distribution is the replica distribution node ID corresponding to each node;

[0017] Mark the topics to be deleted as deleted status;

[0018] Send topic deletion requests within the node according to each node ID;

[0019] Generate a deletion queue according to the sent topic deletion requests.

[0020] The resident data deletion thread continuously scans the data cache. Since the deletion requests are queued in the deletion queue waiting to be executed, when the topics marked for deletion are scanned again, they are directly skipped. In addition, the topics marked for deletion are prohibited from external access.

[0021] Further, after receiving the topic deletion request, the steps of parsing the actual data blocks within the node, obtaining the actual list to be deleted, and adding the list to be deleted to the deletion queue include:

[0022] After receiving the topic deletion request, obtain the storage path of the topic to be deleted;

[0023] Obtain the data paths of each partition of the topic within the node and obtain the index files and log data files in the data directories of each path;

[0024] Parse all data indexes and position fields below the message consumption progress value, and obtain all corresponding log data files according to the field parsing;

[0025] Put the same log data files in the obtained log data files into the list to be deleted; among them, the same log data file is the log data file that only contains data not greater than the message consumption progress value in the log data file;

[0026] Add the list to be deleted to the deletion queue.

[0027] Ensure that all progress values of the files in the list to be deleted are lower than the message consumption progress value. After traversing the partition data files under all data paths, all index files and log files to be deleted can be parsed.

[0028] Further, the specific steps of adding the list to be deleted to the deletion queue include:

[0029] Judge whether the list to be deleted is empty;

[0030] If the list to be deleted is empty, that is, there are no files to be deleted, skip and enter the deletion queue; execute the steps: sequentially obtain the deletion queue and delete data blocks according to the list;

[0031] If the list to be deleted is not empty, judge whether there is a deletion request for the same topic name in the deletion queue,

[0032] If so, compare whether the list to be deleted is the same. If it is the same, do not update the deletion queue and execute the steps: sequentially obtain the deletion queue and delete data blocks according to the list; if it is not the same, update the request in the deletion queue; execute the steps: sequentially obtain the deletion queue and delete data blocks according to the list;

[0033] If not, add a new topic deletion request to the deletion queue and add the list to be deleted to the deletion queue.

[0034] Further, to avoid the deletion request affecting the read and write performance of the Kafka service itself. The steps of sequentially obtaining the deletion queue and deleting data blocks according to the list include:

[0035] Judge whether the performance of disk I / O is greater than the set threshold;

[0036] If so, sequentially obtain the deletion queue and delete data blocks according to the list;

[0037] If not, wait for the set time and execute the steps again: judge whether the performance of disk I / O is greater than the set threshold.

[0038] Obtain the topic configuration distribution within the cluster, and distribute the topic deletion request according to the actual status of the topic. After receiving the topic deletion request, parse the actual data blocks within the node, obtain the actual list of data blocks to be deleted, and update the queue according to the condition judgment, waiting for subsequent actual deletion. Obtain the deletion queue in sequence, delete the data blocks according to the list, and update the topic status and synchronize the topic status with other nodes after completion.

[0039] In a second aspect, the technical solution of the present invention also provides a message queue cache automatic deletion device, including a topic distribution acquisition module, a replica data parsing module, and a node data deletion module;

[0040] The topic distribution acquisition module is used to obtain the configuration and distribution information of topics within the cluster, distribute the topic deletion request according to the topic status, and generate a deletion queue;

[0041] The replica data parsing module is used to parse the actual data blocks within the node after receiving the topic deletion request, obtain the actual list to be deleted, and add the list to be deleted to the deletion queue;

[0042] The node data deletion module is used to obtain the deletion queue in sequence and delete the data blocks according to the list.

[0043] Further, the device also includes a broadcast update module;

[0044] The broadcast update module is used to update the topic status and synchronize the topic status with other nodes after deletion is completed.

[0045] Further, the topic distribution acquisition module includes a topic information acquisition unit, a marking setting unit, and a request sending unit;

[0046] The topic information acquisition unit is used to obtain the topic list within the cluster; and obtain the number of replicas, replica distribution, and message consumption progress value of each topic in the topic list; among them, the replica distribution is the replica distribution node ID corresponding to each node;

[0047] The marking setting unit is used to mark the deletion status of the topic to be deleted;

[0048] The request sending unit is used to send the topic deletion request within the node according to each node ID; and generate a deletion queue according to the sent topic deletion request.

[0049] The resident data deletion thread continuously scans the data cache. Since the deletion request is queued in the deletion queue waiting to be executed, when the marked deleted topic is scanned again, it is directly skipped. In addition, the marked deleted topic is prohibited from external access.

[0050] Further, the replica data parsing module includes a data distribution acquisition unit, a parsing unit, and a processing unit;

[0051] A data distribution acquisition unit, configured to, after receiving a topic deletion request, acquire the storage path of the topic to be deleted; acquire the data paths of each partition of the topic within the node and acquire the index file and the log data file in the data directory of each path;

[0052] A parsing unit, configured to parse all data indexes and position fields lower than the message consumption progress value, and acquire all corresponding log data files according to the field parsing;

[0053] A processing unit, configured to put the same log data file in the acquired log data files into a deletion list; wherein, the same log data file is a log data file that only contains data not greater than the message consumption progress value in the log data file; and add the deletion list to a deletion queue.

[0054] Further, the processing unit is specifically configured to, when the deletion list is not empty, determine whether there is already a topic deletion request with the same name in the deletion queue. If so, compare whether the deletion lists are the same. If they are the same, do not update the deletion queue; if they are different, update the request of the deletion queue; if not, add a new topic deletion request to the deletion queue and add the deletion list to the deletion queue.

[0055] Further, the node data deletion module is specifically configured to, when it is determined that the performance of the disk IO is greater than a set threshold, sequentially acquire the deletion queue and delete data blocks according to the list.

[0056] In a third aspect, the technical solution of the present invention provides a computer device, including a processor and a memory. The processor and the memory complete communication with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute the message queue cache automatic deletion method as described in the first aspect by invoking the program instructions.

[0057] In a fourth aspect, the technical solution of the present invention further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the message queue cache automatic deletion method as described in the first aspect.

[0058] It can be seen from the above technical solutions that the present invention has the following advantages: real-time parsing of the topic message consumption status, issuing processing tasks in each node, and realizing real-time deletion of the topic cache data through data status parsing within the node, reducing the pressure on the cluster storage resources, improving the utilization rate of the cluster resources, and reducing the operation and maintenance problems of insufficient cluster storage resources.

[0059] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect.

[0060] Thus, compared with the prior art, the present invention has prominent substantive features and remarkable progress, and the beneficial effects of its implementation are also obvious. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0062] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention.

[0063] Figure 2 is a schematic block diagram of the device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0065] As Figure 1 shown, an embodiment of the present invention provides a method for automatically deleting a message queue cache, including the following steps:

[0066] Step 1: Obtain the configuration and distribution information of the topics in the cluster, distribute the topic deletion requests according to the topic status, and generate a deletion queue;

[0067] Step 2: After receiving the topic deletion request, parse the actual data blocks within the node, obtain the actual list to be deleted, and add the list to be deleted to the deletion queue;

[0068] Step 3: Obtain the deletion queue in sequence, and delete the data blocks according to the list.

[0069] By setting up a resident deletion process, parsing the topic message consumption status in real time, and issuing processing tasks to each node, and through parsing the data status within the node, real-time deletion of the topic cache data is realized, reducing the pressure on the cluster storage resources, improving the utilization rate of the cluster resources, and reducing the operation and maintenance problems of insufficient cluster storage resources.

[0070] Step 4: After the deletion is completed, update the topic status and synchronize the topic status with other nodes.

[0071] If all nodes have completed the deletion request, the deletion flag of the topic is removed, the restriction on deleting the topic is released, and the topic is allowed to be added to the deletion queue again.

[0072] In some embodiments, for the Kafka cluster, to ensure that data caches can be deleted in real time, a permanent data deletion thread needs to be created to perform data deletion in real time without affecting the Kafka service itself. The steps of obtaining the configuration and distribution information of topics in the cluster in Step 1, distributing topic deletion requests according to the topic status, and generating a deletion queue include:

[0073] Step 11: Obtain the list of topics in the cluster;

[0074] Step 12: Obtain the number of replicas, replica distribution, and message consumption progress value of each topic in the topic list; where the replica distribution is the replica distribution node ID corresponding to each node;

[0075] Obtain the number of replicas of each topic and the node information of its distribution through the Kafka service. The fields in the returned detailed topic configuration information include the number of replicas: Replications, replica distribution: Replicas, message consumption status: Offset. The information in the Replicas field is the replica distribution node ID corresponding to each Broker node, and Offset is the message consumption progress value;

[0076] Step 13: Mark the topic to be deleted as the deletion status;

[0077] Mark the topic to be deleted as the deletion status. Here, it can be implemented by setting variable values. The marked deleted topic cannot be added to the queue again and can only be added again after the previous deletion request is processed. That is, the permanent data deletion thread continuously scans the data cache. Since the deletion request is queued in the deletion queue waiting to be executed, when the marked deleted topic is scanned again, it is directly skipped. In addition, the marked deleted topic is prohibited from external access.

[0078] Step 14: Send the topic deletion request within the node according to each node ID;

[0079] It should be noted that deleting a topic deletes all replicas of the topic. Since the replicas exist on different nodes, a deletion request needs to be sent to each node where the topic replicas exist.

[0080] Step 15: Generate a deletion queue according to the sent topic deletion request.

[0081] In some embodiments, the steps of parsing the actual data blocks within the node, obtaining the actual list to be deleted, and adding the list to be deleted to the deletion queue after receiving the topic deletion request in Step 2 include:

[0082] Step 21: After receiving the topic deletion request, obtain the storage path of the topic to be deleted;

[0083] After each node receives the topic deletion request, it is necessary to obtain the detailed data block distribution information of the topic within the node. First, obtain the Kafka service storage path. The storage path is usually multiple paths, e.g., [ / kafka / data1, / kafka / data2, / kafka / data3, / kafka / data4].

[0084] Step 22: Obtain the data paths of each partition of the topic within the node and obtain the index files and log data files in the data directories of each path;

[0085] Topic data is usually distributed by partition to each data directory. Obtain the data distribution of each partition of the topic within the node, e.g., [ / kafka / data1 / topic-0, / kafka / data1 / topic-1, / kafka / data2 / topic-0, / kafka / data2 / topic-1, / kafka / data3 / topic-0, / kafka / data3 / topic-1, / kafka / data4 / topic-0, / kafka / data4 / topic-1]. Taking one of the topic partition data paths as an example, there are multiple xxxxxxxxxx.index and xxxxxxxxxxxx.log data files in the data directory. The log is the actual data file, and the index is the index file of the log data file for fast indexing of the log file. There is a corresponding relationship between the offset and position in the index file, e.g., [offset: 1, position: 2], [offset: 3, position: 5]. The actual relationship depends on the data within the topic. The position of each data is stored in the log, which is consistent with that in the index file index.

[0086] Step 23: Parse all data indexes and position fields below the message consumption progress value, and obtain all corresponding log data files according to the field parsing;

[0087] According to the message consumption progress value offset, parse all index data indexes and position fields below this value, and obtain all corresponding log files according to these fields.

[0088] Step 24: Place the same log data files in the obtained log data files into the list to be deleted; among them, the same log data file is the log data file that only contains data not greater than the message consumption progress value in the log data file;

[0089] It should be noted that there may be data in some of the obtained log files that do not need to be deleted. In this case, the entire log file will be retained. Only when all the data in the log file is data to be deleted, the entire log file will be deleted.

[0090] Step 25: Add the list to be deleted to the deletion queue.

[0091] Ensure that all the progress values of the files in the list to be deleted are lower than the message consumption progress value. After traversing the partition data files under all data paths, all the index files and log files to be deleted can be parsed.

[0092] It should be noted that the specific steps for adding the list to be deleted to the deletion queue in Step 25 include:

[0093] Step 251: Determine whether the list to be deleted is empty; if so, execute Step 252, if not, execute Step 253;

[0094] Step 252: Then skip and enter the deletion queue; execute Step 3;

[0095] Step 253: Determine whether there is already a deletion request for the same named topic in the deletion queue; if so, execute Step 254, if not, execute Step 256;

[0096] Step 254: Compare whether the list to be deleted is the same; if it is the same, do not update the deletion queue and execute Step 3; if it is not the same, execute Step 255;

[0097] Step 255: Update the request of the deletion queue; execute Step 3;

[0098] Step 256: Add a new topic deletion request to the deletion queue and add the list to be deleted to the deletion queue, and execute Step 3.

[0099] It should be noted that the newly created permanent data deletion thread executes the method steps of this application. Since the permanent data deletion thread scans the data cache in real time, after the topic to be deleted is marked and the main body deletion request is sent, due to the fact that the deletion process of the permanent data deletion thread takes a certain amount of time, the scanning process of the permanent data deletion thread may occasionally scan the topic that has been marked for deletion. At this time, as long as the topic deletion request has been sent for this topic, it will not be processed again until the mark is cleared after the deletion is completed and then it can be processed again. In addition, there are also topics for which the topic deletion request has not been initiated and are accessed or processed during the scanning process. In this case, the message consumption progress value will change, and the progress value is mainly updated when the topic is updated.

[0100] In some embodiments, in step 3, to avoid the deletion request affecting the read and write performance of the Kafka service itself. The steps of sequentially obtaining the deletion queue and deleting data blocks according to the list include:

[0101] Step 31: Determine whether the performance of the disk I / O is greater than a set threshold; if so, execute step 33, otherwise, execute step 32;

[0102] Step 32: Wait for a set time and then execute step 31 again;

[0103] Step 33: Sequentially obtain the deletion queue and delete data blocks according to the list.

[0104] Delete according to the data path and the corresponding path file. To avoid the deletion request affecting the read and write performance of the Kafka service itself, delete when the performance of the disk I / O is greater than the set threshold. In specific cases, it can be set that the thread processing only deletes when the disk I / O is low. If the disk I / O is high, the thread waits and does not execute. It will not perform file deletion until the I / O decreases. The level of the disk I / O can be judged by setting corresponding thresholds.

[0105] Obtain the topic configuration distribution in the cluster and distribute the topic deletion request according to the actual status of the topic. After receiving the topic deletion request, parse the actual data blocks within the node, obtain the actual list of data blocks to be deleted, and update the queue according to the condition judgment, and wait for subsequent actual deletion. Sequentially obtain the deletion queue and delete data blocks according to the list. After completion, update the topic status and synchronize the topic status with other nodes.

[0106] As Figure 2 shown, the embodiment of the present invention also provides a message queue cache automatic deletion device, including a topic distribution acquisition module, a replica data parsing module, a node data deletion module, and a broadcast update module;

[0107] A topic distribution acquisition module, which is used to acquire the configuration and distribution information of topics in the cluster, distribute topic deletion requests according to the topic status, and generate a deletion queue;

[0108] A replica data parsing module, which is used to parse the actual data blocks in the node after receiving a topic deletion request, obtain the actual list of data to be deleted, and add the list of data to be deleted to the deletion queue;

[0109] A node data deletion module, which is used to sequentially obtain the deletion queue and delete data blocks according to the list.

[0110] A broadcast update module, which is used to update the topic status and synchronize the topic status with other nodes after the deletion is completed.

[0111] In some embodiments, the topic distribution acquisition module includes a topic information acquisition unit, a mark setting unit, and a request distribution unit;

[0112] The topic information acquisition unit is used to acquire the topic list in the cluster; and acquire the number of replicas, replica distribution, and message consumption progress value of each topic in the topic list; wherein, the replica distribution is the replica distribution node ID corresponding to each node;

[0113] The mark setting unit is used to mark the to-be-deleted topic with a deletion status;

[0114] The request distribution unit is used to distribute topic deletion requests in the node according to each node ID; and generate a deletion queue according to the distributed topic deletion requests.

[0115] The resident data deletion thread continuously scans the data cache. Since the deletion requests are queued in the deletion queue waiting to be executed, when the marked-deleted topic is scanned again, it is directly skipped. In addition, the marked-deleted topic is prohibited from external access.

[0116] In some embodiments, the replica data parsing module includes a data distribution acquisition unit, a parsing unit, and a processing unit;

[0117] The data distribution acquisition unit is used to acquire the storage path of the to-be-deleted topic after receiving a topic deletion request; acquire the data paths of each partition of the topic in the node and acquire the index file and log data file in each path data directory;

[0118] The parsing unit is used to parse all data indexes and position fields lower than the message consumption progress value, and obtain all corresponding log data files according to the field parsing;

[0119] A processing unit, configured to place the same log data files in the obtained log data files into a to-be-deleted list; wherein, the same log data files are log data files that only contain data not greater than the message consumption progress value in the log data files; and add the to-be-deleted list to a deletion queue.

[0120] In some embodiments, the processing unit is specifically configured to, when the to-be-deleted list is not empty, determine whether there is already a deletion request for the same topic name in the deletion queue. If so, compare whether the to-be-deleted list is the same. If they are the same, do not update the deletion queue; if they are different, update the request in the deletion queue; if not, add a new topic deletion request to the deletion queue and add the to-be-deleted list to the deletion queue.

[0121] A node data deletion module, specifically configured to, when it is determined that the performance of disk I / O is greater than a set threshold, sequentially obtain the deletion queue and delete data blocks according to the list.

[0122] The node data deletion module sequentially takes out requests from the deletion queue and executes file deletion operations, deletes according to the data path and the corresponding path files. To avoid the impact of deletion requests on the read and write performance of the Kafka service itself, the thread processing only performs deletion when the disk I / O is low. If the disk I / O is high, the thread waits and does not execute. It waits until the I / O decreases and then performs file deletion. After completing all data file deletion requests in the list, broadcast the information that its deletion request has been processed to other nodes. After all nodes have completed the deletion requests, remove the topic data deletion mark, release the restriction on the deletion of this topic, and allow this topic to be added to the deletion queue again.

[0123] A computer device provided by an embodiment of the present invention, the device may include: a processor, a communication interface, a memory, and a bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the bus. The bus can be used for information transmission between the electronic device and the sensor. The processor can call the logical instructions in the memory to execute the following methods: Step 1: Obtain the configuration and distribution information of topics in the cluster, distribute topic deletion requests according to the topic status, and generate a deletion queue; Step 2: After receiving a topic deletion request, perform parsing of the actual data blocks within the node, obtain the actual to-be-deleted list, and add the to-be-deleted list to the deletion queue; Step 3: Sequentially obtain the deletion queue and delete data blocks according to the list. Step 4: After the deletion is completed, update the topic status and synchronize the topic status with other nodes.

[0124] In addition, when the logical instructions in the above-mentioned memory can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0125] An embodiment of the present invention provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and these computer instructions cause the computer to execute the method provided by the above method embodiment. For example, it includes: Step 1: Obtain the configuration and distribution information of topics in the cluster, distribute topic deletion requests according to the topic status, and generate a deletion queue; Step 2: After receiving a topic deletion request, parse the actual data blocks within the node, obtain the actual list to be deleted, and add the list to be deleted to the deletion queue; Step 3: Sequentially obtain the deletion queue and delete data blocks according to the list. Step 4: After the deletion is completed, update the topic status and synchronize the topic status with other nodes.

[0126] In some specific embodiments, the program instructions executed by the processor in the readable storage medium can specifically implement the following steps: Step 11: Obtain the topic list in the cluster; Step 12: Obtain the number of replicas, replica distribution, and message consumption progress value of each topic in the topic list; where the replica distribution is the replica distribution node ID corresponding to each node; Step 13: Mark the topic to be deleted with a deletion status; Step 14: Send topic deletion requests within the node according to each node ID; Step 15: Generate a deletion queue according to the sent topic deletion requests.

[0127] In some specific embodiments, the program instructions executed by the processor in the readable storage medium may specifically implement the following steps: Step 21: After receiving a topic deletion request, obtain the storage path of the topic to be deleted; Step 22: Obtain the data paths of each partition of the topic within the node and obtain the index files and log data files in the data directories of each path; Step 23: Parse all data indexes and position fields below the message consumption progress value, and obtain all corresponding log data files according to the field parsing; Step 24: Put the same log data file in the obtained log data files into the deletion list; wherein, the same log data file is a log data file that only contains data not greater than the message consumption progress value in the log data file; Step 25: Add the deletion list to the deletion queue.

[0128] In some specific embodiments, the program instructions executed by the processor in the readable storage medium may specifically implement the following steps: Step 31: Determine whether the performance of the disk IO is greater than a set threshold; if so, execute Step 33, otherwise, execute Step 32; Step 32: Wait for a set time and execute Step 31 again; Step 33: Obtain the deletion queue in sequence and delete data blocks according to the list.

[0129] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0130] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for automatically deleting a message queue cache, characterized in that It includes the following steps: Obtain the configuration and distribution information of topics in the cluster, distribute topic deletion requests according to the topic status, and generate a deletion queue; After receiving a topic deletion request, parse the actual data blocks within the node, obtain the actual list of items to be deleted, and add the list of items to be deleted to the deletion queue; Obtain the deletion queue in sequence and delete data blocks according to the list; The step of obtaining the configuration and distribution information of topics in the cluster, distributing topic deletion requests according to the topic status, and generating a deletion queue includes: Obtain the topic list in the cluster; Obtain the replication factor, replica distribution, and message consumption progress value of each topic in the topic list; among them, the replica distribution is the replica distribution node ID corresponding to each node; Mark the topics to be deleted with a deletion status; Send topic deletion requests within the node according to each node ID; Generate a deletion queue according to the sent topic deletion requests; The step of, after receiving a topic deletion request, parsing the actual data blocks within the node, obtaining the actual list of items to be deleted, and adding the list of items to be deleted to the deletion queue includes: After receiving a topic deletion request, obtain the storage path of the topic to be deleted; Obtain the data paths of each partition of the topic within the node and obtain the index files and log data files in the data directories of each path; Parse all data indexes and position fields lower than the message consumption progress value, and obtain all corresponding log data files according to the field parsing; Put the same log data file in the obtained log data files into the list of items to be deleted; among them, the same log data file is the log data file that only contains data not greater than the message consumption progress value in the log data file; Add the list of items to be deleted to the deletion queue.

2. The method for automatically deleting a message queue cache according to claim 1, wherein, After the step of obtaining the deletion queue in sequence and deleting data blocks according to the list, it further includes: After deletion is completed, update the topic status and synchronize the topic status with other nodes.

3. The method for automatically deleting a message queue cache according to claim 2, wherein The specific steps of adding the list of items to be deleted to the deletion queue include: Judge whether the list of items to be deleted is empty; If the list of items to be deleted is empty, that is, there are no files to be deleted, skip and enter the deletion queue; execute the step: obtain the deletion queue in sequence and delete data blocks according to the list; If the list of items to be deleted is not empty, judge whether there is already a topic deletion request with the same name in the deletion queue, If so, compare whether the list of items to be deleted is the same. If they are the same, do not update the deletion queue, and execute the step: obtain the deletion queue in sequence and delete data blocks according to the list; if they are different, update the request of the deletion queue; execute the step: obtain the deletion queue in sequence and delete data blocks according to the list; If not, add a new topic deletion request to the deletion queue and add the list of items to be deleted to the deletion queue.

4. The method for automatically deleting a message queue cache according to claim 3, wherein The step of obtaining the deletion queue in sequence and deleting data blocks according to the list includes: Judge whether the performance of disk I / O is greater than the set threshold; If so, obtain the deletion queue in sequence and delete data blocks according to the list; If not, wait for the set time and execute the step again: judge whether the performance of disk I / O is greater than the set threshold.

5. A message queue cache automatic deletion device, characterized in that It includes a topic distribution acquisition module, a replica data parsing module, and a node data deletion module; A topic distribution acquisition module, configured to acquire the configuration and distribution information of topics in a cluster, distribute topic deletion requests according to topic statuses, and generate a deletion queue; A replica data parsing module, configured to parse actual data blocks within a node after receiving a topic deletion request, obtain an actual list of data to be deleted, and add the list of data to be deleted to the deletion queue; A node data deletion module, configured to sequentially obtain the deletion queue and delete data blocks according to the list.

6. The message queue cache automatic deletion device according to claim 5, wherein The apparatus further includes a broadcast update module; The broadcast update module is configured to update the topic status and synchronize the topic status with other nodes after deletion is completed.

7. A computer device, characterized in that, It includes a processor and a memory, and the processor and the memory communicate with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute the message queue cache automatic deletion method described in any one of claims 1 to 4 by invoking the program instructions.

8. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the message queue cache automatic deletion method described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Message subscription processing device, system and method

    CN106657349A

  • A method and system for sequential consumption data

    CN109002484A