Distributed cache database processing method, device and equipment and computer readable storage medium

By adopting the distributed cache database processing method in Redis, data is read and sorted in batches, and data whose memory and access frequency exceeds the threshold are determined, the problem of not being able to effectively identify large keys and hot keys in the prior art is solved, and a solution of data skew and performance degradation is achieved.

CN120045591APending Publication Date: 2025-05-27CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510037795.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art cannot effectively determine the large key and hot key in Redis, resulting in problems such as data skew, performance degradation and call timeouts.

Method used

A distributed cache database processing method is provided. By receiving selection instructions, the data to be analyzed is determined from the distributed cache database, and the target algorithm is used to read and sort the data in batches to determine the first type of data whose memory exceeds the threshold and the second type of data whose access frequency exceeds the threshold.

Benefits of technology

Accurate identification of large and hot keys in Redis is achieved, avoiding data skew and performance degradation, and improving system stability and response speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045591A_ABST
    Figure CN120045591A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a distributed cache database processing method. The method comprises the steps that a selection instruction for data in a distributed cache database is received; based on the selection instruction, determining to-be-analyzed data from the distributed cache database; determining a first type of data and a second type of data from the to-be-analyzed data by adopting a target algorithm; wherein the target algorithm is an algorithm for reading data in batches and sorting the data under the condition that the read data meets a target condition; the first type of data is data of which the memory exceeds a first threshold value; the second type of data is data of which the data access frequency exceeds a second threshold value; the first type of data and the second type of data comprise various types of data. The embodiment of the invention further discloses a distributed cache database processing device and equipment and a computer readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, device, and computer-readable storage medium for processing a distributed cache database. Background Art

[0002] As the most popular cache database at present, Redis has been widely used in large application scenarios such as e-commerce and games. Due to the widespread use of Redis, the problems of large keys and hot keys in Redis have gradually emerged. If the problems of large keys and hot keys cannot be discovered and solved in time, it will lead to problems such as data skew, performance degradation, and call timeout in Redis; therefore, how to determine large keys and hot keys in Redis has become an urgent problem to be solved. Summary of the Invention

[0003] To solve the above technical problems, embodiments of this application are expected to provide a method, apparatus, device, and computer-readable storage medium for processing a distributed cache database, which solves the problem that large keys and hot keys in Redis cannot be determined in the prior art.

[0004] The technical solution of this application is implemented as follows:

[0005] A method for processing a distributed cache database, the method includes:

[0006] Receiving a selection instruction for data in the distributed cache database;

[0007] Based on the selection instruction, determining data to be analyzed from the distributed cache database;

[0008] Using a target algorithm to determine first-type data and second-type data from the data to be analyzed; wherein, the target algorithm is an algorithm that reads data in batches and sorts the data when the read data meets the target condition; the first-type data is data whose memory exceeds a first threshold; the second-type data is data whose access frequency exceeds a second threshold; the first-type data and the second-type data include multiple types of data.

[0009] In the above solution, the determining data to be analyzed from the distributed cache database based on the selection instruction includes:

[0010] When the selection instruction includes a shard selection instruction, determining a target shard corresponding to the shard selection instruction from the distributed cache database, and determining the data in the target shard as the data to be analyzed; wherein, the distributed cache database includes multiple shards;

[0011] When the selection instruction includes a database selection instruction, determine the data in the distributed cache database as the data to be analyzed.

[0012] In the above solution, the determining the first type of data and the second type of data from the data to be analyzed by using the target algorithm includes:

[0013] Obtain multiple first data to be screened from the data to be analyzed, and store the multiple first data to be screened into a target array;

[0014] Obtain multiple second data to be screened from the data to be analyzed, and store the multiple second data to be screened into the target array;

[0015] Obtain multiple third data to be screened from the data to be analyzed, and when the amount of data in the target array meets the target condition, determine the first type of data and the second type of data from the target array.

[0016] In the above solution, the determining the first type of data and the second type of data from the target array includes:

[0017] Classify the data to be screened in the target array based on the data type to obtain multiple types of data to be screened; wherein, the data to be screened includes the multiple first data to be screened, the multiple second data to be screened, and the multiple third data to be screened;

[0018] For each type of data to be screened, sort the data to be screened based on the memory size of the data to be screened, and determine the first type of data from the sorted data to be screened;

[0019] For each type of data to be screened, sort the data to be screened based on the access frequency of the data to be screened, and determine the second type of data from the sorted data to be screened.

[0020] In the above solution, after determining the first type of data and the second type of data from the target array, the method further includes:

[0021] Perform an association analysis on the first type of data and the second type of data to determine whether there is target data in the first type of data in the second type of data;

[0022] When there is target data in the second type of data, determine the target identifier of the target data; wherein, the target identifier indicates that the target data is both the first type of data and the second type of data.

[0023] In the above solution, after determining the target identifier of the target data, the method further includes:

[0024] Store the first description information of each first piece of data in the first type of data into the target database; wherein, the first description information includes a first index, a first name, a first type, a memory size, and a target identifier;

[0025] Store the second description information of each second piece of data in the second type of data into the target database; wherein, the second description information includes a second index, a second name, a second type, and an access frequency.

[0026] In the above solution, the distributed cache database processing method further includes:

[0027] Receive a visualization instruction for the target database;

[0028] Based on the visualization instruction, perform visualization processing on the first description information in the target database to obtain first visualization information, and perform visualization processing on the second description information in the target database to obtain second visualization information.

[0029] A distributed cache database processing device, the device includes:

[0030] A receiving unit, configured to receive a selection instruction for data in the distributed cache database;

[0031] A determining unit, configured to determine data to be analyzed from the distributed cache database based on the selection instruction;

[0032] A processing unit, configured to determine a first type of data and a second type of data from the data to be analyzed by using a target algorithm; wherein, the target algorithm is an algorithm for reading data in batches and sorting when the read data meets a target condition; the first type of data is data whose memory exceeds a first threshold; the second type of data is data whose access frequency exceeds a second threshold; the first type of data and the second type of data include multiple types of data.

[0033] A distributed cache database processing device, the device includes: a processor, a memory, and a communication bus;

[0034] The communication bus is used to implement a communication connection between the processor and the memory;

[0035] The processor is configured to execute a distributed cache database processing program in the memory to implement the steps of the above distributed cache database processing method.

[0036] A computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the above-mentioned distributed cache database processing method.

[0037] The distributed cache database processing method, device, equipment and computer-readable storage medium provided by the embodiments of the present application first receive a selection instruction for data in the distributed cache database; then, based on the selection instruction, determine the data to be analyzed from the distributed cache database; and then use a target algorithm to determine the first type of data and the second type of data from the data to be analyzed, and the target algorithm is an algorithm that reads data in batches and sorts the read data when the read data meets the target conditions; the first type of data is data whose memory exceeds the first threshold; the second type of data is data whose access frequency exceeds the second threshold; the first type of data and the second type of data include multiple types of data. In this way, by using the batch sorting comparison algorithm, the read data that meets the target conditions is sorted to obtain the first type of data including different types and the second type of data including different types. That is to say, both the first type of data in the data to be analyzed and the second type of data in the data to be analyzed are obtained, rather than only being able to obtain the first type of data as in the related art, thus solving problems such as data skew, performance degradation, and access timeout caused by the first type of data and the second type of data. Description of the Drawings

[0038] Figure 1 It is a schematic flowchart of a distributed cache database processing method provided by an embodiment of the present application;

[0039] Figure 2 It is a schematic diagram of module interaction in a distributed cache database processing method provided by an embodiment of the present application;

[0040] Figure 3 It is a schematic flowchart of another distributed cache database processing method provided by an embodiment of the present application;

[0041] Figure 4 It is a Redis cluster architecture diagram in a distributed cache database processing method provided by an embodiment of the present application;

[0042] Figure 5 It is a schematic diagram of a backup module in a distributed cache database processing method provided by an embodiment of the present application;

[0043] Figure 6 It is a schematic diagram of a data processing module in a distributed cache database processing method provided by an embodiment of the present application;

[0044] Figure 7Schematic diagram of a storage module in a distributed cache database processing method provided by an embodiment of the present application;

[0045] FIG. 8(a) is a schematic diagram of a kind of visualization information in a distributed cache database processing method provided by an embodiment of the present application;

[0046] FIG. 8(b) is a schematic diagram of another kind of visualization information in a distributed cache database processing method provided by an embodiment of the present application;

[0047] Figure 9 Schematic structural diagram of a distributed cache database processing device provided by an embodiment of the present application;

[0048] Figure 10 Schematic structural diagram of a distributed cache database processing device provided by an embodiment of the present application. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.

[0050] It should be understood that "the embodiments of the present application" or "the foregoing embodiments" mentioned throughout the specification means that specific features, structures, or characteristics related to the embodiments are included in at least one embodiment of the present application. Therefore, the appearances of "in the embodiments of the present application" or "in the foregoing embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.

[0051] Without special instructions, when an electronic device executes any step in the embodiments of the present application, it may be the processor of the electronic device that executes the step. It is also worth noting that the embodiments of the present application do not limit the order of execution of the following steps by the electronic device. In addition, the ways of processing data in different embodiments may be the same method or different methods. It should also be noted that any step in the embodiments of the present application can be independently executed by the electronic device, that is, when the electronic device executes any step in the following embodiments, it may not depend on the execution of other steps.

[0052] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] It should be noted that Redis, as the most popular cache database at present, has been widely used in large application scenarios such as e-commerce and games. Due to the wide use of Redis, the problems of large keys and hot keys in Redis have gradually emerged. If not discovered and solved in time, it will lead to problems such as Redis data skew, performance degradation, and call timeout.

[0054] Currently, the native method of Redis can obtain the largest key in Redis by executing the redis-cli - bigkeys command, but only a limited number of large keys can be displayed. When there are multiple large keys in production, this method is not suitable for large key analysis. Currently, native Redis only provides commands for large keys and does not have corresponding commands for analyzing hot keys. In addition, the Redis scan command can traverse all keys in the cache by setting parameters to obtain information about all keys. However, this command accesses Redis in an online manner, which will block normal access. Seriously, it will cause problems such as Redis connection timeout, and this solution cannot determine whether a key is a large key or a hot key.

[0055] However, the above existing technical solutions have the following disadvantages: (1) For the native Redis - bigkeys command method, during the analysis of large keys, only the largest key of each type is displayed, but the entire Redis needs to be traversed, which will cause resource waste and there is no hot key query command; (2) When analyzing through the Redis scan command, different scan parameters need to be configured according to different Redis specifications. And after testing, although the scan is a non - blocking command, during the scan process, it will indeed affect the stability of Redis itself, reduce the Queries Per Second (QPS) of the current Redis, with a significant performance reduction, and the analysis process is relatively slow. Moreover, when the database is very large, all scanned results need to be stored, which is likely to cause storage waste and program crashes; and neither of the above two methods can distinguish large keys or hot keys; (3) The existing solutions using the scan command and the native Redis - bigkeys command both take a relatively long time, especially the native solution. In summary, the current existing solutions cannot achieve real - time performance, which is also the traditional drawback of key analysis.

[0056] Based on this, the embodiments of this application provide a method for processing a distributed cache database. This method can be applied to a distributed cache database processing device. Referring to Figure 1 as shown, this method includes the following steps:

[0057] Step 101, receive a selection instruction for the data in the distributed cache database.

[0058] In the embodiment of the present application, the distributed cache database may specifically refer to, for example, Figure 2 the Redis cluster database deployed on the cloud as shown, and Redis is an open-source, network-supported, memory-based, distributed, optionally persistent key-value storage database written in NASI C; the selection instruction may be sent by the user and includes a shard selection instruction and a database selection instruction, and the selection instruction is specifically used for the user to select data from the selected distributed cache database according to their own will; the distributed cache database processing device can receive the selection instruction sent by the user in real time.

[0059] Step 102: Based on the selection instruction, determine the data to be analyzed from the distributed cache database.

[0060] In the embodiment of the present application, the distributed cache database may correspond to multiple shards, and the data to be analyzed may refer to the backup files (i.e., relational database, rdb) of the data in each shard. Therefore, the data to be analyzed is offline data; the data to be analyzed may refer to the data selected from the backup files corresponding to the distributed cache database according to the user's selection; when the selection instruction includes a shard selection instruction, the data to be analyzed can be determined from a certain shard corresponding to the distributed cache database according to the shard selection instruction; when the selection instruction includes a database selection instruction, the data to be analyzed can be determined from the entire distributed cache database according to the database selection instruction.

[0061] Step 103: Use the target algorithm to determine the first type of data and the second type of data from the data to be analyzed.

[0062] Among them, the target algorithm is an algorithm that reads data in batches and sorts the data when the read data meets the target conditions; the first type of data is the data whose memory exceeds the first threshold; the second type of data is the data whose access frequency exceeds the second threshold; the first type of data and the second type of data include multiple types of data.

[0063] In the embodiments of the present application, the first type of data may refer to data whose memory exceeds a first threshold, and specifically, the first type of data may refer to large key data; the second type of data may refer to data whose access frequency exceeds a second threshold, and specifically, the second type of data may refer to hot key data; the first threshold and the second threshold may be data determined based on historical experimental data; the multiple types of data may include data of the string (STRING) type, data of the list (LIST) type, data of the sorted set (ZSET) type, and data of the hash (HASH) type. Determining the first type of data and the second type of data from the data to be analyzed using the target algorithm may be to read the data in batches from the data to be analyzed, and determine the first type of data and the second type of data from the read data when the read data meets the target conditions, which can avoid the problem of program crash caused by excessive memory occupation during the subsequent parsing of the data to be analyzed.

[0064] In a feasible implementation manner, that the memory of the data exceeds the first threshold may specifically be: for a data of the STRING type, its value is 5MB (i.e., the data is too large); for a data of the LIST type, the number of its lists is 20,000 (i.e., the number of lists is too large); for a data of the ZSET type, the number of its members is 10,000 (i.e., the number of members is too large); for a data of the HASH type, although the number of its members is only 1,000, the total size of the values of these members is 100MB (i.e., the volume of the members is too large); the second type of data may specifically be: when the QPS reaches a certain value, the access to a certain data accounts for 80% of the entire QPS, that is, the access frequency exceeds the second threshold.

[0065] The distributed cache database processing method provided by the embodiments of the present application uses a batch sorting and comparison algorithm to sort the data that meets the target conditions read, and obtains the first type of data including different types and the second type of data including different types. That is to say, both the first type of data in the data to be analyzed and the second type of data in the data to be analyzed are analyzed, rather than only being able to obtain the first type of data as in the related art, thereby solving problems such as data skew, performance degradation, and access timeout caused by the first type of data and the second type of data.

[0066] Based on the foregoing embodiments, the embodiments of the present application provide another distributed cache database processing method. Referring to Figure 3 as shown, the method includes the following steps:

[0067] Step 201, the distributed cache database processing device receives a selection instruction for the data in the distributed cache database.

[0068] Step 202: When the selection instruction includes a shard selection instruction, the distributed cache database processing device determines the target shard corresponding to the shard selection instruction from the distributed cache database, and determines the data in the target shard as the data to be analyzed.

[0069] Among them, the distributed cache database includes multiple shards.

[0070] In the embodiment of the present application, the distributed cache database includes a shard layer, and the shard layer is the place where the entire Redis cluster finally stores data and executes commands. It exists in the form of nodes (pods) in the containerized application management platform (kubernetes, k8s), and the shard layer includes multiple shards; the shard selection instruction corresponds to shard-level data selection, and the target shard refers to a shard selected from multiple shards based on the shard selection instruction; as Figure 2 shown, when receiving the shard selection instruction sent by the user, only the data in the target shard corresponding to the shard selection instruction needs to be backed up (i.e., the backup file (rdb)) to obtain the data to be analyzed, and then the latest offline data in this target shard is analyzed offline. In this way, the efficiency of subsequent analysis is greatly improved, and it is very useful in real scenarios; it should be noted that each shard corresponds to a backup file, that is, each shard gets a backup file.

[0071] It should be noted that, as Figure 4 shown, in addition to the above-mentioned shard layer, the distributed cache database also includes a service layer and a proxy layer; among them, the service layer is the entrance of the entire Redis cluster, and a unified entrance is exposed through the nodeport method using containerization technology (i.e., k8s) for invocation; the proxy layer is also called the Redis cluster computing layer, and k8s maps each request to the corresponding proxy through load balancing (such as Figure 4 shown as proxy-0, proxy-1, proxy-2...), and then the proxy calculates the slot (converts the key to a 16-bit integer using the CRC16 algorithm according to the key value of the operation, and then performs a modulo (mod) operation on the key according to the slot numbers of all shards obtained in advance) to obtain which shard the current key belongs to, and then maps the request to the corresponding shard and executes; in addition, the number of replicas of the proxy (proxy-N) and the number of shards are in a one-to-one relationship.

[0072] Step 203: When the selection instruction includes a database selection instruction, the distributed cache database processing device determines the data in the distributed cache database as the data to be analyzed.

[0073] In the embodiment of the present application, the database selection instruction is the cluster selection instruction, and the cluster selection instruction corresponds to cluster-level data selection; as Figure 2As shown, when a cluster selection instruction sent by a user is received, all data in the Redis cluster can be backed up to obtain the latest data (i.e., the data to be analyzed) of the Redis cluster and perform offline analysis. It should be noted that this application supports two types of selections: shard-level selection and cluster-level selection. In this way, the corresponding data to be analyzed can be determined and analyzed according to the user's selection, which not only solves practical problems but also improves the analysis efficiency. Moreover, through offline analysis, the stability of Redis is not affected at all.

[0074] It should be noted that as Figure 5 shown, after obtaining the backup file, the backup module can mount the backup file (rdb) of each shard to the fixed object storage directory of the corresponding instance through the pod mounting method to achieve data persistence. The process of data persistence is as Figure 5 shown, and during the entire backup process: manually trigger the backup or trigger the backup through the Redis backup policy, save the data in the Redis memory, and map the saved backup file (rdb) to the directory corresponding to this shard in the object storage.

[0075] It should be noted that after obtaining the data to be analyzed, the data to be analyzed can be processed through the data processing module as Figure 2 shown, and specifically, it can be analyzed through the target algorithm as Figure 6 shown. Specifically, it is as follows:

[0076] Step 204, the distributed cache database processing device obtains multiple first data to be screened from the data to be analyzed and stores the multiple first data to be screened into the target array.

[0077] Step 205, the distributed cache database processing device obtains multiple second data to be screened from the data to be analyzed and stores the multiple second data to be screened into the target array.

[0078] Step 206, the distributed cache database processing device obtains multiple third data to be screened from the data to be analyzed. Until the amount of data in the target array meets the target condition, the first type of data and the second type of data are determined from the target array.

[0079] In the embodiments of the present application, the first data to be screened may refer to the first batch of data obtained from the data to be analyzed using a target algorithm; the second data to be screened may refer to the second batch of data obtained from the data to be analyzed using the target algorithm; the third data to be screened may refer to the third batch of data obtained from the data to be analyzed using the target algorithm; that the data volume in the target array meets the target condition may refer to that the data volume is greater than or equal to a preset threshold; the target array may refer to the array corresponding to the GO parsing tool, and when the data volume in the array is greater than or equal to the preset threshold, the target array may start to perform a sort operation, so as to determine the first type of data and the second type of data from the target array; specifically, using the target algorithm, the first type of data and the second type of data are determined from the target array based on the data type. In a feasible implementation manner, the preset threshold may be set to 2000, that is, a sort operation is performed every 2000 pieces of data to determine a large key data and a hot key data, so that the first type of data and the second type of data can be determined periodically.

[0080] It should be noted that step 206 can be implemented in the following manner:

[0081] Step 206A: The distributed cache database processing device classifies the data to be screened in the target array based on the data type to obtain multiple types of data to be screened.

[0082] Among them, the data to be screened includes multiple first data to be screened, multiple second data to be screened, and multiple third data to be screened.

[0083] In the embodiments of the present application, the data type may include four types: String, Hash, Set, and Zest. Then, the multiple types of data to be screened may include data of the String type, data of the Hash type, data of the Set type, and data of the Zest type; the multiple first data to be screened, multiple second data to be screened, and multiple third data to be screened obtained can be classified according to these four data types to obtain data of the String type, data of the Hash type, data of the Set type, and data of the Zest type.

[0084] Step 206B: For each type of data to be screened, the distributed cache database processing device sorts the data to be screened based on the memory size of the data to be screened, and determines the first type of data from the sorted data to be screened.

[0085] In an embodiment of the present application, after obtaining each type of data to be screened, for each type of data to be screened, the data to be screened can be sorted in ascending or descending order based on the memory size of each data to be screened. Then, according to the order of memory from large to small or from small to large, a target number of data to be screened are obtained from the sorted data to be screened. Then, the first type of data is composed of the target number of data to be screened corresponding to each type of data to be screened. That is to say, the first type of data includes data of multiple data types, which can reflect the situation of large keys of each type in the cache, thereby solving the data skew problem.

[0086] In a feasible implementation, the target number can be set to 20, that is, the first 20 data are obtained from the data to be screened of the String type, the first 20 data are obtained from the data to be screened of the Hash type, the first 20 data are obtained from the data to be screened of the Set type, and the first 20 data are obtained from the data to be screened of the Zest type to form the first type of data.

[0087] Step 206C: For each type of data to be screened, the distributed cache database processing device sorts the data to be screened based on the access frequency of the data to be screened, and determines the second type of data from the sorted data to be screened.

[0088] In an embodiment of the present application, after obtaining each type of data to be screened, for each type of data to be screened, the data to be screened can be sorted in ascending / descending order based on the size of the access frequency of each data to be screened. Then, according to the order of access frequency from large to small, a target number of data to be screened are obtained from the sorted data to be screened, that is, the second type of data. That is to say, the second type of data includes data of multiple data types, which can reflect the situation of hot keys of each type in the cache, thereby solving the data skew problem. In a feasible implementation, the target number can be set to 20.

[0089] It should be noted that the above embodiments may further include the following steps:

[0090] Step 207: The distributed cache database processing device performs correlation analysis on the first type of data and the second type of data to determine whether there is target data in the first type of data in the second type of data.

[0091] Step 208: When there is target data in the second type of data, the distributed cache database processing device determines the target identifier of the target data.

[0092] Among them, the target identifier indicates that the target data is both the first type of data and the second type of data.

[0093] In an embodiment of the present application, after obtaining the first type of data and the second type of data, correlation analysis can also be performed on the first type of data and the second type of data, that is, traverse each large key data in the hot key data. When it is determined that the target data in the large key data appears in the hot key data, it means that this target data is both large key data and hot key data, and then a target identifier (hotFlag) of the hot key data can be set for the target data. It should be noted that by performing correlation analysis on each type of large key data and hot key data, the call frequency of the large key data can be reflected, thereby reducing performance loss.

[0094] Step 209, the distributed cache database processing device stores the first description information of each first data in the first type of data into the target database.

[0095] Among them, the first description information includes a first index, a first name, a first type, a memory size, and a target identifier.

[0096] Step 210, the distributed cache database processing device stores the second description information of each second data in the second type of data into the target database.

[0097] Among them, the second description information includes a second index, a second name, a second type, and an access frequency.

[0098] In an embodiment of the present application, the first type of data may include multiple first data; the second type of data may include multiple second data; the first description information may refer to the description information of the first data, and the second description information may refer to the description information of the second data; the target database may specifically refer to a MySQL database. After obtaining the first type of data and the second type of data, the first type of data and the second type of data can also be persisted to the MySQL database through the storage module as shown in Figure 2 shown. Specifically, as shown in Figure 7 shown, the MySQL database mainly stores the first description information of the first type of data and the second description information of the second type of data. Among them, for the first type of data, it mainly stores the first index (db), the first name (key), the first type (type), the memory size (size), and the target identifier (hotFlag) corresponding to each first data; for the second type of data, it mainly stores the second index (db), the second name (key), the second type (type), and the access frequency (freq) of each second data.

[0099] Step 211, the distributed cache database processing device receives a visualization instruction for the target database.

[0100] Step 212: The distributed cache database processing device performs visualization processing on the first description information in the target database based on the visualization instruction to obtain first visualization information, and performs visualization processing on the second description information in the target database to obtain second visualization information.

[0101] In the embodiments of the present application, the visualization instruction refers to an instruction for visually displaying the information of the data in the target database; the first visualization information may refer to the information obtained by performing visualization processing on the first description information of each first data in the first type of data, and the second visualization information may refer to the information obtained by performing visualization processing on the second description information of each second data in the second type of data. After persisting the first description information of each first data and the second description information of each second data to the target database, the visualization module as shown in Figure 2 obtains the first description information of each first data and the second description information of each second data persisted in the MySQL database through the target interface (such as the Restful interface), and then renders them to the page through the visualization module, and displays the first visualization information (as shown in Figure 8(a)) and the second visualization information (as shown in Figure 8(b)) in the form of a pop-up window; in this way, through the visualization page, the user can clearly see the information of the top 20 key data of each current type and the information of the hot key data.

[0102] In other embodiments of the present application, the embodiments of the present application further provide a high-performance cache association analysis system for the Redis cluster version on the cloud. This system mainly includes a database module, a backup module, a data processing module, a storage module, and a visualization module, and the specific functions of each module can refer to the descriptions in the foregoing embodiments.

[0103] It should be noted that the descriptions of the same steps and the same content in this embodiment and other embodiments can refer to the descriptions in other embodiments, and will not be repeated here.

[0104] The distributed cache database processing method provided by the embodiments of the present application uses a batch sorting and comparison algorithm to sort the data read that meets the target conditions, and obtains the first type of data including different types and the second type of data including different types. That is to say, both the first type of data and the second type of data in the data to be analyzed are analyzed, rather than only being able to obtain the first type of data as in the related art, thereby solving problems such as data skew, performance degradation, and access timeout caused by the first type of data and the second type of data.

[0105] Based on the foregoing embodiments, the embodiments of the present application provide a distributed cache database processing device, and this distributed cache database processing device can be applied toFigure 1 and Figure 3 In the distributed cache database processing method provided by the corresponding embodiment, with reference to Figure 9 as shown, the distributed cache database processing apparatus 3 may include: a receiving unit 31, a determining unit 32, and a processing unit 33, where:

[0106] The receiving unit 31 is configured to receive a selection instruction for data in the distributed cache database;

[0107] The determining unit 32 is configured to determine data to be analyzed from the distributed cache database based on the selection instruction;

[0108] The processing unit 33 is configured to determine first-class data and second-class data from the data to be analyzed by using a target algorithm; where the target algorithm is an algorithm that reads data in batches and sorts the read data when the read data meets the target condition; the first-class data is data whose memory exceeds a first threshold; the second-class data is data whose access frequency exceeds a second threshold; the first-class data and the second-class data include multiple types of data.

[0109] In other embodiments of the present application, the determining unit 32 is specifically configured to perform the following steps:

[0110] When the selection instruction includes a shard selection instruction, determine a target shard corresponding to the shard selection instruction from the distributed cache database, and determine the data in the target shard as the data to be analyzed; where the distributed cache database includes multiple shards;

[0111] When the selection instruction includes a database selection instruction, determine the data in the distributed cache database as the data to be analyzed.

[0112] In other embodiments of the present application, the processing unit 33 is specifically configured to perform the following steps:

[0113] Obtain multiple first data to be screened from the data to be analyzed, and store the multiple first data to be screened in a target array;

[0114] Obtain multiple second data to be screened from the data to be analyzed, and store the multiple second data to be screened in the target array;

[0115] Obtain multiple third data to be screened from the data to be analyzed, and when the data volume in the target array meets the target condition, determine the first-class data and the second-class data from the target array.

[0116] In other embodiments of the present application, the processing unit 33 is specifically configured to perform the following steps:

[0117] Classify the data to be filtered in the target array based on the data type to obtain multiple categories of data to be filtered; wherein, the data to be filtered includes multiple first data to be filtered, multiple second data to be filtered, and multiple third data to be filtered;

[0118] For each category of data to be filtered, sort the data to be filtered based on the memory size of the data to be filtered, and determine the first category of data from the sorted data to be filtered;

[0119] For each category of data to be filtered, sort the data to be filtered based on the access frequency of the data to be filtered, and determine the second category of data from the sorted data to be filtered.

[0120] In other embodiments of the present application, the processing unit 33 is specifically configured to perform the following steps:

[0121] Perform an association analysis on the first category of data and the second category of data to determine whether there is target data in the first category of data in the second category of data;

[0122] In the case where there is target data in the second category of data, determine the target identifier of the target data; wherein, the target identifier indicates that the target data is both the first category of data and the second category of data.

[0123] In other embodiments of the present application, the processing unit 33 is specifically configured to perform the following steps:

[0124] Store the first description information of each first data in the first category of data into the target database; wherein, the first description information includes a first index, a first name, a first type, a memory size, and a target identifier;

[0125] Store the second description information of each second data in the second category of data into the target database; wherein, the second description information includes a second index, a second name, a second type, and an access frequency.

[0126] In other embodiments of the present application, the processing unit 33 is specifically configured to perform the following steps:

[0127] Receive a visualization instruction for the target database;

[0128] Based on the visualization instruction, perform visualization processing on the first description information in the target database to obtain first visualization information, and perform visualization processing on the second description information in the target database to obtain second visualization information.

[0129] It should be noted that the specific descriptions of the steps executed by each unit can be referred to Figure 1 and Figure 3 In the distributed cache database processing method provided in the corresponding embodiments, details are not described herein again.

[0130] The distributed cache database processing device provided by the embodiment of the present application adopts a batch sorting comparison algorithm to sort the data read that meets the target conditions, and obtains the first type of data including different types and the second type of data including different types. That is to say, both the first type of data in the data to be analyzed and the second type of data in the data to be analyzed are analyzed, rather than only obtaining the first type of data as in the related art, thus solving the problems of data skew, performance degradation, access timeout, etc. caused by the first type of data and the second type of data.

[0131] Based on the foregoing embodiment, an embodiment of the present application provides a distributed cache database processing device, which can be applied to Figure 1 and Figure 3 the distributed cache database processing method provided by the corresponding embodiment. As shown in Figure 10 , the distributed cache database processing device 4 may include: a processor 41, a memory 42, and a communication bus 43, where:

[0132] The communication bus 43 is used to implement the communication connection between the processor 41 and the memory 42;

[0133] The processor 41 is used to execute the distributed cache database processing program in the memory 42 to implement the following steps:

[0134] Receive a selection instruction for the data in the distributed cache database;

[0135] Based on the selection instruction, determine the data to be analyzed from the distributed cache database;

[0136] Use a target algorithm to determine the first type of data and the second type of data from the data to be analyzed; where the target algorithm is an algorithm that reads data in batches and sorts the data when the read data meets the target conditions; the first type of data is the data whose memory exceeds the first threshold; the second type of data is the data whose access frequency exceeds the second threshold; the first type of data and the second type of data include multiple types of data.

[0137] In other embodiments of the present application, the processor 41 is used to execute the distributed cache database processing program in the memory 42 based on the selection instruction to determine the data to be analyzed from the distributed cache database to implement the following steps:

[0138] When the selection instruction includes a shard selection instruction, determine the target shard corresponding to the shard selection instruction from the distributed cache database, and determine the data in the target shard as the data to be analyzed; where the distributed cache database includes multiple shards;

[0139] When the selection instruction includes a database selection instruction, determine the data in the distributed cache database as the data to be analyzed.

[0140] In other embodiments of the present application, the processor 41 is used to execute the distributed cache database processing program in the memory 42 to determine the first type of data and the second type of data from the data to be analyzed by using the target algorithm, so as to implement the following steps:

[0141] Obtain multiple first data to be screened from the data to be analyzed, and store the multiple first data to be screened into the target array;

[0142] Obtain multiple second data to be screened from the data to be analyzed, and store the multiple second data to be screened into the target array;

[0143] Obtain multiple third data to be screened from the data to be analyzed, and when the amount of data in the target array meets the target condition, determine the first type of data and the second type of data from the target array.

[0144] In other embodiments of the present application, the processor 41 is used to execute the distributed cache database processing program in the memory 42 to determine the first type of data and the second type of data from the target array, so as to implement the following steps:

[0145] Classify the data to be screened in the target array based on the data type to obtain multiple types of data to be screened; wherein, the data to be screened includes multiple first data to be screened, multiple second data to be screened, and multiple third data to be screened;

[0146] For each type of data to be screened, sort the data to be screened based on the memory size of the data to be screened, and determine the first type of data from the sorted data to be screened;

[0147] For each type of data to be screened, sort the data to be screened based on the access frequency of the data to be screened, and determine the second type of data from the sorted data to be screened.

[0148] In other embodiments of the present application, the processor 41 is used to execute the distributed cache database processing program in the memory 42 for the distributed cache database processing method, so as to implement the following steps:

[0149] Perform correlation analysis on the first type of data and the second type of data to determine whether there is target data in the first type of data in the second type of data;

[0150] When there is target data in the second type of data, determine the target identifier of the target data; wherein, the target identifier represents that the target data is both the first type of data and the second type of data.

[0151] In other embodiments of the present application, the processor 41 is used to execute the distributed cache database processing method of the distributed cache database processing program in the memory 42 to implement the following steps:

[0152] Store the first description information of each first data in the first type of data into the target database; wherein, the first description information includes a first index, a first name, a first type, a memory size, and a target identifier;

[0153] Store the second description information of each second data in the second type of data into the target database; wherein, the second description information includes a second index, a second name, a second type, and an access frequency.

[0154] In other embodiments of the present application, the processor 41 is used to execute the distributed cache database processing method of the distributed cache database processing program in the memory 42 to implement the following steps:

[0155] Receive a visualization instruction for the target database;

[0156] Based on the visualization instruction, perform visualization processing on the first description information in the target database to obtain first visualization information, and perform visualization processing on the second description information in the target database to obtain second visualization information.

[0157] It should be noted that the specific description of the steps executed by the processor can be referred to Figure 1 and Figure 3 In the distributed cache database processing method provided in the corresponding embodiments, details are not described herein again.

[0158] The distributed cache database processing device provided by the embodiments of the present application adopts a batch sorting comparison algorithm to sort the read data that meets the target conditions, and obtains the first type of data including different types and the second type of data including different types. That is to say, both the first type of data in the data to be analyzed and the second type of data in the data to be analyzed are obtained, rather than only being able to obtain the first type of data as in the related art, thus solving problems such as data skew, performance degradation, and access timeout caused by the first type of data and the second type of data.

[0159] Based on the foregoing embodiments, an embodiment of the present application provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement Figure 1 and Figure 3 The steps of the distributed cache database processing method provided in the corresponding embodiments.

[0160] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0161] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0162] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0163] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0164] As mentioned above, it is only a preferred embodiment of the present application, and is not used to limit the protection scope of the present application.

Claims

1. A distributed cache database processing method, characterized in that: The method comprises: Receiving a selection instruction for data in the distributed cache database; Based on the selection instruction, determining the data to be analyzed from the distributed cache database; A target algorithm is used to determine the first category of data and the second category of data from the data to be analyzed; wherein the target algorithm is an algorithm that reads data in batches and sorts the data when the read data meets the target condition; the first category of data is data whose memory exceeds a first threshold; the second category of data is data whose access frequency exceeds a second threshold; the first category of data and the second category of data include multiple types of data.

2. The method according to claim 1, characterized in that The determining the data to be analyzed from the distributed cache database based on the selection instruction includes: In the case where the selection instruction includes a shard selection instruction, determining a target shard corresponding to the shard selection instruction from the distributed cache database, and determining that the data in the target shard is the data to be analyzed; wherein the distributed cache database includes a plurality of shards; In a case where the selection instruction includes a database selection instruction, the data in the distributed cache database is determined to be the data to be analyzed.

3. The method according to claim 1, characterized in that The step of using a target algorithm to determine the first type of data and the second type of data from the data to be analyzed includes: Acquire a plurality of first data to be screened from the data to be analyzed, and store the plurality of first data to be screened in a target array; Acquire a plurality of second data to be screened from the data to be analyzed, and store the plurality of second data to be screened in the target array; A plurality of third data to be screened are obtained from the data to be analyzed until the amount of data in the target array meets the target condition, and then the first category of data and the second category of data are determined from the target array.

4. The method according to claim 3, characterized in that The determining the first category of data and the second category of data from the target array comprises: Classifying the data to be screened in the target array based on data types to obtain multiple types of data to be screened; wherein the data to be screened includes the multiple first data to be screened, the multiple second data to be screened, and the multiple third data to be screened; For each category of data to be screened, sorting the data to be screened based on the memory size of the data to be screened, and determining the first category of data from the sorted data to be screened; For each category of data to be screened, the data to be screened are sorted based on the access frequency of the data to be screened, and the second category of data is determined from the sorted data to be screened.

5. The method according to claim 4, characterized in that After determining the first category of data and the second category of data from the target array, the method further includes: Performing association analysis on the first category of data and the second category of data to determine whether the target data in the first category of data exists in the second category of data; In the case that the target data exists in the second category of data, a target identifier of the target data is determined; wherein the target identifier indicates that the target data is both the first category of data and the second category of data.

6. The method according to claim 5, characterized in that: After determining the target identifier of the target data, the method further includes: Storing first description information of each first data in the first category of data in a target database; wherein the first description information includes a first index, a first name, a first type, a memory size, and a target identifier; The second description information of each second data in the second category of data is stored in the target database; wherein the second description information includes a second index, a second name, a second type and an access frequency.

7. The method according to claim 6, characterized in that The method further comprises: receiving a visualization instruction for the target database; Based on the visualization instruction, the first description information in the target database is visualized to obtain first visualization information, and the second description information in the target database is visualized to obtain second visualization information.

8. A distributed cache database processing device, characterized in that: The device comprises: A receiving unit, configured to receive a selection instruction for data in the distributed cache database; A determination unit, configured to determine the data to be analyzed from the distributed cache database based on the selection instruction; A processing unit, used for determining a first category of data and a second category of data from the data to be analyzed by adopting a target algorithm; wherein the target algorithm is an algorithm for reading data in batches and sorting the data when the read data meets a target condition; the first category of data is data whose memory exceeds a first threshold; the second category of data is data whose access frequency exceeds a second threshold; the first category of data and the second category of data include multiple types of data.

9. A distributed cache database processing device, characterized in that: The device comprises: a processor, a memory and a communication bus; The communication bus is used to realize the communication connection between the processor and the memory; The processor is used to execute the distributed cache database processing program in the memory to implement the steps of the distributed cache database processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the distributed cache database processing method according to any one of claims 1 to 7.