A hotspot recognition method and a current limiting method
By using hot spot identification tree and current limiting methods in distributed databases, the problem of hot spot identification and processing in distributed databases with large amounts of data stored by a single physical machine is solved, and the stability of the system and memory utilization efficiency are improved.
Patent Information
- Application Number
- CN202210289116.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-03-22
AI Technical Summary
In distributed databases, when the amount of data stored by a single physical machine is large, it is difficult for the existing technology to effectively identify and deal with hot issues, resulting in system downtime and affecting the stability of database services.
A hot spot identification method and current limiting method are proposed. By pre-initializing the hot spot index count value in the hot spot identification tree to 0, and upon receiving a data access request, the hot spot primary key is identified by calculating the hot spot index count value corresponding to the primary key, and the current limiting process is performed according to the preset conditions.
While ensuring a certain degree of accuracy, it occupies a small amount of memory to identify hot spots, avoid hot spots from occupying large memory or affecting the normal operation of database services, and improves the stability of the system.
Smart Images

Figure CN114756544B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computer application technologies, and in particular, to a hot spot identification method and a current limiting method. Background Art
[0002] In a distributed database, the data of a relatively large database is divided into multiple parts and stored on different physical machines. If a piece of data in the database is frequently accessed within a certain period of time, since the access service of a piece of data is generally borne by a single physical machine, when the access volume of this piece of data is relatively large, a hot spot will be formed, that is, in a distributed system, access requests are concentrated on a single physical machine. For example, in scenarios such as Weibo hot searches and e-commerce flash sales, hot spot problems are likely to occur. In the case of a hot spot problem, the physical machine where the hot spot is located may crash due to excessive traffic, affecting the stability of the distributed database service.
[0003] For a non-relational database, each piece of data generally corresponds to a primary key (key). In related technologies, some distributed databases generally identify hot spots by counting the access times of each primary key. However, for a database with a large amount of data stored on a single physical machine, this method will occupy a large amount of memory and affect the normal operation of the database.
[0004] It can be seen that for a distributed database with a large amount of data on a single physical machine (and it is a non-relational database), there is a lack of a hot spot identification method that can avoid system crashes. Summary of the Invention
[0005] In view of this, one or more embodiments of this specification provide a hot spot identification method and a current limiting method.
[0006] According to the first aspect of one or more embodiments of this specification, a hot spot identification method is proposed. The count values corresponding to multiple hot spot indexes included in the hot spot identification tree are initialized to 0 in advance; the number of hot spot indexes in the hot spot identification tree is less than the number of data in the database; the method includes:
[0007] When a data access request is received, determine the primary key of the data accessed by the data access request;
[0008] Calculate the hot spot index corresponding to the primary key, and determine the corresponding count value in the hot spot identification tree according to the hot spot index;
[0009] In response to the count value satisfying a preset hot spot condition, determine that the primary key is a hot spot primary key; in response to the count value not satisfying the preset hot spot condition, increment the count value by one.
[0010] According to the second aspect of one or more embodiments of this specification, a current limiting method is proposed, including:
[0011] Receive a data access request for target data in a database, where the data access request carries a primary key corresponding to the target data;
[0012] In response to the primary key belonging to a hot primary key, block the data access request for the target data; the hot primary key is obtained by the foregoing hot spot identification method.
[0013] According to a third aspect of the embodiments of the present specification, a hot spot identification device is provided, and the count values corresponding to a plurality of hot spot indexes included in a hot spot identification tree are initialized to 0 in advance; the number of hot spot indexes in the hot spot identification tree is less than the number of data in the database; the device includes:
[0014] A primary key determination module, configured to determine a primary key of data accessed by the data access request when receiving the data access request;
[0015] A count value determination module, configured to calculate a hot spot index corresponding to the primary key, and determine a corresponding count value in the hot spot identification tree according to the hot spot index;
[0016] A hot spot determination module, configured to determine that the primary key is a hot primary key in response to the count value satisfying a preset hot spot condition; and increment the count value by 1 in response to the count value not satisfying the preset hot spot condition.
[0017] According to a fourth aspect of the embodiments of the present specification, a current limiting device is provided, including:
[0018] A request receiving module, configured to receive a data access request for target data in a database, where the data access request carries a primary key corresponding to the target data;
[0019] A request blocking module, configured to block the data access request for the target data in response to the primary key belonging to a hot primary key; the hot primary key is obtained by the foregoing hot spot identification method.
[0020] According to a fifth aspect of the embodiments of the present specification, an electronic device is provided, including:
[0021] A processor;
[0022] A memory for storing executable instructions of the processor;
[0023] Wherein, the processor realizes the foregoing hot spot identification method or the foregoing current limiting method by running the executable instructions.
[0024] According to a sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided. Computer instructions are stored on the computer-readable storage medium, and when the computer instructions are executed by a processor, the foregoing hotspot recognition method or the foregoing current limiting method is implemented.
[0025] The present specification provides a hotspot recognition method and a current limiting method. Initially, the count values corresponding to multiple hotspot indexes included in the hotspot recognition tree are set to 0; the number of hotspot indexes in the hotspot recognition tree is less than the number of data in the database; when a data access request is received, determine the primary key of the data accessed by the data access request; calculate the hotspot index corresponding to the primary key, and determine the corresponding count value in the hotspot recognition tree according to the hotspot index; the count value is used to represent the number of times the data corresponding to the hotspot index has been accessed; in response to the count value satisfying a preset hotspot condition, determine that the primary key is a hotspot primary key; in response to the count value not satisfying the preset hotspot condition, increment the count value by one.
[0026] By the above method, by counting the count values of the hotspot indexes corresponding to each primary key, since the number of hotspot indexes is less than the number of primary keys in the database, it is possible to identify hotspots with a relatively small amount of memory while ensuring a certain degree of accuracy, and process the hotspots, avoiding the situation where hotspots occupy a large amount of memory or affect the normal operation of the database service.
[0027] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present specification, and are used together with the specification to explain the principles of the present specification.
[0029] Figure 1 is a flowchart of a hotspot recognition method shown according to an exemplary embodiment of the present specification.
[0030] Figure 2 is a schematic structural diagram of a hotspot recognition tree shown according to an exemplary embodiment of the present specification.
[0031] Figure 3 is a flowchart of a current limiting method shown according to an exemplary embodiment of the present specification.
[0032] Figure 4A is a schematic structural diagram of a sketch shown according to a specific embodiment of the present specification.
[0033] Figure 4B is a schematic structural diagram of a hotspot recognition tree shown according to a specific embodiment of the present specification.
[0034] Figure 5 This is a block diagram of a hotspot recognition device shown in accordance with an exemplary embodiment of this specification.
[0035] Figure 6 This is a block diagram of a current limiting device shown in accordance with an exemplary embodiment of this specification.
[0036] Figure 7 This is a hardware structure diagram of an electronic device where a hotspot recognition device or a current limiting device is located, shown in accordance with an exemplary embodiment of this specification. Detailed implementation manners
[0037] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with one or more embodiments of this specification. On the contrary, they are merely examples of devices and methods that are consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0038] It should be noted that: in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0039] In a distributed database, generally, the user's data is sorted in byte array order, and the sorted data is segmented. The different data obtained by segmentation is stored in different modules (regions) of the database, and this region provides access services for this part of the data. This design is generally called range sharding design, which is a common design method in the field of distributed systems. Under this design, a piece of user data can definitely be uniquely located to a range shard.
[0040] Distributed systems generally obtain great scalability through horizontal scaling. By simply adding more machines, the overall throughput can be continuously improved almost indefinitely. However, due to the range sharding design, in some cases, due to business characteristics (such as Weibo hot searches, e-commerce flash sales, where the data corresponding to a Weibo post or a commodity is a single piece of data), most of the traffic may be concentrated on one or a few pieces of data in a short period of time. That is, most of the traffic will hit a specific one or several regions in a short time, which will impose a huge load on the machines where these regions are located, thus forming a hotspot problem. The hotspot problem may lead to abnormal situations such as hardware crashes and process exits. In many distributed systems, there are disaster recovery designs for single-machine anomalies. When a machine crashes, other machines will be responsible for taking over the access services on the crashed machine. However, the hotspot problem is not a hardware problem. As machines crash and other machines take over the services, they will be overwhelmed by the hotspot traffic and crash again, easily causing an avalanche effect on the entire cluster and seriously affecting the stability of the service. It can be seen that limited by the finite single-machine hardware resources, the processing capacity of a single machine is also necessarily limited. In the face of the hotspot problem, the horizontal scalability of distributed systems is useless. Therefore, other methods need to be sought to solve the hotspot problem.
[0041] In related technologies, the common means to address the hotspot problem is usually traffic limiting. However, the difficulty of the hotspot problem lies in the identification and discovery of the hotspot itself. To identify hotspots, considering that each piece of data corresponds to a primary key (key), some in-memory databases will count the access volumes of all primary keys to determine the hot primary keys.
[0042] However, the above method is only suitable for scenarios with limited data in single-machine storage such as in-memory databases. For some persistent databases with a large amount of data stored in a single machine (data in the order of several terabytes, with a throughput of hundreds of thousands of QPS per second), counting the access volume of each key will consume a large amount of memory, making this method inapplicable and inefficient.
[0043] Based on this, this specification provides a hotspot identification method and a traffic limiting method. Initialize the count values corresponding to multiple hotspot indexes included in the hotspot identification tree to 0 in advance; the number of hotspot indexes in the hotspot identification tree is less than the number of data in the database; when a data access request is received, determine the primary key of the data accessed by the data access request; calculate the hotspot index corresponding to the primary key, and determine the corresponding count value in the hotspot identification tree according to the hotspot index; the count value is used to represent the number of times the data corresponding to the hotspot index has been accessed; in response to the count value meeting a preset hotspot condition, determine that the primary key is a hot primary key; in response to the count value not meeting the preset hotspot condition, increment the count value by one.
[0044] The above method identifies hotspots by counting the count values of hot indexes corresponding to each primary key. Since the number of hot indexes is less than the number of primary keys in the database, it is possible to occupy less memory to identify hotspots while ensuring a certain degree of accuracy, and process the hotspots based on the identified hotspots, avoiding the problem that hotspots occupy a large amount of memory or affect the normal operation of the database service.
[0045] Next, a method for identifying hotspots shown in this specification will be described.
[0046] As Figure 1 shown, Figure 1 is a flowchart of a method for identifying hotspots shown in this specification according to an exemplary embodiment, including the following steps:
[0047] Step 103, when a data access request is received, determine the primary key of the data accessed by the data access request.
[0048] It should also be noted that before the execution of this method, the count values corresponding to multiple hot indexes included in the hotspot identification tree need to be initialized to 0; the number of hot indexes in the hotspot identification tree is less than the number of data in the database.
[0049] The reason for determining the primary key in step 103 is that the data access request is a request for reading or writing data. In this specification, the database targeted by the data access request is a non-relational database, and each piece of data corresponds to a primary key. Different primary keys can be used to distinguish access requests for different data. Therefore, in order to count the fuzzy access volume for different data (the specific meaning will be described in detail below), it is necessary to first determine the primary key of the data accessed by the data access request.
[0050] The reason for initializing the count values corresponding to multiple hot indexes to 0 before the execution of this method is that in this specification, the fuzzy access volume (that is, the count value) for different data is counted through the hotspot identification tree, and then hotspots are identified. Therefore, it is necessary to initialize the hotspot identification tree to a situation where no access volume is recorded in advance, that is, to clear the count values corresponding to multiple hot indexes of the hotspot identification tree, so as to facilitate the counting of the fuzzy access volume.
[0051] The reason why the number of hot indexes in the hotspot identification tree is less than the number of data in the database is to save memory. To save memory, compared with the related technology where each primary key counts an access volume, it is necessary to make a counted fuzzy access volume correspond to multiple primary keys. Therefore, it is necessary to ensure that the number of hot indexes is less than the number of primary keys in the database, so as to save memory space, and thus enable the method in this specification to support the identification of hotspots in a database with a large amount of data.
[0052] In addition, the hotspot identification tree is a tree structure that records the count values of each hotspot index. It can have only one layer or multiple layers. In the case of one layer, multiple hotspot indexes are stored in one layer of the hotspot identification tree, and the count values corresponding to the multiple hotspot indexes are also stored.
[0053] In the case of multiple layers, multiple hotspot indexes and their corresponding count values are stored in each layer of the hotspot identification tree. For a multi-layer hotspot identification tree, one hotspot index in the upper layer corresponds to multiple hotspot indexes in the lower layer. The reason for this correspondence is that the methods for calculating hotspot indexes in the two layers are different. Therefore, multiple primary keys corresponding to one hotspot index in the first layer will be calculated into multiple hotspot indexes by a different calculation method in the lower layer compared to the upper layer. In other words, one hotspot index in the upper layer and the multiple hotspot indexes in the lower layer corresponding to this upper-layer hotspot index represent the same batch of data. In this way, under the multi-layer structure, as the number of layers increases, the data corresponding to each hotspot index gradually decreases.
[0054] It is also necessary to explain the method for expanding the hotspot identification tree. When the count value of a hotspot index in a certain layer reaches its maximum (i.e., the maximum value that the count value can reach), the hotspot indexes in the lower layer corresponding to this hotspot index will be expanded to determine which data among the data corresponding to this hotspot index is / are the hotspot.
[0055] In addition, the advantages and other details of the multi-layer structure hotspot identification tree will be described in detail below and will not be elaborated here for the time being.
[0056] Step 105: Calculate the hotspot index corresponding to the primary key, and determine the corresponding count value in the hotspot identification tree according to the hotspot index.
[0057] First, the meanings of the various terms involved in step 105 will be explained.
[0058] The hotspot index is an index value calculated through methods similar to hash functions, etc. That is, taking the primary key as the input, the hotspot index is calculated through methods such as hash functions.
[0059] The count value corresponding to the hotspot index in the hotspot identification tree represents the magnitude of the access volume of the data corresponding to the hotspot index. In the case where the hotspot identification tree has only one layer, the count value corresponding to the hotspot index represents the sum of the access volumes of the multiple primary keys / data corresponding to the hotspot index within a certain period of time. That is, the count value represents the fuzzy access volume of the data corresponding to the hotspot index.
[0060] When the hotspot recognition tree has multiple layers, the count value corresponding to any hotspot index in each layer is used to characterize the access level of the data corresponding to the hotspot index (for example, if the hotspot recognition tree has two layers, the data corresponding to any hotspot index in the second layer refers to the data with the index value of the second layer hotspot index among the data corresponding to the first layer hotspot data corresponding to the second layer hotspot index).
[0061] For example, when the hotspot recognition tree has two layers, if a certain hotspot index only corresponds to the first layer hotspot index, or the count values of some of the second layer hotspot indexes corresponding to this hotspot index are 0, it proves that the data corresponding to these hotspot indexes (the first layer hotspot index in the former case, some of the second layer hotspot indexes in the latter case) has a relatively low access volume; if there are second layer hotspot indexes corresponding to the first layer hotspot index corresponding to a certain hotspot index, and the corresponding second layer hotspot indexes are not 0, it proves that the access volume of the data corresponding to this second layer hotspot index is medium; if the count value of the second layer hotspot index corresponding to this primary key (data) is full, it proves that the access volume of this primary key (data) is relatively large and it belongs to a hotspot primary key.
[0062] Step 107, in response to the count value satisfying the preset hotspot condition, determine that the primary key is a hotspot primary key; in response to the count value not satisfying the preset hotspot condition, increment the count value by one.
[0063] In other words, since the recognition condition for a hotspot primary key is that the access volume within a period of time is greater than a certain value, so if the statistical quantity of the count value within a certain time is greater than a certain value, or the proportion of the fuzzy access volume characterized by the count value in the total access volume within a certain time is greater than a certain value, then it is considered that the primary key is a hotspot primary key. And it is necessary to continue to count the count value of the hotspot index corresponding to this primary key when the primary key is not a hotspot primary key.
[0064] Therefore, step 107 may specifically include: in response to the proportion of the count value in the sum of all count values exceeding the preset hotspot proportion, determine that the primary key is a hotspot primary key.
[0065] In other words, it can be determined whether it is hotspot data by calculating whether the proportion of the count value is greater than a certain value, which is a relatively convenient method for a single-layer hotspot recognition tree. For a multi-layer hotspot recognition tree, in order to calculate accurate hotspots, the count value can be the sum of the count values of the last layer (the count value of each layer will be accumulated to the next layer, for example, when the count value of the first layer is 15 and the next layer is expanded, the count value of the second layer starts counting from 15). The sum of all count values can also be the sum of access numbers (because each access request will increment a certain count value by 1 when it arrives, so the sum of access numbers is the sum of all count values).
[0066] In addition, in addition to the above method, step 107 can also be implemented by the following method.
[0067] The method further includes: when receiving a data access request for any data, incrementing the access request count by 1; when the access request count is greater than a preset access request threshold, resetting the access request count to 0 and setting the count values corresponding to multiple hot index included in the hot spot identification tree to 0. And step 107 includes: in response to the count value reaching a preset hot spot threshold, determining the primary key as a hot spot primary key.
[0068] In other words, when the total access volume for all data in the database is greater than a certain value, reset the hot spot identification tree and start counting again. In the above case, since the hot spot identification tree will be reset, the preset hot spot condition can be that the count value exceeds a preset hot spot threshold, so that the proportion of the access traffic of hot spot data in the access traffic of all data within a certain period of time can be counted.
[0069] In addition, the hot spot identification tree can be reset every once in a while (i.e., re-initialized), and in this case, the hot spot identification condition can be that the count value exceeds a preset hot spot threshold. In this way, it can be counted whether the access volume within a certain period of time exceeds a predetermined threshold, so as to determine whether there is a hot spot.
[0070] It should also be noted that when a hot spot primary key is identified, the hot spot primary key can also be recorded and the specific access volume of the hot spot primary key can be started to be counted, mainly to facilitate observing the traffic situation and observing the processing situation of the hot spot.
[0071] In other words, the method further includes: when determining the primary key as a hot spot primary key, recording the primary key and recording the access volume of the primary key.
[0072] In the above case, in order to prevent the system from recording the access volume of the primary key without limit and occupying too much memory, a cache eviction policy based on Least Recently Used (LRU) or Least Frequency Used (LFU) can be used to delete a part of the recorded primary keys.
[0073] Specifically, when the number of recorded primary keys exceeds a certain threshold, the primary key with the lowest access frequency (the primary key with the lowest access frequency may be data mis-identified as a hot spot because the number of hot indexes is less than the number of primary keys) can be deleted, or when the number of primary keys exceeds a certain threshold, the primary key that has not been accessed for a certain period of time can be deleted.
[0074] For the former case of the previous paragraph, that is: The method further includes: when the number of recorded primary keys exceeds a preset primary key quantity threshold, deleting the primary key with the lowest access frequency according to the access volumes of the respective primary keys of the records.
[0075] After the method shown in this specification is described, the hotspot recognition tree will be further described in detail below.
[0076] As previously mentioned in the above text, there is a case where the hotspot recognition tree has multiple layers. The purpose of setting multiple layers for the hotspot recognition tree is to avoid hash collisions and further save memory. Specifically, the hotspot index is calculated through a certain calculation method, and since the number of hotspot indexes is less than the number of primary keys, it results in multiple primary keys corresponding to one hotspot index. To distinguish different primary keys corresponding to the same hotspot index as much as possible, when the hotspot recognition tree has only one layer, it can be achieved by increasing the number of hotspot indexes. However, increasing the number of hotspot indexes will also increase the amount of memory occupied (it is very likely that the count values of some hotspot indexes are 0 or very small). Therefore, to avoid the above contradiction, it can be solved by setting the hotspot recognition tree with multiple layers.
[0077] Specifically, when the hotspot recognition tree is set with multiple layers, its structure can be as Figure 2 shown. The number of first-layer hotspot indexes is N, each first-layer hotspot index corresponds to M second-layer hotspot indexes, and each second-layer hotspot index can correspond to Q third-layer hotspot indexes (N, M, and Q are all preset positive integers, and the sizes of N, M, and Q can be the same or different), and so on. Different layers of hotspot indexes correspond to different index calculation methods (which can be hash functions), so as to avoid hash collisions (in the case where the index calculation method is a hash function) and reduce the possibility of misidentifying some data as hotspots. And to reduce memory, it can be set that when the count value of the hotspot index in the upper layer is not full, the hotspot index of the lower layer is not established. In other words, the initial hotspot recognition tree has only one layer. When the count value corresponding to any hotspot index in this layer exceeds the preset expansion threshold, the hotspot index of the next layer corresponding to this hotspot index is newly established. In this way, in the case where only a small number of data are hotspots, most of the hotspot indexes stop at the previous layers, and very few hotspot indexes can reach the last layer. Compared with the scheme with only one layer of hotspot indexes, hash collisions can be avoided with the same storage space size.
[0078] It should also be noted that Figure 2 although the structure with corresponding hotspot indexes in multiple layers is shown in
[0079] It can be seen that the multi-layer structure of the hotspot recognition tree can reduce the occupied memory space while avoiding hash collisions.
[0080] Next, a two-layer hotspot recognition tree will be used as an example to illustrate the above solution.
[0081] The hotspot recognition tree includes a first layer and a second layer. Step 103 specifically includes: calculating the first-layer hotspot index corresponding to the primary key according to the hash function corresponding to the first layer of the hotspot recognition tree; in response to the non-existence of the second-layer hotspot index corresponding to the first-layer hotspot index in the hotspot recognition tree, determining the technical value corresponding to the first-layer hotspot index as the corresponding count value; in response to the existence of the second-layer hotspot index corresponding to the first-layer hotspot index in the hotspot recognition tree, determining the second-layer hotspot index corresponding to the primary key according to the hash function corresponding to the second layer, and determining the count value corresponding to the second-layer hotspot index as the corresponding count value. Among them, the hash function corresponding to the first layer is different from the hash function corresponding to the second layer.
[0082] In other words, when the hotspot recognition tree includes two layers, the determined count value is the count value of the last layer where the primary key exists. Specifically, when there is a second hotspot index corresponding to the first-layer hotspot index of the primary key, the count value of the second-layer hotspot index corresponding to the primary key is used as the determined count value; when there is no second-layer hotspot index, the count value of the first-layer hotspot index is used as the determined count value.
[0083] In addition to determining the count value, there is also a process of expanding the next layer. Specifically: initializing the count values corresponding to multiple hotspot indexes included in the hotspot recognition tree to 0 in advance, including: initializing the count values corresponding to multiple hotspot indexes included in the first layer of the hotspot recognition tree to 0 in advance. The method further includes: in response to the count value corresponding to the first-layer hotspot index after incrementing exceeding a preset expansion threshold, creating multiple second-layer hotspot indexes corresponding to the first-layer hotspot index, and initializing the count values corresponding to the multiple second-layer hotspot indexes to 0.
[0084] In other words, only the first-layer hotspot indexes are initialized during the pre-initialization. When the count value of the first layer exceeds the expansion threshold, the next layer of hotspot indexes is expanded and the next layer of hotspot indexes is initialized.
[0085] By the above method, memory is saved. Since the number of hot indexes is fixed and less than the number of primary keys, it can be better applied to databases with massive data. On the premise of low memory overhead, the false positive rate can be calculated according to the number of layers of the hot spot recognition tree, the maximum number of theoretical hot indexes per layer, the maximum value that the count value can reach (if the hot spot primary key is recognized by recognizing that the count value exceeds the predetermined hot spot threshold), and the access request threshold (in the case of an access request threshold). The hot spot recognition requirements can be met by reasonably setting each threshold. And since it does not occupy memory, the performance overhead is not large either.
[0086] After the hot spot primary key is recognized, it is also necessary to process the hot spot primary key to avoid the impact of the hot spot primary key on the single machine operation. There are various methods for processing hot spots. Next, the process of processing the hot spot primary key will be illustrated by taking rate limiting as an example.
[0087] Such as Figure 3 shown, Figure 3 is a flowchart of a rate limiting method shown in this specification according to an exemplary embodiment, including:
[0088] Step 301, receive a data access request for target data in the database, and the data access request carries the primary key corresponding to the target data.
[0089] Step 303, in response to the primary key belonging to the hot spot primary key, block the data access request for the target data.
[0090] Among them, the hot spot primary key is obtained by the foregoing hot spot recognition method. In this way, the data access request of the hot spot primary key is blocked, thereby achieving the purpose of protecting the system itself.
[0091] It should be noted that in step 303, the implementation manner of blocking the access request for the target data can be that after the hot spot primary key is recognized by the foregoing hot spot recognition method, if it is a hot spot primary key, the access request is blocked. In addition, in order to further reduce the impact of the hot spot problem on the system operation, after the hot spot is discovered, the hot spot primary key can be added to a single machine rate limiter (for example: RateLimiter in guava. Of course, the rate limiter can also be other rate limiters, and this specification does not limit the type of the rate limiter). When the system further processes the access request (for example, determines whether it is hot spot data by the foregoing method), the access request is blocked to save the operation performance.
[0092] In addition, to ensure the accuracy of identification, in addition to identifying the hot primary key through the above methods, when the access volume corresponding to the hot primary key is recorded, the hot primary key can also be identified through the recorded access volume. When the access volume of a certain primary key recorded is greater than a certain value, the data access request for this primary key is blocked, which can improve the accuracy of traffic limiting.
[0093] It should also be noted that when the access volume corresponding to the hot primary key is recorded, after traffic limiting the hot primary key, the access volume of the traffic-limited hot primary key can be observed for a period of time, which can feedback the traffic limiting situation of the hot primary key and facilitate the adjustment of traffic limiting strategies, etc.
[0094] In addition, in the above situation, if there are also LRF and LFU cache eviction policies, the hot primary keys to be observed can be locked to avoid the data of the hot primary keys to be observed being evicted by the LRF or LFU policy in the case of successful traffic limiting.
[0095] Next, a specific embodiment will be used to illustrate a hot spot identification method and a traffic limiting method shown in this specification.
[0096] Next, taking 4 layers and the number of hot spot indexes in each group of each layer being 256, the structure of the hot spot identification tree, the update, reset, etc. of the hot spot identification tree will be described.
[0097] First, introduce the concept of Sketch, which refers to a group of hot spot indexes and their corresponding count values (a group includes 256 hot spot indexes), and its structure is as Figure 4A shown. Each Sketch includes 256 hot spot indexes, 256 count values, as well as the maximum value that each count value can count to, the layer to which the Sketch belongs, etc.
[0098] After describing the Sketch, next, the structure of the hot spot identification tree will be introduced, as Figure 4B shown. The hot spot identification tree includes 4 layers of structure. The first layer includes 1 Sketch, the second layer includes the Sketches corresponding to each hot spot index in the first layer, that is, the second layer includes at most 256 Sketches, the third layer includes at most 256 * 256 Sketches, and the fourth layer includes at most 256 * 256 * 256 Sketches (for the convenience of showing only 1 Sketch for each layer). Since each Sketch has 256 hot spot indexes, the maximum number of hot spots that the hot spot identification tree can identify is the fourth power of 256 (that is, the number of hot spot indexes included in the fourth layer). It should also be noted that different layers correspond to different Hash functions.
[0099] There are two operations for the hotspot identification tree. The first is to update the hotspot identification tree, that is, when a data access request is received, the hotspot identification tree is updated, and the data accessed by the data access request is determined to be a hotspot. The second is to reset the hotspot identification tree.
[0100] Next, we will first describe the process of updating the hotspot identification tree.
[0101] First, the hash function corresponding to the first layer of the hotspot identification tree is used to calculate the first-layer hotspot index of the primary key of the data targeted by the data access request. In order to limit the number of hotspot indexes to 256, the remainder of 256 can be taken from the solved index value to obtain the first-layer hotspot index.
[0102] After obtaining the first-level hot index, determine whether the first-level hot index corresponds to a second-level hot index. If there is a second-level hot index, calculate the second-level hot index corresponding to the primary key, and determine whether the second-level hot index has a third-level hot index, and so on, until the last level of hot index corresponding to the primary key is determined (the last level refers to the one with the largest number of levels. It should be noted that the last level of hot index here is the last level of hot index where the primary key exists).
[0103] After determining the last layer of hot index, add 1 to the count value corresponding to the last layer of hot index, and determine whether the count value after adding 1 exceeds the maximum count value (14). If it exceeds, expand the sketch of the next layer corresponding to the hot index (i.e., initialize the sketch of the next layer). If there is no expandable next layer sketch, determine that the primary key is the hot primary key.
[0104] This completes the identification of the hotspot primary key. In the above process, a hash function is used to disperse data access requests into multiple hotspot indexes for accumulation. A higher accumulated value represents a higher frequency (this effect should be combined with the hotspot identification tree reset process). Because the number of hotspot indexes in each sketch (256) is much smaller than the data distribution to be accessed by the data access request, low-frequency data may be falsely reported due to hash conflicts and high-frequency data being dispersed on the same hotspot index. Therefore, a multi-layer design is added to the hotspot identification tree, and each layer corresponds to a separate hash function. In this way, data that is mixed together due to hash conflicts in the first layer is more likely to be dispersed in the second layer using another hash function. Through 4 layers and 4 hash functions, false alarms caused by hash conflicts can be greatly reduced.
[0105] In addition, it should be noted that in a non-relational database (noSQL), each piece of data has a primary key that can uniquely identify it. In the read-write link, for each piece of data accessed, its primary key will be used as an input parameter to call the above method to determine whether the accessed data is a hot spot. Each time the above method is called, a return value of boolean type (that is, it can only return true or false) will be returned. When the return value is true, there is a high probability that the data corresponding to the input primary key is a hot spot data that is being frequently accessed.
[0106] In addition, for any primary key that returns true, its each read-write access will be recorded. There are two purposes for this: First, the above method cannot inform the specific access volume of the data access request. Through recording, the actual access volume of the hot primary key can be obtained so that the operation and maintenance personnel can intuitively know the access volume. Second, because the above method has a low probability of returning low-frequency data (mistaking low-frequency accessed data for hot spot data), through recording, it can be known which are the truly high-frequency hot spots, so as to delete the low-frequency data and prevent mis-limiting the flow of some low-frequency data, causing troubles to users.
[0107] In addition, the number of continuously recorded primary keys should be limited within 1000. Once it exceeds 1000, the data will be eliminated according to the LRU principle. At the same time, data that has not been accessed for a certain period of time (such as 5 minutes) will also be eliminated. Doing so is on the one hand to filter out the low-frequency data returned with low probability faster, and on the other hand to limit the memory overhead and only track and record the data of the highest-frequency hot key access.
[0108] Next, the process of resetting the hot spot identification tree will be described in detail. Regarding the above process as a process of expanding the hot spot identification tree, the reset process is a process of shrinking the hot spot identification tree.
[0109] Since each time a data access request is received, the above method needs to be used to update the hot spot identification tree, the number of times the above method is executed can be used as the number of data access requests received (data access requests not limited by flow). When the number of received data access requests exceeds the preset number threshold, the hot spot identification tree is reset, that is, the second to fourth layers of the hot spot identification tree are deleted, and all the count values of the first layer are reset to 0. The purpose of doing this is that the traffic is constantly changing, and the data with high-frequency access may only last for a while and then no longer be high-frequency. Therefore, a periodic shrinking mechanism needs to be added to filter out the historical high-frequency data, so that the above hot spot identification method can always correctly judge the data that is being accessed with high frequency under limited memory overhead.
[0110] It should also be noted that the addition of the reset method can filter out the hot primary keys whose access frequency (i.e., the access volume of the data / the access volume of all data) is greater than a certain value. Specifically, the hot identification tree is reset every P updates. Then, by setting the maximum value of the count value for each layer, the hot primary keys above the predetermined access frequency can be obtained.
[0111] In addition, after discovering the hot primary keys, the true read and write access throughput can be obtained by comprehensively recording the primary keys continuously. Then, the operation and maintenance personnel can perform flow limiting. For example, a single-machine flow limiter (such as RateLimiter in Guava) can be used for flow limiting, and the requests being flow-limited adopt the fail-fast strategy to return exceptions to the client, realizing self-protection of the system.
[0112] The above method has a very good effect of saving memory. For each table in the database, using the parameters in the above example, a single table only occupies 500 bytes of resident memory without hot spots. There is a theoretical upper limit for the memory occupation of each table, which is 30k (each table is a hot spot and each table has multiple hot keys). It can be seen that the memory overhead required by the above method has nothing to do with the actual amount of stored data and can be applied to the database scenario of mass storage.
[0113] On the premise of achieving extremely low memory overhead, using the above parameters, the scheme can detect the keys with a frequency of more than 3.5% at least, and count qps and rt, which can fully meet the requirements of identifying hot spots.
[0114] The above method has a cpu calculation overhead with a complexity of log(N) and has a low impact on the response of read and write requests. In the case of no hot spots, a personal computer can support the determination of millions of hot keys by a single core. (In the case of hot spots, the sketch will be expanded and it is slightly slower by 1 - 2 times), which reflects the advantage of small performance overhead.
[0115] In addition, the false positive rate can be calculated through the above parameters, and the above parameters can be reasonably set to reduce the false positive rate.
[0116] Corresponding to the embodiments of the foregoing method, this specification also provides embodiments of a device and the electronic device to which it is applied.
[0117] As Figure 5 shown, Figure 5 is a block diagram of a hot spot identification device shown in this specification according to an exemplary embodiment. The device includes:
[0118] An initialization module 500, configured to initialize the count values corresponding to multiple hot spot indexes included in the hot spot identification tree to 0 in advance; the number of hot spot indexes in the hot spot identification tree is less than the number of data in the database.
[0119] The primary key determination module 510 is configured to determine the primary key of the data accessed by the data access request when receiving the data access request.
[0120] The count value determination module 520 is configured to calculate the hot index corresponding to the primary key, and determine the corresponding count value in the hot identification tree according to the hot index.
[0121] The hot key determination module 530 is configured to determine that the primary key is a hot primary key in response to the count value satisfying a preset hot condition; and increment the count value by one in response to the count value not satisfying the preset hot condition.
[0122] In an alternative embodiment, the hot identification tree includes a first layer and a second layer. The count value determination module 520 is specifically configured to: calculate the first-layer hot index corresponding to the primary key according to the hash function corresponding to the first layer of the hot identification tree; in response to the non-existence of the second-layer hot index corresponding to the first-layer hot index in the hot identification tree, determine the technical value corresponding to the first-layer hot index as the corresponding count value; in response to the existence of the second-layer hot index corresponding to the first-layer hot index in the hot identification tree, determine the second-layer hot index corresponding to the primary key according to the hash function corresponding to the second layer, and determine the count value corresponding to the second-layer hot index as the corresponding count value.
[0123] In an alternative embodiment, the initialization module 500 is specifically configured to: initialize the count values corresponding to multiple hot indexes included in the first layer of the hot identification tree to 0 in advance. In addition, the apparatus further includes: an expansion module 521 (not shown in the figure), configured to, in response to the count value corresponding to the first-layer hot index after incrementing by 1 exceeding a preset expansion threshold, create multiple second-layer hot indexes corresponding to the first-layer hot index, and initialize the count values corresponding to the multiple second-layer hot indexes to 0.
[0124] In an alternative embodiment, the apparatus further includes: a reset module 523 (not shown in the figure), configured to increment the access request count by 1 when receiving a data access request for any data; and in the case where the access request count is greater than a preset access request threshold, reset the access request count to 0, and set the count values corresponding to multiple hot indexes included in the hot identification tree to 0. In this case, the hot key determination module 530 is specifically configured to determine that the primary key is a hot primary key in response to the count value reaching a preset hot threshold.
[0125] In an alternative embodiment, the hot key determination module 530 is specifically configured to determine that the primary key is a hot primary key in response to the ratio of the count value to the sum of all count values exceeding a preset hot ratio.
[0126] In an alternative embodiment, the device further comprises a recording module 531 (not shown in the figure), configured to record the primary key and the access volume of the primary key when it is determined that the primary key is a hot primary key.
[0127] In an alternative embodiment, the device further comprises a deletion module 532, configured to delete the primary key with the lowest access frequency according to the access volumes of the recorded primary keys when the number of recorded primary keys exceeds a preset primary key quantity threshold.
[0128] As Figure 6 shown, Figure 6 is a block diagram of a current limiting device shown in accordance with an exemplary embodiment of the present specification. The device comprises:
[0129] A request receiving module 610, configured to receive a data access request for target data in a database, where the data access request carries a primary key corresponding to the target data;
[0130] A request blocking module 620, configured to block the data access request for the target data in response to the primary key belonging to a hot primary key; the hot primary key is obtained by the foregoing hot spot identification method.
[0131] For the implementation processes of the functions and roles of each module in the above device, refer specifically to the implementation processes of the corresponding steps in the above method, which will not be elaborated herein.
[0132] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present specification. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0133] As Figure 7 shown, Figure 7 shows a hardware structure diagram of an electronic device where the hot spot identification device or the current limiting device of the embodiment is located. The device may include: a processor 1010, a memory 1020 for storing computer instructions, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0134] The processor 1010 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute computer instructions to implement the above-mentioned hotspot recognition method or traffic limiting method.
[0135] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and called and executed by the processor 1010.
[0136] The input / output interface 1030 is used to connect to an input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0137] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. Among them, the communication module can achieve communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).
[0138] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0139] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0140] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the hotspot recognition method and / or the current limiting method as described above.
[0141] Computer-readable media include both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVDs) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0142] The embodiments of this specification also provide a computer program, which implements the hotspot recognition method or the current limiting method as described above when it runs.
[0143] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent in such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.
[0144] The specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
Claims
1. A hotspot recognition method, which initializes the count values corresponding to multiple hotspot indexes included in a hotspot recognition tree to 0 in advance; the number of hotspot indexes in the hotspot recognition tree is less than the number of data in the database; the hotspot recognition tree is a tree structure with at least one layer, and each layer of the hotspot recognition tree stores at least one hotspot index and its corresponding count value; the method includes: Upon receiving a data access request, determine the primary key of the data accessed by the data access request; Calculate the hot index corresponding to the primary key, and determine the corresponding count value in the hot identification tree according to the hot index; In response to the count value satisfying a preset hot condition, determine that the primary key is a hot primary key; in response to the count value not satisfying the preset hot condition, increment the count value by one.
2. The method according to claim 1, wherein the hotspot recognition tree includes a first layer and a second layer; Calculating the hotspot index corresponding to the primary key and determining the corresponding count value in the hotspot recognition tree according to the hotspot index includes: According to the hash function corresponding to the first layer of the hot identification tree, calculate the first-layer hot index corresponding to the primary key; In response to the non-existence of the second-layer hot index corresponding to the first-layer hot index in the hot identification tree, determine the count value corresponding to the first-layer hot index as the corresponding count value; In response to the existence of the second-layer hot index corresponding to the first-layer hot index in the hot identification tree, determine the second-layer hot index corresponding to the primary key according to the hash function corresponding to the second layer, and determine the count value corresponding to the second-layer hot index as the corresponding count value.
3. The method according to claim 2, wherein initializing the count values corresponding to multiple hotspot indexes included in the hotspot recognition tree to 0 in advance includes: Initialize the count values corresponding to multiple hot indexes included in the first layer of the hot identification tree to 0 in advance; The method further includes: In response to the count value corresponding to the incremented first-layer hot index exceeding a preset expansion threshold, create multiple second-layer hot indexes corresponding to the first-layer hot index, and initialize the count values corresponding to the multiple second-layer hot indexes to 0.
4. The method according to claim 1, the method further includes: Upon receiving a data access request for any data, increment the access request count by one; In the case where the access request count is greater than a preset access request threshold, reset the access request count to 0, and set the count values corresponding to multiple hot indexes included in the hot identification tree to 0; The step of, in response to the count value satisfying a preset hot condition, determining that the primary key is a hot primary key, includes: In response to the count value reaching a preset hot threshold, determine that the primary key is a hot primary key.
5. The method according to claim 1, wherein in response to the count value satisfying a preset hotspot condition, determining that the primary key is a hotspot primary key includes: In response to the proportion of the count value in the sum of all count values exceeding a preset hot proportion, determine that the primary key is a hot primary key.
6. The method according to claim 1, the method further includes: In the case where the primary key is determined to be a hot primary key, record the primary key and record the access volume of the primary key.
7. The method according to claim 6, the method further includes: In the case where the number of recorded primary keys exceeds a preset primary key quantity threshold, delete the primary key with the lowest access frequency according to the access volumes of the recorded primary keys.
8. A flow limiting method, including: Receive a data access request for target data in the database, where the data access request carries the primary key corresponding to the target data; In response to the primary key belonging to a hot primary key, block the data access request for the target data; the hot primary key is obtained by the hot identification method according to any one of claims 1-7.
9. A hotspot recognition device that initializes the count values corresponding to multiple hotspot indexes included in a hotspot recognition tree to 0 in advance; the number of hotspot indexes in the hotspot recognition tree is less than the number of data in the database; the hotspot recognition tree is a tree structure with at least one layer, and each layer of the hotspot recognition tree stores at least one hotspot index and its corresponding count value; the device includes: A primary key determination module, configured to determine the primary key of the data accessed by the data access request upon receiving the data access request; A count value determination module, configured to calculate the hot index corresponding to the primary key, and determine the corresponding count value in the hot identification tree according to the hot index; A hot determination module, configured to determine that the primary key is a hot primary key in response to the count value satisfying a preset hot condition; In response to the count value not satisfying the preset hot condition, increment the count value by one.
10. A current limiting device, including: A request receiving module, configured to receive a data access request for target data in a database, where the data access request carries a primary key corresponding to the target data; A request blocking module, configured to block the data access request for the target data in response to the primary key belonging to a hot primary key; the hot primary key is obtained by the hot spot identification method according to any one of claims 1-7.
11. An electronic device, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor realizes the hot spot identification method according to any one of claims 1-7 or the current limiting method according to claim 8 by running the executable instructions.
12. A computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the hotspot recognition method according to any one of claims 1-7 or the current limiting method according to claim 8 is implemented.
13. A computer program, which when run, implements the hotspot recognition method according to any one of claims 1-7 or the current limiting method according to claim 8.
Citation Information
Patent Citations
Data flow frequency estimation method and system based on double-layer structure and medium
CN111782700A
Traffic control method, cluster resource guarantee method, equipment and storage medium
CN113965519A