Cache pool management method and device, processor and storage medium

By using a unified cache coordination mechanism in a distributed database, the problem of duplicate cached data is solved, resulting in higher cache hit rates and query efficiency.

CN121301412APending Publication Date: 2026-01-09CHINA CONSTRUCTION BANK +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511523080.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

In distributed databases, the cache status of each node is independent, resulting in a lot of duplicate cached data, which affects cache hit rate and query efficiency.

Method used

By setting up a cache coordination mechanism in the distributed database, the caches of different types of database nodes are unified into an organic whole, avoiding duplicate cached data and ensuring that valid, non-duplicate data is cached in the cache pool.

Benefits of technology

This improved cache hit rate, which in turn improved query efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301412A_ABST
    Figure CN121301412A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cache pool management method and device, a processor and a storage medium. The method comprises the following steps: determining target query data corresponding to a query request under the condition of determining that a target service layer node obtains the query request from a client; determining first data and second data stored in a disk of a distributed database on the basis of a cache distribution table of the service layer node under the condition of determining that the target query data is not completely cached in a first cache region; the first data comprises uncached data which is not cached in the first cache region in the target query data, and the first data, the second data and currently cached data in the cache pool are not overlapped; determining a target cache service layer node corresponding to the first data and a target cache storage layer node corresponding to the second data; and caching the first data into a first cache region of the target cache service layer node, caching the second data into a second cache region of the target cache storage layer node, and updating a corresponding cache distribution table.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a cache pool management method, a cache pool management device, a processor and a machine readable storage medium. BACKGROUND

[0002] In a distributed database, each node has its own local cache area. The node usually temporarily stores high-frequency access data in the local cache to reduce the IO overhead of the node local to the disk / remote node, thereby improving the response speed.

[0003] However, the cache states of each node are completely independent and do not know what data the other nodes have cached. This information island cannot fully play the role of the cache, and there are too many repeated data in the cache, resulting in less effective non-repeated data stored in the cache pool, which ultimately affects the cache hit rate and query efficiency. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a cache pool management method, a cache pool management device, a processor and a storage medium.

[0005] In order to achieve the above-mentioned purpose, the first aspect of the present application provides a cache pool management method applied to a distributed database, the distributed database comprising a cache pool and a plurality of database nodes; the database nodes are divided into service layer nodes and storage layer nodes, and the cache pool comprises a first cache area of each service layer node and a second cache area of each storage layer node; the method comprises: In the case where it is determined that a target service layer node obtains a query request from a client, determining target query data corresponding to the query request; Based on the cache distribution table of the service layer node, in the case where it is determined that the target query data is not all cached in the first cache area, determining first data and second data stored in the disk of the distributed database; wherein the first data comprises uncached data in the target query data which is not cached in the first cache area, and the first data, the second data and the current cached data in the cache pool are all non-overlapping; Determining a target cache service layer node corresponding to the first data and a target cache storage layer node corresponding to the second data; Caching the first data into the first cache area of the target cache service layer node and caching the second data into the second cache area of the target cache storage layer node, and updating the corresponding cache distribution table.

[0006] In the embodiments of the present application, the data size of the first data is an integer multiple of the preset cache unit size; determining the first data comprises: determining, in a case where it is determined that the uncached data has been completely cached in the second cache area, a set of all data cached in the cache unit in the second cache area for caching the uncached data as the first data; determining, in a case where it is determined that the uncached data has been partially cached in the second cache area, first sub-data of the uncached data that has been cached in the second cache area and second sub-data of the uncached data that has not been cached in the second cache area, determining a set of all data cached in the cache unit in the second cache area for caching the first sub-data as first to-be-cached data, and determining second to-be-cached data stored in the disk according to a data size of the second sub-data and the preset cache unit size; wherein the second to-be-cached data comprises the second sub-data, and the first data comprises the first to-be-cached data and the second to-be-cached data; determining, in a case where it is determined that the second cache area does not cache any of the uncached data, the first data stored in the disk according to a data size of the uncached data and the preset cache unit size.

[0007] In the embodiments of the present application, determining the target cache service layer node corresponding to the first data comprises: determining a data size of the first data; wherein the data size of the first data is a positive integer multiple of the preset cache unit size; if it is determined that the free space size of the first cache area is greater than or equal to the data size of the first data, determining the target cache service layer node from the service layer node corresponding to the first cache area in which the free space currently exists; if it is determined that the free space size of the first cache area is less than the data size of the first data, determining the target cache service layer node based on a preset eviction policy.

[0008] In the embodiments of the present application, caching the first data to the first cache area of the target cache service layer node comprises: migrating, according to the cache distribution table of the storage layer node, all data in the cache unit in the second cache area for caching the first data to the first cache area of the target cache service layer node in a case where it is determined that the first data has been completely cached in the second cache area; migrating, in a case where it is determined that the first data has been partially cached in the second cache area, all data cached in the cache unit in the second cache area for caching the first to-be-cached data to the first cache area of the target cache service layer node, and caching the second to-be-cached data from the disk to the first cache area of the target cache service layer node; migrating, in a case where it is determined that the second cache area does not cache any of the first data, the first data from the disk to the first cache area of the target cache service layer node.

[0009] In the embodiment of the present application, after determining the target query data corresponding to the query request, the method further comprises: According to the cache distribution table of the service layer node, it is determined whether the target query data is cached in the first cache area of the target service layer node, and if so, the target query data is obtained from the first cache area of the target service layer node; In the case where it is determined that the target query data is not cached in the first cache area of the target service layer node, it is determined whether the target query data is cached in the first cache area of other service layer nodes except the target service layer node in the distributed database, and if so, the target query data is obtained from the first cache area of the other service layer nodes; In the case where it is determined that the target query data is not cached in the first cache area of each service layer node, according to the cache distribution table of the storage layer node, it is determined whether the target query data is cached in the second cache area, and if so, the target query data is obtained from the second cache area; In the case where it is determined that the target query data is not cached in the cache pool, the target query data is obtained from the disk; The target query data obtained is sent to the client through the target service layer node.

[0010] In the embodiment of the present application, the method further comprises: In the case where it is determined that the target service layer node obtains the write request from the client, according to the write request, a target identifier value set corresponding to the write request is determined, and target write data is generated; wherein the target identifier value set corresponds to the target write data; In the case where it is determined that the target identifier value set is recorded in the cache distribution table of the service layer node, a target write database node corresponding to the target write data is determined; wherein the target write database node includes a target write service layer node and a target write storage layer node; The target write data is cached to the first cache area of the target write service layer node and the second cache area of the target write storage layer node, and the corresponding cache distribution table is updated; In the case where it is determined that the target write data has been cached to the second cache area of the target write storage layer node, the target write data is written into the disk, and the target write data is deleted from the second cache area, and the corresponding cache distribution table is updated.

[0011] The second aspect of the present application provides a cache pool management device applied to a distributed database, wherein the distributed database comprises a cache pool and a plurality of database nodes; the database nodes are divided into service layer nodes and storage layer nodes, and the cache pool comprises first cache areas of the service layer nodes and second cache areas of the storage layer nodes; the device comprises: The query data determination module is used to determine the target query data corresponding to the query request when the target service layer node obtains the query request from the client. The module for determining cacheable data is used to determine, based on the cache distribution table of the service layer nodes, the first data and the second data stored on the disk of the distributed database when it is determined that not all of the target query data is cached in the first cache area. The first data includes the uncached data in the target query data that is not cached in the first cache area, and the first data, the second data, and the currently cached data in the cache pool do not overlap. The node to be operated module is used to determine the target cache service layer node corresponding to the first data and the target cache storage layer node corresponding to the second data. The cache pool management module is used to cache the first data in the first cache area of ​​the target cache service layer node, cache the second data in the second cache area of ​​the target cache storage layer node, and update the corresponding cache distribution table.

[0012] A third aspect of this application provides a processor configured to perform the cache pool management method described above.

[0013] A fourth aspect of this application provides a machine-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the cache pool management method described above.

[0014] The fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described cache pool management method.

[0015] The above technical solution, by setting up a cache coordination mechanism among database nodes in a distributed database, unifies the caches of different types of database nodes into an organic whole. This effectively avoids duplication of cached data in the cache pool, thereby enabling the caching of as much valid, non-duplicate data as possible. By caching as much valid, non-duplicate data as possible in the cache pool, the cache hit rate is effectively improved, thus enhancing query efficiency.

[0016] Other features and advantages of the embodiments of this application will be described in detail in the following detailed description section. Attached Figure Description

[0017] The accompanying drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the following detailed description to explain the embodiments of this application, but do not constitute a limitation on the embodiments of this application. In the drawings: Figure 1 This illustration schematically shows an application environment diagram of the cache pool management method according to an embodiment of this application; Figure 2 A flowchart of a cache pool management method according to an embodiment of the present application is shown schematically; Figure 3 A structural block diagram of a cache pool management apparatus according to an embodiment of the present application is shown schematically; Figure 4 An internal structural diagram of a computer device according to an embodiment of the present application is shown schematically.

[0018] Explanation of reference numerals 102 - terminal; 104 - server; 402 - query data determination module; 404 - to-be-cached data determination module; 406 - to-be-operated node determination module; 408 - cache pool management module; A01 - processor; A02 - network interface; A03 - internal memory; A04 - display screen; A05 - input device; A06 - nonvolatile storage medium; B01 - operating system; B02 - computer program. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It should be understood that the specific embodiments described herein are merely used to explain and illustrate the embodiments of the present application, and should not be used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort belong to the scope of protection of the present application.

[0020] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are merely used to explain the relative positional relationship, movement condition, etc. between components in a certain posture (as shown in the drawings), and if the certain posture changes, the directional indications also change accordingly.

[0021] In addition, if the embodiments of the present application involve descriptions such as “first”, “second”, etc., the descriptions of “first”, “second”, etc. are merely for description purposes, and should not be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by “first”, “second” can explicitly or implicitly include at least one of the features. In addition, the technical solutions of the various embodiments can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can implement it, and when the combination of technical solutions appears to be contradictory or unimplementable, it should be considered that the combination of technical solutions does not exist, and is also not within the scope of protection claimed by the present application.

[0022] The acquisition, transmission, storage, use, processing and the like of data in the technical solutions of the present application comply with relevant provisions of national laws and regulations. In addition, it should be noted that in the embodiments of the present application, some industry existing solutions of software, components, models and the like may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solutions of the present application, but does not mean that the applicant has or will necessarily use the solutions.

[0023] The cache pool management method provided by the present application can be applied to an application environment as shown in Figure 1 The terminal 102 communicates with the server 104 through the network. In the case where it is determined that the target service layer node obtains a query request from the client, the method first determines the target query data corresponding to the query request, and then determines the first data and the second data stored in the disk of the distributed database based on the cache distribution table of the service layer node in the case where it is determined that the target query data is not all cached in the first cache area, wherein the first data includes uncached data in the target query data that is not cached in the first cache area, the first data, the second data and the currently cached data in the cache pool do not overlap, then the corresponding target cache service layer node and target cache storage layer node are determined based on the first data and the second data, and finally the first data is cached in the first cache area of the target cache service layer node, the second data is cached in the second cache area of the target cache storage layer node, and the corresponding cache distribution table is updated. The method provided by the present application can effectively avoid the repetition of the cache data in the cache pool by setting the cache area coordination mechanism between the database nodes of the distributed database, and unify the cache areas of different types of database nodes into an organic whole, so as to realize the caching of as much effective non-repetitive data as possible in the cache pool. By realizing the caching of as much effective non-repetitive data as possible in the cache pool, the cache hit rate is effectively improved, and the query efficiency is improved. The terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0024] Figure 2 The flowchart of the cache pool management method according to the embodiments of the present application is schematically shown. As shown in Figure 2 In an embodiment of the present application, a cache pool management method is provided, and the present embodiment mainly applies the method to the application environment as shown in Figure 1The method is applied to a distributed database, and the distributed database comprises a cache pool, a plurality of database nodes and a cache distribution table of each database node. The database nodes are divided into service layer nodes and storage layer nodes, and the cache pool comprises a first cache area of each service layer node and a second cache area of each storage layer node. The cache pool management method provided in the embodiment of the application comprises the following steps.

[0025] In step 200, when it is determined that the target service layer node obtains the query request from the client, the target query data corresponding to the query request is determined.

[0026] It should be understood that the structured query language (SQL) is a de facto standard for interacting with a database, and by executing the SQL, the client can query and operate the data in the database and the properties of the database. The database is a warehouse for organizing, storing and managing data according to a data structure, that is, a collection of a large amount of data stored in a computer system, and has the characteristics of long-term storage, organization, sharing and unified management.

[0027] In the embodiment of the application, when the query request from the client is obtained, the target query data corresponding to the query request can be determined by generating and executing the corresponding SQL.

[0028] In the embodiment of the application, the hierarchical structure of the distributed database comprises a service layer and a storage layer, the service layer comprises a plurality of independently running service layer nodes, and similarly, the storage layer comprises a plurality of independently running storage layer nodes. The service layer nodes are mainly responsible for processing the requests of the client, analyzing the client commands and interacting with the storage layer nodes to obtain data, and returning the combined data to the client; and the storage layer nodes are mainly responsible for storing and organizing data.

[0029] In the embodiment of the application, the remaining memory of the database node except for the memory space necessary for itself to run is added to the cache pool. That is, the memory area in the database node used for storing data to accelerate reading is the cache area of the database node. The cache pool is a total area composed of a plurality of cache areas, and is uniformly managed and allocated by the database node.

[0030] In the embodiment of the present application, the disk of the distributed database stores a plurality of unit data, each unit data has a corresponding identification value, and the cache distribution table records the identification value of the unit data stored in the cache area of the database node and the data structure of the database node. Specifically, the unit data stored in the distributed database is distinguished according to the primary key. For example, the primary key of a certain database table is id, and the value range of id in the table is 1 to 1 billion. For a continuous id value range, it can be understood as an "identification value range", that is, an "identification value set".

[0031] It is worth mentioning that in the embodiment of the present application, the unit data in the distributed database is loaded and stored into the cache pool as an indivisible whole.

[0032] In step 300, based on the cache distribution table of the service layer node, in the case that it is determined that the target query data is not all cached in the first cache area, the first data and the second data stored in the disk of the distributed database are determined.

[0033] The first data includes the uncached data in the target query data that is not cached in the first cache area, and the first data, the second data, and the currently cached data in the cache pool do not overlap.

[0034] It should be understood that the target query data includes at least part of the data contained in at least one unit data. After the target query data is determined by SQL, the cache distribution table of the storage layer node is queried to determine whether the target query data has been all cached in the first cache area. In addition, by querying the corresponding cache distribution table, it can be determined which unit data has been cached in the current first cache area and / or second cache area, and then the first data and the second data are determined.

[0035] It should be noted that for each database node in a layer, if the identification value of all unit data cached in the cache area of the database node in the cache distribution table of each database node in the layer is recorded, then by querying the cache distribution table of any database node in the layer, it can be known which unit data is cached in the cache area of the database node in the layer. If only the identification value of all unit data cached in the cache area of the database node in the cache distribution table of each database node in the layer is recorded, then by querying the cache distribution table of each database node in the layer, it can be known which unit data is cached in the cache area of the database node in the layer.

[0036] In step 400, the target cache service layer node corresponding to the first data and the target cache storage layer node corresponding to the second data are determined.

[0037] It should be noted that in the embodiments of the present application, the target cache service layer node includes at least one service layer node, when the target cache service layer node includes multiple service layer nodes, it is represented that the determined first data is cached through the first cache area of the multiple service layer nodes, that is, the first cache area of each service layer node in the target cache service layer node respectively caches part of the first data.

[0038] Similarly, the target cache storage layer node also includes at least one storage layer node, when the target cache storage layer node includes multiple storage layer nodes, it is represented that the determined second data is cached through the second cache area of the multiple storage layer nodes, that is, the second cache area of each storage layer node in the target cache storage layer node respectively caches part of the second data.

[0039] Step 500, cache the first data into the first cache area of the target cache service layer node, cache the second data into the second cache area of the target cache storage layer node, and update the corresponding cache distribution table.

[0040] In the embodiments of the present application, after the data is cached into the corresponding cache area, the updated cache distribution table includes at least the cache distribution table of the database node whose cache area has changed, that is, it can also include the cache distribution table of all database nodes in the layer where the database node whose cache area has changed is located.

[0041] It can be seen that the cache pool management method provided by the embodiments of the present application unifies the cache areas of different types of database nodes into an organic whole through the cache area coordination mechanism between the database nodes of the distributed database, which can effectively avoid the repetition of cache data in the cache pool, so as to realize the caching of as much effective non-repeated data as possible in the cache pool. By caching as much effective non-repeated data as possible in the cache pool, the query ratio (i.e. cache hit rate) of returning the target query data corresponding to the query request by querying the cache pool is effectively improved, and the query efficiency is improved.

[0042] In an embodiment, the first data and the second data each include a plurality of unit data, and the data size of the first data and the second data is a positive integer multiple of a preset cache unit size. The preset cache unit size is the space size of each cache unit contained in each cache area in the cache pool.

[0043] It should be noted that in this embodiment, when the target cache service layer node includes multiple service layer nodes, the data size of the first data is M (M is a positive integer, and M>1) times of the preset cache unit size, that is, the data size of the data respectively cached by the first cache area of each service layer node in the target cache service layer node is a positive integer times of the preset cache unit size. Similarly, when the target cache storage layer node includes multiple storage layer nodes, the data size of the second data is N (N is a positive integer, and N>1) times of the preset cache unit size, that is, the data size of the data respectively cached by the second cache area of each storage layer node in the target cache storage layer node is a positive integer times of the preset cache unit size.

[0044] It is worth mentioning that the "integer times" mentioned in this embodiment does not require a strict integer times relationship between the data size and the preset cache unit size. When the ratio of the data size to the preset cache unit size is close to an integer (for example, 0.99 times when close to 1 times, 1.98 times when close to 2 times, etc.), it can be determined to meet the integer times requirement. In an ideal scenario, the data size and the preset cache unit size are in a completely integer times relationship.

[0045] In an embodiment, the step 300 of determining the first data comprises the following steps.

[0046] Step 311, in the case where it is determined that the uncached data has been completely cached in the second cache area, determining the set of all data cached by the cache unit in the second cache area for caching the uncached data as the first data.

[0047] Step 312, in the case where it is determined that the uncached data has been partially cached in the second cache area, determining first sub-data in the uncached data that has been cached in the second cache area and second sub-data in the uncached data that has not been cached in the second cache area, determining the set of all data cached by the cache unit in the second cache area for caching the first sub-data as the first to-be-cached data, and determining the second to-be-cached data stored in the disk according to the data size of the second sub-data and the preset cache unit size.

[0048] Among them, the second to-be-cached data includes the second sub-data, and the first data includes the first to-be-cached data and the second to-be-cached data.

[0049] In this embodiment, the step 312 first determines the set of identification values corresponding to the currently cached data in the cache pool based on the cache distribution table, and then determines the second to-be-cached data based on the set of identification values corresponding to the second sub-data using the nearest principle or the correlation principle. Among them, the data size of the second to-be-cached data is a positive integer times of the preset cache unit size, and the currently cached data and the first sub-data in the cache pool are not overlapped with the second to-be-cached data.

[0050] It needs to be understood that, in the distributed database, the principle of proximity refers to that the identification value of the data in the table is close to the reference identification value / reference identification value set (such as the identification value set corresponding to the second sub-data) in the value dimension; the principle of relevance refers to that the data are closely associated in the business logic (such as belonging to the same business process, the same entity), and the association can be determined by the business association field (such as the common business identification) in the table.

[0051] In a specific example, based on the identification value set corresponding to the second sub-data, the specific operation of determining the second to-be-cached data based on the principle of proximity is as follows: It is determined whether the data size of the second sub-data is an integer multiple of the preset cache unit size, if yes, the second sub-data is determined as the second to-be-cached data, if not, the identification value set corresponding to the second sub-data is determined, and based on the identification value set and the principle of proximity, the first identification value set is determined, and the second to-be-cached data is determined according to the first identification value set. The first identification value set includes the identification value set corresponding to the second sub-data, and the first identification value set does not overlap with the identification value recorded in the cache distribution table.

[0052] In step 313, in the case that no un-cached data is cached in the second cache area, the first data stored in the disk is determined according to the data size of the un-cached data and the preset cache unit size.

[0053] In this embodiment, the step 313 first determines the identification value set corresponding to the currently cached data in the cache pool based on the cache distribution table, and then determines the first data based on the principle of proximity or the principle of relevance based on the identification value set corresponding to the un-cached data. The data size of the first data is a positive integer multiple of the preset cache unit size, and the currently cached data in the cache pool does not overlap with the first data.

[0054] In a specific example, in the case that the preset cache unit size is set to 1M and no un-cached data is cached in the second cache area, if the data size of the un-cached data is less than 1M, the un-cached data is expanded, and the first data obtained after the expansion is controlled to be 1M; if the data size of the un-cached data is a positive integer multiple of 1M, the un-cached data is directly determined as the first data; if the data size of the un-cached data is greater than 1M and is not an integer multiple of 1M, the un-cached data is expanded, and the first data obtained after the expansion is controlled to be a positive integer multiple of 1M (in this case, the un-cached data needs to be cached by multiple cache units).

[0055] It can be seen that, in the case that it is determined that the uncached data is all cached in the second cache area, the embodiment directly migrates all data cached in the cache unit where the uncached data is located to the first cache area of the target cache service layer node, effectively avoiding the repetition of cached data in the cache pool, while ensuring high space utilization of the cache unit; in the case that it is determined that the uncached data is not all cached in the second cache area, by filling the cache unit occupied by the data cached from the disk to the target cache service layer node, high space utilization of the occupied cache unit is achieved.

[0056] In an embodiment, the step 300 of determining the second data comprises the following steps.

[0057] Step 321, determining the data size of the second data.

[0058] In a specific example, if it is determined that the uncached data is partially cached in the second cache area, the space size of the cache unit in the second cache area used for caching the first sub-data (i.e., the data in the uncached data that has been cached in the second cache area) is determined as the data size of the second data; if it is determined that the second cache area does not cache any of the uncached data, the data size of the second data is set to be a preset cache unit size or a multiple of the preset cache unit size (at this time, the data size of the second data is a fixed value).

[0059] Step 322, determining the second data stored in the disk according to the data size of the second data.

[0060] In this embodiment, the step 322 first determines a set of identification values corresponding to the currently cached data in the cache pool based on the cache distribution table, and then determines the second data based on the set of identification values corresponding to the target query data using the nearest principle. The second data does not overlap with the currently cached data in the cache pool and the first data.

[0061] It can be seen that, in the case that it is determined that the target query data is not all cached in the first cache area, by filling the cache unit occupied by the data cached from the disk to the target cache storage layer node, high space utilization of the occupied cache unit is achieved.

[0062] In an embodiment, the step 400 of determining the target cache service layer node corresponding to the first data comprises the following steps.

[0063] Step 411, determining the data size of the first data.

[0064] The data size of the first data is a positive integer multiple of a preset cache unit size.

[0065] Step 412, if it is determined that the free space size of the first cache area is greater than or equal to the data size of the first data, a target cache service layer node is determined from the service layer node corresponding to the first cache area in which the free space currently exists; if it is determined that the free space size of the first cache area is less than the data size of the first data, the target cache service layer node is determined based on the preset eviction policy.

[0066] It should be understood that after the service layer node is started, the first cache area thereof will be gradually filled with data through interaction with the client. Similarly, after the storage layer node is started, the second cache area thereof will be gradually filled with data through interaction with the service layer node. In this embodiment, if it is determined that the free space size of the first cache area is greater than or equal to the data size of the first data, the target cache service layer node (i.e., the cache area for caching the first data) is directly selected from the service layer node corresponding to the first cache area in which the free space currently exists, and it is only required that the free space size of the selected target cache service layer node is greater than or equal to the data size of the first data, and the first data is directly cached from the second cache area and / or the disk to the first cache area of the target cache service layer node.

[0067] In this embodiment, the access times and / or the last access times of each cache unit of the database node are recorded in the cache distribution table, and the preset eviction policy adopts the least recently used principle.

[0068] Specifically, when a cache area needs to cache new data, if the cache area is full or the free space of the cache area is insufficient to accommodate the new data, the cache data that is least frequently accessed in the recent period of time or the cache data whose last access time is the earliest will be preferentially evicted to create space, while the data that is likely to be frequently used in the future is preserved as much as possible, so as to improve the utilization rate and the hit rate of the cache area.

[0069] In this embodiment, the specific implementation method of step 400 for determining the target cache storage layer node corresponding to the second data is similar to the specific implementation method of determining the target cache service layer node corresponding to the first data, and will not be described herein again.

[0070] It can be seen that the embodiment formulates the update mechanism of the cache pool based on the preset eviction policy, which can not only ensure that the data to be cached can be completely cached to the corresponding cache area, but also improve the hit rate of the cache pool, thereby improving the query efficiency.

[0071] In an embodiment, step 500 of caching the first data to the first cache area of the target cache service layer node includes the following steps.

[0072] Step 511, according to the cache distribution table of the storage layer node, in the case that it is determined that the first data is completely cached in the second cache area, migrating all data in the cache unit for caching the first data in the second cache area to the first cache area of the target cache service layer node.

[0073] Step 512, in the case that it is determined that the first data is partially cached in the second cache area, migrating all data cached in the cache unit for caching the first data to be cached in the second cache area to the first cache area of the target cache service layer node, and caching the second data to be cached from the disk to the first cache area of the target cache service layer node.

[0074] Step 513, in the case that it is determined that no first data is cached in the second cache area, caching the first data from the disk to the first cache area of the target cache service layer node.

[0075] It should be noted that in this embodiment, after the data migration operation is completed, the migrated data will be deleted from the original location to avoid duplicate caching of data.

[0076] In a specific example, the step 500, after successfully caching the first data to the first cache area of the target cache service layer node and caching the second data to the second cache area of the target cache storage layer node, the specific operation of updating the cache distribution table is as follows: The target cache service layer node records the identification value of the data newly cached to the first cache area of itself in the cache distribution table of itself, and broadcasts a first cache message (the message content at least includes an identification value set corresponding to the operation of caching the first data to the first cache area of the target cache service layer node) in the service layer node, and other service layer nodes receiving the first cache message save the message to the cache distribution table of themselves; It can be seen that, in the case that the first data is completely cached in the second cache area, the embodiment directly migrates it from the second cache area to the cache area of the service layer node, and in the case that the first data is not completely cached in the second cache area, caches the first data from the second cache area and / or the disk to the first cache area, effectively avoiding the existence of duplicate cached data in the cache area of the service layer node and the storage layer node, and ensuring that as much effective non-duplicate data as possible can be cached in the cache pool.

[0077] In an embodiment, after the step 200, the method further includes the following steps.

[0078] Step 610, according to the cache distribution table of the service layer node, determining whether the target query data is cached in the first cache area of the target service layer node, if yes, obtaining the target query data from the first cache area of the target service layer node, and executing step 650.

[0079] Step 620, in the case of determining that the target query data is not cached in the first cache area of the target service layer node, it is determined whether the target query data is cached in the first cache area of other service layer nodes in the distributed database except the target service layer node, if yes, the target query data is obtained from the first cache area of the other service layer nodes, and step 650 is executed.

[0080] Step 630, in the case of determining that the target query data is not cached in the first cache area of each service layer node, it is determined whether the target query data is cached in the second cache area according to the cache distribution table of the storage layer node, if yes, the target query data is obtained from the second cache area, and step 650 is executed.

[0081] Step 640, in the case of determining that the target query data is not cached in the cache pool, the target query data is obtained from the disk, and step 650 is executed.

[0082] Step 650, the target query data obtained is sent to the client through the target service layer node.

[0083] In this embodiment, whether the target query data is cached in the first cache area of other service layer nodes except the target service layer node can be determined by querying the cache distribution table of the target service layer node, or by querying the cache distribution table of other service layer nodes in the distributed database except the target service layer node.

[0084] In a specific example, after the service layer node n1 (i.e. the target service layer node) receives the query request of the client, the cache distribution table is queried. If it is found that the data corresponding to the query request (i.e. the target query data) is cached in the cache area of the service layer node n1, the corresponding data is obtained from the cache area of the service layer node n1 and returned to the client. If it is found that the data corresponding to the query request is cached in the cache area of the service layer node n2, communication is initiated to the service layer node n2 to obtain the data and return to the client. If it is found that no service layer node caches the data, communication is initiated to the storage layer node to obtain the data. If the cache area of the storage layer node n3 has cached the data, the data is directly returned from the cache pool, otherwise the storage layer node follows the normal process, queries the disk data to return the data to the service layer node n1, returns the data sent by the storage layer node to the client through the service layer node n1, caches the data in the cache area of the service layer node, and updates the corresponding cache distribution table.

[0085] It can be seen that, in this embodiment, whether the target query data is cached in the cache pool is judged, and in the case of determining that the target query data is cached in the cache pool, the target query data is directly sent to the client from the cache pool, which can effectively improve the query efficiency.

[0086] In an embodiment, the method further comprises the following steps.

[0087] At step 710, in a case where it is determined that the target service layer node obtains the write request from the client, a target identifier value set corresponding to the write request is determined according to the write request, and target write data is generated.

[0088] The target identifier value set corresponds to the target write data.

[0089] At step 720, in a case where it is determined that the target identifier value set is recorded in the cache distribution table of the service layer node, a target write database node corresponding to the target write data is determined.

[0090] The target write database node comprises a target write service layer node and a target write storage layer node.

[0091] In this embodiment, the target write service layer node is the service layer node in which the identifier value contained in the target identifier value set is recorded in the corresponding cache distribution table; the target write storage layer node is determined according to the data size of the target write data. The specific implementation method of determining the target write storage layer node at step 720 is similar to the specific implementation method of determining the target cache service layer node corresponding to the first data, which will not be described here.

[0092] At step 730, the target write data is cached to the first cache area of the target write service layer node and the second cache area of the target write storage layer node, and the corresponding cache distribution table is updated.

[0093] For example, after the service layer node n1 (i.e., the target service layer node) receives the write request from the client, it queries the cache distribution table. If it is found that the identifier value range (i.e., the target identifier value set) corresponding to the write request is recorded in the cache distribution table of the service layer node n1, the cache area of the service layer node n1 is directly updated according to the content contained in the write request, and a write request is initiated to the storage layer node. If it is found that the cache distribution table of the service layer node n2 records the identifier value range corresponding to the write request, a communication is initiated to the service layer node n2 to update the cache area of the service layer node n2 according to the content contained in the write request, and a write request is initiated to the storage layer node. The storage layer node (i.e., the target write storage layer node) receiving the write request writes the content contained in the write request into the cache area of the storage layer node according to the usual process. It is worth mentioning that the write process is completed only when the two operations of updating the first cache area of the service layer node and the second cache area of the write storage layer node are both completed.

[0094] Step 740, in the case of determining that the target write data has been cached to the second cache area of the target write storage layer node, writing the target write data into the disk, and deleting the target write data from the second cache area, and updating the corresponding cache distribution table.

[0095] In particular, when the target identification value set is not recorded in the cache distribution table of the service layer node, the target write data is written into the disk according to the general process, and the target write service layer node is determined to cache the target write data into the first cache area.

[0096] It can be seen that the embodiment updates the cache areas of the corresponding service layer node and storage layer node when receiving the write request, and finally updates the data in the disk to ensure data persistence. In terms of write efficiency, only when there is a corresponding identification value set in the cache distribution table of the service layer node, the network interaction time is increased, and in other cases, no time is increased, maintaining the write efficiency basically unchanged.

[0097] In an embodiment, the method further comprises the following steps.

[0098] Step 810, according to the cache distribution table, listening to whether the unit data cached to the database node in the disk has changed, if so, updating the changed data to the corresponding cache area.

[0099] It can be seen that the embodiment ensures the accuracy of returning data from the cache pool to the client by updating the data in the cache pool in a timely manner.

[0100] Figure 2 The flowchart of the cache pool management method in an embodiment. It should be understood that although Figure 2 The steps in the flowchart are displayed in sequence according to the direction of the arrow, but these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise stated in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other order. Moreover, Figure 2 At least part of the steps in the flowchart can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed with at least part of other steps or sub-steps or stages of other steps.

[0101] In an embodiment, as Figure 3As shown, a cache pool management apparatus is provided, applied to a distributed database, the distributed database comprising a cache pool and a plurality of database nodes; the database nodes are divided into service layer nodes and storage layer nodes, and the cache pool comprises a first cache area of each service layer node and a second cache area of each storage layer node. The apparatus comprises a query data determination module 402, a data to be cached determination module 404, a node to be operated determination module 406, and a cache pool management module 408, wherein: The query data determination module 402 is configured to determine target query data corresponding to a query request from a client in a case where it is determined that the target service layer node acquires the query request.

[0102] The data to be cached determination module 404 is configured to determine first data and second data stored in a disk of the distributed database based on a cache distribution table of the service layer node in a case where it is determined that the target query data is not all cached in the first cache area. The first data comprises uncached data in the target query data that is not cached in the first cache area, and the first data, the second data, and currently cached data in the cache pool are all non-overlapping.

[0103] The node to be operated determination module 406 is configured to determine a target cache service layer node corresponding to the first data and a target cache storage layer node corresponding to the second data.

[0104] The cache pool management module 408 is configured to cache the first data in the first cache area of the target cache service layer node, cache the second data in the second cache area of the target cache storage layer node, and update the corresponding cache distribution table.

[0105] In an embodiment, the data size of the first data is a positive integer multiple of a preset cache unit size. In this embodiment, the data to be cached determination module 404 comprises a first data determination sub-module, which comprises: A first determination unit is configured to determine, in a case where it is determined that the uncached data is all cached in the second cache area, a set of all data cached by a cache unit in the second cache area for caching the uncached data as the first data.

[0106] A second determination unit is configured to determine, in a case where it is determined that the uncached data is partially cached in the second cache area, first sub-data in the uncached data that is cached in the second cache area and second sub-data in the uncached data that is not cached in the second cache area, determine a set of all data cached by a cache unit in the second cache area for caching the first sub-data as first data to be cached, and determine second data to be cached stored in the disk according to the data size of the second sub-data and the preset cache unit size. The second data to be cached comprises the second sub-data, and the first data comprises the first data to be cached and the second data to be cached.

[0107] The third determining unit is configured to, in a case where it is determined that none of the uncached data is cached in the second cache area, determine the first data stored in the disk according to a data size of the uncached data and a preset cache unit size.

[0108] In an embodiment, the node to be operated determining module 406 comprises a service layer node determining sub-module, which comprises: The data size determining unit is configured to determine a data size of the first data. The data size of the first data is a positive integer multiple of the preset cache unit size.

[0109] The service layer node determining unit is configured to, if it is determined that the free space size of the first cache area is greater than or equal to the data size of the first data, determine the target cache service layer node from the service layer node corresponding to the first cache area in which the free space currently exists; and if it is determined that the free space size of the first cache area is less than the data size of the first data, determine the target cache service layer node based on a preset eviction policy.

[0110] In an embodiment, the cache pool management module 408 comprises a first data caching sub-module, which comprises: The first cache unit is configured to, according to the cache distribution table of the storage layer node, in a case where it is determined that the first data has been completely cached in the second cache area, migrate all data in the cache unit for caching the first data in the second cache area to the first cache area of the target cache service layer node.

[0111] The second cache unit is configured to, in a case where it is determined that the first data has been partially cached in the second cache area, migrate all data cached in the cache unit for caching the first data to be cached in the second cache area to the first cache area of the target cache service layer node, and cache the second data to be cached from the disk to the first cache area of the target cache service layer node.

[0112] The third cache unit is configured to, in a case where it is determined that none of the first data is cached in the second cache area, cache the first data from the disk to the first cache area of the target cache service layer node.

[0113] In an embodiment, the apparatus further comprises a data returning module, which is configured to: determine, according to the cache distribution table of the service layer node, whether the target query data is cached in the first cache area of the target service layer node, and if so, acquire the target query data from the first cache area of the target service layer node; In a case where it is determined that the target query data is not cached in the first cache area of the target service layer node, it is determined whether the target query data is cached in the first cache area of another service layer node in the distributed database except the target service layer node, and if so, the target query data is acquired from the first cache area of the another service layer node. In a case where it is determined that the target query data is not cached in the first cache area of each service layer node, it is determined whether the target query data is cached in the second cache area according to the cache distribution table of the storage layer node, and if so, the target query data is acquired from the second cache area. In a case where it is determined that the target query data is not cached in the cache pool, the target query data is acquired from the disk. The target query data acquired is sent to the client through the target service layer node.

[0114] In an embodiment, the apparatus further includes: The write data determination module is configured to, in a case where it is determined that the target service layer node acquires a write request from the client, determine a target identifier value set corresponding to the write request according to the write request, and generate target write data. The target identifier value set corresponds to the target write data.

[0115] The to-be-written node determination module is configured to, in a case where it is determined that the target identifier value set is recorded in the cache distribution table of the service layer node, determine a target write database node corresponding to the target write data. The target write database node includes a target write service layer node and a target write storage layer node.

[0116] The first data write module is configured to cache the target write data to the first cache area of the target write service layer node and the second cache area of the target write storage layer node, and update the corresponding cache distribution table.

[0117] The second data write module is configured to, in a case where it is determined that the target write data has been cached to the second cache area of the target write storage layer node, write the target write data to the disk, delete the target write data from the second cache area, and update the corresponding cache distribution table.

[0118] The cache pool management apparatus includes a processor and a memory. The cache pool management module, the cache pool management module, the cache pool management module, and the cache pool management module are all stored in the memory as program units, and the corresponding functions are realized by the processor executing the above-mentioned program modules stored in the memory.

[0119] The processor includes a kernel, and the corresponding program units are called from the memory by the kernel. The kernel can be set to one or more, and the cache pool management method is realized by adjusting the kernel parameters.

[0120] The memory can include non-persistent memory in a computer readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory, including at least one memory chip.

[0121] Embodiments of the present application provide a machine readable storage medium, which stores a program, and the program is executed by a processor to implement the cache pool management method.

[0122] Embodiments of the present application provide a processor, which is used to run a program, and the program is executed to implement the cache pool management method.

[0123] In one embodiment, a computer device is provided, which can be a terminal, and an internal structure diagram of the computer device can be as shown in Figure 4 The computer device includes a processor A01, a network interface A02, a display screen A04, an input device A05 and a memory (not shown in the figure) connected through a system bus. The processor A01 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes an internal memory A03 and a non-volatile storage medium A06. The non-volatile storage medium A06 stores an operating system B01 and a computer program B02. The internal memory A03 provides an environment for the operating system B01 and the computer program B02 in the non-volatile storage medium A06 to run. The network interface A02 of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor A01 to implement a cache pool management method. The display screen A04 of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device A05 of the computer device can be a touch layer overlaid on the display screen, or can be a key, trackball or touchpad arranged on the shell of the computer device, or can be an external keyboard, touchpad or mouse, etc.

[0124] Those skilled in the art can understand that Figure 4 the structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0125] In one embodiment, the cache pool management apparatus provided by the present application can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 4 The memory of the computer device can store various program modules constituting the cache pool management apparatus, such as Figure 3The first data determining module and the cache pool management module are shown. The computer program composed of various program modules enables the processor to execute the steps in the cache pool management method of various embodiments of the present application described in the specification.

[0126] Figure 4 The computer device shown can execute the steps of the method by, for example, Figure 3 The query data determining module 402 in the cache pool management apparatus shown executes step 200, the data to be cached determining module 404 executes step 300, the node to be operated determining module 406 executes step 400, and the cache pool management module 408 executes step 500.

[0127] The present application also provides a computer program product adapted to execute the program initialized with the following method steps when executed on a data processing device: In the case where it is determined that the target service layer node obtains the query request from the client, the target query data corresponding to the query request is determined; Based on the cache distribution table of the service layer node, in the case where it is determined that the target query data is not all cached in the first cache area, the first data and the second data stored in the disk of the distributed database are determined; wherein the first data includes the uncached data in the target query data that is not cached in the first cache area, the first data, the second data, and the currently cached data in the cache pool do not overlap; The target cache service layer node corresponding to the first data and the target cache storage layer node corresponding to the second data are determined; The first data is cached in the first cache area of the target cache service layer node, the second data is cached in the second cache area of the target cache storage layer node, and the corresponding cache distribution table is updated.

[0128] The program is applied to a distributed database, and the distributed database includes a cache pool and a plurality of database nodes; the database nodes are divided into service layer nodes and storage layer nodes, and the cache pool includes first cache areas of the service layer nodes and second cache areas of the storage layer nodes.

[0129] In an embodiment, the data size of the first data is a positive integer multiple of the preset cache unit size, and the method for determining the first data includes: In the case where it is determined that the uncached data is all cached in the second cache area, a set of all data cached by the cache unit in the second cache area for caching the uncached data is determined as the first data; determining, in a case where it is determined that the uncached data is partially cached in the second cache area, first sub-data of the uncached data that is cached in the second cache area and second sub-data of the uncached data that is not cached in the second cache area, determining a set of all data cached in a cache unit of the second cache area for caching the first sub-data as first to-be-cached data, and determining second to-be-cached data stored in the disk according to a data size of the second sub-data and the preset cache unit size; wherein the second to-be-cached data includes the second sub-data, and the first data includes the first to-be-cached data and the second to-be-cached data; determining, in a case where it is determined that the second cache area does not cache any of the uncached data, first data stored in the disk according to a data size of the uncached data and the preset cache unit size.

[0130] In an embodiment, the method determines a target cache service layer node corresponding to the first data, including: determining a data size of the first data; wherein the data size of the first data is a positive integer multiple of the preset cache unit size; if it is determined that the free space size of the first cache area is greater than or equal to the data size of the first data, determining the target cache service layer node from the service layer node corresponding to the first cache area in which the free space currently exists; if it is determined that the free space size of the first cache area is less than the data size of the first data, determining the target cache service layer node based on a preset eviction policy.

[0131] In an embodiment, the method caches the first data in the first cache area of the target cache service layer node, including: according to the cache distribution table of the storage layer node, in a case where it is determined that the first data is completely cached in the second cache area, migrating all data in the cache unit of the second cache area for caching the first data to the first cache area of the target cache service layer node; in a case where it is determined that the first data is partially cached in the second cache area, migrating all data cached in the cache unit of the second cache area for caching the first to-be-cached data to the first cache area of the target cache service layer node, and caching the second to-be-cached data from the disk to the first cache area of the target cache service layer node; in a case where it is determined that the second cache area does not cache any of the first data, caching the first data from the disk to the first cache area of the target cache service layer node.

[0132] In an embodiment, after the method determines the target query data corresponding to the query request, the method steps further include: According to the cache distribution table of the service layer node, it is determined whether the target query data is cached in the first cache area of the target service layer node, and if yes, the target query data is obtained from the first cache area of the target service layer node; In a case where it is determined that the target query data is not cached in the first cache area of the target service layer node, it is determined whether the target query data is cached in the first cache area of another service layer node except the target service layer node in the distributed database, and if yes, the target query data is obtained from the first cache area of the other service layer node; In a case where it is determined that the target query data is not cached in the first cache area of each service layer node, according to the cache distribution table of the storage layer node, it is determined whether the target query data is cached in the second cache area, and if yes, the target query data is obtained from the second cache area; In a case where it is determined that the target query data is not cached in the cache pool, the target query data is obtained from the disk; The target query data obtained is sent to the client through the target service layer node.

[0133] In an embodiment, the method further includes: In a case where it is determined that the target service layer node obtains the write request from the client, according to the write request, a target identifier value set corresponding to the write request is determined, and target write data is generated; wherein the target identifier value set corresponds to the target write data; In a case where it is determined that the target identifier value set is recorded in the cache distribution table of the service layer node, a target write database node corresponding to the target write data is determined; wherein the target write database node includes a target write service layer node and a target write storage layer node; The target write data is cached to the first cache area of the target write service layer node and the second cache area of the target write storage layer node, and the corresponding cache distribution table is updated; In a case where it is determined that the target write data has been cached to the second cache area of the target write storage layer node, the target write data is written into the disk, and the target write data is deleted from the second cache area, and the corresponding cache distribution table is updated.

[0134] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0135] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0136] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0137] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks.

[0138] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0139] The memory can include non-persistent memory, Random Access Memory (RAM), and / or non-volatile memory, e.g., Read Only Memory (ROM) or flash memory, among others in a computer readable medium. The memory is an example of computer readable media.

[0140] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.

[0141] It should also be noted that the terms "comprising", "containing", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article or apparatus that includes the element.

[0142] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of claims of the present application.

Claims

1. A cache pool management method characterized by, The method is applied to a distributed database, the distributed database comprising a cache pool and a plurality of database nodes; the database nodes are divided into service layer nodes and storage layer nodes, and the cache pool comprises a first cache area of each service layer node and a second cache area of each storage layer node; the method comprises: In a case where it is determined that a target service layer node obtains a query request from a client, determining target query data corresponding to the query request; In a case where it is determined that the target query data is not all cached in the first cache area, determining first data and second data stored in a disk of the distributed database based on a cache distribution table of the service layer nodes; wherein the first data comprises uncached data in the target query data which is not cached in the first cache area, and the first data, the second data and currently cached data in the cache pool are all non-overlapping; Determining a target cache service layer node corresponding to the first data and a target cache storage layer node corresponding to the second data; Caching the first data into the first cache area of the target cache service layer node and caching the second data into the second cache area of the target cache storage layer node, and updating the corresponding cache distribution table.

2. The method of claim 1, wherein, The data size of the first data is a positive integer multiple of a preset cache unit size; the first data is determined by: In a case where it is determined that the uncached data is all cached in the second cache area, determining a set of all data cached by a cache unit in the second cache area for caching the uncached data as the first data; In a case where it is determined that the uncached data is partially cached in the second cache area, determining first sub-data in the uncached data which is cached in the second cache area and second sub-data in the uncached data which is not cached in the second cache area, determining a set of all data cached by a cache unit in the second cache area for caching the first sub-data as first to-be-cached data, and determining second to-be-cached data stored in the disk according to the data size of the second sub-data and the preset cache unit size; wherein the second to-be-cached data comprises the second sub-data, and the first data comprises the first to-be-cached data and the second to-be-cached data; In a case where it is determined that none of the uncached data is cached in the second cache area, determining the first data stored in the disk according to the data size of the uncached data and the preset cache unit size.

3. The method of claim 1, wherein, The target cache service layer node corresponding to the first data is determined by: Determining the data size of the first data; wherein the data size of the first data is a positive integer multiple of a preset cache unit size; If it is determined that the idle space size of the first cache area is greater than or equal to the data size of the first data, the target cache service layer node is determined from the service layer nodes corresponding to the first cache areas which currently have idle space; If it is determined that the idle space size of the first cache area is less than the data size of the first data, the target cache service layer node is determined based on a preset eviction policy.

4. The method of claim 2, wherein, The first data is cached into the first cache area of the target cache service layer node by: In a case where it is determined that the first data is completely cached in the second cache area, all data in the cache unit for caching the first data in the second cache area is migrated to the first cache area of the target cache service layer node according to the cache distribution table of the storage layer node; In a case where it is determined that the first data is partially cached in the second cache area, all data cached in the cache unit for caching the first data to be cached in the second cache area is migrated to the first cache area of the target cache service layer node, and the second data to be cached is cached from the disk to the first cache area of the target cache service layer node; In a case where it is determined that no first data is cached in the second cache area, the first data is cached from the disk to the first cache area of the target cache service layer node.

5. The method of claim 1, wherein, After determining the target query data corresponding to the query request, the method further comprises: According to the cache distribution table of the service layer node, it is determined whether the target query data is cached in the first cache area of the target service layer node, and if so, the target query data is obtained from the first cache area of the target service layer node; In a case where it is determined that the target query data is not cached in the first cache area of the target service layer node, it is determined whether the target query data is cached in the first cache area of other service layer nodes except the target service layer node in the distributed database, and if so, the target query data is obtained from the first cache area of the other service layer nodes; In a case where it is determined that the target query data is not cached in the first cache area of each service layer node, it is determined whether the target query data is cached in the second cache area according to the cache distribution table of the storage layer node, and if so, the target query data is obtained from the second cache area; In a case where it is determined that the target query data is not cached in the cache pool, the target query data is obtained from the disk; The target query data obtained is sent to the client through the target service layer node.

6. The method of claim 1, wherein, The method further comprises: In a case where it is determined that the target service layer node obtains a write request from the client, a target identifier value set corresponding to the write request is determined according to the write request, and target write data is generated; wherein the target identifier value set corresponds to the target write data; In a case where it is determined that the target identifier value set is recorded in the cache distribution table of the service layer node, a target write database node corresponding to the target write data is determined; wherein the target write database node includes a target write service layer node and a target write storage layer node; The target write data is cached to the first cache area of the target write service layer node and the second cache area of the target write storage layer node, and the corresponding cache distribution table is updated; In a case where it is determined that the target write data is cached to the second cache area of the target write storage layer node, the target write data is written to the disk, and the target write data is deleted from the second cache area, and the corresponding cache distribution table is updated.

7. A cache pool management apparatus characterized by comprising: The distributed database comprises a cache pool and a plurality of database nodes; the database nodes are divided into service layer nodes and storage layer nodes, and the cache pool comprises first cache areas of the service layer nodes and second cache areas of the storage layer nodes; the apparatus comprises: In a case where it is determined that the first data is completely cached in the second cache area, all data in the cache unit for caching the first data in the second cache area is migrated to the first cache area of the target cache service layer node according to the cache distribution table of the storage layer node; In a case where it is determined that the first data is partially cached in the second cache area, all data cached in the cache unit for caching the first data to be cached in the second cache area is migrated to the first cache area of the target cache service layer node, and the second data to be cached is cached from the disk to the first cache area of the target cache service layer node; In a case where it is determined that no first data is cached in the second cache area, the first data is cached from the disk to the first cache area of the target cache service layer node. After determining the target query data corresponding to the query request, the method further comprises: According to the cache distribution table of the service layer node, it is determined whether the target query data is cached in the first cache area of the target service layer node, and if so, the target query data is obtained from the first cache area of the target service layer node; In a case where it is determined that the target query data is not cached in the first cache area of the target service layer node, it is determined whether the target query data is cached in the first cache area of other service layer nodes except the target service layer node in the distributed database, and if so, the target query data is obtained from the first cache area of the other service layer nodes; In a case where it is determined that the target query data is not cached in the first cache area of each service layer node, it is determined whether the target query data is cached in the second cache area according to the cache distribution table of the storage layer node, and if so, the target query data is obtained from the second cache area; In a case where it is determined that the target query data is not cached in the cache pool, the target query data is obtained from the disk; The target query data obtained is sent to the client through the target service layer node. The method further comprises: In a case where it is determined that the target service layer node obtains a write request from the client, a target identifier value set corresponding to the write request is determined according to the write request, and target write data is generated; wherein the target identifier value set corresponds to the target write data; In a case where it is determined that the target identifier value set is recorded in the cache distribution table of the service layer node, a target write database node corresponding to the target write data is determined; wherein the target write database node includes a target write service layer node and a target write storage layer node; The target write data is cached to the first cache area of the target write service layer node and the second cache area of the target write storage layer node, and the corresponding cache distribution table is updated; In a case where it is determined that the target write data is cached to the second cache area of the target write storage layer node, the target write data is written to the disk, and the target write data is deleted from the second cache area, and the corresponding cache distribution table is updated. The query data determination module is configured to determine target query data corresponding to the query request from the client when it is determined that the target service layer node obtains the query request from the client. The to-be-cached data determination module is configured to determine first data and second data stored in a disk of the distributed database based on a cache distribution table of the service layer node when it is determined that the target query data is not all cached in the first cache area; the first data includes uncached data in the target query data that is not cached in the first cache area, and the first data, the second data, and currently cached data in the cache pool do not overlap. The to-be-operated node determination module is configured to determine a target cache service layer node corresponding to the first data and a target cache storage layer node corresponding to the second data. The cache pool management module is configured to cache the first data into the first cache area of the target cache service layer node, cache the second data into the second cache area of the target cache storage layer node, and update the corresponding cache distribution table.

8. A processor, comprising: The computer program is configured to implement the cache pool management method according to any one of claims 1 to 6 when executed by a processor.

9. A machine-readable storage medium having stored thereon instructions, the instructions being executable by a machine to cause the machine to perform operations comprising: The computer program is configured to implement the cache pool management method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product comprising a computer program, characterized in that, The computer program is configured to implement the cache pool management method according to any one of claims 1 to 6 when executed by a processor.