Data reading method and device, electronic equipment, storage medium and computer program product

By using preset information sets in the database cluster to manage data within the nodes and identifying the target nodes to read the target data, the inefficient reading problem caused by node independence in the database cluster is solved, and the data reading efficiency is significantly improved.

CN119938759APending Publication Date: 2025-05-06CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411999653.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Each node in the database cluster is relatively independent and cannot form a cluster effect, resulting in low reading efficiency.

Method used

The permanent and non-resident data of each node are controlled by preset information sets, and the target node is determined based on the target identification to read the target data.

Benefits of technology

The island state between each node is broken, the correlation between nodes is established, the cache ratio and cache hit rate of user data are improved, and the data reading efficiency is greatly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938759A_ABST
    Figure CN119938759A_ABST
Patent Text Reader

Abstract

The invention provides a data reading method and device, electronic equipment, a storage medium and a computer program product, and relates to the technical field of databases, and the method comprises the steps that a user request is acquired; wherein the user request comprises a target identifier of the target data; determining a target node in a preset information set based on the target identifier; wherein the preset information set is used for managing and controlling the resident data and the non-resident data in each node, and is also used for explaining the distribution condition of the data in each node; and forwarding the user request to the target node to read the target data. Through the scheme in the embodiment of the invention, the data reading efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of database technology, and in particular to a data reading method, device, electronic device, storage medium and computer program product. Background Art

[0002] In the related technologies, with the rapid development of the information technology industry, massive amounts of data are being generated in various fields. In today's information age, high-speed data access has become an inevitable problem that must be solved by the database system (one master node and multiple standby nodes). This requires the deployment of a database cluster to provide multi-node read services to meet the demand for high-speed data access. The deployment architecture of the database cluster usually adopts a one-master-multiple-standby architecture, where the master node undertakes read and write requests, and the standby node undertakes read requests. The master and standby nodes synchronize data through streaming replication, and each node stores a full copy of the data to provide independent services. Each node in the database cluster provides read requests relatively independently, without interaction with each other, and cannot form a cluster effect due to their independence. The performance improvement is limited, resulting in low reading efficiency. Summary of the invention

[0003] A data reading method, device, electronic device, storage medium and computer program product provided in the embodiments of the present application can improve data reading efficiency.

[0004] The technical solution of this application is implemented as follows:

[0005] The present application provides a data reading method, including:

[0006] Obtaining a user request; wherein the user request includes: a target identifier of the target data;

[0007] Determine the target node in a preset information set based on the target identifier; wherein the preset information set is used to manage the resident data and non-resident data in each node, and is also used to describe the distribution of data in each node;

[0008] The user request is forwarded to the target node to read the target data.

[0009] In the above scheme, the preset information set includes: a first information set and a second information set; the first information set is used to manage the resident data and the non-resident data in each of the nodes; the second information set is used to describe the distribution of data in each of the nodes;

[0010] The first information set includes: a first mapping relationship between a first identifier of each data and a second identifier of a storage database, a second mapping relationship between each first identifier and a third identifier of a storage node, a third mapping relationship between each first identifier and a first mark of the resident data or a second mark of the non-resident data, and a fourth mapping relationship for determining whether the data mark corresponding to each first identifier is effective;

[0011] The second information set includes: a fifth mapping relationship between the first identifier of each data and the third identifier of the storage node.

[0012] In the above scheme, the method further comprises:

[0013] In response to a configuration operation by a user, forming the first information set;

[0014] Each of the nodes periodically reads the first information set so that each of the nodes can manage the resident data and the non-resident data in the node based on the first information set.

[0015] In the above solution, each of the nodes performs management and control of the resident data in the node based on the first information set, including:

[0016] The node traverses the first identifier of each of the data in the first information set, and if the node traverses and determines that the first identifier of the first data corresponds to the third identifier of the node, determines a mark validity detection result based on the fourth mapping relationship corresponding to the first identifier; wherein the mark validity detection result is used to indicate whether the mark for the first data is valid;

[0017] The node completes the marking of the first data based on the marking effectiveness detection result, and traverses the next first identifier until the traversal of the first identifier for each of the data is completed, thereby realizing the management and control of the resident data in the node.

[0018] In the above solution, the step of completing the marking of the first data based on the marking effectiveness detection result includes:

[0019] If the mark validity detection result indicates that the mark for the first data is not valid, the node completes the marking for the first data based on the first mark corresponding to the first identifier in the third mapping relationship, or the second mark.

[0020] In the above scheme, the method further comprises:

[0021] Each of the nodes periodically scans each of the cached data, and updates the local second information set based on the correspondence between the first identifier of each of the cached data and the third identifier of the node;

[0022] Each of the nodes broadcasts the updated second information set to other nodes, so that other nodes can update the corresponding second information set.

[0023] In the above solution, the step of determining the target node in the preset information set based on the target identifier includes:

[0024] Based on the target identifier, query the third mapping relationship in the first information set to determine a query result; wherein the query result is used to indicate whether the target data is the resident data;

[0025] If the query result indicates that the target data is the resident data, a target node having the largest storage proportion for the target data is determined in the second mapping relationship based on the target identifier.

[0026] In the above solution, after searching the first information set based on the target identifier and determining the query result, the method further includes:

[0027] If the query result indicates that the target data is the non-resident data, a target node having the largest storage proportion for the target data is determined in the second information set based on the target identifier.

[0028] In the above scheme, the method further comprises:

[0029] Each of the nodes dynamically eliminates the non-resident data stored therein.

[0030] The present application also provides a data reading device, including:

[0031] A request acquisition unit, used to acquire a user request; wherein the user request includes: a target identifier of the target data;

[0032] A determination unit, configured to determine a target node in a preset information set based on the target identifier; wherein the preset information set is used to manage resident data and non-resident data in each node, and is also used to describe the distribution of data in each node;

[0033] The forwarding unit is used to forward the user request to the target node to read the target data.

[0034] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps in the above method when executing the computer program.

[0035] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.

[0036] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps in the above method when executed by a processor.

[0037] In the embodiment of the present application, a user request is obtained; wherein the user request includes: a target identifier of the target data; a target node is determined in a preset information set based on the target identifier; wherein the preset information set is used to control the resident data and non-resident data in each node, and is also used to explain the distribution of data in each node; the user request is forwarded to the target node to read the target data. In this way, by controlling the resident data and non-resident data of each node through the preset information set, and explaining the data distribution of each node, the island state between each node can be broken, and the association between each node can be established. From the perspective of the entire database cluster, the cache ratio of user data is greatly improved, the cache hit rate when the user accesses can be improved, and the data reading efficiency can be greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 An optional flow chart of a data reading method provided in an embodiment of the present application;

[0039] Figure 2 An optional effect schematic diagram of the data reading method provided in an embodiment of the present application;

[0040] Figure 3 An optional flow chart of a data reading method provided in an embodiment of the present application;

[0041] Figure 4 An optional flow chart of a data reading method provided in an embodiment of the present application;

[0042] Figure 5 An optional flow chart of a data reading method provided in an embodiment of the present application;

[0043] Figure 6 An optional flow chart of a data reading method provided in an embodiment of the present application;

[0044] Figure 7 An optional flow chart of a data reading method provided in an embodiment of the present application;

[0045] Figure 8 An optional effect schematic diagram of the data reading method provided in an embodiment of the present application;

[0046] Fig. 9 An optional flow chart of a data reading method provided in an embodiment of the present application;

[0047] Fig.10 An optional flow chart of a data reading method provided in an embodiment of the present application;

[0048] Fig.11 A schematic diagram of the structure of a data reading device provided in an embodiment of the present application;

[0049] Fig.12 A hardware entity schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions of the present application are further elaborated in detail below in conjunction with the drawings and embodiments. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0051] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0052] If similar descriptions of "first / second" appear in the application documents, the following instructions are added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0054] In related technologies, with the rapid development of the information technology industry, massive amounts of data are being generated in various fields. In today's information age, high-speed data access has become an inevitable problem that must be solved by database systems. This requires the deployment of database clusters to provide multi-node read services to meet the demand for high-speed data access.

[0055] The deployment architecture of a database cluster usually adopts a one-master-multiple-standby architecture. The master node accepts read and write requests, and the standby node accepts read requests. Data is synchronized between the master and the standby through streaming replication. Each node stores a copy of the full data and provides independent services. When each database node starts, it will apply for a section of memory to cache hot data to accelerate data reading. This part of memory is called shared_buffers. The data stored in the database node is much larger than the size of shared_buffers, so the hot data in shared_buffers changes dynamically. The least recently used algorithm (Least Recently Used, LRU) and its variants are usually used for data elimination. The business system accesses the cluster to read data through the middleware, and usually uses random polling to access it to ensure that the number of connections to each node is in a balanced state.

[0056] If user requests can be classified by time or other dimensions, you can also organize business data by using separate databases and tables, and distribute the data originally stored in one database to multiple databases or tables for storage and query. And forward requests by time or other dimensions through middleware, so that different instances can process different data.

[0057] However, each node in the database cluster (one master node and multiple backup nodes) provides read requests relatively independently, without interaction with each other. They are independent of each other and cannot form a cluster effect, so the performance improvement is limited. The shared_buffers memory space of each database node is limited. The LRU (Least Recently Used) algorithm and its variant algorithms are used to eliminate hot data. All types of data may be eliminated from the memory, which is not conducive to more efficient management of data stored in the shared_buffers memory space. Once the accessed data does not hit the shared_buffers memory, the reading speed will be significantly affected. Sub-library and table sharding is a physical means that can only logically control the distribution rules of data, but cannot control the distribution of data in shared_buffers. In addition, the rules for sub-library and table sharding will generally not change after they are set. The applicable scenarios are relatively limited and lack flexibility.

[0058] In order to solve the above technical problems, the present application embodiment provides a data reading method, see Figure 1 , is an optional flow chart of the data reading method provided in the embodiment of the present application, which will be combined with Figure 1 The steps shown are explained:

[0059] S101. Obtain a user request; wherein the user request includes: a target identifier of target data.

[0060] In the embodiment of the present application, the reading device obtains a user request sent by other devices or a client, wherein the user request includes: a target identifier of the target data.

[0061] In the embodiment of the present application, the reading device may be a master node in the database system, or may be a server or a cloud server for managing the database system.

[0062] S102. Determine a target node in a preset information set based on the target identifier; wherein the preset information set is used to manage resident data and non-resident data in each node, and is also used to illustrate the distribution of data in each node.

[0063] In an embodiment of the present application, the reading device can be pre-configured to form a preset information set. The preset information set is used to manage the resident data and non-resident data in each node, and is also used to illustrate the distribution of data in each of the nodes. After receiving a user request, the reading device can determine whether the target data is resident data based on the target identifier in combination with the preset information set. If the target data is resident data, the target node with the largest proportion of target data storage is determined in the preset information set. If the target data is not a non-resident number, the target node with the largest proportion of target data storage can also be determined in the preset information set. Among them, the data in each node includes resident data and non-resident data.

[0064] In the embodiment of the present application, the preset information set includes: a first information set and a second information set; the first information set is used to manage the resident data and the non-resident data in each of the nodes; the second information set is used to describe the distribution of data in each of the nodes;

[0065] The first information set includes: a first mapping relationship between a first identifier of each data and a second identifier of a storage database, a second mapping relationship between each first identifier and a third identifier of a storage node, a third mapping relationship between each first identifier and a first mark of the resident data or a second mark of the non-resident data, and a fourth mapping relationship for determining whether the data mark corresponding to each first identifier is effective;

[0066] The second information set includes: a fifth mapping relationship between the first identifier of each data and the third identifier of the storage node.

[0067] In the embodiment of the present application, when the target data is resident data, the reading device can determine the target node with the largest storage ratio for the target data in the first information set based on the target identifier. When the target data is not resident data, the reading device can determine the target node with the largest storage ratio for the target data in the second information set.

[0068] In an embodiment of the present application, each of the nodes in the database system can read the first information set, and each node marks resident data and non-resident data based on the first information set. Each node can also dynamically eliminate the stored non-resident data. Exemplarily, each node can use LRU and its variant algorithms to eliminate hot data for data stored in the shared_buffers cache space.

[0069] Combination Figure 2 , the database cluster system adopts a master-slave architecture and may include a master node and slave nodes 1, 2, and 3. Data synchronization is performed between the master and slave nodes through stream replication. Each node has an independent shared_buffers cache space for caching user data. The cached data is managed in two ways, namely dynamic planning and resident memory. The dynamic planning management method refers to the use of LRU and its variant algorithms to eliminate hot data. The resident memory management method refers to specifying that part of the data resides in the shared_buffers cache space by setting a rule table (first information set), which is not affected by the elimination algorithm and only changes with the change of the rules. In addition to the shared_buffers cache space, each node also maintains a global cache space memory map hash table (GSMM: global shared_buffersmemory map) (second information set) in memory, which is used to record the metadata information of all nodes cached data in the shared_buffers cache space. The key of the hash table is the page identifier, and the value is the node identifier set that caches this page.

[0070] S103: Forward the user request to the target node to read the target data.

[0071] In the embodiment of the present application, the reading device forwards the user request to the determined target node, so that the target node searches for the corresponding target data based on the target identifier and feeds back the target data to the user.

[0072] In an embodiment of the present application, a user request is obtained; wherein the user request includes: a target identifier of the target data; based on the target identifier, the target node with the largest storage ratio for the target data is determined in the preset information set; wherein the preset information set is used to control the resident data and non-resident data in each node, and is also used to explain the distribution of data in each node; the user request is forwarded to the target node to read the target data. In this way, by controlling the resident data and non-resident data of each node through the preset information set, and explaining the data distribution of each node, the island state between each node can be broken, and the association between each node can be established. From the perspective of the entire database cluster, the cache ratio of user data is greatly improved, the cache hit rate when the user accesses can be improved, and the data reading efficiency can be greatly improved.

[0073] See also Figure 3 , is an optional flow chart of the data reading method provided in the embodiment of the present application, which will be described in combination with the steps:

[0074] S201. In response to a configuration operation by a user, form the first information set.

[0075] In the embodiment of the present application, the reading device may form a first information set based on the human-computer interaction device responding to the user's configuration operation. The first information set is used to manage the resident data and the non-resident data in each of the nodes.

[0076] Each standby node in the database cluster system manages the resident data and non-resident data in the shared_buffers cache through the second information set (hot_data_rules table). The hot_data_rules table design is shown in Table 1, where the databasename and tablename fields are used to specify the first mapping relationship, the nodename is used to specify the second mapping relationship, the iscache field is used to specify the third mapping relationship, and the isset field is used to specify the fourth mapping relationship.

[0077] Column Tуre Nullable databasename name not null tablename name not null nodename name not null iscache boolean not null isset boolean

[0078] Table 1

[0079] S202. Each of the nodes periodically reads the first information set, so that each of the nodes can manage the resident data and the non-resident data in the node based on the first information set.

[0080] In the embodiment of the present application, each standby node can periodically read the first information set from the master node, and then mark the data in the node as resident data or non-resident data according to the mapping relationship corresponding to the first identifier of each data in the first information set.

[0081] In an embodiment of the present application, each backup node can search for the first identifier of the data stored in this node in the second mapping relationship in the first information set, and then mark the data corresponding to the first identifier as resident data or non-resident data based on the third mapping relationship and the fourth mapping relationship corresponding to the found first identifier.

[0082] In the embodiment of the present application, a first information set is formed in response to the user's configuration operation. Each node periodically reads the first information set so that each node can manage and control the resident data and non-resident data in the node based on the first information set. In this way, the data in each spare node can be marked and divided into resident data and non-resident data based on the first information, eliminating the problem of mistakenly eliminating important data in the related technology, so that each spare node can manage the stored data more effectively, and can make the main node hit the spare node faster for user requests, which is convenient for hitting the spare node to read data and improves data reading efficiency.

[0083] See also Figure 4 , is an optional flow chart of a data reading method provided in an embodiment of the present application, Figure 3 S202 in the embodiment can also be implemented through S301 to S302, which will be described in combination with the steps:

[0084] S301. The node traverses the first identifier of each data in the first information set. If the node traverses and determines that the first identifier of the first data corresponds to the third identifier of the node, the mark effectiveness detection result is determined based on the fourth mapping relationship corresponding to the first identifier.

[0085] In the embodiment of the present application, the backup node traverses each first identifier in the second mapping relationship in the first information set. If the backup node determines that the first identifier of the first data corresponds to the third identifier of the node, it determines whether the mark for the first data is effective based on the fourth mapping relationship corresponding to the first identifier of the first data.

[0086] The mark validity detection result is used to indicate whether the mark for the first data is valid.

[0087] S302. The node completes the marking of the first data based on the marking effectiveness detection result, and traverses the next first identifier until the traversal of the first identifier for each data is completed, thereby realizing the management and control of the resident data and the non-resident data in the node.

[0088] In an embodiment of the present application, if the result of the mark effectiveness detection indicates that the mark for the first data is effective, each standby node performs a traversal of the next first identifier until the traversal of the first identifier for each of the data is completed, thereby realizing the control of the resident data and the non-resident data in the node. If the result of the mark effectiveness detection indicates that the mark for the first data is not effective, each standby node completes the marking of the first data based on the first identifier in the third mapping relationship corresponding to the first identifier of the first data, or the second identifier, and performs a traversal of the next first identifier until the traversal of the first identifier for each of the data is completed, thereby realizing the control of the resident data and the non-resident data in the node.

[0089] If the first data is marked with a first tag, it indicates that the first data is resident data. If the first data is marked with a second tag, it indicates that the first data is non-resident data.

[0090] In the embodiment of the present application, the solution for each standby node to perform resident data and non-resident data management based on the first information set can be combined with Table 1 for reference. Figure 5 The solution in the example will be described in combination with the steps:

[0091] S11. Traverse the hot_data_rules table row by row.

[0092] In the embodiment of the present application, each standby node traverses each row in Table 1 row by row.

[0093] S12: Is there any data with nodename being the node?

[0094] In the embodiment of the present application, the standby node searches the second mapping relationship in the row corresponding to nodename to see whether data of the first identifier corresponding to the third identifier of the node exists. If so, S13 is executed.

[0095] S13: Check whether the mark is effective.

[0096] In the embodiment of the present application, the standby node can determine whether the mark of the first data is valid according to the isset field in the fourth mapping relationship corresponding to the first data. If it is valid, S11 is executed. If it is not valid, S14 is executed.

[0097] S14, iscache mark whether it is resident.

[0098] In the embodiment of the present application, the standby node can determine whether the mark of the first data is resident data according to the iscache field in the third mapping relationship corresponding to the first identifier of the first data. If yes, execute S15. If no, execute S16.

[0099] S15. Read the data into the cache and mark it as resident.

[0100] In the embodiment of the present application, if the first data is marked as resident data, the first data is stored in the shared_buffers cache and marked as non-eliminable.

[0101] S16. Mark the first data in the cache as non-resident.

[0102] In the embodiment of the present application, if the first data is marked as non-resident data, the first data is marked as being obsolete.

[0103] S17. Set the isset value of this row in hot_data_rules to true.

[0104] In the embodiment of the present application, after the marking of the first data is completed, the isset in the fourth mapping relationship corresponding to the first identifier of the first data in Table 1 can be modified to true, indicating that the marking is effective.

[0105] In an embodiment of the present application, the first information set (hot_data_rules table) is set manually. Each standby node can perceive the changes in the rules in the rule table hot_data_rules. Each standby node periodically reads the rule table hot_data_rules (every T seconds) and decides whether to cache and cancel the cache based on the iscache field and the isset field. If the table data needs to be cached, the table data is read into the shared_buffers cache and marked as non-eliminable. If the table data resident cache needs to be canceled, the table data in the shared_buffers cache is marked as eliminable. This part of the data will later participate in the elimination process in the dynamic planning module. When setting rules, follow two principles: 1. Set the tables that need to be accessed frequently to resident memory. 2. Evenly set the tables that need to be resident in memory according to the size of the cluster, and try to ensure that the amount of data cached by each node is in a balanced state. For data in the shared_buffers cache that is not marked as non-eliminable, dynamic programming is used for management. The so-called dynamic programming is based on the LRU (least recently used algorithm). When the shared_buffers cache space is full, the data pages that have not been used recently or are least used are eliminated first to free up cache space for reading new data from the disk.

[0106] In an embodiment of the present application, the node traverses the first identifier of each data in the first information set. If the node traverses and determines that the first identifier of the first data corresponds to the third identifier of the node, then based on the fourth mapping relationship corresponding to the first identifier, the mark effectiveness detection result is determined. Based on the mark effectiveness detection result, the marking of the first data is completed, and the next first identifier is traversed until the traversal of the first identifier of each data is completed, thereby realizing the management and control of resident data and non-resident data in the node. In this way, the data in each backup node can be marked and divided into resident data and non-resident data based on the first information, eliminating the problem of mistakenly eliminating important data in the related technology, so that each backup node can manage the stored data more effectively, and the main node can hit the backup node faster for user requests, which is convenient for hitting the backup node to read data and improves data reading efficiency.

[0107] See also Figure 6 , is an optional flow chart of the data reading method provided in the embodiment of the present application, which will be described in combination with the steps:

[0108] S401: Each of the nodes periodically scans each cached data, and updates the local second information set based on the correspondence between the first identifier of each cached data and the third identifier of the node.

[0109] In the embodiment of the present application, the standby node periodically scans the first identifier of each cached data, establishes a corresponding relationship between each first identifier obtained by scanning and a local third identifier, and uses the corresponding relationship to update the local second information set.

[0110] S402: Each of the nodes broadcasts the updated second information set to other nodes, so that other nodes can update corresponding second information sets.

[0111] In the embodiment of the present application, each standby node broadcasts the updated second information set to other nodes, so that other nodes can update the corresponding second information set.

[0112] In the embodiment of the present application, in addition to the shared_buffers cache space, each standby node also maintains a global cache space memory map hash table (GSMM: global shared_buffers memory map) in memory, which is the second information set. It is used to record the mapping relationship between the data cached in the shared_buffers cache space of each node in the database cluster and the nodes. That is, through the first identifier of the data, it can be retrieved whether there is a node that has cached this data in the shared_buffers cache space. If so, the specific node information is given. Combined with Figure 7 Each node (primary node, standby node 1, standby node 2 and standby node 3) will scan the shared_buffers cache space of its own node regularly (every T seconds), record the correspondence between the first tag of the data in the shared_buffers cache space and the third tag of the node (primary, slave1, slave2, slave3) in the GSMM hash table, and synchronize the content of the GSMM hash table of the node to other nodes (including the primary node and standby node) by broadcasting. After each node receives the GSMM hash table of other nodes, it will merge it with the GSMM hash table of the node. Finally, the GSMM hash table of each node records the data page distribution of the shared_buffers cache space of each node in the entire database cluster.

[0113] See also Figure 8 , is an optional flow chart of the data reading method provided in the embodiment of the present application, which will be described in combination with the steps:

[0114] S18. Traverse the cache space of this node.

[0115] In the embodiment of the present application, standby node 1 traverses the data in the cache space of the node, and establishes a corresponding relationship based on the first identifier of the data cached by the node and the local third identifier. Standby node 2 traverses the data in the cache space of the node, and establishes a corresponding relationship based on the first identifier of the data cached by the node and the local third identifier.

[0116] S19. Update the GSMM hash table of this node.

[0117] In the embodiment of the present application, the standby node 1 can update the GSMM1 hash table of the node using the corresponding relationship formed in S18. The standby node 2 can update the GSMM2 hash table of the node using the corresponding relationship formed in S18.

[0118] S20. Broadcast the GSMM hash table of this node to other nodes.

[0119] In the embodiment of the present application, the standby node 1 broadcasts the GSMM1 hash table of the node to other nodes. The standby node 2 broadcasts the GSMM2 hash table of the node to other nodes.

[0120] S21. Update the GSMM hash table content.

[0121] In the embodiment of the present application, the standby node 1 can perform a union operation on GSMM1 and GSMM2 to update GSMM1. The standby node 2 can perform a union operation on GSMM1 and GSMM2 to update GSMM2.

[0122] S22, sleep (T).

[0123] In the embodiment of the present application, the standby node 1 and the standby node 2 wait for a period of time T before executing steps S18 to S21.

[0124] In the embodiment of the present application, each node periodically scans each cached data, and updates the local second information set based on the correspondence between the first identifier of each cached data and the third identifier of the node. Each node broadcasts the updated second information set to other nodes for other nodes to update the corresponding second information set. In this way, the distribution of data in each node can be synchronized to any other node, breaking the state of each node island, making each node related to each other. From the perspective of the entire database cluster, the cache ratio of user data is greatly improved, the cache hit rate when the user accesses can be improved, and the data access efficiency is greatly improved.

[0125] See also Fig. 9 , is an optional flow chart of a data reading method provided in an embodiment of the present application, Figure 1 The illustrated S102 can also be implemented through S501 to S502, which will be described in conjunction with the steps:

[0126] S501. Query the third mapping relationship in the first information set based on the target identifier to determine a query result; wherein the query result is used to indicate whether the target data includes the resident data.

[0127] In the embodiment of the present application, the target identifier may be a target identifier corresponding to at least one target data. The reading device may search for the corresponding third mapping relationship in the first information set based on the at least one target identifier, and determine the query result of whether the target data corresponding to each target identifier is resident data based on the corresponding mark in the third mapping relationship. If the target identifier in the third mapping relationship corresponds to the first mark, it is determined to be resident data. If the target identifier in the third mapping relationship corresponds to the second mark, it is determined not to be resident data.

[0128] S502: If the query result indicates that the target data includes the resident data, determine the target node with the largest storage proportion for the target data in the second mapping relationship based on the target identifier.

[0129] In an embodiment of the present application, if the query result indicates that at least one of the target data has the resident data, matching is performed in the second mapping relationship of the first information set based on at least one target identifier, and the backup node storing the most resident target data is determined as the target node.

[0130] In the embodiment of the present application, if the query result indicates that the target data does not include the resident data, a target node having the largest storage proportion for the target data is determined in the second information set based on the target identifier.

[0131] In an embodiment of the present application, the query result indicates that at least one target data does not include resident data. Then, based on at least one target identifier, the storage node corresponding to each target data is searched in the second information set, and the backup node that stores the most target data is determined as the target node.

[0132] In the embodiment of the present application, the query result is determined by querying in the third mapping relationship in the first information set based on the target identifier; wherein the query result is used to characterize whether the target data includes resident data; if the query result characterizes that the target data includes resident data, the target node with the largest storage ratio for the target data is determined in the second mapping relationship based on the target identifier. In this way, through the setting of the first information set and the second information set, each node can obtain the situation that the shared_buffers cache space of each node caches different user data, break the island state of each node, and make each node related to each other. From the perspective of the entire database cluster, the data cache ratio is greatly improved, the cache hit rate when the user accesses can be improved, and the data access efficiency is greatly improved.

[0133] See also Fig.10 , is an optional flow chart of the data reading method provided in the embodiment of the present application, which will be described in combination with the steps:

[0134] S22. Obtain user request.

[0135] S23, middleware.

[0136] In the embodiment of the present application, each node in the database cluster can accept user read requests, and the application accesses the database cluster through the middleware.

[0137] S24. The master node parses the user request to determine the target identifier.

[0138] In the embodiment of the present application, when the middleware receives a user request, it first obtains the target identifier of the target data that the user needs to access.

[0139] S25. Connect to the master node and scan the rule table hot_data_rules.

[0140] S26. Count the cache quantity of the target data in the cache space of each node.

[0141] In the embodiment of the present application, the master node queries the rule table hot_data_rules according to the target identifier, and determines whether the target data has been set as resident data according to the rule table hot_data_rules.

[0142] S27: Whether the target data is resident data.

[0143] In the embodiment of the present application, if the target data has been set as resident data, S29 is executed. If the target data has not been set as resident data, S28 is executed.

[0144] S28, GSMM hash table statistics.

[0145] In an embodiment of the present application, if the tables that the user wants to access are not resident tables in the memory, the number of tables that the user wants to access cached in the shared_buffers cache space of each node is counted according to the GSMM hash table, and the node with the highest cache number is selected as the service node, and the user connection is forwarded to the corresponding target node.

[0146] S29. Select a target node.

[0147] In the embodiment of the present application, if there is a table set to be resident in memory, the node with the most cached tables is selected as the service node, and the user connection is forwarded to the corresponding target node.

[0148] S30: Forward the user request for processing.

[0149] In the embodiment of the present application, when processing user requests to select a database service node, the hot_data_rules table and the GSMM table shown in Table 1 are relied upon. The distribution of resident memory tables in each database node can be sensed through the metadata table hot_data_rules. The data page distribution of the shared_buffers cache space of each node in the entire database cluster can be sensed through the GSMM module. The most suitable database service node can be selected through the distribution of resident memory tables and the data page distribution of the shared_buffers cache space of each node. Through the design ideas of this proposal, a database cluster system based on federal memory pool technology is proposed. Compared with the traditional database cluster system, this proposal realizes the controllability of cached data of each node by introducing a resident memory management module in the shared_buffers cache space. Through reasonable organization, different nodes cache different table data. From a global perspective, the data cached in the shared_buffers cache space of the entire cluster has a reduced duplication of data, and more table data is involved, which can provide more effective data access services. By introducing the GSMM hash table, the shared_buffers cache space of each node is logically connected. Each node can sense the data cache status of the shared_buffers cache space of each node in the entire cluster, provide node selection for user requests, select the node with the most cached data as the service node, and maximize the shared_buffers cache space hit rate of target data access, thereby improving data access efficiency. The cached data of each node can change dynamically according to the rules, which has very good flexibility compared with the sharding of libraries and tables, and can dynamically control the data in shared_buffers, which is more suitable for business usage scenarios and more conducive to improving data access efficiency.

[0150] See also Fig.11 , is a schematic diagram of the structure of a data reading device provided in an embodiment of the present application.

[0151] The embodiment of the present application provides a data reading device 800 , including: a request acquisition unit 801 , a determination unit 802 , and a forwarding unit 803 .

[0152] The request acquisition unit 801 is used to acquire a user request; wherein the user request includes: a target identifier of target data;

[0153] A determination unit 802 is used to determine a target node in a preset information set based on the target identifier; wherein the preset information set is used to manage resident data and non-resident data in each node, and is also used to describe the distribution of data in each node;

[0154] The forwarding unit 803 is used to forward the user request to the target node to read the target data.

[0155] In the embodiment of the present application, the preset information set includes: a first information set and a second information set; the first information set is used to manage the resident data and the non-resident data in each of the nodes; the second information set is used to describe the distribution of data in each of the nodes;

[0156] The first information set includes: a first mapping relationship between a first identifier of each data and a second identifier of a storage database, a second mapping relationship between each first identifier and a third identifier of a storage node, a third mapping relationship between each first identifier and a first mark of the resident data or a second mark of the non-resident data, and a fourth mapping relationship for determining whether the data mark corresponding to each first identifier is effective;

[0157] The second information set includes: a fifth mapping relationship between the first identifier of each data and the third identifier of the storage node.

[0158] In the embodiment of the present application, the data reading device 800 is further used to form the first information set in response to the configuration operation of the user;

[0159] Each of the nodes periodically reads the first information set so that each of the nodes can manage the resident data and the non-resident data in the node based on the first information set.

[0160] In the embodiment of the present application, the data reading device 800 is further used for the node to traverse the first identifier of each of the data in the first information set, and if the node traverses and determines that the first identifier of the first data corresponds to the third identifier of the node, then based on the fourth mapping relationship corresponding to the first identifier, determine the mark validity detection result; wherein the mark validity detection result is used to indicate whether the mark for the first data is effective;

[0161] The marking of the first data is completed based on the marking effectiveness detection result, and the next first identification is traversed until the first identification traversal of each data is completed, thereby realizing the management and control of the resident data and the non-resident data in the node.

[0162] In an embodiment of the present application, the data reading device 800 is also used to complete the marking of the first data based on the first mark corresponding to the first identifier in the third mapping relationship, or the second mark if the mark validity detection result indicates that the mark for the first data is not valid.

[0163] In the embodiment of the present application, the data reading device 800 is further used for each node to periodically scan each cached data, and update the local second information set based on the correspondence between the first identifier of each cached data and the third identifier of the node;

[0164] Each of the nodes broadcasts the updated second information set to other nodes, so that other nodes can update the corresponding second information set.

[0165] In the embodiment of the present application, the determination unit 802 in the data reading device 800 is further used to query the third mapping relationship in the first information set based on the target identifier to determine a query result; wherein the query result is used to indicate whether the target data includes the resident data;

[0166] If the query result indicates that the target data includes the resident data, a target node having the largest storage proportion for the target data is determined in the second mapping relationship based on the target identifier.

[0167] In the embodiment of the present application, the determination unit 802 in the data reading device 800 is also used to determine the target node with the largest storage proportion for the target data in the second information set based on the target identifier if the query result indicates that the target data does not include the resident data.

[0168] In the embodiment of the present application, the data reading device 800 is used for each of the nodes to dynamically eliminate the non-resident data stored therein.

[0169] It should be noted that in the embodiments of the present application, if the above-mentioned method for processing item information is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly reflected in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium and includes several instructions for enabling an item information processing device (which can be a personal computer, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0170] Correspondingly, an embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method on the data reading device side are implemented.

[0171] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.

[0172] It should be noted that Fig.12 A hardware entity diagram of an electronic device provided in an embodiment of the present application, such as Fig.12 As shown, an embodiment of the present application provides an electronic device 900, including a memory 902 and a processor 901, wherein the memory 902 stores a computer program that can be run on the processor 901, and the processor 901 implements the steps in the above method when executing the program, wherein;

[0173] The processor 901 generally controls the overall operation of the electronic device 900 .

[0174] The memory 902 is configured to store instructions and applications executable by the processor 901, and can also cache data to be processed or processed by the processor 901 and various modules in the electronic device 900 (for example, image data, audio data, voice communication data, and video communication data), which can be implemented through flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0175] Correspondingly, an embodiment of the present application further provides a computer program product, including a computer program, which can be executed by a processor 901 of an electronic device 900 to complete the steps in the method on the data reading device 800 side.

[0176] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0177] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0178] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0179] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0180] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0181] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: a mobile storage device, a read-only memory (ROM), a magnetic disk or an optical disk, and other media that can store program codes.

[0182] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a mobile storage device, a ROM, a magnetic disk, or an optical disk.

[0183] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application.

Claims

1. A data reading method, characterized in that: include: Obtaining a user request; wherein the user request includes: a target identifier of the target data; Determine the target node in a preset information set based on the target identifier; wherein the preset information set is used to manage the resident data and non-resident data in each node, and is also used to describe the distribution of data in each node; The user request is forwarded to the target node to read the target data.

2. The data reading method according to claim 1, characterized in that: The preset information set includes: a first information set and a second information set; the first information set is used to manage the resident data and the non-resident data in each of the nodes; the second information set is used to describe the distribution of data in each of the nodes; The first information set includes: a first mapping relationship between a first identifier of each data and a second identifier of a storage database, a second mapping relationship between each first identifier and a third identifier of a storage node, a third mapping relationship between each first identifier and a first mark of the resident data or a second mark of the non-resident data, and a fourth mapping relationship for determining whether the data mark corresponding to each first identifier is effective; The second information set includes: a fifth mapping relationship between the first identifier of each data and the third identifier of the storage node.

3. The data reading method according to claim 2, characterized in that: The method further comprises: In response to a configuration operation by a user, forming the first information set; Each of the nodes periodically reads the first information set so that each of the nodes can manage the resident data and the non-resident data in the node based on the first information set.

4. The data reading method according to claim 3, characterized in that: Each of the nodes performs management and control of the resident data and the non-resident data in the node based on the first information set, including: The node traverses the first identifier of each of the data in the first information set, and if the node traverses and determines that the first identifier of the first data corresponds to the third identifier of the node, determines a mark validity detection result based on the fourth mapping relationship corresponding to the first identifier; wherein the mark validity detection result is used to indicate whether the mark for the first data is valid; The node completes the marking of the first data based on the marking effectiveness detection result, and traverses the next first identifier until the traversal of the first identifier for each data is completed, thereby realizing the management and control of the resident data and the non-resident data in the node.

5. The data reading method according to claim 4, characterized in that: The node completes marking of the first data based on the marking validity detection result, including: If the mark validity detection result indicates that the mark for the first data is not valid, the node completes the marking for the first data based on the first mark corresponding to the first identifier in the third mapping relationship, or the second mark.

6. The data reading method according to any one of claims 2 to 5, characterized in that: The method further comprises: Each of the nodes periodically scans each of the cached data, and updates the local second information set based on the correspondence between the first identifier of each of the cached data and the third identifier of the node; Each of the nodes broadcasts the updated second information set to other nodes, so that other nodes can update the corresponding second information set.

7. The data reading method according to any one of claims 2 to 5, characterized in that: The determining the target node in the preset information set based on the target identifier includes: Based on the target identifier, query the third mapping relationship in the first information set to determine a query result; wherein the query result is used to indicate whether the target data includes the resident data; If the query result indicates that the target data includes the resident data, a target node having the largest storage proportion for the target data is determined in the second mapping relationship based on the target identifier.

8. The data reading method according to claim 7, characterized in that: After searching in the first information set based on the target identifier and determining the query result, the method further includes: If the query result indicates that the target data does not include the resident data, a target node having the largest storage proportion for the target data is determined in the second information set based on the target identifier.

9. The data reading method according to any one of claims 2 to 5, characterized in that: The method further comprises: Each of the nodes dynamically eliminates the non-resident data stored therein.

10. A data reading device, characterized in that: include: A request acquisition unit, used to acquire a user request; wherein the user request includes: a target identifier of the target data; A determination unit, configured to determine a target node in a preset information set based on the target identifier; wherein the preset information set is used to manage resident data and non-resident data in each node, and is also used to describe the distribution of data in each node; The forwarding unit is used to forward the user request to the target node to read the target data.

11. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 9 when executing the computer program.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 9 are implemented.

13. A computer program product, comprising a computer program, characterized in that The computer program implements the steps of the method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Cited By

  • Cloud native database memory pooling method and device, electronic equipment and storage medium

    CN121681600A