Database Cache Management Method, Database, Device, and Storage Medium

During the database brushing cycle, the target brushing parameters are calculated based on the number of data blocks in the buffer data space and historical brushing data, and the memory space and read and write performance of the database are optimized, and the problem of fluctuations in the brushing process is solved, achieving more stable database performance and file integrity.

CN119248833BActive Publication Date: 2025-06-24本原数据(北京)信息技术有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411331423.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-24
Publication Date
2025-06-24
Estimated Expiration
2044-09-24

AI Technical Summary

Technical Problem

During the process of brushing dirty, the existing database determines the brushing speed based on the number of dirty pages in memory, resulting in large fluctuations in read and write speeds, affecting the stable operation of other database services.

Method used

By obtaining the number of the first and second data blocks in the buffer data space during the dirty brushing cycle, combining the historical dirty brushing number and the data sequence number, the target dirty brushing parameters are calculated, the target data block is determined, and written to the local data space.

Benefits of technology

It realizes optimization of memory space and read and write performance, reduces unnecessary data writing operations, improves dirty brushing efficiency, provides more stable database performance, and avoids file corruption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119248833B_ABST
    Figure CN119248833B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a database cache management method, a database, a device, and a storage medium. The method includes: in the current dirty flushing period, obtaining a first data block and a second data block in the buffer data space; obtaining a first data quantity of the first data block and a second data quantity of the second data block, and determining a first dirty flushing parameter based on the first data quantity and the second data quantity; obtaining historical dirty flushing quantities of multiple historical periods, and determining a second dirty flushing parameter based on the historical dirty flushing quantities; obtaining data sequence numbers of each first data block, and determining a third dirty flushing parameter based on the data sequence numbers; determining a target dirty flushing parameter according to the first dirty flushing parameter, the second dirty flushing parameter, and the third dirty flushing parameter; determining a target data block according to the target dirty flushing quantity, and writing the target data block into the local data space. The embodiment of the present application can optimize the memory space and the read and write performance, and avoid file damage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to, but is not limited to, the field of database management technology, and particularly relates to a database cache management method, a database, a device, and a storage medium. Background Art

[0002] The database stores user data using a file system. To improve read and write performance, current mainstream databases first load data blocks into memory, and a data block in memory is called a page. When the database performs a write operation, the modified data block is called a dirty page. A dirty page is only considered to be successfully written when it is persisted to the data file, and data loss due to system crashes can be avoided. The process of persisting dirty pages is called flushing dirty pages.

[0003] In related technologies, most databases adopt an incremental dirty page flushing mode. This mode avoids the read and write peaks caused by flushing dirty pages and reduces the impact of background dirty page flushing on other database operations to a certain extent. However, the database usually determines the flushing speed based on the number of dirty pages in the current memory, and the fluctuating read and write speeds are not conducive to the stable operation of other database operations. Summary of the Invention

[0004] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.

[0005] Embodiments of this application provide a database cache management method, a database, a device, and a storage medium, which can optimize the memory space and read and write performance and avoid file corruption.

[0006] To achieve the above object, a first aspect of an embodiment of this application proposes a database cache management method, including: in the current dirty page flushing cycle, obtaining a first data block and a second data block in the buffer data space, where the first data block represents data that has not been written to the local data space, and the second data block represents data that has been written to the local data space; obtaining a first data quantity of the first data block and a second data quantity of the second data block, and determining a first dirty page flushing parameter based on the first data quantity and the second data quantity; obtaining the historical dirty page flushing quantities of multiple historical cycles, and determining a second dirty page flushing parameter based on the historical dirty page flushing quantities; obtaining the data sequence numbers of each of the first data blocks, and determining a third dirty page flushing parameter based on the data sequence numbers; determining a target dirty page flushing parameter according to the first dirty page flushing parameter, the second dirty page flushing parameter, and the third dirty page flushing parameter; determining a target data block according to the target dirty page flushing quantity, and writing the target data block to the local data space.

[0007] In some embodiments, obtaining the first data quantity of the first data block and the second data quantity of the second data block, and determining a first dirty brush parameter based on the first data quantity and the second data quantity includes: obtaining the first data quantity of the first data block and the second data quantity of the second data block, and determining the dirty brush ratio of the buffer data space according to the first data quantity and the second data quantity; determining a first dirty brush parameter based on the dirty brush ratio and the first data quantity.

[0008] In some embodiments, determining a first dirty brush parameter based on the dirty brush ratio and the first data quantity includes: obtaining the historical dirty brush speed of the previous dirty brush cycle, and determining a first dirty brush limit value based on the historical dirty brush speed; when the first data quantity is greater than or equal to the first dirty brush limit value, or the dirty brush ratio is greater than or equal to a preset second dirty brush limit value, obtaining a third dirty brush limit value, and determining a first dirty brush parameter according to the third dirty brush limit value, where the third dirty brush limit value represents the maximum quantity of data blocks for dirty brushing in a single dirty brush cycle of the database; when the first data quantity is less than the first dirty brush limit value and the dirty brush ratio is less than the second dirty brush limit value, determining a first dirty brush parameter according to the preset second dirty brush limit value.

[0009] In some embodiments, obtaining the historical dirty brush quantities of multiple historical cycles and determining a second dirty brush parameter based on all the historical dirty brush quantities includes: obtaining the historical dirty brush quantities of multiple historical cycles and determining an average dirty brush quantity based on all the historical dirty brush quantities; determining a second dirty brush parameter based on the historical dirty brush quantities.

[0010] In some embodiments, obtaining the data sequence numbers of each of the first data blocks and determining a third dirty brush parameter based on the data sequence numbers includes: obtaining the data sequence numbers of each of the first data blocks, determining a first minimum sequence number and a first maximum sequence number in the first data block, and determining a sequence difference based on the first minimum sequence number and the first maximum sequence number; determining a third dirty brush parameter according to the sequence difference.

[0011] In one embodiment, determining a third dirty brush parameter according to the sequence difference includes: when the sequence difference is greater than or equal to a preset difference threshold, obtaining a first difference factor and determining a third dirty brush parameter according to the first difference factor and the historical dirty brush quantity; when the sequence difference is less than the difference threshold, obtaining a second difference factor and determining the third dirty brush parameter according to the second difference factor and the historical dirty brush quantity, where the first difference factor is greater than the second difference factor.

[0012] In some embodiments, determining the target data block according to the target dirty write quantity and writing the target data block into the local data space includes: determining the target data block according to the target dirty write quantity, and obtaining the second smallest sequence number and the second largest sequence number in the target data block; based on the data sequence numbers, starting from the target data block corresponding to the second smallest sequence number, sequentially writing the target data blocks into the local data space; when the first data quantity is greater than or equal to a preset fourth dirty write limit value and when the sequence difference is greater than or equal to the difference threshold, continuously obtaining the dirty write running time of the current dirty write period; obtaining the periodic dirty write time, and when the target data block corresponding to the second largest sequence number is written into the local data space and the dirty write running time is less than the periodic dirty write time, re-obtaining the first data block and the second data block in the current buffer data space to enter the next dirty write period.

[0013] To achieve the above object, a second aspect of the present application provides a database, which is configured to execute the database cache management method described in the first aspect.

[0014] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the database cache management method described in the first aspect is implemented.

[0015] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, which is a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the database cache management method described in the first aspect is implemented.

[0016] The embodiments of the present application at least include the following beneficial effects: In the current cycle, by differentiating the first data block and the second data block and counting their quantities, it is possible to more precisely evaluate which data needs to be flushed preferentially, thereby reducing unnecessary data write operations and improving the flushing efficiency; the target flushing quantity represents the quantity of the first data block written into the local data space in the current cycle, and the target flushing quantity is affected by the density of the data in the database buffer data space and the database write and read capabilities. When calculating the target flushing quantity of the current cycle, first, the capacity of the buffer data space is constant. The space occupied by the second data block can be flushed when needed, while the first data block needs to be persistently processed. By obtaining the first data quantity of the first data block and the second data quantity of the second data block, the congestion degree of the buffer data space at the start of the current cycle can be known. When the first data quantity is greater than a certain threshold, it indicates that the first data block occupies a relatively large capacity in the buffer data space, there are more files in the buffer data space that have not been written into the local data space, new files cannot be written, and the first data block in the buffer data space is also prone to loss. Therefore, by determining the first flushing parameter based on the first data quantity and the second data quantity, the first data block can be written into the local data block as soon as possible, enhancing the write ability of the buffer data space; then, when the target flushing quantity fluctuates greatly in adjacent cycles, it may cause an instantaneous decline in the database performance. Therefore, by obtaining the historical flushing quantities of multiple historical cycles, where the historical flushing quantity represents the target flushing quantity of the corresponding cycle, determining the second flushing parameter based on the historical flushing quantity, and determining the target flushing quantity of the current cycle according to the second flushing parameter, it is possible to smooth the target flushing quantities of adjacent cycles, reduce the competition of different threads for simultaneous write operations to the local data space, and improve the overall write efficiency. By reducing the fluctuation of the target flushing quantity in adjacent cycles, a more stable database performance can be provided; finally, in the buffer data space, the data sequence number of the first data block is determined based on the write time. The data sequence numbers of the first data blocks of the same file are consecutive, but the first data blocks with consecutive data sequence numbers are not physically consecutive in the buffer data space. When writing the first data block into the local data space, for the sake of read and write efficiency, usually, the first data blocks adjacent in physical space are read and written in one cycle. Therefore, by obtaining the data sequence number of the first data block, when the difference in the data sequence numbers of the first data blocks is too large, it indicates that there are first data blocks that have been stored for a long time. Determining the third flushing parameter based on the data sequence number can thus preferentially write the first data block that most needs to be persisted into the local data space according to the importance of the first data block data, avoid file damage, and also optimize the buffer data space and the local data space, enhancing the database write and read efficiency;In summary, the database cache management method proposed in this application determines the target dirty flush quantity through the above method, determines the target data blocks according to the target dirty flush quantity, can optimize the memory space and read / write performance, and can also avoid file damage.

[0017] Other features and advantages of this application will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing this application. The objectives and other advantages of this application can be achieved and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings are used to provide a further understanding of the technical solutions of this application, and constitute a part of the specification. Together with the embodiments of this application, they are used to explain the technical solutions of this application, and do not constitute a limitation to the technical solutions of this application.

[0019] Figure 1 It is an optional flowchart of the database cache management method provided by an embodiment of this application;

[0020] Figure 2 It is an optional flowchart of calculating the first dirty flush parameter provided by an embodiment of this application;

[0021] Figure 3 It is an optional specific flowchart of calculating the first dirty flush parameter provided by an embodiment of this application;

[0022] Figure 4 It is an optional flowchart of calculating the second dirty flush parameter provided by an embodiment of this application;

[0023] Figure 5 It is an optional flowchart of calculating the third dirty flush parameter provided by an embodiment of this application;

[0024] Figure 6 It is an optional specific flowchart of calculating the third dirty flush parameter provided by an embodiment of this application;

[0025] Figure 7 It is an optional flowchart of entering the next dirty flush cycle provided by an embodiment of this application;

[0026] Figure 8 It is an optional flowchart of determining the target dirty flush quantity according to the database capabilities provided by an embodiment of this application;

[0027] Figure 9 It is an optional hardware structure diagram of the electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Other features and advantages of the present application will be set forth in the following description, and in part will be obvious from the description, or can be learned by practicing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0030] In the description of the present application, the meaning of "several" is one or more, the meaning of "multiple" is more than two, and understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.

[0031] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the flowchart in the flowchart. Terms such as "first", "second", etc. in the description, claims or the above drawings are used to distinguish similar objects and do not have to be used to describe a specific order or sequence.

[0032] To facilitate the understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application will be explained here:

[0033] Flushing Dirty Pages is an operation in database memory management, which refers to writing the dirty pages (data pages that have been modified but not yet written to disk) in memory to disk to ensure data consistency. This process is part of the database buffer pool management and is usually triggered by the database's Checkpoint mechanism or automatically when the number of dirty pages in memory reaches a certain threshold. Flushing dirty ensures that data changes in memory are synchronized to disk, but does not itself guarantee the durability of transactions.

[0034] Persistence means that data is stored on a persistent storage medium, such as a hard disk or solid-state drive, so that the data will not be lost even in the event of a system crash or power failure. In a database system, persistence usually relies on transaction logs (such as redo logs) to achieve. When a transaction is committed, its changes are recorded in the log, and these log records can be used to recover data to ensure the "D" (Durability) in the ACID properties of database operations. Persistence is a broader concept that includes not only the flushing dirty operation, but also log recording, transaction management, and other mechanisms to ensure data is not lost.

[0035] In the related art, most databases adopt the incremental dirty page flushing mode. This mode avoids the read-write peaks caused by dirty page flushing and reduces the impact of background dirty page flushing on the rest of the database operations to a certain extent. However, the database usually determines the dirty page flushing speed based on the number of dirty pages in the current memory, and the fluctuating read-write speed is not conducive to the stable operation of the rest of the database operations.

[0036] Based on this, the embodiments of the present application provide a database cache management method, a database, a device, and a storage medium, which can optimize the memory space and read-write performance and avoid file damage.

[0037] The database cache management method, database, device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the database cache management method in the embodiments of the present application is described.

[0038] The database cache management method provided by the embodiments of the present application relates to the field of computer technology. The database cache management method provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the database cache management method, etc., but is not limited to the above forms.

[0039] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0040] The following further elaborates on the embodiments of the present application in conjunction with the accompanying drawings.

[0041] As Figure 1 shown, Figure 1 is an alternative flowchart of the database cache management method provided by the embodiments of the present application. This database cache management method can be executed by a server, or by a terminal, or by a server in cooperation with a terminal. The database cache management method includes, but is not limited to, the following steps S110 to S160:

[0042] Step S110: In the current dirty flush cycle, obtain the first data block and the second data block in the buffer data space. Among them, the first data block represents the data that has not been written into the local data space, and the second data block represents the data that has been written into the local data space;

[0043] Step S120: Obtain the first data quantity of the first data block and the second data quantity of the second data block, and determine the first dirty flush parameter based on the first data quantity and the second data quantity;

[0044] Step S130: Obtain the historical dirty flush quantities of multiple historical cycles, and determine the second dirty flush parameter based on the historical dirty flush quantities;

[0045] Step S140: Obtain the data sequence numbers of each first data block, and determine the third dirty flush parameter based on the data sequence numbers;

[0046] Step S150: Determine the target dirty flush parameter according to the first dirty flush parameter, the second dirty flush parameter, and the third dirty flush parameter;

[0047] Step S160: Determine the target data block according to the target dirty flush quantity, and write the target data block into the local data space.

[0048] It can be understood that within the current cycle, by differentiating between the first data block and the second data block and counting their quantities, it is possible to more precisely evaluate which data needs to be flushed dirty first, thereby reducing unnecessary data write operations and improving the dirty flushing efficiency. The target dirty flushing quantity represents the number of first data blocks written to the local data space within the current cycle. The target dirty flushing quantity is affected by the density of data in the database buffer data space and the database write and output capabilities. When calculating the target dirty flushing quantity for the current cycle, first, the capacity of the buffer data space is constant. The space occupied by the second data block can be flushed when needed, while the first data block needs to be persisted. By obtaining the first data quantity of the first data block and the second data quantity of the second data block, the congestion degree of the buffer data space at the start of the current cycle can be known. When the first data quantity is greater than a certain threshold, it indicates that the first data block occupies a relatively large capacity in the buffer data space, there are many files in the buffer data space that have not been written to the local data space, new files cannot be written, and the first data blocks in the buffer data space are also prone to loss. Therefore, by determining the first dirty flushing parameter based on the first data quantity and the second data quantity, the first data block can be written to the local data block as soon as possible, enhancing the write ability of the buffer data space. Then, when the target dirty flushing quantity fluctuates significantly in adjacent cycles, it may cause an instantaneous decline in database performance. Therefore, by obtaining the historical dirty flushing quantities of multiple historical cycles, where the historical dirty flushing quantity represents the target dirty flushing quantity for the corresponding cycle, determining the second dirty flushing parameter based on the historical dirty flushing quantity, and determining the target dirty flushing quantity for the current cycle according to the second dirty flushing parameter, it is possible to smooth the target dirty flushing quantities in adjacent cycles, reduce the competition among different threads for simultaneous write operations to the local data space, and improve the overall write efficiency. By reducing the fluctuations in the target dirty flushing quantity within adjacent cycles, a more stable database performance can be provided. Finally, within the buffer data space, the data sequence number of the first data block is determined based on the write time. The data sequence numbers of the first data blocks of the same file are consecutive, but the first data blocks with consecutive data sequence numbers are not physically contiguous in the buffer data space. When writing the first data block to the local data space, for the sake of read and write efficiency, within one cycle, usually the first data blocks adjacent in physical space are read and written. Therefore, by obtaining the data sequence number of the first data block, when the difference in the data sequence numbers of the first data blocks is too large, it indicates that there are first data blocks that have been stored for a long time. By determining the third dirty flushing parameter based on the data sequence number, it is possible to preferentially write the first data block that most needs to be persisted to the local data space according to the importance of the first data block data, avoid file corruption, and also optimize the buffer data space and the local data space, enhancing the database write and output efficiency.In summary, the database cache management method proposed in this application determines the target dirty flush quantity in the above manner, determines the target data blocks according to the target dirty flush quantity, can optimize the memory space and read-write performance, and can also avoid file corruption.

[0049] It should be noted that the first data block can be understood as a dirty page, the second data block can be understood as a clean page, the buffered data space can be understood as the memory in the database, the local data space can be understood as the local hard disk in the database, and the operation of writing the first data block from the buffered data space to the local data space can be understood as flushing or persistence. In the buffered data space, the file is divided into multiple data blocks of 8 kb in size for storage, that is, the sizes of both the first data block and the second data block are 8 kb.

[0050] In a specific embodiment, the target dirty flush quantity is equal to the average of the first dirty flush parameter, the second dirty flush parameter, and the third dirty flush parameter.

[0051] In a specific embodiment, referring to Figure 8 as shown, when the target dirty flush quantity is greater than the maximum read-write capacity of the database, the target dirty flush quantity is set to the maximum read-write capacity of the database; when the target dirty flush quantity is less than the minimum single read-write quantity of the database, the target dirty flush quantity is set to the minimum single read-write quantity.

[0052] In addition, referring to Figure 2 as shown, in some embodiments of this application, Figure 1 step S120 in

[0053] includes but is not limited to the following steps S210 to S220:

[0054] Step S210: Obtain the first data quantity of the first data block and the second data quantity of the second data block, and determine the to-be-dirty-flushed ratio of the buffered data space according to the first data quantity and the second data quantity;

[0054] Step S220: Determine the first dirty flush parameter based on the to-be-dirty-flushed ratio and the first data quantity.

[0055] It can be understood that by calculating the to-be-dirty-flushed ratio, it is possible to avoid over-flushing data blocks that have not been modified significantly, thereby reducing unnecessary read-write operations and improving the overall efficiency of the system. Under different workloads, the calculation of the to-be-dirty-flushed ratio and the determination of the dirty flush parameter can adapt to different data access patterns, enabling the database system to maintain good performance when facing different workloads.

[0056] In addition, referring to Figure 3As shown, some embodiments of the present application Figure 2 Step S220 in includes, but is not limited to, the following steps S310 to S330:

[0057] Step S310: Obtain the historical dirty brushing speed of the previous dirty brushing cycle, and determine the first dirty brushing limit value based on the historical dirty brushing speed;

[0058] Step S320: When the first data quantity is greater than or equal to the first dirty brushing limit value, or the dirty brushing ratio to be processed is greater than or equal to a preset second dirty brushing limit value, obtain the third dirty brushing limit value, and determine the first dirty brushing parameter according to the third dirty brushing limit value, where the third dirty brushing limit value represents the maximum quantity of data block dirty brushing in a single dirty brushing cycle of the database;

[0059] Step S330: When the first data quantity is less than the first dirty brushing limit value and the dirty brushing ratio to be processed is less than the second dirty brushing limit value, determine the first dirty brushing parameter according to the preset second dirty brushing limit value.

[0060] In a specific embodiment, the first dirty brushing limit value is equal to the target dirty brushing quantity corresponding to the previous dirty brushing cycle, the dirty brushing ratio to be processed is equal to the ratio of the first data quantity to the sum of the first data quantity and the second data quantity, the second dirty brushing limit value is equal to 75%, and the third dirty brushing limit value represents the maximum read / write capacity of the database.

[0061] The value of the first dirty brushing parameter is a multi-level step function. When the first data quantity is greater than or equal to the first dirty brushing limit value, or the dirty brushing ratio to be processed is greater than or equal to the preset second dirty brushing limit value, the value of the first dirty brushing parameter is 75% of the third dirty brushing limit value. When the first data quantity is less than the first dirty brushing limit value and the dirty brushing ratio to be processed is less than the second dirty brushing limit value, the value of the first dirty brushing parameter is the preset low value.

[0062] In a specific embodiment, when the first data quantity is less than the first dirty brushing limit value and the dirty brushing ratio to be processed is less than the second dirty brushing limit value, the first dirty brushing parameter is equal to the target dirty brushing quantity corresponding to the previous dirty brushing cycle.

[0063] In addition, referring to Figure 4 As shown, some embodiments of the present application Figure 1 Step S130 in includes, but is not limited to, the following steps S410 to S420:

[0064] Step S410: Obtain the historical dirty brushing quantities of multiple historical cycles, and determine the average dirty brushing quantity based on all the historical dirty brushing quantities;

[0065] Step S420: Determine the second dirty brushing parameter based on the historical dirty brushing quantities.

[0066] It can be understood that when the number of target dirty pages fluctuates greatly in adjacent cycles, it may cause an instantaneous decrease in database performance. Therefore, by obtaining the historical dirty page numbers of multiple historical cycles, where the historical dirty page numbers represent the target dirty page numbers of the corresponding cycles, determining the second dirty page parameter based on the historical dirty page numbers, and determining the target dirty page number of the current cycle according to the second dirty page parameter, it is possible to smooth the target dirty page numbers in adjacent cycles, reduce the competition among different threads for simultaneous write operations on the local data space, improve the overall write efficiency, and provide more stable database performance by reducing the fluctuations of the target dirty page numbers in adjacent cycles.

[0067] In a specific embodiment, the second dirty page parameter is equal to the average value of the historical dirty page numbers of 30 historical cycles.

[0068] In addition, referring to Figure 5 as shown, some embodiments of the present application Figure 1 step S140 in includes but is not limited to the following steps S510 to step S520:

[0069] Step S510: Obtain the data sequence numbers of each first data block, determine the first minimum sequence number and the first maximum sequence number in the first data block, and determine the sequence difference based on the first minimum sequence number and the first maximum sequence number;

[0070] Step S520: Determine the third dirty page parameter according to the sequence difference.

[0071] It can be understood that in the buffered data space, that is, in the memory of the database, the data sequence numbers of the first data blocks are determined based on the write time. The data sequence numbers of the first data blocks of the same file are continuous, but the first data blocks with continuous data sequence numbers are not continuous in the physical space of the buffered data space. When writing the first data block into the local data space, for the sake of read and write efficiency, usually in one cycle, the first data blocks adjacent in physical space are read and written. Therefore, by obtaining the data sequence numbers of the first data blocks, when the data sequence numbers of the first data blocks differ greatly, it indicates that there are first data blocks that have been stored for a long time. Determining the third dirty page parameter based on the data sequence numbers can thus preferentially write the first data blocks that most need to be persisted into the local data space according to the importance of the first data block data, avoid file damage, and also optimize the buffered data space and the local data space, improving the write and read efficiency of the database.

[0072] In addition, referring to Figure 6 as shown, some embodiments of the present application Figure 5 step S520 in includes but is not limited to the following steps S610 to step S620:

[0073] Step S610: When the sequence difference is greater than or equal to a preset difference threshold, obtain a first difference factor, and determine a third dirty brush parameter according to the first difference factor and the historical dirty brush quantity.

[0074] Step S620: When the sequence difference is less than the difference threshold, obtain a second difference factor, and determine a third dirty brush parameter according to the second difference factor and the historical dirty brush quantity, where the first difference factor is greater than the second difference factor.

[0075] In a specific embodiment, the first minimum sequence number represents the first data block earliest written into the buffer data space, and the second maximum sequence number represents the first data block latest written into the buffer data space. When the sequence difference is greater than or equal to the preset difference threshold, the third dirty brush parameter is equal to the product of the second dirty brush parameter and the first difference factor. When the sequence difference is less than the difference threshold, the third difference factor is equal to the product of the second dirty brush parameter and the second difference factor.

[0076] In a specific embodiment, when the sequence difference is greater than or equal to the preset difference threshold, the first difference factor is equal to 3, and when the sequence difference is less than the difference threshold, the second difference factor is equal to 1.

[0077] It can be understood that when the sequence difference is large, it indicates that there are data blocks in the buffer data space that were written earlier but have not yet been written into the local data space for persistence. In the related art, when a new data block is written into the buffer data space, the database preferentially clears the data block with a smaller data sequence number, that is, the earlier written data block. Therefore, when the sequence difference is large, a larger first difference factor is given to increase the third dirty brush parameter to increase the target dirty brush quantity, which is beneficial to writing the first data block with a longer storage time into the local data space in a timely manner and avoiding file data damage. At the same time, determining the third dirty brush parameter based on the average value of the first difference factor or the second difference factor and multiple historical dirty brush quantities can also smooth the target dirty brush quantity in adjacent cycles, reduce the competition for writing operations on the local data space by different threads, improve the overall writing efficiency, and provide more stable database performance by reducing the fluctuation of the target dirty brush quantity in adjacent cycles.

[0078] In addition, referring to Figure 7 as shown, some embodiments of the present application Figure 1 Step S150 in includes but is not limited to the following steps S710 to S740:

[0079] Step S710: Determine target data blocks according to the target dirty brush quantity, and obtain the second minimum sequence number and the second maximum sequence number in the target data blocks.

[0080] Step S720: Starting from the target data block corresponding to the second smallest sequence number based on the data sequence number, write the target data blocks into the local data space in sequence;

[0081] Step S730: When the first data quantity is greater than or equal to a preset fourth dirty-writing limit value, and when the sequence difference is greater than or equal to the difference threshold, continuously obtain the dirty-writing running time of the current dirty-writing cycle;

[0082] Step S740: Obtain the periodic dirty-writing time. When the target data block corresponding to the second largest sequence number is written into the local data space and the dirty-writing running time is less than the periodic dirty-writing time, re-obtain the first data block and the second data block in the current buffer data space to enter the next dirty-writing cycle.

[0083] Specifically, after determining the target dirty-writing quantity, based on the database persistence strategy in the related technology, the database determines the target dirty-writing quantity, the target data blocks to be written into the local data space in the current cycle, and their read-write order. The second smallest sequence number represents the target data block that is read and written first, and the second largest sequence number represents the second data block that is read and written last.

[0084] In the related technology, when the database performs the persistence strategy, the target time for each dirty-writing cycle is usually fixed. When the first data quantity is less than the preset fourth dirty-writing limit value, or when the sequence difference is less than the difference threshold, the database performs dirty-writing according to the preset target time. Even when the target data block corresponding to the second largest sequence number is written into the local data space and the dirty-writing running time is less than the periodic dirty-writing time, the database still waits until the dirty-writing time is equal to the target time before performing the dirty-writing of the next dirty-writing cycle.

[0085] When the target data block corresponding to the second largest sequence number is written into the local data space and the dirty-writing running time is less than the periodic dirty-writing time, when the database completes the read and write of the target data block corresponding to the second largest sequence number, the database automatically completes and stops the current dirty-writing cycle, re-obtains the target data block and the second data block to enter the next dirty-writing cycle, that is, when the target data block corresponding to the second largest sequence number is completed with persistence, regardless of whether the dirty-writing time is equal to the target time specified by the persistence strategy, the database exits the current dirty-writing cycle and enters the next dirty-writing cycle.

[0086] In a specific implementation manner, the fourth dirty-writing limit value is equal to the first dirty-writing limit value, that is, the fourth dirty-writing limit value is equal to the target dirty-writing quantity corresponding to the previous dirty-writing cycle.

[0087] It can be understood that when the number of the first data is greater than or equal to a preset fourth dirty limit value, and when the sequence difference is greater than or equal to the difference threshold, it can directly indicate that there are more target data blocks in the buffer data space, indirectly indicate that the target dirty quantity is large, and the task of persistence is heavy. At this time, after the dirty task in the current dirty cycle is completed, completing and stopping the current dirty cycle in time and quickly entering the next dirty cycle can reduce the time when the database is in the idle state, effectively improve the read and write efficiency to write the target data block into the local data space in time, and ensure the integrity of the file.

[0088] In addition, the present application also provides a database, which is used to execute the database cache management method in the above embodiments.

[0089] It can be understood that the specific implementation manner of this database is basically the same as the specific embodiments of the above database cache management method, and will not be elaborated here.

[0090] In addition, referring to Figure 9 , Figure 9 illustrates the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0091] A processor 901, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0092] A memory 902, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and are called by the processor 901 to execute the database cache management method of the embodiments of the present application. For example, execute the method steps S110 to step S160 described above in Figure 1 , Figure 2 the method steps S210 to step S220 in Figure 3 the method steps S310 to step S330 in Figure 4 the method steps S410 to step S420 in Figure 5 the method steps S510 to step S520 in Figure 6The method steps S610 to S620 in Figure 7 The method steps S710 to S740 in

[0093] The input / output interface 903 is used to implement information input and output.

[0094] The communication interface 904 is used to implement communication interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0095] The bus 905 transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904).

[0096] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.

[0097] The embodiment of the present application also provides a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and one or more programs can be executed by one or more processors to implement the above database cache management method. For example, execute the Figure 1 The method steps S110 to S160 in Figure 2 The method steps S210 to S220 in Figure 3 The method steps S310 to S330 in Figure 4 The method steps S410 to S420 in Figure 5 The method steps S510 to S520 in Figure 6 The method steps S610 to S620 in Figure 7 The method steps S710 to S740 in

[0098] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0099] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0100] Those skilled in the art can understand that Figures 1 to 8 the technical solutions shown in do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown, or combine certain steps, or different steps.

[0101] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0102] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0103] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0104] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0105] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0106] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0107] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0108] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0109] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.

Claims

1. A database cache management method, characterized in that: include: In the current dirty flushing cycle, a first data block and a second data block in the buffer data space are obtained, wherein the first data block represents data that has not been written into the local data space, and the second data block represents data that has been written into the local data space; Acquire a first data quantity of the first data block and a second data quantity of the second data block, and determine a dirty ratio to be flushed of the buffer data space according to the first data quantity and the second data quantity; Acquire a historical dirty-cleaning speed of a previous dirty-cleaning cycle, and determine a first dirty-cleaning limit value based on the historical dirty-cleaning speed; When the first data quantity is greater than or equal to the first dirty flushing limit, or the proportion of dirty data to be flushed is greater than or equal to a preset second dirty flushing limit, a third dirty flushing limit is obtained, and a first dirty flushing parameter is determined according to the third dirty flushing limit, wherein the third dirty flushing limit represents the maximum number of dirty flushing of data blocks performed by the database in a single dirty flushing cycle; When the first data quantity is less than the first dirty brushing limit, and the proportion of dirty data to be brushed is less than the second dirty brushing limit, the first dirty brushing parameter is equal to the target dirty brushing quantity corresponding to the previous dirty brushing cycle; Obtaining historical dirty flushing numbers of multiple historical periods, and determining a second dirty flushing parameter based on the historical dirty flushing numbers; Obtaining a data sequence number of each of the first data blocks, and determining a third dirty flushing parameter based on the data sequence number; Determine a target dirty flushing parameter according to the first dirty flushing parameter, the second dirty flushing parameter and the third dirty flushing parameter, wherein the target dirty flushing parameter represents the target dirty flushing quantity of the first data block written into the local data space in the current dirty flushing cycle; A target data block is determined according to the target dirty flushing parameter, and the target data block is written into the local data space.

2. The database cache management method according to claim 1, characterized in that: The obtaining of historical dirty flushing numbers of multiple historical periods and determining a second dirty flushing parameter based on all the historical dirty flushing numbers includes: Obtain historical dirty flushing numbers of multiple historical periods, and determine an average dirty flushing number based on all the historical dirty flushing numbers; A second dirty brushing parameter is determined based on the historical dirty brushing quantity.

3. The database cache management method according to claim 2, characterized in that: The obtaining of the data sequence number of each of the first data blocks and determining the third dirty flushing parameter based on the data sequence number includes: Obtaining a data sequence number of each of the first data blocks, determining a first minimum sequence number and a first maximum sequence number in the first data blocks, and determining a sequence difference based on the first minimum sequence number and the first maximum sequence number; A third dirty flushing parameter is determined according to the sequence difference.

4. The database cache management method according to claim 3, characterized in that: The determining a third dirty flushing parameter according to the sequence difference includes: When the sequence difference is greater than or equal to a preset difference threshold, a first difference factor is obtained, and a third dirty flushing parameter is determined according to the first difference factor and the historical dirty flushing quantity; When the sequence difference is less than the difference threshold, a second difference factor is obtained, and the third dirty flushing parameter is determined according to the second difference factor and the historical dirty flushing quantity, wherein the first difference factor is greater than the second difference factor.

5. The database cache management method according to claim 4, characterized in that: The step of determining a target data block according to the target dirty flushing quantity and writing the target data block into the local data space includes: Determine a target data block according to the target dirty flushing quantity, and obtain a second minimum sequence number and a second maximum sequence number in the target data block; Based on the data sequence number, starting from the target data block corresponding to the second smallest sequence number, the target data blocks are sequentially written into the local data space; When the first data quantity is greater than or equal to a preset fourth dirty flushing limit value, and when the sequence difference value is greater than or equal to the difference threshold value, continuously acquiring the dirty flushing running time of the current dirty flushing cycle; Obtain the periodic dirty flushing time. When the target data block corresponding to the second maximum sequence number is written into the local data space and the dirty flushing running time is less than the periodic dirty flushing time, re-acquire the first data block and the second data block in the current buffer data space to enter the next dirty flushing cycle.

6. A database, characterized in that: The database is used to execute the database cache management method described in any one of claims 1 to 5.

7. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the database cache management method according to any one of claims 1 to 5 when executing the computer program.

8. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the database cache management method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Method and device for refreshing page from solid storage device

    CN107870732A

  • Data synchronization method, device and system, apparatus, storage medium and program product

    CN113434476A