Write memory allocation model training method and write memory allocation method and system
By using a write memory allocation model trained through reinforcement learning, the write memory allocation of the log-structured merge tree is dynamically adjusted, which solves the performance and efficiency problems caused by static allocation in the existing technology and achieves more efficient memory utilization and throughput.
Patent Information
- Application Number
- CN202510570845.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-26
AI Technical Summary
In the prior art, the write memory allocation method of the log-structured merge tree is too static, which affects the system performance and efficiency and makes it difficult to adapt to dynamic environments.
A reinforcement learning-based method is used to train the write memory allocation model. The deep Q-network (DQN) is used to optimize the memory allocation strategy through interactive learning with the environment, and the write memory allocation of the log-structured merge tree is dynamically adjusted.
Improves the flexibility and diversity of write memory allocation, reduces overhead, and improves memory utilization and overall throughput.
Smart Images

Figure CN120706495A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of artificial intelligence technology, and in particular to a training method for a write memory allocation model, a write memory allocation method, and a system. Background Art
[0002] Log-Structured Merge Tree (LSM-tree) is a data structure widely used in storage systems (such as databases and file systems) that require efficient write operations. Efficient write memory management is crucial for storage systems to achieve better performance.
[0003] In related technologies, a fixed or even write memory allocation method can be used to allocate write memory to multiple log-structured merge trees. For example, the write memory size of each log-structured merge tree can be fixed, or the total write memory size can be evenly distributed to each log-structured merge tree.
[0004] However, this allocation method is too static and may have a negative impact on system performance and efficiency.
[0005] It is worth noting that the content of the above-mentioned related technologies is only information known to the inventor personally, and does not mean that the above-mentioned information has entered the public domain before the filing date of this specification, nor does it mean that it can become the prior art of this specification. Summary of the Invention
[0006] This specification provides a training method for a write memory allocation model, a write memory allocation method and a system to avoid at least one of the above-mentioned technical problems.
[0007] In a first aspect, this specification provides a training method for a write memory allocation model, wherein the training method is implemented based on reinforcement learning, wherein the reinforcement learning includes an evaluation Q network and a target Q network, and in the Nth training, where N is an integer greater than or equal to 1, including:
[0008] In response to each query request being obtained and meeting a preset requirement, obtaining a current state and a current action, wherein the current state represents the write memory size corresponding to each current log structure merge tree, and the current action represents the current write memory adjustment allocation information corresponding to each log structure merge tree;
[0009] Determining a current reward corresponding to the evaluation Q network executing the current action, wherein the current reward represents performance change information before and after the write memory adjustment corresponding to each log structure merge tree based on the write memory adjustment allocation information; and
[0010] A current sub-sample including the current state, the current action, the current reward, and the determined next state is constructed, wherein the current sub-sample is used to update the target Q network to a write memory allocation model.
[0011] In a second aspect, this specification provides a write memory allocation method, comprising:
[0012] In response to each target query request being obtained meeting a preset requirement, the write memory size of each log structure merge tree and the write operation ratio corresponding to each target query request are obtained respectively;
[0013] The write memory size and the write operation ratio are input into a write memory allocation model to obtain the write memory allocated to each log structure merge tree, wherein the write memory allocation model is trained based on the training method described in the first aspect.
[0014] In a third aspect, this specification provides a training system for writing a memory allocation model, including:
[0015] At least one storage medium storing at least one instruction set for training a write memory allocation model;
[0016] At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, it reads the at least one instruction set and executes the training method as described in the first aspect according to the instructions of the at least one instruction set.
[0017] In a fourth aspect, this specification provides a write memory allocation system, comprising:
[0018] at least one storage medium storing at least one instruction set for allocating write memory;
[0019] At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, it reads the at least one instruction set and executes the write memory allocation method as described in the second aspect according to the instructions of the at least one instruction set.
[0020] In a fifth aspect, this specification provides a computer-readable non-temporary storage medium, wherein the computer-readable non-temporary storage medium stores at least one instruction set, and the at least one instruction set is executed by at least one processor to implement the method described in the first aspect or the second aspect.
[0021] As can be seen from the above technical solutions, the training method, write memory allocation method, and system for the write memory allocation model provided in this specification automatically learn write memory allocation strategies through reinforcement learning, reducing the need for manual intervention and allowing for rapid allocation of write memory to log-structured merge trees requiring more write memory. This system can also adapt to dynamic environments to increase the flexibility and diversity of write memory allocation. Furthermore, by intelligently adjusting the write memory allocation of log-structured merge trees based on performance changes, it is possible to effectively reduce overhead, improve memory utilization, and improve overall throughput.
[0022] Other features of the training method, writing memory allocation method, and system provided in this specification are partially listed in the following description. The creative aspects of the training method, writing memory allocation method, and system provided in this specification can be fully explained by practicing or using the methods, devices, and combinations described in the following detailed examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of this specification, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 A schematic diagram of an application scenario of the training method for the write memory allocation model provided in an embodiment of this specification;
[0025] Figure 2 A schematic diagram of the structure of a training system for a write memory allocation model provided in an embodiment of this specification;
[0026] Figure 3 A flowchart of a method for training a write memory allocation model provided in one embodiment of this specification;
[0027] Figure 4 A flowchart of a method for training a write memory allocation model provided in another embodiment of this specification;
[0028] Figure 5 A schematic diagram illustrating the principles of a method for training a write memory allocation model according to one embodiment of this specification;
[0029] Figure 6 A schematic diagram illustrating the principles of a method for training a write memory allocation model provided in another embodiment of this specification;
[0030] Figure 7 This is a flowchart of the write memory allocation method provided in the embodiments of this specification. DETAILED DESCRIPTION
[0031] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this specification. Rather, they are merely examples of apparatus and methods consistent with certain aspects of this specification, as detailed in the appended claims.
[0032] It should be understood that the terms "including" and "having" and any variations thereof in the embodiments of this specification are intended to cover but not exclude inclusion. For example, a product or device that includes a series of components is not necessarily limited to those components explicitly listed, but may include other components that are not explicitly listed or are inherent to these products or devices.
[0033] In the embodiments of this specification, the term "and / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0034] The term "plurality" in the embodiments of this specification refers to two or more than two, and other quantifiers are similar to it.
[0035] The terms "first," "second," "third," etc., in this specification are used to distinguish between similar or similar objects or entities and are not necessarily intended to limit a particular order or precedence, unless otherwise indicated. It should be understood that the terms used in this manner are interchangeable where appropriate, e.g., capable of being implemented in an order other than that shown or described in the embodiments of this specification.
[0036] The term "unit / module" as used in this specification refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware and / or software code that is capable of performing the functions associated with that element.
[0037] As mentioned above, the log-structured merge tree data structure has been widely used in non-relational database management systems (such as NoSQL systems) in recent years, including LevelDB, RocksDB, Cassandra, HBase, and X-Engine. Unlike traditional in-place update structures, this data structure employs an out-of-place update design.
[0038] The log-structured merge tree is designed to optimize write performance by first writing data to an in-memory data structure (such as a memtable). When the data reaches a certain threshold, it is then written sequentially to an immutable file on disk (such as an SSTable or sorted string table). A background process then merges and compresses these files to maintain data consistency and query efficiency.
[0039] For example, in a log-structured merge-tree, all write operations are first cached in memory and then flushed to disk to form immutable disk components. Disk components play a key role in log-structured merge-trees, performing periodic compaction operations to improve query performance and reclaim space occupied by outdated records. This design enables log-structured merge-trees to better handle large numbers of write operations while also effectively reclaiming space and improving storage efficiency.
[0040] Furthermore, the log-structured merge-tree design provides better data durability and consistency. In traditional in-place update structures, a system failure can result in data loss. In a log-structured merge-tree, all write operations are first recorded in an in-memory log, allowing data recovery even in the event of a system failure.
[0041] In general, the design of the log-structured merge tree gives it great advantages in handling a large number of write operations and improving storage efficiency.
[0042] To avoid at least one of the technical problems mentioned in the background, in some embodiments, machine learning methods can be used to tune memory parameters as database configurations. For example, Gaussian process regression and Bayesian optimization can be used to recommend various database parameter options, including memory options, to allocate appropriate write memory for each log-structured merge tree.
[0043] Among them, machine learning is a multidisciplinary cross-disciplinary major that covers knowledge of probability theory, statistics, approximate theory and complex algorithms. It uses computers as tools and is committed to simulating human learning methods in real time, and divides existing content into knowledge structures to effectively improve learning efficiency.
[0044] However, using machine learning methods to tune databases requires certain domain knowledge and expert experience; in addition, using machine learning may fall into local optimality, and sometimes the results are difficult to reproduce.
[0045] In other embodiments, a rule-based multi-log-structured merge tree allocation method can be used to allocate corresponding write memory to each log-structured merge tree. Unlike machine learning, which treats the memory allocation process as a black box process, the rule-based memory allocation method uses a white box approach to allocate write memory to multiple trees.
[0046] For example, we model the input / output (I / O) cost of each log-structured merge-tree and allocate memory based on the ratio of the write rate of each log-structured merge-tree to the write rate of all log-structured merge-trees. Relatively speaking, hotter log-structured merge-trees receive more write memory than cooler log-structured merge-trees.
[0047] In this embodiment, I / O overhead (hereinafter referred to as overhead) can be understood as resource consumption related to disk reading and writing, for example, the resources consumed by all disk reading and writing operations involved in data processing using a log-structured merge tree.
[0048] However, rule-based approaches only analyze overhead, without considering the additional system overhead (such as CPU overhead) associated with actual write operations and merging log-structured merge trees. Furthermore, such approaches struggle to comprehensively consider performance metrics such as system throughput, latency, and read / write overhead.
[0049] Based on the above analysis, this specification proposes a technical conception that has been creatively developed: a write memory allocation model is obtained based on reinforcement learning training. If reinforcement learning is based on a deep learning framework, it can be specifically implemented as a deep Q-Learning network (DQN). The deep Q network includes an evaluation Q network and a target Q network. The training data set is determined based on the evaluation Q network, and the target Q network is updated based on the training data set to obtain a write memory allocation model. Each sample in the training data set includes the current state, current action, current reward, and next state. The current state is used to represent the write memory size corresponding to each current log structure merge tree; the current action is used to represent the current write memory adjustment allocation information corresponding to each log structure merge tree; the current reward is used to represent the performance change information before and after the write memory adjustment corresponding to each log structure merge tree based on the write memory adjustment allocation information.
[0050] The write memory allocation model can dynamically allocate the corresponding write memory to each log structure merge tree based on performance.
[0051] In addition, by applying reinforcement learning (such as deep Q-network) to memory allocation, which relies mainly on its ability to learn better strategies through interaction with the environment, reinforcement learning can optimize resource allocation in the memory allocation scenario, improve memory usage efficiency and overall system performance.
[0052] For example, in the Linux kernel, memory allocation is typically page-based. Each block of memory is allocated per page. This mechanism facilitates the allocation of large blocks of memory and reduces memory fragmentation. Reinforcement learning can learn when to allocate more pages and how to most efficiently release pages, thereby reducing memory fragmentation and optimizing memory usage.
[0053] For example, in some high-level programming languages, such as Java, memory allocation is object-oriented. Objects are allocated on the heap, while primitive variables are allocated on the stack. Reinforcement learning can learn how to optimize stack allocation strategies based on actual object usage, reducing memory waste and improving memory utilization.
[0054] For example, non-contiguous memory allocation, such as segment management and page management, allows programs to use non-contiguous physical address spaces, improving memory utilization and management flexibility. Reinforcement learning can learn when to use page or segment management and how to dynamically adjust memory allocation strategies to adapt to different workloads.
[0055] Reinforcement learning is an important branch of machine learning that focuses on enabling intelligent agents to learn how to complete specific tasks through interaction with their environment. This learning method simulates the learning process of humans and animals, which involves adjusting behavioral strategies by observing the consequences of their actions (rewards or punishments).
[0056] The Deep Q-Network is a reinforcement learning algorithm that combines Q-learning and deep neural networks. It is designed to solve decision-making problems in high-dimensional state spaces. By using deep neural networks to approximate the action-value function (i.e., the Q-function), the Deep Q-Network enables intelligent agents to learn relatively optimal strategies in complex environments.
[0057] In a deep Q-network, the evaluation Q-network can be used to predict the reward obtained after taking a certain action, thereby guiding the strategy to select the optimal action; the target Q-network is the Q-network used for evaluation, which is usually updated periodically based on the evaluation Q-network to provide a more stable learning goal.
[0058] To facilitate readers' understanding of this manual, the application scenarios of this manual are now introduced.
[0059] The training method for a write memory allocation model provided in this specification (hereinafter referred to as the training method) is applicable to scenarios where a write memory allocation model needs to be trained. For example, the training method provided in this specification can be applied to scenarios such as NoSQL databases, time series databases, search engine indexes, and real-time analysis platforms.
[0060] Take the above NoSQL database scenario as an example:
[0061] NoSQL databases can manage their data storage using log-structured merge trees. Using the training method provided in this manual, a write memory allocation model can be trained. Accordingly, the write memory allocation model can be used to allocate appropriate write memory for each log-structured merge tree in the NoSQL database.
[0062] For descriptions of other application scenarios, please refer to the above description of NoSQL database scenarios, which are not listed here one by one.
[0063] Figure 1 The training method of the embodiment of this specification is applied to the following scenarios: Figure 1 The scene shown is 100. Figure 1 As shown, scenario 100 may include a target user 101 , a client 102 , a server 103 , and a network 104 .
[0064] The target user 101 may be a user who triggers the training to obtain the write memory allocation model. For example, the target user 101 may perform a target operation on the client 102 to trigger the training to obtain the write memory allocation model.
[0065] The client 102 may be an electronic device that provides interactive functions to the target user 101. For example, the client 102 may provide an interactive interface to the target user 101, and the target user 101 may perform interactive operations in the interactive page. In some embodiments, the client 102 executes the training method described in this specification in response to detecting the target operation triggered by the target user 101. At this time, the client 102 may store data or instructions for executing the training method described in this specification and may execute or be used to execute the data or instructions. In some embodiments, the client 102 may include a hardware device with data information processing capabilities and the necessary programs required to drive the hardware device to execute the training method described in this specification.
[0066] In some embodiments, client 102 may include a mobile device, a tablet computer, a laptop computer, a built-in device in a motor vehicle, or the like, or any combination thereof. In some embodiments, the mobile device may include a smart home device, a smart mobile device, a virtual reality device, an augmented reality device, or the like, or any combination thereof. In some embodiments, the smart home device may include a smart television, a desktop computer, or the like, or any combination thereof. In some embodiments, the smart mobile device may include a smartphone, a personal digital assistant, a gaming device, a navigation device, or the like, or any combination thereof. In some embodiments, the built-in device in a motor vehicle may include an onboard computer, an onboard television, or the like.
[0067] In some embodiments, the client 102 may be installed with one or more applications (APPs). APPs can provide the target user 101 with the ability and interface to interact with the outside world via the network 104. APPs include, but are not limited to, web browser APPs, search APPs, chat APPs, shopping APPs, video APPs, financial management APPs, instant messaging tools, email clients, social networking platform software, and the like.
[0068] like Figure 1 As shown, the client 102 can be in communication connection with the server 103. The server 103 can be in communication connection with one client 102 or with multiple clients 102. In some embodiments, the client 102 can interact with the server 103 via the network 104 to receive or send messages, etc.
[0069] The server 103 may be a server that provides various services. For example, the server 103 may be a cloud server or a local server. The server 103 may be connected to a client 102 and receive data sent by the client 102, or may be connected to multiple clients 102 and receive data sent by each client 102.
[0070] In some embodiments, the training methods described herein can be executed on server 103. In this case, server 103 can store data or instructions for executing the training methods described herein and can execute or be used to execute the data or instructions. Server 103 can include hardware devices capable of data information processing and the necessary programs to operate the hardware devices.
[0071] The network 104 is a medium for providing a communication connection between the client 102 and the server 103. The network 104 can facilitate the exchange of information or data. Figure 1As shown, the client 102 and the server 103 can be connected to the network 104 respectively, and transmit information or data to each other through the network 104.
[0072] In some embodiments, the network 104 can be any type of wired or wireless network, or a combination thereof. For example, the network 104 can include a cable network, a wired network, a fiber optic network, a telecommunications network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth™ network, a ZigBee™ network, a near field communication (NFC) network, or the like.
[0073] In some embodiments, network 104 may include one or more network access points. For example, network 104 may include a wired or wireless network access point, such as a base station or an Internet exchange point, through which one or more components of client 102 and server 103 can connect to network 104 to exchange data or information.
[0074] It is worth mentioning that Figure 1 The number of clients 102, servers 103, and networks 104 in the examples is merely illustrative. Any number of clients 102, servers 103, and networks 104 may be used as needed. Furthermore, the training method provided herein may be executed entirely on the client 102, entirely on the server 103, or partially on the client 102 and partially on the server 103.
[0075] That is to say, Figure 1 and targeting Figure 1 The above description is only used to illustrate the application scenarios to which the training method of this specification may be applicable, and should not be understood as a limitation on the application scenarios.
[0076] Figure 2The hardware structure diagram of a training system 200 provided according to an embodiment of this specification is shown. Training system 200 can execute the training method described in this specification. The training method is described elsewhere in this specification. When the training method is executed on client 102, training system 200 can be client 102. When the training method is executed on server 103, training system 200 can be server 103. When the training method is partially executed on client 102 and partially executed on server 103, training system 200 can be a system including client 102 and server 103.
[0077] like Figure 2 As shown, training system 200 may include at least one storage medium 203 and at least one processor 202. In some embodiments, training system 200 may further include a communication port 204 and an internal communication bus 201. Training system 200 may further include an I / O component 205.
[0078] The internal communication bus 201 can connect different system components. For example, the internal communication bus 201 can connect the storage medium 203, the processor 202, the communication port 204 and the I / O component 205.
[0079] I / O components 205 support input / output between training system 200 and other components.
[0080] Communication port 204 is used for data communication between training system 200 and the outside world. For example, communication port 204 can be used for data communication between training system 200 and network 104. Communication port 204 can be a wired communication port or a wireless communication port.
[0081] Storage medium 203 may include a data storage device. The data storage device may be a non-transitory storage medium or a temporary storage medium. For example, the data storage device may include one or more of a disk 2031, a read-only storage medium (ROM) 2032, or a random access storage medium (RAM) 2033. Storage medium 203 also includes at least one instruction set stored in the data storage device. The instruction set includes computer program code, which may include programs, routines, objects, components, data structures, processes, modules, etc. that execute the training methods provided in this specification.
[0082] At least one processor 202 can be communicatively connected to at least one storage medium 203. The at least one processor 202 is configured to execute the at least one instruction set described above. When the training system 200 is running, the at least one processor 202 reads the at least one instruction set and, according to the instructions of the at least one instruction set, executes the training method provided herein. The processor 202 can perform all steps included in the training method. The processor 202 can be in the form of one or more processors. In some embodiments, the processor 202 can include one or more hardware processors, such as a microcontroller, a microprocessor, a reduced instruction set computer (RISC), an application-specific integrated circuit (ASIC), an application-specific instruction set processor (ASIP), a central processing unit (CPU), a graphics processing unit (GPU), a physical processing unit (PPU), a microcontroller unit, a digital signal processor (DSP), a field programmable gate array (FPGA), an advanced RISC machine (ARM), a programmable logic device (PLD), any circuit or processor capable of performing one or more functions, or any combination thereof.
[0083] For illustrative purposes only, only one processor 202 is shown in the training system 200 in the accompanying drawings. However, it should be noted that the training system 200 described herein may also include multiple processors. Therefore, the operations and / or method steps disclosed herein may be performed by a single processor or jointly by multiple processors. For example, if the training system 200 is described herein as processor 202 performing steps A and B, it should be understood that steps A and B may also be performed jointly or separately by two different processors 202 (e.g., a first processor performing step A and a second processor performing step B, or a first and a second processor performing steps A and B together).
[0084] See also Figure 3 , Figure 3 This is a flow chart of a method for training a write memory allocation model according to one embodiment of this specification. Figure 3 The training method shown may be performed by a training system. For a description of the training system, please refer to the above examples and will not be repeated here.
[0085] in addition, Figure 3 The training method shown is based on reinforcement learning. Combined with the above analysis, it can be seen that reinforcement learning includes evaluation Q network and target Q network. Reinforcement learning includes multiple iterative training. In the Nth training, N is an integer greater than or equal to 1, such as Figure 3 As shown, it includes the following S301 to S303:
[0086] S301: In response to each query request obtained meeting the preset requirements, obtain the current state S tand the current action a t , where the current state represents the write memory size corresponding to each log structure merge tree, and the current action represents the current write memory adjustment allocation information corresponding to each log structure merge tree.
[0087] A query request may be a query request initiated by a user or a system to obtain data or perform corresponding operations, such as a write data operation (referred to as a write operation for short).
[0088] The preset requirements can be determined by the training system based on requirements, historical records, experiments, etc., and are not limited in this embodiment. For example, the preset requirements can be a limit on the amount of query requests or a limit on the content of query requests.
[0089] For example, the training system can detect the query request obtained to determine whether it meets the preset requirements, and if the preset requirements are met, perform the operation of obtaining the current state and current action.
[0090] The current state can be understood as the current state of the application scenario system. Taking the application scenario system as an example, the current state can be understood as the write memory size corresponding to each log structure merge tree in the database at the current moment.
[0091] The current action can be understood as a control behavior for the database. For example, at the current moment, it indicates the adjustment information for the write memory allocation corresponding to each log-structured merge tree in the database. Specifically, for example, the write memory allocation for a log-structured merge tree is adjusted from allocation size A to allocation size B. In other words, the current action indicates how the write memory allocation for each log-structured merge tree is adjusted at the current moment.
[0092] Correspondingly, this step can be understood as, when each query request at the current moment meets the preset requirements, the training system can obtain the write memory size corresponding to each log structure merge tree at the current moment, and obtain information on how to adjust the write memory size allocated to each log structure merge tree at the current moment.
[0093] S302: Determine the current reward r for evaluating the Q network's execution of the current action t , where the current reward representation is based on the write memory adjustment allocation information, and the performance change information before and after the write memory adjustment corresponding to each log structure merge tree is obtained.
[0094] The current reward can be understood as a measure of how good the current action is based on the performance change of the database after the current action is taken.
[0095] Accordingly, the evaluation Q network can execute the current action to obtain the corresponding current reward. That is, by executing the information on how to adjust the write memory size allocated to each log structure merge tree at the current moment, the performance change of the data system before and after the write memory adjustment can be obtained.
[0096] S303: Construct the current state, current action, current reward, and the next state S t+1 The current subsample is used to update the target Q network to the write memory allocation model.
[0097] It is understood that, to avoid tedious descriptions, the present embodiment will not further elaborate on the same or similar features already described in the above examples. For example, the understanding of the next state can refer to the current state. Specifically, the next state can be understood as the new state of the database after the current action is executed.
[0098] The next state may be determined by the current state and the current action performed in the current state. The next state may reflect, for example, how the storage system evaluates changes in the Q network after taking the current action.
[0099] In some embodiments, a pre-built state transition function can be used to calculate the next state. For example, a training system uses a state transition function to calculate the next state based on the current state and the current action. The state transition function can be pre-deployed in the training system or learned, and this embodiment does not limit this.
[0100] After the training system executes S301 to S303 once, the corresponding subsample (S t ,a t ,r t ,S t+1 The training system can obtain samples corresponding to each step by executing S301 to S303 multiple times. The training system can update the target Q network based on each sample to obtain a write memory allocation model. For example, the training system can use a gradient descent algorithm to update the evaluation Q network based on each sample and update the target Q network at corresponding steps.
[0101] Accordingly, a training system or other system can utilize the write memory allocation model to allocate write memory for each log-structured merge tree. In this embodiment, by automatically learning a relatively optimal write memory allocation strategy using reinforcement learning, the need for manual intervention is reduced, allowing for rapid allocation of write memory for log-structured merge trees that require more write memory. This can also adapt to dynamic environments to increase the flexibility and diversity of write memory allocation. Furthermore, by intelligently adjusting the write memory allocation of log-structured merge trees based on performance changes, overhead can be effectively reduced, improving memory utilization and overall throughput.
[0102] In order to make readers more deeply understand the technical principles of the training method provided in this manual, Figures 4 to 6 The training method provided in this manual is explained in more detail.
[0103] in, Figure 4 This is a flow chart of a method for training a write memory allocation model provided in another embodiment of this specification. Figure 4 As shown, in the Nth training, the method includes:
[0104] S401: Execute the workload. If the key-value pairs (Key-Value Pair,<K,V> When a pair of log structures reaches a preset queue size, the write ratio, current status, and request type of each query request corresponding to each log structure merge tree are obtained. The request type includes at least one of a null point type, a non-null point type, a range type, and a write type.
[0105] Similarly, in order to avoid tedious description, the present embodiment will not repeat the same or similar technical features as those in the above examples. For example, regarding the current state of the present embodiment, reference can be made to the description in the above examples.
[0106] For example, Figure 5 As shown in Figure 2, workload can be understood as a series of operations or tasks that the training system needs to process. It can be query requests initiated by users.
[0107] Continue reading Figure 5 , the data in the table can be understood as specific tasks in the workload. Figure 5 As shown, the data in the table may include fields such as time (time), column family (cf), identifier (ID), size (size), and query operation (query_operation).
[0108] In addition, the number of log structure merge trees can be n+1, where n is an integer greater than or equal to 1. Figure 5 As shown, the n+1 log structure merge trees include: Tree_0, Tree_1 to Tree_n.
[0109] Accordingly, continue to combine the above examples and Figure 5 , the training system can allocate corresponding write memory for n+1 log structure merge trees from the write memory pool (Write MemoryPool).
[0110] A key is a unique identifier used to index data. A value is the actual data associated with a key. Key-value pairs are the fundamental data units of many modern storage systems (such as NoSQL databases and distributed storage systems). Based on the above analysis, in a log-structured merge tree, written key-value pairs are typically first stored in the in-memory MemTable and then flushed to the on-disk SSTable.
[0111] Each query request has a corresponding key-value pair. The key-value pairs can be cached in a queue until a certain size (such as a preset queue size) is reached.
[0112] The preset queue size may be determined by the training system based on demand, historical records, experiments, etc., and is not limited in this embodiment.
[0113] Accordingly, in conjunction with the above example, if the key-value pairs corresponding to each query request reach the preset queue size, the training system determines that each query request meets the preset requirements and can proceed with subsequent training operations. This queue-based training operation improves training efficiency and resource utilization, ensuring effective training.
[0114] The write operation ratio, also known as the write rate or write request ratio, can be understood as the ratio of write operations to the total number of operations (including read and write operations) in a database or storage system.
[0115] Different query requests may correspond to different types.
[0116] The null point (query) type can be understood as an operation to find a key-value pair that does not exist in the database. This type usually requires traversing the entire data structure (including the MemTable in memory and the SSTable file on disk), but the final result returned is "not found".
[0117] The non-empty point (query) type can be understood as involving finding key-value pairs that actually exist in the database. This type may also require accessing multiple layers of data structures, but since the target data does exist, the query can terminate at the location where the data is found.
[0118] Range queries are queries that retrieve a series of consecutive key-value pairs in a certain order (e.g., key sorting). When executing range queries in a log-structured merge tree, multiple SSTables need to be scanned simultaneously, and data from different levels may need to be merged.
[0119] The write type can be understood as that the new write or update is first added to the MemTable in memory. Once the MemTable reaches the preset size, it will be refreshed to an immutable state and written to disk to form a new SSTable.
[0120] In some embodiments, an initialization operation is also included before training.
[0121] For example, Figure 6 As shown, the initialization operation may include: initializing the experience pool, the evaluation Q network, the target Q network, and the state.
[0122] For example, initialize the capacity of the experience pool; initialize the evaluation Q network and randomly generate weights ω; initialize the target Q network and randomly generate weights ω \ , and ω ' =ω; the initialization state can generate the state randomly.
[0123] S402: Determine the current action according to the current state and the write operation ratio.
[0124] For example, the training system can generate a current action related to the write operation ratio based on the current state and the ε-greedy policy, and the training system can execute the current action based on the evaluation Q network.
[0125] Relatively speaking, the training system determines the current action by combining the current action and the write operation ratio, which can make the write memory allocation highly consistent with the corresponding write memory size and corresponding write memory demand of each log structure merge tree, thereby improving the effectiveness and reliability of training.
[0126] In some embodiments, the current action is one of the following:
[0127] a) Allocate write memory for log-structured merge-trees that increase write operations.
[0128] b) Allocate corresponding write memory to each log structure merge tree based on the write operation ratio.
[0129] c) Allocate corresponding write memory to each log structure merge tree based on the write operation ratio, and adjust the allocated write memory based on the average memory utilization of each log structure merge tree.
[0130] The current action a) can be understood as requesting a memory page from the memory pool when a write operation arrives on each log structure merge tree.
[0131] The current action (b) can be understood as allocating write memory to each log-structured merge tree based on the write operation ratio corresponding to each log-structured merge tree during the queue size period. For example, for a particular log-structured merge tree, the corresponding memory page can be allocated from the memory pool based on the write operation ratio of the log-structured merge tree.
[0132] The current action c) can be understood as first allocating write memory to each log-structured merge tree based on the total write operation ratio. Then, the memory utilization of each log-structured merge tree can be calculated when the key-value pairs reach 1 / 4, 1 / 2, 3 / 4, and 1, respectively. When the ratio of the actual occupied write memory to the allocated write memory is less than the preset threshold, the allocated write memory is reduced; otherwise, the current write memory size is maintained.
[0133] Similarly, the preset threshold value can be determined by the training system based on demand, historical records, experiments, etc., and this embodiment does not limit it.
[0134] In this embodiment, the current action may be one of a), b), or c), so as to perform different actions in different states, thereby improving the effectiveness and reliability of training.
[0135] S403: Obtain the total cost corresponding to each query request.
[0136] The total cost can be understood as the sum of the costs corresponding to each query request. For example, the total cost is accumulated based on the costs of different query requests.
[0137] Based on the above analysis, different query requests may correspond to different request types. Therefore, the training system can determine the total cost based on the request type corresponding to each query request. For example, the training system can use different methods to determine the corresponding cost for different request types and determine the total cost based on the cost corresponding to each request type.
[0138] In other words, when calculating the total cost, the training system can first determine the request type corresponding to each query request. Then, based on the characteristics of each request type, it uses a targeted method to calculate the query cost corresponding to each request type. Finally, the query cost of each request type is added together to obtain the final total cost.
[0139] In contrast, by using different methods to determine the cost of query requests of different request types, it is possible to achieve diversity and flexibility in cost determination, thereby improving the accuracy and reliability of the determined total cost.
[0140] In some embodiments, different methods are used to determine corresponding overheads for different request types, including the following four cases:
[0141] Case 1: If it is a null point type, the corresponding overhead is determined based on the number of layers of the log structure merge tree and the false positive rate of the Bloom filter at each layer.
[0142] A Bloom filter can be thought of as a relatively space-efficient probabilistic data structure. It can be used to test whether an element belongs to a set. It can quickly determine whether an element is likely to be in the set (a true positive or false positive) or not (a true negative).
[0143] An important feature of a Bloom filter is that it can produce false positives (i.e. elements that are not actually in the set are incorrectly marked as possibly present), but it cannot produce false negatives (if the Bloom filter says an element does not exist, it does not exist).
[0144] Therefore, the false positive rate can be understood as the probability that an element is considered to be possible when checking it using a Bloom filter when it is not actually part of the set. The false positive rate is related to the size of the Bloom filter (the length of the bit array) and the number of hash functions.
[0145] Generally speaking, the larger the Bloom filter, the more appropriate the number of hash functions, and the lower the false positive rate. The false positive rate of each layer of Bloom filter In a log-structured merge tree, data is distributed across multiple levels, and each level contains one or more SSTables. In order to speed up queries, each SSTable is usually associated with a Bloom filter. This means that when trying to query whether a key value exists, you can first query the corresponding Bloom filter to quickly determine whether the key value is likely to exist in the corresponding SSTable, thereby avoiding unnecessary disk I / O operations.
[0146] In other words, the false positive rate of the Bloom filter at each level of the log-structured merge tree can be used to describe the probability that the Bloom filter designed for each level in the log-structured merge tree architecture may make a false positive when identifying non-member elements.
[0147] In some embodiments, the training system may determine the overhead corresponding to Case 1 based on Equation 1:
[0148]
[0149] Where L(T) is the total number of layers of the log structure merge tree, f i (T) is the false positive rate of the Bloom filter at each layer.
[0150] Case 2: If it is a non-empty point type, the corresponding overhead is determined based on the number of layers of the log structure merge tree, the layer capacity ratio, the total write memory, and the size of the query request.
[0151] The tier capacity ratio can be understood as the capacity ratio between different tiers. This ratio defines the amount of data that each tier can store relative to other tiers.
[0152] In some embodiments, the training system may determine the overhead corresponding to Case 2 based on Equation 2:
[0153]
[0154] Where T is the level capacity ratio, m is the number of entries (key-value pairs) that fill the total level in the log structure merge tree, and buf is the total write memory size, E is the size of a query request, L(T) is the total number of layers of the log structure merge tree, and f i (T) is the false positive rate of the Bloom filter at each layer.
[0155] For example, the calculation of the overhead corresponding to the non-empty point type can be divided into two parts:
[0156] Part 1: Assuming that the probability of a point query finding a non-empty result in each level of the log-structured merge tree is proportional to the size of the level, then the unit I / O operation cost at level i satisfies The query probability.
[0157] Part II: Assume that all levels before level i will trigger I / O operations with a probability equivalent to the false positive rate of the Bloom filter at these levels. Similar to the null point query, the cost of the I / O operation that fails to query at the above levels is
[0158] Accordingly, by combining the above two parts, the overhead calculated as shown in Formula 2 can be obtained.
[0159] Case 3: If it is a range type, the corresponding overhead is determined based on the number of disk seeks, the cumulative number of pages obtained, the preset page size, and the total number of pages.
[0160] For example, the range type may issue L(T) disk seeks (one per run) at the level, each followed by a sequential scan. The cumulative number of pages scanned in all runs is S RQ , where S RQ is the average proportion of all entries included in the range lookup.
[0161] In some embodiments, the training system may determine the overhead corresponding to Case 3 based on Equation 3:
[0162]
[0163] Among them, L(T) is the number of disk seeks, S RQ is the cumulative number of pages, N is the total number of pages, and B is the preset page size.
[0164] Case 4: If it is a write type, the corresponding overhead is determined based on the number of levels of the log structure merge tree, the level capacity ratio, the preset page size, and the preset storage asymmetry parameters.
[0165] In some embodiments, the training system may determine the overhead corresponding to Case 4 based on Equation 4:
[0166]
[0167] Where L(T) is the total number of layers in the log structure merge tree, B is the preset page size, T is the layer capacity ratio, and A rw Stores asymmetric parameters for presets.
[0168] For example, a log-structured merge tree can use solid-state storage with asymmetric read and write costs. This storage asymmetry can be expressed as A rw If a write operation costs twice as much as a read operation, then A rw =2.
[0169] Based on the above analysis, we can see that by using different calculation methods to determine the overhead of different request types, the accuracy and reliability of the total overhead can be improved. In addition, the overhead before and after write memory adjustment can also be accurately and reliably analyzed.
[0170] S404: Determine a current reward corresponding to the current action performed by the Q network based on the total overhead, wherein the current reward represents performance change information before and after the write memory adjustment corresponding to each log structure merge tree based on the write memory adjustment allocation information, and the performance change information includes overhead change information.
[0171] Combining the above examples and Figure 6 , after obtaining the current state and determining the current action based on the current state, the network can perform the current action based on the evaluation Q. Accordingly, the training system can obtain the reward for the current action (i.e., the current reward).
[0172] For example, the total cost can be understood as the total cost before executing the current action. After the evaluation network executes the current action, the training system can still use the above method to obtain the corresponding total cost. Accordingly, the training system can determine the current reward based on the total cost before and after the evaluation network executes the current action.
[0173] Continuing with the above example and Figure 5 , the evaluation Q network can execute the current action for writing to the memory pool. Correspondingly, the training system will receive the evaluation effect (i.e., the current reward) of the evaluation Q network executing the current action.
[0174] Combined with the above analysis of S403 and S404, we can see that in this example, by determining the current reward by evaluating the total cost before and after the Q network executes the current action, the trained write memory allocation model can be highly aligned with the cost dimension to allocate the corresponding write memory to each log-structured merge tree. This ensures that the cost meets the scenario requirements, such as ensuring that the storage system cost is as low as possible, while quickly allocating the corresponding write memory to each log-structured merge tree.
[0175] In some embodiments, the performance change information also includes memory utilization change information and / or throughput change information.
[0176] For example, taking the performance change information including memory utilization change information as an example:
[0177] The training system can determine the current reward by evaluating the memory utilization of the Q network before and after executing the current action. This allows the trained write memory allocation model to closely align with the memory utilization dimension and allocate the corresponding write memory to each log-structured merge tree. This ensures that memory utilization meets scenario requirements while quickly allocating the corresponding write memory to each log-structured merge tree, such as ensuring the highest possible memory utilization of the storage system.
[0178] Similarly, taking performance change information including throughput change information as an example:
[0179] The training system can determine the current reward by evaluating the throughput of the Q network before and after executing the current action. This allows the trained write memory allocation model to closely align with the throughput dimension and allocate the corresponding write memory to each log-structured merge tree. This ensures that the memory throughput meets the scenario requirements while quickly allocating the corresponding write memory to each log-structured merge tree, such as ensuring the highest possible memory utilization of the storage system.
[0180] That is to say, in this embodiment, based on the write memory allocation model obtained by training the training system, the size of the write memory allocated to the log-structured merge tree can be adjusted as the workload changes; or the size of the write memory allocated to the log-structured merge tree can be adjusted as the memory utilization changes; or the size of the write memory allocated to the log-structured merge tree can be adjusted as the throughput changes; or the size of the write memory allocated to the log-structured merge tree can be adjusted as multiple dimensions (such as at least two of the overhead, memory utilization, and throughput) change.
[0181] In some embodiments, the current reward can represent the normalized difference between the cost, memory utilization, and throughput before and after adjusting the memory write for the current action. That is, the current reward obtained by the training system is determined by combining the cost, memory utilization, and throughput metrics.
[0182] For example, the current reward Reward can be expressed based on Formula 5:
[0183]
[0184] Where i is the log structure merge tree, n is the total number of log structure merge trees, c is the overhead before write memory adjustment, c ' is the overhead after write memory adjustment, u is the memory utilization before write memory adjustment, and u ' is the memory utilization after write memory adjustment, Th is the throughput before write memory adjustment, and Th ' Throughput adjusted for write memory.
[0185] In this embodiment, especially when determining the current reward based on the three dimensions of overhead, memory utilization, and throughput, it is possible to quickly allocate corresponding write memory to log-structured merge trees that require more memory. Furthermore, this allows for high overall throughput and low read and write latency while ensuring high memory utilization and minimal I / O overhead for storage systems.
[0186] S405: Construct a current sub-sample including the current state, current action, current reward, and the determined next state, wherein the current sub-sample is used to update the target Q network to a write memory allocation model.
[0187] Similarly, regarding the implementation principle of S405, reference can be made to the description of S303 in the above example, which will not be repeated here.
[0188] In some embodiments, continue to combine the examples and Figure 5 The training system can store the current sample in the experience replay buffer, which is equivalent to the experience pool. For example, the experience replay buffer is used to store historical interaction data (such as state, action, reward, and next state). In other words, the experience replay buffer stores the corresponding batch (buffer) of samples obtained from each training iteration, such as the current batch of samples (current state, current action, current reward, and next state).
[0189] Accordingly, the training system can (randomly) draw samples from the experience replay buffer to update the target Q network to the write memory allocation model.
[0190] For example, continuing with the above example and Figure 6, the training system can randomly extract a sample data set (including multiple batches of samples) from the experience pool. The training system can determine whether the time step j+1 of the current sample is a terminal state.
[0191] Continue reading Figure 6 , if the judgment result is yes, then let the corresponding immediate reward r j (i.e., the reward after executing the corresponding action at time step j) = target Q value y j (Representing the expected future return under the current strategy.) Furthermore, on the one hand, the training system returns to determine whether a preset number of iterations (or rounds) has been reached; on the other hand, the training system can update the target Q network at a preset number of steps to keep the network parameters of the target Q network consistent with those of the evaluation Q network.
[0192] Conversely, if the judgment result is negative, the training system can use the gradient descent algorithm to update the evaluation Q network. Similarly, after the update is completed, the training system can update the target Q network at a preset number of steps to keep the network parameters of the target Q network consistent with the parameters of the evaluation Q network.
[0193] In some embodiments, the gradient algorithm can be expressed by Equation 6:
[0194] (x j -Q(S j ,a j ;ω)) 2
[0195] Among them, y j is the target Q value, and y j =r j +γmax a′ Q(S t+1 ,a′);r j is the immediate reward, γ is the preset discount factor, S t+1 is the state of the next step, a′ is all possible actions; Q() is the Q function; S j is the state corresponding to time step j (which can be called the instantaneous state); a j is the action corresponding to time step j (which can be called immediate action).
[0196] Based on the training of the write memory allocation model, this specification also provides a write memory allocation method. Regarding the application scenario of the write memory allocation method, please refer to the description of the application scenario of the training method in the above example, which will not be repeated here.
[0197] Similarly, the system for executing the write memory allocation method may be a write memory allocation system. Regarding the structure of the write memory allocation system, reference may be made to the structural description of the training system in the above example, which will not be repeated here.
[0198] It is understood that the write memory allocation system and the training system can be the same system or different systems. If they are different systems, the training system can transfer the trained write memory allocation model to the write memory allocation system. Accordingly, the write memory allocation system can dynamically allocate the corresponding write memory for each log structure merge tree based on the write memory allocation model.
[0199] Alternatively, the training system can also provide an interface for invoking the write memory allocation service of the write memory allocation model. Accordingly, the write memory allocation system can invoke the write memory allocation service of the write memory allocation model through this interface and dynamically allocate the corresponding write memory for each log structure merge tree based on the write memory allocation service.
[0200] See also Figure 7 , Figure 7 This is a flowchart of the write memory allocation method provided in the embodiment of this specification. Figure 7 As shown, the method includes the following S701 and S702:
[0201] S701: In response to each target query request being obtained meeting a preset requirement, the write memory size of each log structure merge tree and the write operation ratio corresponding to each target query request are obtained respectively.
[0202] Similarly, regarding the same or similar technical features of this embodiment and the above example, please refer to the description of the above example, which will not be repeated here. For example, regarding the preset requirements, write operation ratio, etc., please refer to the above example.
[0203] S702: Input the write memory size and the write operation ratio into the write memory allocation model to obtain the write memory allocated to each log structure merge tree, wherein the write memory allocation model is trained based on the training method described in any of the above embodiments.
[0204] Exemplarily, the input of the write memory allocation model includes the write memory size and the write operation ratio; the output of the write memory allocation model is the write memory allocation information, that is, the write memory allocated to each log structure merge tree.
[0205] For example, referring to the above example, the allocation system obtains the write memory size corresponding to Tree_0, Tree_1, through Tree_n, and obtains the write operation ratio. The allocation system inputs the write memory size and write operation ratio into the write memory allocation model deployed within it. The write memory allocation model dynamically determines the write memory allocation information for Tree_0, Tree_1, through Tree_n based on actual production environment workloads, high memory utilization, throughput, overhead, etc., such as the write memory size allocated to Tree_0, Tree_1, through Tree_n.
[0206] This allows for quick allocation of write memory for Tree_0, Tree_1, and finally Tree_n, resulting in higher throughput, lower read and write latency, and at the same time ensuring higher memory utilization and lower overhead.
[0207] It is worth noting that the above examples are only used to illustrate the possible implementation methods of the training method and the write memory allocation method of this specification, and should not be understood as limiting the implementation methods of the training method and the write memory allocation method of this specification. For example, based on the above technical concept, some of the above technical features can be combined to obtain a new embodiment; new technical features can be added on the basis of the above examples to obtain a new embodiment; some technical features can be reduced on the basis of the above examples to obtain a new embodiment; some technical features in the above examples can be replaced with other technical features; some technical features and their order in the above examples can be adjusted to obtain a new embodiment, etc., which will not be listed here one by one.
[0208] Based on the above technical concept, this specification also provides a computer-readable non-temporary storage medium, which stores at least one instruction set. When the at least one instruction set is executed by the processor, the steps of the training method and the write memory allocation method described in this specification are implemented.
[0209] In some possible implementations, various aspects of this specification may also be implemented in the form of a program product comprising program code. Taking a training method as an example, when the program product is executed on training system 200, the program code is used to cause training system 200 to perform the steps of the training method described in this specification. The program product for implementing the above method may utilize a portable compact disc read-only memory (CD-ROM) comprising program code and may be executed on training system 200. However, the program product of this specification is not limited thereto. In this specification, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system. The program product may utilize any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or any combination thereof. More specific examples of computer-readable storage media include: an electrical connection having one or more conductors, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable storage medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing. Program code for performing the operations described herein may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the training system 200, partially on the training system 200, as a stand-alone software package, partially on the training system 200 and partially on a remote training system, or entirely on the remote training system 200.
[0210] In addition, when writing the memory allocation method in the form of a program product, please refer to the description of the training method, which will not be listed here one by one.
[0211] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0212] In summary, after reading this detailed disclosure, those skilled in the art will appreciate that the foregoing detailed disclosure may be presented by way of example only and may not be limiting. Although not expressly stated herein, those skilled in the art will understand that this specification encompasses various reasonable changes, improvements, and modifications to the embodiments. Such changes, improvements, and modifications are intended to be suggested by this specification and are within the spirit and scope of the exemplary embodiments of this specification.
[0213] Furthermore, certain terms in this specification have been used to describe embodiments of this specification. For example, “one embodiment,” “an embodiment,” and / or “some embodiments” mean that a particular feature, structure, or characteristic described in connection with that embodiment may be included in at least one embodiment of this specification. Therefore, it is emphasized and should be understood that two or more references to “an embodiment,” “one embodiment,” or “an alternative embodiment” in various parts of this specification do not necessarily refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be appropriately combined in one or more embodiments of this specification.
[0214] It should be understood that in the foregoing descriptions of the embodiments of this specification, to facilitate understanding of a feature and to simplify this specification, various features are combined in a single embodiment, figure, or description thereof. However, this does not necessarily mean that these features are combined. When reading this specification, a person skilled in the art may label some of the devices as separate embodiments. In other words, the embodiments of this specification can also be understood as the integration of multiple sub-embodiments. The content of each sub-embodiment is also valid even when it includes fewer than all the features of a single previously disclosed embodiment.
[0215] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, and the like, cited herein (excluding any historical review documents related thereto) is hereby incorporated by reference for all purposes relevant to this document, such as within the specification and claims herein. However, if there is any inconsistency or conflict between the descriptions, definitions, and / or terminology of such materials and the descriptions, definitions, and / or terminology used herein, the descriptions, definitions, and / or terminology used herein shall control.
[0216] Finally, it should be understood that the embodiments of the application disclosed herein are illustrative of the principles of the embodiments of this specification. Other modified embodiments are also within the scope of this specification. Therefore, the embodiments disclosed in this specification are merely examples and not limitations. Those skilled in the art can adopt alternative configurations based on the embodiments in this specification to implement the application in this specification. Therefore, the embodiments of this specification are not limited to the embodiments precisely described in the application.
Claims
1. A method for training a write memory allocation model, the method being implemented based on reinforcement learning, the reinforcement learning including an evaluation Q network and a target Q network, wherein in an Nth training run, where N is an integer greater than or equal to 1, the method comprises: In response to each query request being obtained and meeting a preset requirement, obtaining a current state and a current action, wherein the current state represents the write memory size corresponding to each current log structure merge tree, and the current action represents the current write memory adjustment allocation information corresponding to each log structure merge tree; Determining a current reward for the evaluation Q network to execute the current action, wherein the current reward represents performance change information before and after the write memory adjustment corresponding to each log structure merge tree based on the write memory adjustment allocation information; and A current sub-sample including the current state, the current action, the current reward, and the determined next state is constructed, wherein the current sub-sample is used to update the target Q network to a write memory allocation model.
2. The method according to claim 1, wherein The method further comprises: In response to each query request meeting the preset requirement, obtaining the write operation ratio corresponding to each log structure merge tree; And, obtaining the current action includes: determining the current action based on the current state and the write operation ratio.
3. The method according to claim 2, wherein the current action is one of the following: Allocate write memory for log-structured merge trees that increase write operations; Allocate corresponding write memory to each log structure merge tree according to the write operation ratio; A corresponding write memory is allocated to each log structure merge tree according to the write operation ratio, and the allocated write memory is adjusted according to the average memory utilization of each log structure merge tree.
4. The method according to any one of claims 1 to 3, wherein The method further comprises: In response to each query request meeting the preset requirement, obtaining a total cost corresponding to each query request; And, determining the current reward corresponding to the evaluation Q network performing the current action includes: determining the current reward corresponding to the evaluation Q network performing the current action based on the total overhead, wherein the performance change information includes overhead change information.
5. The method according to claim 4, wherein Obtaining the total cost corresponding to each query request includes: Determine a request type corresponding to each query request, where the request type includes at least one of a null point type, a non-null point type, a range type, and a write type; Different methods are used to determine the corresponding overhead for different request types; and The total cost is determined according to the costs corresponding to each request type.
6. The method according to claim 5, wherein: Different methods are used to determine the corresponding overhead for different request types, including: If it is the null point type, the corresponding overhead is determined according to the number of layers of the log structure merge tree and the false positive rate of the Bloom filter of each layer; If it is the non-empty point type, the corresponding overhead is determined based on the number of layers of the log structure merge tree, the layer capacity ratio, the total write memory, and the size of the query request; If it is the range type, then determine the corresponding overhead based on the number of disk seeks, the cumulative number of pages obtained, the preset page size, and the total number of pages; If it is the write type, the corresponding overhead is determined according to the number of layers of the log structure merge tree, the layer capacity ratio, the preset page size, and the preset storage asymmetry parameters.
7. The method according to claim 4, wherein: The performance change information also includes memory utilization change information and / or throughput change information.
8. A write memory allocation method, comprising: In response to each target query request being obtained meeting a preset requirement, the write memory size of each log structure merge tree and the write operation ratio corresponding to each target query request are obtained respectively; The write memory size and the write operation ratio are input into a write memory allocation model to obtain the write memory allocated to each of the log structure merge trees, wherein the write memory allocation model is trained based on the training method described in any one of claims 1 to 7.
9. A training system for a write memory allocation model, comprising: At least one storage medium storing at least one instruction set for training a write memory allocation model; At least one processor is communicatively connected to the at least one storage medium, wherein when the at least one processor is running, it reads the at least one instruction set and executes the training method as described in any one of claims 1 to 7 according to the instructions of the at least one instruction set.
10. A write memory allocation system comprising: at least one storage medium storing at least one instruction set for allocating write memory; At least one processor is communicatively connected to the at least one storage medium, wherein the at least one processor reads the at least one instruction set when running and executes the write memory allocation method as described in claim 8 according to the instructions of the at least one instruction set.