Storage architecture optimization methods, equipment, media and products based on reinforcement learning
By applying the storage architecture optimization method based on reinforcement learning in multimode databases, the table connection operation is dynamically determined, and the problem of storage architecture optimization in multimode databases is solved, and more efficient query performance and response capabilities are achieved.
Patent Information
- Application Number
- CN202411975070.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Under the framework of multi-mode database, how to optimize the storage architecture used to store multiple modal data remains a technical problem that needs to be solved urgently.
The storage architecture optimization method based on reinforcement learning is adopted. By determining the environment and agent in reinforcement learning, using the strategy model to make action decisions, and dynamically decide whether to connect different tables to optimize the storage architecture.
Through reinforcement learning, the agent can optimize the storage architecture according to specific query needs, shorten query execution time, improve responsiveness, avoid management complexity, and automatically balance the selection of "wide table" and "multi-table JOIN" solutions.
Smart Images

Figure CN119377448B_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of the present specification relate to the field of storage technology, and in particular, to a storage architecture optimization method based on reinforcement learning, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the advent of the big data era, data types and sources are becoming increasingly diverse, showing a situation where structured data, semi-structured data, and unstructured data coexist. In order to process these data of different modes, traditional methods often require the configuration of multiple storage engines, each engine is responsible for a specific mode of data. This method increases the complexity of the system and leads to an increase in operation and maintenance, data migration, and development and maintenance costs. In order to solve this problem, multi-modal databases came into being. Multi-modal databases are designed to support the storage and management of data of multiple modes through a unified database management platform. However, how to optimize the storage architecture for storing data of multiple modes under the framework of multi-modal databases is still a technical problem that needs to be solved urgently. Summary of the invention
[0003] In view of this, one or more embodiments of the present specification provide a storage architecture optimization method based on reinforcement learning, an electronic device, a computer-readable storage medium, and a computer program product.
[0004] To achieve the above objectives, one or more embodiments of this specification provide the following technical solutions:
[0005] According to a first aspect of one or more embodiments of this specification, a storage architecture optimization method based on reinforcement learning is proposed, including:
[0006] Determine the environment in the reinforcement learning according to the storage architecture to be optimized, and determine the agent in the reinforcement learning according to the policy model used to optimize the storage architecture; wherein the storage architecture includes N tables, different tables in the N tables are used to store different modal data, N>1;
[0007] The actions in the reinforcement learning are determined by running the policy model, the actions including connecting different tables and not connecting different tables, the state in the reinforcement learning is determined based on the unique identifiers of all tables included in the storage architecture, the reward in the reinforcement learning is determined based on the query duration of the query service provided by the storage architecture to the user, and the policy model is trained for reinforcement learning with maximizing the reward in the reinforcement learning as the optimization goal, so as to optimize the storage architecture using the policy model.
[0008] According to a second aspect of an embodiment of this specification, a storage architecture optimization device based on reinforcement learning is provided, including:
[0009] An environment and agent determination module, used to determine the environment in the reinforcement learning according to the storage architecture to be optimized, and to determine the agent in the reinforcement learning according to the policy model used to optimize the storage architecture; wherein the storage architecture includes N tables, different tables in the N tables are used to store different modal data, N>1;
[0010] A reinforcement learning training module is used to determine the actions in the reinforcement learning by running the strategy model, wherein the actions include connecting different tables and not connecting them, determining the state in the reinforcement learning based on the unique identifiers of all tables included in the storage architecture, determining the reward in the reinforcement learning based on the query duration of the query service provided by the storage architecture to the user, and performing reinforcement learning training on the strategy model with maximizing the reward in the reinforcement learning as the optimization goal, so as to optimize the storage architecture using the strategy model.
[0011] According to a third aspect of the embodiments of this specification, there is provided an electronic device, including:
[0012] processor;
[0013] a memory for storing processor-executable instructions;
[0014] When the processor executes the executable instructions, it is used to implement the method described in the first aspect.
[0015] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0016] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program, which implements the steps of the method described in the first aspect when executed by a processor.
[0017] The technical solutions provided by the embodiments of this specification may have the following beneficial effects:
[0018] In the embodiments of this specification, through reinforcement learning, the intelligent agent can continuously optimize the storage architecture according to specific query requirements (determined by the query duration of the query service provided to users based on the storage architecture based on the rewards in reinforcement learning), and dynamically decide whether to connect different tables, which is beneficial to shorten the query execution time of the optimized query architecture and improve responsiveness. In addition, the storage architecture of this solution supports storing multimodal data through different tables, avoiding the management complexity introduced by directly merging multimodal data into a single table. Through dynamic decision-making in the reinforcement learning process, it automatically balances the selection of "wide table" and "multi-table JOIN" solutions to meet different query requirements.
[0019] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 is a flowchart of a storage architecture optimization method based on reinforcement learning provided by an exemplary embodiment;
[0021] Figure 2 is a schematic diagram of a reinforcement learning architecture provided by an exemplary embodiment;
[0022] Figure 3 is a schematic diagram of a storage architecture provided by an exemplary embodiment;
[0023] Figure 4 is a flowchart of a reinforcement learning training provided by an exemplary embodiment;
[0024] Figure 5 It is a schematic structural diagram of an electronic device provided by an exemplary embodiment. DETAILED DESCRIPTION
[0025] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with one or more embodiments of this specification. Instead, they are merely examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0026] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0027] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0028] The following is an explanation of the relevant terms in the embodiments of this specification:
[0029] 1. Reinforcement Learning (RL) is a machine learning method, the core of which is to enable the agent to optimize its behavior strategy in a specific environment through continuous trial and learning, in order to maximize the cumulative reward.
[0030] An agent is an entity that makes decisions in a reinforcement learning scenario. It can be a robot, a software algorithm, or any system that can perceive the environment and take actions.
[0031] The environment is the external world or situation in which the agent is located, which gives feedback (usually rewards or penalties) based on the agent's actions.
[0032] State (S) is a feature or information that describes the environment at any moment, and the agent decides actions based on the current state.
[0033] Action (A) is an operation or decision that an agent can perform in a given state. Actions affect the environment and lead to state transitions.
[0034] Reward (R) is the feedback given by the environment to the agent after it takes an action. It is a direct indicator of the quality of the action. Rewards can be given immediately or delayed to a certain point in the future.
[0035] The policy model (Policy, π) is the rule or strategy for the agent to choose actions, which defines what action the agent should choose to achieve the goal in a given state. The policy model can be deterministic or stochastic (i.e., policy π(s) = a, meaning take action a in state s).
[0036] 2. Multi-model data: refers to data that uses multiple storage methods, such as the following storage methods:
[0037] (1) Relational storage: Relational databases store data using table structures, and relationships between data are established using foreign keys. Each table consists of rows and columns, and tables are connected through JOIN operations.
[0038] (2) Key-value pair storage: Data is stored in the form of "key" and "value" pairs. Each key is unique, and the corresponding value can be quickly accessed through the key.
[0039] (3) Document storage: Data is stored in documents as the basic unit. Documents are usually in JSON (JavaScript Object Notation), BSON (Binary JSON) or XML (Extensible Markup Language) format, with flexible data structures. Each document is a self-contained data unit that can store complex structures and data types.
[0040] (4) Graph database storage: Graph database is a data model based on graph theory. Data is stored and operated with nodes (entities) and edges (relationships) as basic units. Graph database is particularly suitable for representing and operating data with complex relationships.
[0041] (5) Column family storage: Data is stored in the table in the form of columns. The data in each column family can be stored independently, allowing efficient read and write operations and is suitable for storing sparse data.
[0042] (6) Geographic Information System (GIS) storage: used to process geographic spatial data. Common formats include point, line, surface and other spatial data models. The data usually includes information such as geographic location, boundaries, and area.
[0043] (7) Media file storage: used to store multimedia data, including images, audio, video and other files.
[0044] With the advent of the big data era, data types and sources are becoming increasingly diverse, showing a situation where structured data, semi-structured data and unstructured data coexist. In order to process these different modal data, multimodal databases have emerged. Multimodal databases are designed to support the storage and management of data in multiple modalities through a unified database management platform. However, how to optimize the storage architecture for storing data in multiple modalities under the framework of multimodal databases is still a technical problem that needs to be solved urgently.
[0045] For example, if you store data of multiple modalities directly in one table, because the data of each modality may correspond to a different data format or type, in order to accommodate these different data types, you need to design a "wide table" containing a large number of columns. Each column corresponds to a data field or data type, which may generate a large number of redundant fields. Some fields may be redundant for data of a specific modality, resulting in a waste of storage space. Moreover, as the number of columns increases, the query performance of the wide table may drop sharply. Even if you only retrieve certain data items, more columns need to be scanned during the query due to the large table structure, resulting in performance bottlenecks. Once the number of columns in a table increases to a certain scale, the management and maintenance of the database becomes extremely difficult, especially when the amount of data is extremely large. Wide tables will face performance bottlenecks and are difficult to adapt to horizontal expansion.
[0046] Based on the problems in the related art, the embodiments of this specification provide a storage architecture optimization method based on reinforcement learning, which uses the decision-making advantages of the reinforcement learning algorithm to dynamically optimize the storage architecture selection of multi-modal data. The storage architecture optimization method based on reinforcement learning provided in the embodiments of this specification can be executed by any electronic device with computing power, including but not limited to physical servers, server clusters, cloud servers, virtual servers, smart phones / mobile phones, tablet computers, personal digital assistants (PDAs), laptop computers, and desktop computers.
[0047] See also Figure 1 and Figure 2 , Figure 1 A flowchart of a storage architecture optimization method based on reinforcement learning is shown. Figure 2 A schematic diagram of a training framework for reinforcement learning is shown. The method includes:
[0048] In S101, the environment in reinforcement learning is determined according to the storage architecture to be optimized, and the agent in reinforcement learning is determined according to the policy model used to optimize the storage architecture; wherein the storage architecture includes N tables, different tables among the N tables are used to store different modal data, and N>1.
[0049] Exemplarily, the storage architecture to be optimized can be deployed in a relational database. There are several reasons for storing multimodal data based on a relational database. First, relational databases dominate the current market, and the cost of replacing them with other types of databases is high. Second, after years of theoretical and technological development, relational databases have advantages in security, query optimization, and transaction processing, and a large amount of existing data has been stored in them.
[0050] Of course, other types of databases may also be selected, and this embodiment does not impose any limitation on this.
[0051] Exemplarily, the environment is an object that the agent interacts with in reinforcement learning. In this method, the environment is determined by the storage architecture to be optimized. The storage architecture includes N tables, where different tables store data of different modalities, such as JSON documents, XML documents, graph data, multimedia data (such as images, audio or video), etc.
[0052] The table structure is as follows: each table contains several records, each record contains at least two columns, one of which stores the unique identifier of each record (such as id) to indicate the uniqueness of the record; the other columns store data content related to the modality (such as JSON documents, XML documents, graph relationships, etc.).
[0053] The agent is trained and optimized based on the policy model in reinforcement learning. The policy model is responsible for the agent's decision-making and is the core for optimizing the storage architecture. The agent runs the policy model to decide whether to perform a JOIN operation on the table or keep the current architecture unchanged.
[0054] In S102, actions in reinforcement learning are determined by running the policy model, the actions include connecting different tables and not connecting them, the state in reinforcement learning is determined based on the unique identifiers of all tables included in the storage architecture, the reward in reinforcement learning is determined based on the query duration of the query service provided by the storage architecture to the user, and reinforcement learning training is performed on the policy model with maximizing the reward in reinforcement learning as the optimization goal, so as to optimize the storage architecture using the policy model.
[0055] Regarding the determination of actions: The action space defines all possible actions of the agent, including: (1) JOIN operation: select two tables for JOIN operation to generate a new table; (2) No JOIN operation: keep the current table state unchanged. By running the policy model, the agent selects an action from the action space. The action space A is defined as follows: The total number of actions is |A| = 1 + C(N, 2) = 1 + N * (N - 1) / 2; where action 0 means no JOIN operation, actions 1 to C(N, 2) mean JOIN operations on different tables, and C(N, 2) means the total number of combinations of any two tables selected from N tables.
[0056] For example, when N=4, the total number of action spaces |A| =1 + 4 * (4 - 1) / 2=7. The actions include the following 7 types: (Table 1, Table 2), (Table 1, Table 3), (Table 1, Table 4), (Table 2, Table 3), (Table 2, Table 4), (Table 3, Table 4), and no connection operation. (Table 1, Table 2) means that Table 1 and Table 2 are connected to obtain a new table. The meanings of other brackets are similar and will not be repeated here.
[0057] Regarding the determination of the state: The state space describes the current state of the storage architecture. The initial state of the storage architecture can be determined by the unique identifiers (such as table names) of the N tables included in the storage architecture. Assume that Figure 3 As shown, assuming that there are three tables in the storage architecture, the table names are json_table (for storing json data), xml_table (for storing xml data) and csv_table (for storing csv data), then the state string [json_table, xml_table, csv_table] can be determined, and the state string can be further mapped into a feature vector for input into the policy model for learning.
[0058] Regarding the determination of rewards: Rewards reflect the quality of actions, with query efficiency and storage overhead as the core optimization goals. The rewards in reinforcement learning are determined by the query duration of providing query services to users based on the storage architecture, which comprehensively considers the actual query needs of users and is conducive to optimizing query duration.
[0059] Regarding the training of the policy model: the agent uses reinforcement learning to train the policy model with the optimization goal of maximizing rewards. The adopted policy model continuously optimizes the decision-making ability by learning the relationship between state-action-reward, and finally outputs a model that can efficiently optimize the storage architecture, thereby optimizing the storage architecture.
[0060] The storage architecture optimization method based on reinforcement learning provided in this embodiment can optimize the storage architecture according to specific query requirements through reinforcement learning, dynamically decide whether to perform JOIN operations on tables, shorten query execution time, and improve response capabilities. In addition, this solution supports storing multimodal data through different tables, avoiding the management complexity introduced by directly merging multimodal data into a single table. Through dynamic decision-making during reinforcement learning, the selection of "wide table" and "multi-table JOIN" solutions is automatically balanced to meet different query requirements. Furthermore, the reward in reinforcement learning is determined by the query duration, so that the intelligent agent can continuously adjust the storage architecture as the user's query scenario changes (the change of the query scenario will lead to a change in the reward), so that it is always in a better state, which is particularly suitable for complex and frequently changing database environments.
[0061] In some embodiments, the reinforcement learning training process is exemplarily described below:
[0062] See also Figure 4The electronic device can determine the state of the environment in the first round of iteration based on the unique identifiers of the N tables, and determine the non-connection operation as the action output by the agent in the first round of iteration (S400), and loop the following training process to train the policy model with maximizing the reward in reinforcement learning as the optimization goal:
[0063] In S401, the storage architecture is updated based on the actions output by the agent.
[0064] Exemplarily, if the action output by the intelligent agent is to perform a connection operation on different tables, the electronic device performs a connection operation on the two tables indicated by the action to obtain a new table, and determines the updated storage architecture in this round of iteration based on the new table and all tables included in the storage architecture; for example, assuming that there are 3 tables in the storage architecture, and the table names are json_table, xml_table and csv_table, and the action output by the intelligent agent indicates a connection operation on json_table and xml_table, then the electronic device can, based on the action, perform a connection operation on json_table and xml_table to obtain a new table, then the updated storage architecture can include the original 3 tables and the new table, and the updated storage architecture retains both the original tables and the newly generated tables, which not only avoids the problem of original data loss due to the connection operation, but also provides more optimization selection space for subsequent iterations.
[0065] After obtaining the new table, a unique identifier of the new table can be generated based on the unique identifiers corresponding to the two tables indicated by the action and the preset connection identifier (assuming it is concat). For example, the unique identifier of the new table can be identified as: json_table_concat_xml_table, so that the source and content relationship of the new table are clear at a glance. This design facilitates the management, tracking and optimization of all tables in the storage architecture.
[0066] If the action output by the agent is not to perform a connection operation, the storage architecture obtained in the previous iteration process is determined as the updated storage architecture in this round of iteration. This can avoid unnecessary storage architecture update operations, reduce computing and storage overhead, and improve the efficiency of the optimization process.
[0067] In S402, based on the unique identifiers of all tables included in the updated storage architecture, the state of the environment in this round of iteration is determined.
[0068] Exemplarily, when determining the state of the environment during this round of iterations, the electronic device can combine the unique identifiers of all tables included in the updated storage architecture to obtain a combined result, and then normalize the combined result, and determine the result of the normalization operation as the state of the environment during this round of iterations. By combining and normalizing the unique identifiers of all tables in the storage architecture, complex storage architecture states can be converted into fixed numerical or vector representations, which helps the intelligent agent learn and optimize strategies more efficiently. The normalization operation can map storage architectures of different sizes (such as storage architectures containing different numbers of tables) to a unified numerical range, so that the reinforcement learning model can adapt to storage architectures of different sizes and improve its generalization ability.
[0069] In order to save computational overhead, when the action output by the intelligent agent is to connect different tables, the above-mentioned combination and normalization process can be performed to obtain the state of the environment in this round of iteration; when the action output by the intelligent agent is not to perform a connection operation, the state of the environment in the previous round of iteration is directly determined as the state of the environment in this round of iteration.
[0070] In S403, in the process of providing query services to users based on the updated storage architecture, a reward for the updated storage architecture is determined based on the query duration of the query service.
[0071] In one possible implementation, if the action output by the agent performs a connection operation on different tables, the opposite of the query duration corresponding to the query request for the new table obtained by the connection operation is determined as the reward for the updated storage architecture; in this case, the reward is determined by the query duration of the new table, and the agent can directly perceive the impact of its actions on the query efficiency.
[0072] If the action output by the agent is not to perform a connection operation, the opposite of the query duration corresponding to the query request for a random table is determined as the reward for the updated storage architecture; in this case, the method of randomly selecting a table for query reward enables the agent to still perceive the query efficiency of the current storage architecture, and randomly selecting a table for reward calculation can increase the possibility of exploring different states of the storage architecture.
[0073] In this embodiment, the design of the reward function directly uses the opposite of the query time as the reward, which means that the higher the reward value, the shorter the query time and the more efficient the storage architecture. In this way, the agent is motivated to generate a storage architecture that can respond to query requests faster, thereby optimizing the user experience. The design of the reward function combines randomness and targeting, which can not only enhance the adaptability to different query scenarios, but also avoid overfitting problems in specific query modes, thereby improving robustness and practical application effects.
[0074] In another possible implementation, in addition to considering query performance, the amount of stored data in the storage architecture can be further considered, so as to balance query performance and storage resource utilization while optimizing the storage architecture. For example, the electronic device can determine the reward for the updated storage architecture based on the query duration of the query service and the amount of stored data in the updated storage architecture. In this embodiment, the addition of the amount of stored data guides the agent to minimize unnecessary data redundancy and avoid waste of storage resources caused by generating new tables. After adding the consideration of the amount of stored data, the agent will tend to choose operations that have less impact on the amount of stored data, making the generated storage architecture more lightweight and efficient.
[0075] If the action output by the intelligent agent is to connect different tables, a reward for the updated storage architecture can be obtained by performing a weighted sum based on the inverse of the query duration corresponding to the query request for the new table obtained by the connection operation and the amount of storage data in the updated storage architecture.
[0076] For example, there are three tables named json_table, xml_table, and csv_table. The action output by the agent indicates to connect json_table and xml_table to generate a new table json_table_xml_table_joined. Rewards can be obtained by weighted summing the query time (e.g., 50ms) of the new table json_table_xml_table_joined and the amount of data stored in the updated storage architecture (including 4 tables, such as 300MB of storage data). This reward mechanism ensures the optimization of query performance while constraining the usage of storage resources to avoid generating too large new tables that affect the utilization of storage resources.
[0077] If the action output by the agent is not to perform a connection operation, a weighted sum can be taken based on the inverse of the query duration corresponding to the query request for a random table and the amount of storage data in the updated storage architecture to obtain a reward for the updated storage architecture, thereby enabling the agent to dynamically balance query performance and storage resource utilization during reinforcement learning.
[0078] For example, the weight of query performance and the weight of the amount of stored data in the storage architecture can be set according to actual needs. For example, when query performance is the main optimization goal, the weight corresponding to the opposite of the query time can be set to be greater than the weight corresponding to the amount of stored data. The agent will give priority to actions that can significantly reduce the query time. When storage resources are limited, even if the query performance does not change much, the agent can avoid generating too many new tables that occupy too many storage resources.
[0079] In this implementation, the design of the reward function guides the agent to dynamically adjust the optimization direction according to the actual needs of the storage architecture and query services, thereby achieving comprehensive optimization of query performance and storage resources.
[0080] In S404, it is determined whether the iteration end condition is met.
[0081] Exemplarily, in order to ensure that the reinforcement learning process can be terminated reasonably, the iteration termination condition includes any of the following:
[0082] (1) Reaching the preset number of iterations: The iterative process will automatically end after reaching the preset maximum number of iterations M. The preset number of iterations is to avoid an infinite loop in the training process. Usually, a reasonable upper limit is set according to the complexity of the problem.
[0083] (2) The actions in the action space are empty; the action space is used to define all the actions that the agent can take, and the total number of actions in the action space is determined by the number of combinations of any two tables in the N tables. When all actions in the action space have been executed and there are no new actions available, the iteration process ends.
[0084] (3) The reward variation between multiple adjacent iterations is smaller than a preset range. The reward is calculated by the agent based on the current storage architecture state. If the reward variation between multiple adjacent iterations is continuously smaller than a preset range, it indicates that the optimization of the storage architecture is converging and no further training is required.
[0085] The introduction of the iteration end condition ensures that the training process will not loop infinitely, saving computing resources.
[0086] In S405 , when the iteration end condition is not met, the parameters of the strategy model are updated based on the reward, and the state is input into the updated strategy model to determine the action of the next iteration process through the updated strategy model.
[0087] Exemplarily, the policy model includes a deep Q-network (DQN), which is an important algorithm in reinforcement learning. It combines deep learning and Q learning to solve decision-making problems in high-dimensional state space. DQN approximates the Q-value function by using a deep neural network. When the iteration end condition is not met, the parameters of the deep Q network can be updated based on the reward and the back propagation algorithm of DQN can be run to find a better policy model. The deep Q network (DQN) receives the action in the state and action space as input, and outputs the Q value of each possible action (i.e., the value evaluation of each action). According to the Q value, the agent selects the better action in the current state.
[0088] In order to balance exploration and exploitation, the ε-greedy strategy can be further introduced. The deep Q network guides the agent to choose a better action by estimating the value of each state-action pair. It tells the agent which action to take in each state to maximize the cumulative reward. The ε-greedy strategy controls the agent's strategy for selecting actions in a specific state. The ε-greedy strategy provides the agent with a balance between choosing exploration (random actions) or exploitation (selecting the action with the highest Q value). In the early stages (when the agent knows less about the environment), the agent explores with a higher ε value to gain more experience, that is, to choose a random action without relying on the Q value output by the deep Q network; in the later stages, ε gradually decreases, and the agent relies more on the deep Q network for decision-making, gradually improving the stability and convergence speed of the strategy.
[0089] Of course, other reinforcement learning algorithms may also be used, such as PPO (Proximal Policy Optimization) algorithm, REINFORCE (Monte Carlo Policy Gradient) algorithm, Q-Learning algorithm, etc. This embodiment does not impose any restrictions on this and can be specifically set according to the actual application scenario.
[0090] In this embodiment, through multiple rounds of reinforcement learning iterations, the storage architecture can be gradually adjusted to dynamically form a better storage architecture to meet the performance requirements of user queries. This progressive optimization avoids the overfitting problem that may be caused by one-time optimization. This solution is designed for storing multimodal data (such as JSON, XML, graph data, etc.). The agent can dynamically decide whether to connect the table and optimize the storage architecture according to the actual query requirements, overcoming the complexity of multimodal data integration and query. During the optimization process, the agent dynamically balances the choice of "wide table" and "multi-table JOIN" to avoid the waste of storage space caused by blindly generating wide tables. At the same time, reinforcement learning can guide the agent to perform table connection operations only when the query efficiency is significantly improved, thereby reducing unnecessary new table generation and further reducing storage overhead. Each round of iteration re-determines the action through the updated policy model, making the optimization of the storage architecture a dynamic process. This method not only optimizes the current storage state, but also provides good adaptability to future changes in query requirements.
[0091] The various technical features in the above embodiments can be arbitrarily combined as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, any combination of the various technical features in the above embodiments also falls within the scope of this specification.
[0092] In some embodiments, the embodiments of this specification also provide an electronic device, including: a processor; a memory for storing processor executable instructions; wherein the processor implements any of the above methods by running the executable instructions.
[0093] Figure 5 is a schematic structural diagram of a device provided by an exemplary embodiment. Figure 5 At the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510, and may also include hardware required for other functions. One or more embodiments of this specification may be implemented based on software, such as the processor 502 reading the corresponding computer program from the non-volatile memory 510 into the memory 508 and then running it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementations, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0094] Exemplarily, the storage architecture optimization device based on reinforcement learning can be applied to Figure 5 The device shown in the figure is used to implement the technical solution of this specification. The storage architecture optimization device based on reinforcement learning may include:
[0095] An environment and agent determination module, used to determine the environment in the reinforcement learning according to the storage architecture to be optimized, and to determine the agent in the reinforcement learning according to the policy model used to optimize the storage architecture; wherein the storage architecture includes N tables, different tables in the N tables are used to store different modal data, N>1;
[0096] A reinforcement learning training module is used to determine the actions in the reinforcement learning by running the strategy model, wherein the actions include connecting different tables and not connecting them, determining the state in the reinforcement learning based on the unique identifiers of all tables included in the storage architecture, determining the reward in the reinforcement learning based on the query duration of the query service provided by the storage architecture to the user, and performing reinforcement learning training on the strategy model with maximizing the reward in the reinforcement learning as the optimization goal, so as to optimize the storage architecture using the strategy model.
[0097] In one implementation, the reinforcement learning training module is specifically used to determine the state of the environment in the first round of iterations based on the unique identifiers of the N tables, and to determine not performing a connection operation as the action output by the agent in the first round of iterations, and to loop the following iterative process: based on the action output by the agent, updating the storage architecture; based on the unique identifiers of all tables included in the updated storage architecture, determining the state of the environment in this round of iterations; in the process of providing query services to users based on the updated storage architecture, determining the reward for the updated storage architecture based on the query duration of the query service; if the iteration end condition is not met, updating the parameters of the strategy model based on the reward, and inputting the state into the updated strategy model, so as to determine the action of the next iteration process through the updated strategy model.
[0098] In one implementation, the iteration end condition includes any one of the following: reaching a preset number of iterations, the action in the action space is empty, and the reward change amplitude between adjacent multiple iteration processes is less than a preset amplitude; wherein, the action space is used to define all actions that the agent can take, and the total number of actions in the action space is determined based on the number of combinations of any two tables in the N tables.
[0099] In one implementation, the reinforcement learning training module is specifically used to, if the action output by the agent is to connect different tables, connect the two tables indicated by the action to obtain a new table, and determine the updated storage architecture in this round of iteration based on the new table and all tables included in the storage architecture; if the action output by the agent is not to perform a connection operation, then determine the storage architecture obtained in the previous iteration as the updated storage architecture in this round of iteration.
[0100] In one implementation, if the action output by the agent is to connect different tables, the unique identifier of the new table is generated based on the unique identifiers corresponding to the two tables indicated by the action and a preset connection identifier.
[0101] In one implementation, the reinforcement learning training module is specifically used to combine the unique identifiers of all tables included in the updated storage architecture to obtain a combination result; normalize the combination result, and determine the result of the normalization operation as the state of the environment in this round of iteration.
[0102] In one implementation, the reinforcement learning training module is specifically used to, if the action output by the agent is to connect different tables, determine the opposite of the query time corresponding to the query request for the new table obtained by the connection operation as the reward for the updated storage architecture; if the action output by the agent is not to connect, determine the opposite of the query time corresponding to the query request for a random table as the reward for the updated storage architecture.
[0103] In one implementation, the reinforcement learning training module is specifically used to determine the reward for the updated storage architecture based on the query duration of the query service and the amount of storage data of the updated storage architecture.
[0104] In one implementation, the reinforcement learning training module is specifically used to obtain a reward for the updated storage architecture by performing a weighted summation based on the inverse of the query time corresponding to the query request for the new table obtained by the connection operation and the amount of storage data of the updated storage architecture if the action output by the agent is to connect different tables; and to obtain a reward for the updated storage architecture by performing a weighted summation based on the inverse of the query time corresponding to the query request for a random table and the amount of storage data of the updated storage architecture if the action output by the agent is not to perform a connection operation; wherein the weight corresponding to the inverse of the query time is greater than the weight corresponding to the amount of storage data.
[0105] In one implementation, the storage architecture is deployed in a relational database.
[0106] In one implementation, each of the N tables includes at least two columns, one of which is used to store a unique identifier of each record in the table, and the remaining columns are used to store data related to a specified modality included in each record in the table.
[0107] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0108] Based on the same concept as the above method, this specification also provides a computer-readable storage medium on which computer instructions are stored. When the instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.
[0109] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0110] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instruction, which implements the steps of the method described in any of the above embodiments when executed by a processor.
[0111] The above description is merely a preferred embodiment of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. A storage architecture optimization method based on reinforcement learning, comprising: Determine the environment in the reinforcement learning according to the storage architecture to be optimized, and determine the agent in the reinforcement learning according to the policy model used to optimize the storage architecture; wherein the storage architecture includes N tables, different tables in the N tables are used to store different modal data, N>1; Determine the action in the reinforcement learning by running the policy model, the action includes connecting different tables and not connecting them, determine the state in the reinforcement learning based on the unique identifiers of all tables included in the storage architecture, determine the reward in the reinforcement learning based on the query duration of the query service provided by the storage architecture to the user, and perform reinforcement learning training on the policy model with maximizing the reward in the reinforcement learning as an optimization goal, so as to optimize the storage architecture by using the policy model; Among them, in each round of iteration of the reinforcement learning training, if the action output by the agent is to connect different tables, then the two tables indicated by the action are connected to obtain a new table, and based on the new table and all tables included in the storage architecture, the updated storage architecture in this round of iteration is determined; if the action output by the agent is not to perform a connection operation, then the storage architecture obtained in the previous iteration is determined as the updated storage architecture in this round of iteration.
2. The method according to claim 1, wherein the actions in the reinforcement learning are determined by running the policy model, the actions include connecting different tables and not connecting different tables, determining the state in the reinforcement learning based on the unique identifiers of all tables included in the storage architecture, determining the reward in the reinforcement learning based on the query duration of the query service provided by the storage architecture to the user, and performing reinforcement learning training on the policy model with maximizing the reward in the reinforcement learning as the optimization goal, comprising: The state of the environment in the first round of iteration is determined based on the unique identifiers of the N tables, and the non-connection operation is determined as the action output by the agent in the first round of iteration, and the following iteration process is repeated: Based on the actions output by the agent, updating the storage architecture; Determine the state of the environment in this iteration process based on the unique identifiers of all tables included in the updated storage architecture; In the process of providing a query service to a user based on the updated storage architecture, determining a reward for the updated storage architecture based on the query duration of the query service; When the iteration end condition is not met, the parameters of the policy model are updated based on the reward, and the state is input into the updated policy model to determine the action of the next iteration process through the updated policy model.
3. According to the method of claim 2, the iteration end condition includes any one of the following: reaching a preset number of iterations, the action in the action space is empty, and the reward change range between multiple adjacent iteration processes is less than a preset range; in, The action space is used to define all actions that the agent can take, and the total number of actions in the action space is determined based on the number of combinations of any two tables in the N tables.
4. According to the method of claim 1, if the action output by the agent is to connect different tables, the unique identifier of the new table is generated based on the unique identifiers corresponding to the two tables indicated by the action and a preset connection identifier.
5. The method according to claim 2, wherein determining the state of the environment in this round of iteration based on the unique identifiers of all tables included in the updated storage architecture comprises: Combining the unique identifiers of all tables included in the updated storage architecture to obtain a combination result; A normalization operation is performed on the combination result, and the result of the normalization operation is determined as the state of the environment in this round of iteration.
6. The method according to claim 2, wherein determining the reward for the updated storage architecture based on the query duration of the query service comprises: If the action output by the agent is to perform a connection operation on different tables, the opposite of the query duration corresponding to the query request for the new table obtained by the connection operation is determined as the reward for the updated storage architecture; If the action output by the agent is not to perform a connection operation, the opposite of the query duration corresponding to the query request for a random table is determined as the reward for the updated storage architecture.
7. The method according to claim 2, wherein determining the reward for the updated storage architecture based on the query duration of the query service comprises: Based on the query duration of the query service and the amount of storage data of the updated storage architecture, a reward for the updated storage architecture is determined.
8. The method according to claim 7, wherein determining the reward for the updated storage architecture based on the query duration of the query service and the amount of stored data of the updated storage architecture comprises: If the action output by the agent is to connect different tables, a weighted sum is performed based on the inverse of the query duration corresponding to the query request for the new table obtained by the connection operation and the amount of storage data of the updated storage architecture to obtain a reward for the updated storage architecture; If the action output by the agent is not to perform a connection operation, a weighted sum is performed based on the inverse of the query duration corresponding to the query request for a random table and the amount of storage data of the updated storage architecture to obtain a reward for the updated storage architecture; Among them, the weight corresponding to the inverse of the query duration is greater than the weight corresponding to the amount of stored data.
9. According to the method according to any one of claims 1 to 8, the storage architecture is deployed in a relational database.
10. According to the method according to any one of claims 1 to 8, each of the N tables includes at least two columns, one column is used to store a unique identifier of each record in the table, and the remaining columns are used to store data related to the specified modality included in each record in the table.
11. A storage architecture optimization device based on reinforcement learning, comprising: An environment and agent determination module, used to determine the environment in the reinforcement learning according to the storage architecture to be optimized, and to determine the agent in the reinforcement learning according to the policy model used to optimize the storage architecture; wherein the storage architecture includes N tables, different tables in the N tables are used to store different modal data, N>1; A reinforcement learning training module is used to determine the actions in the reinforcement learning by running the strategy model, wherein the actions include connecting different tables and not connecting them, determining the state in the reinforcement learning based on the unique identifiers of all tables included in the storage architecture, determining the reward in the reinforcement learning based on the query duration of the query service provided by the storage architecture to the user, and performing reinforcement learning training on the strategy model with maximizing the reward in the reinforcement learning as the optimization goal, so as to optimize the storage architecture using the strategy model; wherein, in each round of iteration of the reinforcement learning training, if the action output by the agent is to connect different tables, then a connection operation is performed on the two tables indicated by the action to obtain a new table, and based on the new table and all tables included in the storage architecture, the updated storage architecture in this round of iteration is determined; if the action output by the agent is not to connect, then the storage architecture obtained in the previous iteration is determined as the updated storage architecture in this round of iteration.
12. An electronic device comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 10 by executing the executable instructions.
13. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
14. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Data classified storage method and device based on reinforcement learning
CN117453123A
Data management method and device, program product, equipment and storage medium
CN118733635A