Training Method and Device for Database Request Distribution Model, and Electronic Device

By building a sample pool of the database request distribution model and using self-supervised algorithm training, the problem of unbalanced database request allocation in high-concurrency business scenarios is solved, and efficient operation and second-level adjustment of the database system are achieved.

CN117827438BActive Publication Date: 2025-07-25中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311760637.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-07-25
Estimated Expiration
2043-12-20

AI Technical Summary

Technical Problem

In the high concurrency business scenario, the traffic allocation of database requests in the prior art cannot be guaranteed to be balanced, resulting in unbalanced data load and affecting the processing efficiency of database servers.

Method used

Build a sample pool of the database request distribution model, train the database request distribution model through a self-supervised algorithm, dynamically update the state feature parameters and historical distribution data of the database server, and optimize the request distribution plan.

Benefits of technology

Real-time allocation of database requests is realized, the overall operation efficiency of the database system is improved, the database server can be adjusted in seconds, and historical experience is reused for dynamic decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117827438B_ABST
    Figure CN117827438B_ABST
Patent Text Reader

Abstract

The present application discloses a training method and device for a database request distribution model, and an electronic device. The method includes: constructing a sample pool for the database request distribution model; obtaining training sample data for the database request distribution model from the sample pool, where the training sample data includes state characteristic parameters of each database server at the current moment, a request distribution scheme and the corresponding reward value, and state characteristic parameters at the next moment; training the database request distribution model using a self-supervised algorithm according to the training sample data of the database request distribution model; obtaining historical distribution data and dynamically updating the database request distribution model based on this. The present application uses a self-supervised algorithm to train the database request distribution model, which can perform parameter updates by itself, can adjust the database server in seconds, and the model can dynamically learn from historical distribution data, reuse historical experience, and constantly monitor the current state of the database server to give different distribution schemes, improving the overall operation efficiency of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of database technologies, and in particular, to a method and apparatus for training a database request distribution model, and an electronic device. Background Art

[0002] A database request refers to a request initiated by a user to establish a connection (long connection) with a PostgreSQL server through a database connection pool. A database thread obtains an SQL statement for analysis and parsing, converts the SQL statement into an abstract syntax tree, and then executes a transaction.

[0003] In a high-concurrency business scenario, most of the current request traffic distribution is based on an average distribution form. For example, in the scenario of commemorative coin issuance, database requests for high-concurrency reservations are evenly distributed to multiple database servers.

[0004] The main problem with this method is that the time for the database to process requests cannot be guaranteed to be average. Over time, it will still lead to the problem of unbalanced data load, and the differences in the processing efficiencies of different database servers will also affect the effect of this method. Summary of the Invention

[0005] Embodiments of this application provide a method and apparatus for training a database request distribution model, and an electronic device, to achieve real-time allocation of database requests and improve the overall operation efficiency of a database system.

[0006] Embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, embodiments of this application provide a method for training a database request distribution model. The method for training the database request distribution model includes:

[0008] Construct a sample pool for the database request distribution model;

[0009] Obtain training sample data for the database request distribution model from the sample pool. The training sample data includes state characteristic parameters of each database server at the current moment, a request distribution scheme at the current moment, a reward value corresponding to the request distribution scheme at the current moment, and state characteristic parameters at the next moment;

[0010] Train the database request distribution model using a self-supervised algorithm according to the training sample data for the database request distribution model;

[0011] Obtain historical distribution data, and dynamically update the database request distribution model using the historical distribution data.

[0012] Optionally, constructing the sample pool for the database request distribution model includes:

[0013] Randomly sample the state parameters of each database server at the current moment;

[0014] Convert the randomly sampled state parameters at the current moment into state feature parameters at the current moment;

[0015] Input the state feature parameters at the current moment into the current policy network of the database request distribution model to obtain the request distribution plan at the current moment;

[0016] Execute the request distribution plan at the current moment to obtain the reward value corresponding to the request distribution plan at the current moment and the state feature parameters of each database server at the next moment;

[0017] Construct training sample data from the state feature parameters of each database server at the current moment, the request distribution plan at the current moment and the corresponding reward value, and the state feature parameters at the next moment, and store them in the sample pool.

[0018] Optionally, the reward value corresponding to the request distribution plan at the current moment is obtained by the following method:

[0019] Determine the reward weight of the state feature parameters according to the importance degree of the state feature parameters in the database request distribution scenario;

[0020] Calculate the reward value corresponding to the request distribution plan at the current moment according to the state feature parameters at the current moment and the corresponding reward weight;

[0021] Wherein, the state feature parameters include at least one dimension of the number of transactions processed per second, the transaction execution time per second, the core processor utilization rate, the memory occupancy rate, and the disk storage space.

[0022] Optionally, the calculating the reward value corresponding to the request distribution plan at the current moment according to the state feature parameters at the current moment and the corresponding reward weight includes:

[0023] Calculate the reward value corresponding to each database server according to the state feature parameters at the current moment of each database server and the corresponding reward weight;

[0024] Average the reward values corresponding to each database server as the reward value corresponding to the request distribution plan at the current moment.

[0025] Optionally, training the database request distribution model using a self-supervised algorithm according to the training sample data of the database request distribution model includes:

[0026] Input the state feature parameters at the current moment and the request distribution plan at the current moment into the current value network of the database request distribution model to obtain the value score of the request distribution plan at the current moment;

[0027] Input the state feature parameters at the next moment into the target policy network of the database request distribution model to obtain the request distribution plan at the next moment;

[0028] Input the state feature parameters at the next moment and the request distribution plan at the next moment into the target value network of the database request distribution model to obtain the value score of the request distribution plan at the next moment;

[0029] Determine the loss value of the database request distribution model according to the value score of the request distribution plan at the current moment and the value score of the request distribution plan at the next moment, and update the parameters of the database request distribution model by using the loss value of the database request distribution model.

[0030] In a second aspect, an embodiment of the present application further provides a database request distribution method, and the database request distribution method includes:

[0031] Obtain the state parameters of each database server at the current moment;

[0032] Input the state parameters of each database server at the current moment into the database request distribution model to obtain the request distribution plan at the next moment;

[0033] After receiving the database request at the next moment, distribute the database request at the next moment to the corresponding database server according to the request distribution plan at the next moment;

[0034] Wherein, the database request distribution model is trained by using the training method of any one of the foregoing database request distribution models.

[0035] Optionally, the step of distributing the database request at the next moment to the corresponding database server according to the request distribution plan at the next moment includes:

[0036] Perform risk control processing on the request distribution plan at the next moment by using a preset risk control strategy;

[0037] When the risk control processing is completed, distribute the database request at the next moment to the corresponding database server.

[0038] Optionally, the request distribution plan at the next moment includes the number of requests assigned to each database server at the next moment, and the step of performing risk control inspection on the request distribution plan at the next moment by using a preset risk control strategy includes:

[0039] Compare the number of requests assigned to each database server at the next moment with a preset request quantity threshold respectively;

[0040] If the number of requests assigned to each database server at the next moment is less than the preset request quantity threshold, then according to the request distribution scheme at the next moment, distribute the database requests at the next moment to the corresponding database servers;

[0041] Otherwise, distribute the database requests at the next moment evenly to each database server.

[0042] Optionally, after distributing the database requests at the next moment to the corresponding database servers according to the request distribution scheme at the next moment, the method further includes:

[0043] Store the request distribution scheme at the next moment as historical distribution data in the sample pool, so as to dynamically update the database request distribution model by using the historical distribution data in the sample pool.

[0044] In a third aspect, an embodiment of the present application further provides a training device for a database request distribution model, and the training device for the database request distribution model includes:

[0045] A construction unit, configured to construct a sample pool for the database request distribution model;

[0046] A first acquisition unit, configured to acquire training sample data of the database request distribution model from the sample pool, where the training sample data includes state characteristic parameters of each database server at the current moment, the request distribution scheme at the current moment, the reward value corresponding to the request distribution scheme at the current moment, and the state characteristic parameters at the next moment;

[0047] A training unit, configured to train the database request distribution model by using a self-supervised algorithm according to the training sample data of the database request distribution model;

[0048] An update unit, configured to acquire historical distribution data, and dynamically update the database request distribution model by using the historical distribution data.

[0049] In a fourth aspect, an embodiment of the present application further provides a database request distribution device, and the database request distribution device includes:

[0050] A second acquisition unit, configured to acquire the state parameters of each database server at the current moment;

[0051] An input unit, configured to input the state parameters of each database server at the current moment into the database request distribution model to obtain the request distribution scheme at the next moment;

[0052] A distribution unit, configured to, after receiving the database requests at the next moment, distribute the database requests at the next moment to the corresponding database servers according to the request distribution scheme at the next moment;

[0053] Among them, the database request distribution model is trained based on the training method of any one of the foregoing database request distribution models.

[0054] In a fifth aspect, an embodiment of the present application further provides an electronic device, including:

[0055] A processor; and

[0056] A memory arranged to store computer-executable instructions, and when the executable instructions are executed, the processor executes the training method of any one of the foregoing database request distribution models, or executes the database request distribution method of any one of the foregoing.

[0057] In a sixth aspect, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device executes the training method of any one of the foregoing database request distribution models, or executes the database request distribution method of any one of the foregoing.

[0058] The above at least one technical solution adopted in the embodiment of the present application can achieve the following beneficial effects: The training method of the database request distribution model in the embodiment of the present application first constructs a sample pool of the database request distribution model; then obtains the training sample data of the database request distribution model from the sample pool, and the training sample data in the sample pool includes the state characteristic parameters of each database server at the current moment, the request distribution scheme at the current moment, the reward value corresponding to the request distribution scheme at the current moment, and the state characteristic parameters at the next moment; then trains the database request distribution model according to the training sample data of the database request distribution model by using a self-supervised algorithm; finally, obtains historical distribution data and dynamically updates the database request distribution model by using the historical distribution data. The training method of the database request distribution model in the embodiment of the present application uses a self-supervised algorithm to train the database request distribution model, and can update the parameters by itself. After the model is trained, the parameter recommendation is faster, and the database server can be adjusted in seconds. In addition, the model can dynamically learn from historical distribution data, reuse historical experience, and can always monitor the current state of the database server to give different distribution decision schemes, improving the overall operation efficiency of the database system. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0060] Figure 1Schematic flowchart of a method for training a database request distribution model in an embodiment of the present application;

[0061] Figure 2 Schematic flowchart of a method for distributing database requests in an embodiment of the present application;

[0062] Figure 3 Schematic overall flowchart of database request distribution in an embodiment of the present application;

[0063] Figure 4 Schematic structural diagram of a device for training a database request distribution model in an embodiment of the present application;

[0064] Figure 5 Schematic structural diagram of a device for distributing database requests in an embodiment of the present application;

[0065] Figure 6 Schematic structural diagram of an electronic device in an embodiment of the present application. Detailed implementation manners

[0066] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0067] The following details the technical solutions provided by each embodiment of the present application in conjunction with the drawings.

[0068] The request distribution schemes in the prior art can be mainly divided into two methods. The first method is to set fixed rules based on experience. The specific implementation of this method is that technicians set the number of databases and the distribution volume of database requests based on experience. For example, after technicians estimate the transaction volume at peak hours, they give a solution that 8 database servers need to be deployed, set the specific number of requests received by each database server, and then set some opening rules, such as opening eight at peak hours and six at other times, etc. The defect of this method is that it is too fixed and cannot cope with emergencies. Secondly, the optimal solution cannot be obtained.

[0069] The second method is request distribution implemented based on traditional algorithms. Most current systems adopt this method, and the most widely used one is the MD5 modulo database sharding algorithm, which evenly distributes the total number of requests to each database server. However, a major drawback of this method is that each time the request distribution is treated as the first time, historical experience cannot be reused, and each time it is evenly distributed. Even if there have been cases of lag or excessive database load before, there is no record or reference, and the adjustment is still made according to the original rules. Moreover, this method cannot solve the globally optimal solution.

[0070] Based on this, the embodiments of the present application provide a method for training a database request distribution model, as Figure 1 shown, which provides a schematic flowchart of a method for training a database request distribution model in the embodiments of the present application. The method for training the database request distribution model includes the following steps S110 to S140:

[0071] Step S110, construct a sample pool for the database request distribution model.

[0072] When training the database request distribution model, it is necessary to first construct a sample pool for the database request distribution model. Specifically, training sample data for training the database request distribution model can be constructed. The amount of data in the sample pool can be flexibly set according to the actual business scenario and training requirements, and no specific limitation is made here.

[0073] Step S120, obtain the training sample data for the database request distribution model from the sample pool. The training sample data includes the state characteristic parameters of each database server at the current moment, the request distribution scheme at the current moment, the reward value corresponding to the request distribution scheme at the current moment, and the state characteristic parameters at the next moment.

[0074] Step S130, train the database request distribution model using a self-supervised algorithm according to the training sample data of the database request distribution model.

[0075] After constructing the sample pool, it is necessary to obtain the training sample data from the sample pool to train the database request distribution model. For example, the training sample data can be sampled from the sample pool by batch sampling, and the database request distribution model can be iteratively trained in batches using the batch-sampled training sample data.

[0076] The database request distribution model trained in the embodiments of this application can, based on reinforcement learning techniques, use the DDPG (Deep Deterministic Policy Gradient) algorithm to achieve self-supervised training. Reinforcement learning, also known as reward learning, evaluation learning, or enhanced learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of an agent achieving maximum reward or a specific goal through learning strategies during the interaction with the environment. The DDPG algorithm was proposed to solve the problem of continuous action control. Most previous algorithms only mapped the evaluation function of state-action from a discrete space to a continuous space using neural networks, without solving the problem of action discreteness, while DDPG fundamentally solved this problem. In the business scenario of database request concurrency applied in the embodiments of this application, the real-time distribution of database requests can be regarded as a problem of continuous action control.

[0077] Based on the principle of the DDPG algorithm, each data sample in the sample pool constructed in the embodiments of this application mainly includes several dimensions of information such as the state feature parameter s of each database server at the current moment, the request distribution scheme a at the current moment, the reward value r corresponding to the request distribution scheme at the current moment, and the state feature parameter s' at the next moment.

[0078] Step S140, obtain historical distribution data, and use the historical distribution data to dynamically update the database request distribution model.

[0079] The database request distribution model trained in the foregoing steps can already meet the requirements of database request distribution in a certain business scenario. In business scenarios such as commemorative coin issuance, historical distribution experience also has high reference value for the distribution decision at the current moment.

[0080] Therefore, based on the model trained in the above steps in the embodiments of this application, historical distribution data is further obtained, and the database request distribution model is dynamically updated using the historical distribution data. The historical distribution data can be regarded as the relevant data generated after actually applying the database distribution model to the business environment for real request distribution, so that the model can further consider historical distribution experience on the basis of considering the current state of the database server, thereby further improving the overall operation efficiency of the database system.

[0081] The training method of the database request distribution model according to the embodiments of the present application uses a self-supervised algorithm to train the database request distribution model, which can update parameters by itself. After the model is trained, parameter recommendation is relatively fast, and the database server can be adjusted in seconds. In addition, the model can dynamically learn from historical distribution data, reuse historical experience, and continuously monitor the current state of the database server to give different distribution decision plans, improving the overall operation efficiency of the database system.

[0082] In some embodiments of the present application, the sample pool for constructing the database request distribution model includes: randomly sampling the state parameters of each database server at the current moment; converting the randomly sampled state parameters at the current moment into state feature parameters at the current moment; inputting the state feature parameters at the current moment into the current policy network of the database request distribution model to obtain the request distribution plan at the current moment; executing the request distribution plan at the current moment to obtain the reward value corresponding to the request distribution plan at the current moment and the state feature parameters of each database server at the next moment; and constructing training sample data from the state feature parameters of each database server at the current moment, the request distribution plan at the current moment and the corresponding reward value, and the state feature parameters at the next moment, and storing the training sample data in the sample pool.

[0083] When constructing the sample pool of the database request distribution model, the original state parameters of each database server at the current moment can be randomly sampled first. Here, the state parameters can specifically include the performance parameters and configuration parameters of the database server. The performance parameters can, for example, include parameters such as the core processor usage rate (CPU), memory occupancy rate (Mem), and disk storage space (Disk) of the database server. The configuration parameters can, for example, include connection configuration and network configuration of the database server. Of course, specifically which dimensions of state parameters are included can be flexibly set by those skilled in the art according to actual needs, and no specific limitation is made here.

[0084] Convert the state parameters at the current moment obtained by random sampling into the form of a feature vector, that is, obtain the state feature parameter s at the current moment. Based on the network architecture of the DDPG algorithm, input the state feature parameter s at the current moment into the current policy network of the database request distribution model to obtain the request distribution scheme a at the current moment. Executing the request distribution scheme a at the current moment can obtain the reward value r corresponding to the request distribution scheme a at the current moment and the state feature parameters s' of each database server at the next moment. The reward value can be the reward score assigned to the current distribution scheme based on a certain reward algorithm. The reward algorithm will analyze and calculate based on the representation vector of the database server at the current moment and give guiding opinions. Finally, the state feature parameter s at the current moment of each database server, the request distribution scheme a at the current moment and the corresponding reward value r, and the state feature parameter s' at the next moment are formed into training sample data {s, a, r, s'} and stored in the sample pool.

[0085] In some embodiments of the present application, the reward value corresponding to the request distribution scheme at the current moment is obtained in the following manner: determine the reward weight of the state feature parameter according to the importance degree of the state feature parameter in the database request distribution scenario; calculate the reward value corresponding to the request distribution scheme at the current moment according to the state feature parameter at the current moment and the corresponding reward weight; wherein, the state feature parameter includes at least one dimension of the number of transactions processed per second, the transaction execution time per second, the core processor utilization rate, the memory occupancy rate, and the disk storage space.

[0086] The state feature parameters in the embodiments of the present application mainly can include feature parameters of multiple dimensions such as the number of transactions processed per second (Tps), the transaction execution time per second (ART), the core processor utilization rate (CPU), the memory occupancy rate (Mem), and the disk storage space (Disk). In different business scenarios, the importance degrees of different dimensions of state feature parameters are different. For example, in the scenario of commemorative coin concurrent request distribution, the core processor utilization rate, the memory occupancy rate, and the disk storage space have a greater impact on the database performance and operation efficiency.

[0087] Based on this, the embodiments of the present application can determine the importance degree of state feature parameters of each dimension in combination with the actual database request distribution scenario. The higher the importance degree, the greater the assigned reward weight. On the contrary, the lower the importance degree, the smaller the assigned reward weight. By performing weighted summation on each state feature parameter and the assigned reward weight, the reward value corresponding to the request distribution scheme at the current moment can be obtained. The reward value r(t) at the current moment can be expressed in the following form, for example:

[0088] r(t) = aT + bArt + cCpu + dMem + eDisk

[0089] Among them, a, b, c, d, and e are the reward weights corresponding to each state characteristic parameter.

[0090] Since the setting of the reward weight takes into account the requirements of the actual distribution scenario, a reward value that is more suitable for the requirements of the actual distribution scenario can be obtained.

[0091] In some embodiments of the present application, calculating the reward value corresponding to the request distribution scheme at the current moment according to the state characteristic parameter and the corresponding reward weight at the current moment includes: calculating the reward value corresponding to each database server according to the state characteristic parameter and the corresponding reward weight at the current moment of each database server; averaging the reward values corresponding to each database server as the reward value corresponding to the request distribution scheme at the current moment.

[0092] Since in the scenario of database request distribution, multiple concurrent database requests are distributed to multiple database servers according to the distribution scheme, and each database server has its own current state parameter, the calculation of the reward value of the current distribution scheme needs to comprehensively consider the current state of each database server to achieve the goal of overall optimal system state.

[0093] Therefore, the embodiments of the present application can calculate the reward values corresponding to each database server respectively based on the weighted summation method of the foregoing embodiments, and then average the reward values corresponding to multiple database servers as the final reward value, so that the training process of the model can comprehensively consider the current states of all database servers while focusing on optimizing the state of the core characteristic parameters to achieve the global optimum of the system.

[0094] In some embodiments of the present application, training the database request distribution model using the self-supervised algorithm according to the training sample data of the database request distribution model includes: inputting the state characteristic parameter at the current moment and the request distribution scheme at the current moment into the current value network of the database request distribution model to obtain the value score of the request distribution scheme at the current moment; inputting the state characteristic parameter at the next moment into the target policy network of the database request distribution model to obtain the request distribution scheme at the next moment; inputting the state characteristic parameter at the next moment and the request distribution scheme at the next moment into the target value network of the database request distribution model to obtain the value score of the request distribution scheme at the next moment; determining the loss value of the database request distribution model according to the value score of the request distribution scheme at the current moment and the value score of the request distribution scheme at the next moment, and updating the parameters of the database request distribution model using the loss value of the database request distribution model.

[0095] Based on the network architecture of the DDPG algorithm, when using the training sample data to train the database request distribution model, the state feature parameters s at the current moment and the request distribution scheme a at the current moment are input into the current value network of the database request distribution model, and the value score Q of the request distribution scheme a at the current moment output by the current value network can be obtained. The state feature parameters s' at the next moment are input into the target policy network of the database request distribution model, and the request distribution scheme a' at the next moment can be obtained. The state feature parameters s' at the next moment and the request distribution scheme a' at the next moment are input into the target value network of the database request distribution model to obtain the value score Q' of the request distribution scheme at the next moment.

[0096] On the one hand, the best strategy for DDPG training is to learn a good value network, and the selected database request distribution scheme can maximize the Q value. The purpose of DDPG training is also to solve the database request distribution scheme that maximizes the Q value. Therefore, the gradient for optimizing the policy network is to maximize the Q value. The constructed loss function can be to take a negative sign for Q and put this loss function into the optimizer, and it will automatically minimize the loss, that is, maximize Q. On the other hand, a loss function is constructed based on the value score Q of the request distribution scheme a at the current moment and the value score Q' of the request distribution scheme at the next moment. The constructed loss function can be to directly calculate the mean square error of these two values. After constructing the loss function, put it into the optimizer to let it automatically minimize the loss, so as to obtain the trained database request distribution model.

[0097] The embodiment of the present application also provides a database request distribution method, as Figure 2 shown, which provides a flowchart of a database request distribution method in the embodiment of the present application. The database request distribution method includes the following steps S210 to step S230:

[0098] Step S210, obtain the state parameters of each database server at the current moment;

[0099] Step S220, input the state parameters of each database server at the current moment into the database request distribution model to obtain the request distribution scheme at the next moment;

[0100] Step S230, after receiving the database request at the next moment, according to the request distribution scheme at the next moment, distribute the database request at the next moment to the corresponding database server;

[0101] Among them, the database request distribution model is trained by the training method of any one of the foregoing database request distribution models.

[0102] The database request distribution model trained based on the foregoing embodiments can be applied to an actual database request distribution scenario. Specifically, the status parameters of each database server at the current moment can be obtained first, such as the performance parameters and configuration parameters mentioned in the foregoing embodiments. Then, the status parameters of each database server at the current moment are directly input into the database request distribution model, and the request distribution plan R for the next moment output by the database request distribution model can be obtained. After receiving the database request for the next moment, the request can be allocated according to the request distribution plan R for the next moment.

[0103] In some embodiments of the present application, the distributing the database request for the next moment to the corresponding database server according to the request distribution plan for the next moment includes: performing risk control processing on the request distribution plan for the next moment by using a preset risk control strategy; and when the risk control processing is completed, distributing the database request for the next moment to the corresponding database server.

[0104] Since the training of the database request distribution model in the embodiments of the present application mainly focuses on adjusting the efficiency of the database server, and the recommended request distribution plan is mainly used to make the database system in the best state, the request distribution plan output by the data request distribution model may not be adapted to the actual business scenario requirements. Considering the particularity of the database request distribution scenario, the embodiments of the present application can add a risk control strategy to perform a layer of risk verification on the distribution plan output by the model, and then perform output execution after determining that there is no risk, so as to further improve the applicability of the data request distribution model and the overall availability of the database system.

[0105] In some embodiments of the present application, the request distribution plan for the next moment includes the number of requests allocated to each database server at the next moment. The performing risk control inspection on the request distribution plan for the next moment by using a preset risk control strategy includes: comparing the number of requests allocated to each database server at the next moment with a preset request quantity threshold respectively; if the number of requests allocated to each database server at the next moment is less than the preset request quantity threshold, then distributing the database request for the next moment to the corresponding database server according to the request distribution plan for the next moment; otherwise, distributing the database request for the next moment evenly to each database server.

[0106] One of the main risk situations existing in the request distribution scheme output by the database request distribution model is the extremely unbalanced request allocation. For example, currently there are 100,000 concurrent requests. After calculation and analysis by the model, 90,000 of them are evenly allocated to database server 1, and the remaining 10,000 are evenly allocated to database servers 2 and 3. If database server 1 receives and processes these 90,000 requests simultaneously within a short period of time, it is very likely that the server will crash or go down, resulting in a large number of request processing failures.

[0107] Based on this, the embodiments of the present application can set the traffic limit of a single database server according to the actual business scenario requirements and system operation capabilities, and restrict the maximum number of concurrent requests processed by each database server. Compare the number of requests allocated to each database server in the request distribution scheme for the next moment with this maximum value respectively. If the number of requests allocated to all database servers is less than this maximum value, the distribution scheme is within the safe range and can be executed. If the number of requests allocated to any one database server reaches this maximum value, it is considered that there is a greater risk in the current request distribution scheme and it cannot be executed. At this time, the average distribution strategy can be adopted to evenly allocate all requests to each database server.

[0108] In some embodiments of the present application, after distributing the database requests for the next moment to the corresponding database servers according to the request distribution scheme for the next moment, the method further includes: storing the request distribution scheme for the next moment as historical distribution data in the sample pool, so as to dynamically update the database request distribution model by using the historical distribution data in the sample pool.

[0109] After completing the current request distribution, the request distribution scheme output by the model for the next moment can be used as the action a output by the current policy network, together with other dimensional information such as s, r, and s', to construct historical distribution data and store it in the sample pool in the same way as constructing the training sample described above. Thus, the model can be dynamically updated regularly or in real time by using the historical distribution data in the sample pool, further optimizing the effect of the model.

[0110] For the convenience of understanding the above embodiments, as Figure 3As shown in the figure, a schematic diagram of the overall process of database request distribution in an embodiment of the present application is provided. In the model training stage, first obtain the state parameters of each database server in the database platform at the current moment, including performance parameters, configuration parameters, etc., and then convert the state parameters of each database server into the form of feature vectors to obtain the state feature parameter s at the current moment. Input the state feature parameter s at the current moment into the current policy network of the database request distribution model to obtain the request distribution plan a at the current moment. Executing the request distribution plan a at the current moment can obtain the reward value r corresponding to the request distribution plan a at the current moment and the state feature parameter s' of each database server at the next moment. The reward value can be obtained by analyzing and calculating the representation vector of the database server at the current moment through a certain reward algorithm. Finally, the state feature parameter s of each database server at the current moment, the request distribution plan a at the current moment and the corresponding reward value r, and the state feature parameter s' at the next moment are formed into training sample data {s, a, r, s'} and stored in the sample pool. Through batch sampling, training sample data is obtained from the sample pool to train the database request distribution model, and finally the database request distribution model is dynamically updated in combination with historical distribution data.

[0111] In the model application stage, first obtain the state parameters of each database server at the current moment, and then input the state parameters of each database server at the current moment into the above-trained database request distribution model to obtain the request distribution plan for the next moment. After receiving the database request for the next moment, according to the request distribution plan for the next moment, distribute the database request for the next moment to the corresponding database server, and update the current distribution plan as historical distribution data to the sample pool for subsequent dynamic model update.

[0112] In summary, the training method and database request distribution method of the database request distribution model of the present application have at least achieved the following technical effects:

[0113] 1) Using a self-supervised algorithm, it can automatically update parameters and can always monitor the current state of the database server to give different decision-making schemes;

[0114] 2) The model can learn from historical distribution data and reuse historical distribution experience;

[0115] 3) After the model is trained, the parameter recommendation is fast, and the database can be adjusted in seconds;

[0116] 4) It is not necessary for training samples to be independent of each other. They can be correlated with each other. The sample processing is very simple and the feasibility is strong.

[0117] An embodiment of the present application also provides a training device 400 for a database request distribution model, asFigure 4 As shown in the figure, a structural schematic diagram of a training device for a database request distribution model in an embodiment of the present application is provided. The training device 400 for the database request distribution model includes:

[0118] A construction unit 410, configured to construct a sample pool for the database request distribution model;

[0119] A first acquisition unit 420, configured to acquire training sample data for the database request distribution model from the sample pool. The training sample data includes state characteristic parameters of each database server at the current moment, a request distribution scheme at the current moment, a reward value corresponding to the request distribution scheme at the current moment, and state characteristic parameters at the next moment;

[0120] A training unit 430, configured to train the database request distribution model using a self-supervised algorithm according to the training sample data of the database request distribution model;

[0121] An update unit 440, configured to acquire historical distribution data and dynamically update the database request distribution model using the historical distribution data.

[0122] In some embodiments of the present application, the construction unit 410 is specifically configured to: randomly sample state parameters of each database server at the current moment; convert the randomly sampled state parameters at the current moment into state characteristic parameters at the current moment; input the state characteristic parameters at the current moment into the current policy network of the database request distribution model to obtain a request distribution scheme at the current moment; execute the request distribution scheme at the current moment to obtain a reward value corresponding to the request distribution scheme at the current moment and state characteristic parameters of each database server at the next moment; and form training sample data from the state characteristic parameters of each database server at the current moment, the request distribution scheme at the current moment and the corresponding reward value, and the state characteristic parameters at the next moment, and store the training sample data in the sample pool.

[0123] In some embodiments of the present application, the reward value corresponding to the request distribution scheme at the current moment is obtained in the following manner: determining a reward weight of the state characteristic parameter according to the importance degree of the state characteristic parameter in the database request distribution scenario; calculating the reward value corresponding to the request distribution scheme at the current moment according to the state characteristic parameter at the current moment and the corresponding reward weight; wherein the state characteristic parameter includes at least one dimension of the number of transactions processed per second, the transaction execution time per second, the core processor usage rate, the memory occupancy rate, and the disk storage space.

[0124] In some embodiments of the present application, the reward value corresponding to the request distribution scheme at the current moment is obtained in the following manner: calculate the reward value corresponding to each database server according to the state characteristic parameters and the corresponding reward weights of each database server at the current moment; average the reward values corresponding to each database server to obtain the reward value corresponding to the request distribution scheme at the current moment.

[0125] In some embodiments of the present application, the training unit 430 is specifically configured to: input the state characteristic parameters at the current moment and the request distribution scheme at the current moment into the current value network of the database request distribution model to obtain the value score of the request distribution scheme at the current moment; input the state characteristic parameters at the next moment into the target policy network of the database request distribution model to obtain the request distribution scheme at the next moment; input the state characteristic parameters at the next moment and the request distribution scheme at the next moment into the target value network of the database request distribution model to obtain the value score of the request distribution scheme at the next moment; determine the loss value of the database request distribution model according to the value score of the request distribution scheme at the current moment and the value score of the request distribution scheme at the next moment, and update the parameters of the database request distribution model by using the loss value of the database request distribution model.

[0126] It can be understood that the above training device of the database request distribution model can implement each step of the training method of the database request distribution model provided in the foregoing embodiments. The relevant explanations regarding the training method of the database request distribution model are applicable to the training device of the database request distribution model, and will not be elaborated here.

[0127] Embodiments of the present application further provide a database request distribution device 500, as Figure 5 shown, which provides a structural schematic diagram of a database request distribution device in an embodiment of the present application. The database request distribution device 500 includes:

[0128] A second acquisition unit 510, configured to acquire the state parameters of each database server at the current moment;

[0129] An input unit 520, configured to input the state parameters of each database server at the current moment into the database request distribution model to obtain the request distribution scheme at the next moment;

[0130] A distribution unit 530, configured to, after receiving the database request at the next moment, distribute the database request at the next moment to the corresponding database server according to the request distribution scheme at the next moment;

[0131] Wherein, the database request distribution model is trained based on any one of the foregoing database request distribution model training methods.

[0132] In some embodiments of the present application, the distribution unit 530 is specifically configured to: perform risk control processing on the request distribution plan for the next moment by using a preset risk control strategy; and when the risk control processing is completed, distribute the database request for the next moment to the corresponding database server.

[0133] In some embodiments of the present application, the request distribution plan for the next moment includes the number of requests allocated to each database server at the next moment. The distribution unit 530 is specifically configured to: compare the number of requests allocated to each database server at the next moment with a preset request quantity threshold respectively; if the number of requests allocated to each database server at the next moment is less than the preset request quantity threshold, then distribute the database request for the next moment to the corresponding database server according to the request distribution plan for the next moment; otherwise, distribute the database request for the next moment evenly to each database server.

[0134] In some embodiments of the present application, the device further includes: a storage unit, configured to store the request distribution plan for the next moment as historical distribution data in a sample pool, so as to dynamically update the database request distribution model by using the historical distribution data in the sample pool.

[0135] It can be understood that the above database request distribution device can implement each step of the database request distribution method provided in the foregoing embodiments. The relevant explanations regarding the database request distribution method are applicable to the database request distribution device and will not be elaborated herein.

[0136] Figure 6 It is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 6 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk storage device, etc. Of course, the electronic device may also include other hardware required for other services.

[0137] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a bidirectional arrow is used in

[0138] memory, which is used to store programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0139] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a training device for the database request distribution model at the logical level. The processor executes the program stored in the memory.

[0140] The above, such as in this application Figure 1The method executed by the training device of the database request distribution model disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in software form. The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0141] The electronic device can also execute Figure 1 the method executed by the training device of the database request distribution model in Figure 1 the illustrated embodiment and implement the functions of the training device of the database request distribution model in

[0142] Embodiments of the present application also propose a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 1 the method executed by the training device of the database request distribution model in the illustrated embodiment.

[0143] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0144] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0145] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0147] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0148] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0149] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage space or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0150] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0151] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage space devices, CD-ROMs, optical storage devices, etc.) that contain computer-usable program codes.

[0152] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A training method for a database request distribution model, characterized in that, The training method of the database request distribution model includes: Constructing a sample pool for the database request distribution model; Obtaining training sample data for the database request distribution model from the sample pool, where the training sample data includes the state characteristic parameters of each database server at the current moment, the request distribution scheme at the current moment, the reward value corresponding to the request distribution scheme at the current moment, and the state characteristic parameters at the next moment; Training the database request distribution model using a self-supervised algorithm according to the training sample data of the database request distribution model; Obtaining historical distribution data and dynamically updating the database request distribution model using the historical distribution data; Among them, the state characteristic parameters include at least one dimension of the number of transactions processed per second, the transaction execution time per second, the core processor utilization rate, the memory occupancy rate, and the disk storage space; The reward value is a reward score assigned to the request distribution scheme at the current moment using a reward algorithm based on the state characteristic parameters of the database server at the current moment; The training of the database request distribution model using a self-supervised algorithm according to the training sample data of the database request distribution model includes: Inputting the state characteristic parameters at the current moment and the request distribution scheme at the current moment into the current value network of the database request distribution model to obtain the value score of the request distribution scheme at the current moment; Inputting the state characteristic parameters at the next moment into the target policy network of the database request distribution model to obtain the request distribution scheme at the next moment; Inputting the state characteristic parameters at the next moment and the request distribution scheme at the next moment into the target value network of the database request distribution model to obtain the value score of the request distribution scheme at the next moment; Determining the loss value of the database request distribution model according to the value score of the request distribution scheme at the current moment and the value score of the request distribution scheme at the next moment, and updating the parameters of the database request distribution model using the loss value of the database request distribution model.

2. The training method of the database request distribution model according to claim 1, wherein The construction of the sample pool for the database request distribution model includes: Randomly sampling the state parameters of each database server at the current moment; Converting the randomly sampled state parameters at the current moment into state characteristic parameters at the current moment; Inputting the state characteristic parameters at the current moment into the current policy network of the database request distribution model to obtain the request distribution scheme at the current moment; Executing the request distribution scheme at the current moment to obtain the reward value corresponding to the request distribution scheme at the current moment and the state characteristic parameters of each database server at the next moment; Constructing training sample data from the state characteristic parameters of each database server at the current moment, the request distribution scheme at the current moment and the corresponding reward value, and the state characteristic parameters at the next moment, and storing them in the sample pool.

3. The training method of the database request distribution model according to claim 1, characterized in that The reward value corresponding to the request distribution scheme at the current moment is obtained through the following method: Determining the reward weight of the state characteristic parameters according to the importance of the state characteristic parameters in the database request distribution scenario; Calculating the reward value corresponding to the request distribution scheme at the current moment according to the state characteristic parameters at the current moment and the corresponding reward weight.

4. The training method of the database request distribution model according to claim 3, characterized in that Calculating the reward value corresponding to the request distribution scheme at the current moment according to the state characteristic parameters and corresponding reward weights at the current moment includes: Calculating the reward value corresponding to each database server according to the state characteristic parameters and corresponding reward weights of each database server at the current moment; Averaging the reward values corresponding to each database server as the reward value corresponding to the request distribution scheme at the current moment.

5. A method for distributing database requests, characterized in that, The database request distribution method includes: Obtaining the state parameters of each database server at the current moment; Inputting the state parameters of each database server at the current moment into the database request distribution model to obtain the request distribution scheme for the next moment; After receiving the database request for the next moment, distributing the database request for the next moment to the corresponding database server according to the request distribution scheme for the next moment; Wherein, the database request distribution model is trained based on the training method of the database request distribution model according to any one of claims 1 to 4.

6. The database request distribution method according to claim 5, wherein The distributing the database request for the next moment to the corresponding database server according to the request distribution scheme for the next moment includes: Performing risk control processing on the request distribution scheme for the next moment by using a preset risk control strategy; When the risk control processing is completed, distributing the database request for the next moment to the corresponding database server.

7. The database request distribution method according to claim 6, wherein The request distribution scheme for the next moment includes the number of requests allocated to each database server at the next moment. The performing risk control inspection on the request distribution scheme for the next moment by using a preset risk control strategy includes: Comparing the number of requests allocated to each database server at the next moment with a preset request quantity threshold respectively; If the number of requests allocated to each database server at the next moment is less than the preset request quantity threshold, distributing the database request for the next moment to the corresponding database server according to the request distribution scheme for the next moment; Otherwise, distributing the database request for the next moment evenly to each database server.

8. The database request distribution method according to claim 5, characterized in that After distributing the database request for the next moment to the corresponding database server according to the request distribution scheme for the next moment, the method further includes: Storing the request distribution scheme for the next moment as historical distribution data in the sample pool to dynamically update the database request distribution model by using the historical distribution data in the sample pool.

9. A training device for a database request distribution model, characterized in that, The training device of the database request distribution model includes: A construction unit for constructing a sample pool of the database request distribution model; A first acquisition unit for acquiring the training sample data of the database request distribution model from the sample pool, where the training sample data includes the state characteristic parameters of each database server at the current moment, the request distribution scheme at the current moment, the reward value corresponding to the request distribution scheme at the current moment, and the state characteristic parameters for the next moment; A training unit for training the database request distribution model by using a self-supervised algorithm according to the training sample data of the database request distribution model. An update unit, configured to obtain historical distribution data and dynamically update the database request distribution model by using the historical distribution data; Wherein, the state characteristic parameters include at least one dimension of the number of transactions processed per second, the transaction execution time per second, the core processor utilization rate, the memory occupancy rate, and the disk storage space; The reward value is a reward score assigned to the request distribution scheme at the current moment by using a reward algorithm based on the state characteristic parameters of the database server at the current moment; The training unit is specifically configured to: Input the state characteristic parameters at the current moment and the request distribution scheme at the current moment into the current value network of the database request distribution model to obtain the value score of the request distribution scheme at the current moment; Input the state characteristic parameters at the next moment into the target policy network of the database request distribution model to obtain the request distribution scheme at the next moment; Input the state characteristic parameters at the next moment and the request distribution scheme at the next moment into the target value network of the database request distribution model to obtain the value score of the request distribution scheme at the next moment; Determine the loss value of the database request distribution model according to the value score of the request distribution scheme at the current moment and the value score of the request distribution scheme at the next moment, and update the parameters of the database request distribution model by using the loss value of the database request distribution model; 10. A database request distribution device, characterized in that, The database request distribution device includes: A second acquisition unit, configured to acquire the state parameters of each database server at the current moment; An input unit, configured to input the state parameters of each database server at the current moment into the database request distribution model to obtain the request distribution scheme at the next moment; A distribution unit, configured to, after receiving the database request at the next moment, distribute the database request at the next moment to the corresponding database server according to the request distribution scheme at the next moment; Wherein, the database request distribution model is trained by using the training method of the database request distribution model according to any one of claims 1 to 4; 11. An electronic device, comprising: A processor; And A memory arranged to store computer-executable instructions, which when executed cause the processor to execute the training method of the database request distribution model according to any one of claims 1 to 4, or execute the database request distribution method according to any one of claims 5 to 8; 12. A computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute the training method of the database request distribution model according to any one of claims 1 to 4, or execute the database request distribution method according to any one of claims 5 to 8;