DQN index recommendation method and system based on multi-target reinforcement learning
Through multi-objective reinforcement learning, the problem of slow query speed in the existing technology is solved, and efficient and stable data query is achieved.
Patent Information
- Application Number
- CN202510408756.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-08-01
AI Technical Summary
The prior art lacks consideration of user query mode when recommending data query indexes, resulting in slowing down the query speed.
Through multi-objective reinforcement learning, configure the initial index set for each target, perform reinforcement learning to generate a new index, and select the optimal index according to the user query mode, build the optimal index recommendation model, and recommend the optimal index to the user.
It improves the efficiency and response speed of data queries, avoids query lag caused by insufficient storage space, and ensures the stability of query efficiency.
Smart Images

Figure CN120407856A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of recommendation indexes, and particularly to a DQN index recommendation method and system based on multi-objective reinforcement learning. Background Art
[0002] DQN (Deep Q-network) is an algorithm based on deep learning Q-Learing and reinforcement learning. DQN approximates the Q function (i.e., the state-action value function) through a deep neural network, enabling the agent to gradually learn the optimal strategy in target reinforcement learning, so as to achieve the optimal decision-making of the agent in the replicated environment. The deep neural network can have multiple hidden layers, aiming to approximate the action value of taking action a in a state s. Taking state s as the input of the deep neural network, the deep neural network outputs the action values of each possible action of the agent; if DQN is used to recommend indexes for data queries, the optimal index can be recommended for data queries. In existing related technical solutions, often only consider how to use DQN to recommend the optimal index for data queries. When recommending the optimal index, the factor of the query pattern required by the user is lacking, and it is difficult to recommend the optimal index according to the query pattern required by the user, thus may slow down the user's query speed. Summary of the Invention
[0003] This application provides a DQN index recommendation method and system based on multi-objective reinforcement learning, aiming to improve existing related technical problems.
[0004] In a first aspect, a DQN index recommendation method based on multi-objective reinforcement learning provided by an embodiment of this application may include the following steps:
[0005] Configure an initial index set for each objective in the multi-objective, obtain a feasible new index through reinforcement learning on the initial index set of the objective, and add the new index to the initial index set to generate the index set of the objective;
[0006] Obtain the query pattern input by the user, obtain the selection value when the query pattern selects each index in the index set through DQN, and select the index corresponding to the maximum value of the selection values of all indexes in the index set of each objective as the optimal index of the objective;
[0007] Obtain the optimal indexes of all the objectives in the multi-objective, and obtain the optimal index with the largest selection value among the optimal indexes of all the objectives as the optimal index of the query pattern;
[0008] Recommend the optimal index of the query pattern to the user who inputs the query pattern.
[0009] In the above technical solution, reinforcement learning is performed on the initial index set of the target through multi-objective reinforcement learning to obtain a feasible new index. According to the query pattern required by the user, the optimal index with the largest selection value corresponding to the query pattern is obtained from the optimal indexes of all targets, and is recommended to the user as the optimal index of the query pattern, efficiently recommending the optimal index of the query data for the query pattern required by the user, thereby making the query efficiency of the user's query data very high.
[0010] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0011] The described DQN index recommendation method based on multi-objective reinforcement learning may further include the following steps:
[0012] Based on the query pattern data input by the user and the optimal index data of the query pattern, an optimal index recommendation model is constructed, and the optimal index is recommended to the user through the optimal index recommendation model.
[0013] In the above technical solution, the optimal index is recommended to the user through the constructed optimal index recommendation model, improving the intelligent level of recommending the optimal index, and further improving the efficiency of recommending the optimal index of the query data for the query pattern required by the user.
[0014] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0015] The described DQN index recommendation method based on multi-objective reinforcement learning may further include the following steps:
[0016] When the queried data is updated, an association relationship is established between the updated data and the corresponding index before the data update, and the updated data is queried through the corresponding index before the data update.
[0017] Through the above technical solution, after the data is updated, without waiting for the index of the updated data to be updated, the updated data can be queried through the index before the data update, so the query response speed of the updated data can be improved.
[0018] In a preferred example, the solution of the first aspect of the present application can be further configured as follows:
[0019] The described DQN index recommendation method based on multi-objective reinforcement learning may further include the following steps:
[0020] Sort the selection values of the indexes selected for the query pattern in descending order. When the storage space size required for the current index selected for the query pattern exceeds the preset threshold, change the current index to an index with a selection value smaller than the selection value of the current index, and recommend the indexes with selection values smaller than the selection value of the current index to the user who inputs the query pattern.
[0021] Through the above technical solution, it is possible to avoid the problem that the query freezes when querying using the recommended index due to insufficient storage space used by the index, which slows down the query speed of the query data and thus seriously affects the query efficiency. The above technical solution ensures that the query efficiency will not be greatly affected even when the storage space is not sufficient by changing the current index to the index corresponding to an appropriate size selection value.
[0022] In a preferred example, the solution of the first aspect of the present application can be further configured as:
[0023] The described DQN index recommendation method based on multi-objective reinforcement learning may further include the following steps:
[0024] Test the query response speed of the optimal index recommended for the query pattern through the query pattern test data. When the tested query response speed is lower than the preset speed, strengthen the reinforcement learning of the initial index set of the target and increase the number of feasible new indexes.
[0025] By adopting the above technical solution, it is possible to ensure that the query response speed of the optimal index of the recommended query pattern is not lower than the preset speed, thereby ensuring the query efficiency of the data.
[0026] In a preferred example, the solution of the first aspect of the present application can be further configured as:
[0027] In the step of obtaining the user input query pattern, obtaining the selection value when the query pattern selects each index in the index set through DQN, and selecting the index corresponding to the maximum value of the selection values of all indexes in the index set of each target as the optimal index of the target, the method for obtaining the selection value includes the following expression:
[0028]
[0029] In the formula, C represents the selection value, R represents the expected query response speed, V represents the query speed for querying through the index selected for the query pattern, tan(·) represents the tangent function, and S represents the storage space consumed when querying using the selected index.
[0030] Through the above technical solution, the size of the selection value can be accurately obtained.
[0031] In a second aspect, a DQN index recommendation system based on multi-objective reinforcement learning provided by an embodiment of the present application may include:
[0032] An index set generation module, configured to configure an initial index set for each objective in the multi-objectives, obtain a feasible new index through reinforcement learning on the initial index set of the objective, and add the new index to the initial index set to generate the index set of the objective;
[0033] A target optimal index selection module, configured to obtain the query pattern of the user, obtain the selection value when the query pattern selects each index in the index set through DQN, and select the index corresponding to the maximum value of the selection values of all indexes in the index set of each target as the optimal index of the target;
[0034] A query pattern optimal index acquisition module, configured to obtain the optimal indexes of all the targets in the multi-objectives, and obtain the optimal index with the largest selection value among all the optimal indexes as the optimal index of the query pattern;
[0035] An index recommendation module, configured to recommend the optimal index of the query pattern to the user.
[0036] The solution of the second aspect of the present application may be further configured in a preferred example as follows:
[0037] The above-mentioned DQN index recommendation system based on multi-objective reinforcement learning may further include:
[0038] An optimal index recommendation model construction module, configured to construct an optimal index recommendation model based on the query pattern data input by the user and the optimal index data of the query pattern, and recommend the optimal index to the user through the optimal index recommendation model.
[0039] The solution of the second aspect of the present application may be further configured in a preferred example as follows:
[0040] The above-mentioned DQN index recommendation system based on multi-objective reinforcement learning may further include:
[0041] A current index change module, configured to sort the selection values of the indexes selected for the query pattern in descending order. When the storage space size required for the current index selected for the query pattern exceeds a preset threshold, change the current index to an index with a selection value smaller than the selection value of the current index, and recommend the index with a selection value smaller than the selection value of the current index to the user who inputs the query pattern.
[0042] The solution of the second aspect of the present application can be further configured in a preferred example as follows:
[0043] The described DQN index recommendation system based on multi-objective reinforcement learning may further include:
[0044] A query response speed test module, which is used to test the query response speed of the optimal index recommended for the query pattern by testing data in the query pattern. When the tested query response speed is lower than the preset speed, the reinforcement learning of the initial index set of the target is strengthened, and the number of feasible new indexes is increased.
[0045] Based on the above method item embodiment, the present application correspondingly provides a terminal item embodiment;
[0046] The present application provides a terminal, including a processor, a memory, and a computer program stored in the above memory and configured to be executed by the above processor. When the above processor executes the above computer program, it implements a DQN index recommendation method based on multi-objective reinforcement learning described in any embodiment of the present application.
[0047] Based on the above method item embodiment, the present application correspondingly provides a storage medium item embodiment;
[0048] The present application provides a storage medium, including a processor, a memory, and a computer program stored in the above memory and configured to be executed by the above processor. When the above processor executes the above computer program, it implements a DQN index recommendation method based on multi-objective reinforcement learning described in any embodiment of the present application.
[0049] The present application has at least the following beneficial effects:
[0050] The DQN index recommendation method based on multi-objective reinforcement learning provided by the present application has advantages such as efficiently recommending the optimal index of query data for the query pattern required by the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a flowchart of a DQN index recommendation method based on multi-objective reinforcement learning according to an embodiment of the present application.
[0052] Figure 2 is a block diagram of a DQN index recommendation system structure according to an embodiment of the present application. DETAILED DESCRIPTION
[0053] Next, the technical solutions in this application will be clearly and completely described in conjunction with the accompanying drawings in this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0054] As Figure 1 shown, a DQN index recommendation method based on multi-objective reinforcement learning provided by an embodiment of this application may specifically include the following steps:
[0055] Step S1: Configure an initial index set for each objective in the multi-objective. Through reinforcement learning on the initial index set of the objective, obtain a feasible new index, and add the new index to the initial index set to generate the index set of the objective.
[0056] Step S2: Obtain the query pattern input by the user. Through DQN, obtain the selection value when the query pattern selects each index in the index set, and select the index corresponding to the maximum value of the selection values of all indexes in the index set of each objective as the optimal index of the objective.
[0057] Step S3: Obtain the optimal indexes of all the objectives in the multi-objective, and obtain the optimal index with the largest selection value among all the optimal indexes as the optimal index of the query pattern.
[0058] Step S4: Recommend the optimal index of the query pattern to the user who inputs the query pattern.
[0059] In an embodiment of this application, a DQN index recommendation method based on multi-objective reinforcement learning performs reinforcement learning on the initial index set of the objective through multi-objective reinforcement learning to obtain a feasible new index, and according to the query pattern required by the user, obtains the optimal index with the largest selection value corresponding to the query pattern from the optimal indexes of all objectives as the optimal index of the query pattern and recommends it to the user, efficiently recommending the optimal index of the query data for the query pattern required by the user, thereby making the query efficiency of the user querying data very high.
[0060] In a preferred embodiment, in order to recommend the optimal index for the user through the constructed optimal index recommendation model, improve the intelligent level of recommending the optimal index, and further improve the efficiency of recommending the optimal index of the query data for the query pattern required by the user, the above-mentioned DQN index recommendation method based on multi-objective reinforcement learning may further include the following steps:
[0061] Construct an optimal index recommendation model based on the query pattern data input by the user and the optimal index data of the query pattern, and recommend the optimal index for the user through the optimal index recommendation model.
[0062] In a preferred embodiment, after the data is updated, in order to be able to query the updated data through the index before the data update without waiting for the index to be updated for the updated data, thereby improving the query response speed of the updated data, the DQN index recommendation method based on multi-objective reinforcement learning may further include the following steps:
[0063] When the queried data is updated, establish an association relationship between the updated data and the corresponding index before the data update, and query the updated data through the corresponding index before the data update.
[0064] In a preferred embodiment, in order to avoid the problem that the recommended index may cause a lag when querying due to insufficient storage space used by it, which slows down the query speed of the query data and seriously affects the query efficiency, the above technical solution changes the current index to the index corresponding to the appropriate size of the selection value, so that even when the storage space is not sufficient, the query efficiency can be ensured not to be greatly affected. The DQN index recommendation method based on multi-objective reinforcement learning may further include the following steps:
[0065] Sort the selection values of the indexes selected for the query pattern in descending order. When the storage space size required by the current index selected for the query pattern exceeds the preset threshold, change the current index to an index with a selection value smaller than the selection value of the current index, and recommend the index with a selection value smaller than the selection value of the current index to the user who inputs the query pattern.
[0066] In a preferred embodiment, in order to ensure that the query response speed of the optimal index of the recommended query pattern is not lower than the preset speed, thereby ensuring the query efficiency of the data, the DQN index recommendation method based on multi-objective reinforcement learning may further include the following steps:
[0067] Test the query response speed of the optimal index recommended for the query pattern through the query pattern test data. When the tested query response speed is lower than the preset speed, strengthen the reinforcement learning of the initial index set of the target and increase the number of feasible new indexes.
[0068] In a preferred embodiment, in order to accurately obtain the magnitude of the selection value, in the step of obtaining the query mode of the user input, obtaining the selection value when the query mode selects each index in the index set through DQN, and selecting the index corresponding to the maximum value of the selection values of all indexes in the index set of each target as the optimal index of the target, the method for obtaining the selection value includes the following expression:
[0069]
[0070] In the formula, C represents the selection value, R represents the expected query response speed, V represents the query speed for querying through the index selected for the query mode, tan(·) represents the tangent function, and S represents the storage space consumed when querying using the selected index.
[0071] An embodiment of the present application provides a DQN index recommendation system based on multi-objective reinforcement learning, as Figure 2 shown, and specifically may include:
[0072] An index set generation module, configured to configure an initial index set for each target in the multi-objectives, obtain a feasible new index by performing reinforcement learning on the initial index set of the target, and add the new index to the initial index set to generate the index set of the target;
[0073] A target optimal index selection module, configured to obtain the query mode of the user, obtain the selection value when the query mode selects each index in the index set through DQN, and select the index corresponding to the maximum value of the selection values of all indexes in the index set of each target as the optimal index of the target;
[0074] A query mode optimal index acquisition module, configured to obtain the optimal indexes of all the targets in the multi-objectives, and obtain the optimal index with the largest selection value among the optimal indexes of all the targets as the optimal index of the query mode;
[0075] An index recommendation module, configured to recommend the optimal index of the query mode to the user.
[0076] In a preferred embodiment, the DQN index recommendation system based on multi-objective reinforcement learning may specifically further include:
[0077] An optimal index recommendation model construction module, configured to construct an optimal index recommendation model based on the query mode data input by the user and the optimal index data of the query mode, and recommend the optimal index to the user through the optimal index recommendation model.
[0078] In a preferred embodiment, the DQN index recommendation system based on multi-objective reinforcement learning may specifically further include:
[0079] A current index change module that sorts the selection values of the indexes selected for the query pattern in descending order. When the storage space size required for the current index selected for the query pattern exceeds a preset threshold, the current index is changed to an index with a selection value smaller than the selection value of the current index, and the index with a selection value smaller than the selection value of the current index is recommended to the user who inputs the query pattern.
[0080] In a preferred embodiment, the DQN index recommendation system based on multi-objective reinforcement learning may specifically further include:
[0081] A query response speed test module for testing the query response speed of the optimal index recommended for the query pattern through query pattern test data. When the tested query response speed is lower than the preset speed, the reinforcement learning of the initial index set of the target is strengthened to increase the number of feasible new indexes.
[0082] It should be noted that the system embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the system embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative efforts. The above schematic diagram is only an example of a DQN index recommendation system based on multi-objective reinforcement learning, and does not constitute a limitation on a DQN index recommendation system based on multi-objective reinforcement learning. It may include more or fewer components than shown in the figure, or combine some components, or different components.
[0083] Based on the above method item embodiments, this application correspondingly provides terminal item embodiments.
[0084] Another embodiment of this application provides a terminal, including a processor, a memory, and a computer program stored in the above memory and configured to be executed by the above processor. When the above processor executes the above computer program, it implements the method for DQN index recommendation based on multi-objective reinforcement learning described in any one of the embodiments of this application.
[0085] Exemplarily, in this embodiment, the above computer program can be divided into one or more modules. The above one or more modules are stored in the above memory and executed by the above processor to complete the present application. The above one or more module elements can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the above computer program in the above device;
[0086] The above terminal can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The above device may include, but is not limited to, a processor and a memory.
[0087] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The above processor is the control center of the above device, and connects various parts of the entire device through various interfaces and lines.
[0088] The above memory can be used to store the above computer program and / or module. The above processor realizes various functions of the above device by running or executing the computer program and / or module stored in the above memory, and by calling the data stored in the memory. The above memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; in addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0089] Based on the above method item embodiment, the present application correspondingly provides a storage medium item embodiment.
[0090] Another embodiment of the present application provides a storage medium, which includes a stored computer program. When the computer program runs, it controls the device where the storage medium is located to execute a DQN index recommendation method based on multi-objective reinforcement learning described in any embodiment of the present application.
[0091] In this embodiment, the storage medium is a computer-readable storage medium. The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0092] In the above embodiment of the present application, the internal and external networks of the enterprise are integrated, so that the internal staff of the enterprise participating in the enterprise internal network project can obtain the external network information related to the enterprise internal network project only by logging in to the enterprise internal network, and can very conveniently obtain the relevant information of the enterprise internal network project; by setting up an external network information access account for the internal staff of the enterprise participating in the enterprise internal network project to access the external network information related to the project, and setting different external network information access permissions for different external network information access accounts, the internal staff of the enterprise participating in the enterprise internal network project can conveniently and accurately obtain the external network information related to the enterprise internal network project they participate in.
[0093] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.
Claims
1. A DQN index recommendation method based on multi-objective reinforcement learning, characterized in that, Including the following steps: Configure an initial index set for each target among multiple targets. Through reinforcement learning on the initial index set of the target, obtain a feasible new index, and add the new index to the initial index set to generate the index set of the target; Obtain the query pattern input by the user. Through DQN, obtain the selection value when the query pattern selects each index in the index set, and select the index corresponding to the maximum value among the selection values of all indexes in the index set of each target as the optimal index of the target; Obtain the optimal indexes of all the targets among the multiple targets, and obtain the optimal index with the largest selection value among the optimal indexes of all the targets as the optimal index of the query pattern; Recommend the optimal index of the query pattern to the user who inputs the query pattern.
2. The DQN index recommendation method based on multi-objective reinforcement learning according to claim 1, characterized in that, It further includes the following steps: Based on the query pattern data input by the user and the optimal index data of the query pattern, construct an optimal index recommendation model, and recommend the optimal index to the user through the optimal index recommendation model.
3. A DQN index recommendation method based on multi-objective reinforcement learning according to claim 1, characterized in that It further includes the following steps: When the queried data is updated, establish an association relationship between the updated data and the corresponding index before the data update, and query the updated data through the corresponding index before the data update.
4. The DQN index recommendation method based on multi-objective reinforcement learning according to claim 1, wherein It further includes the following steps: Sort the selection values of the indexes selected for the query pattern in descending order. When the storage space size required for the current index selected for the query pattern exceeds the preset threshold, change the current index to an index with a selection value smaller than the selection value of the current index, and recommend the index with a selection value smaller than the selection value of the current index to the user who inputs the query pattern.
5. The DQN index recommendation method based on multi-objective reinforcement learning according to claim 4, characterized in that, It further includes the following steps: Through the query pattern test data, test the query response speed of the optimal index recommended for the query pattern. When the tested query response speed is lower than the preset speed, strengthen the reinforcement learning of the initial index set of the target and increase the number of feasible new indexes.
6. The DQN index recommendation method based on multi-objective reinforcement learning according to claim 1, wherein In the step of obtaining the query pattern input by the user, obtaining the selection value when the query pattern selects each index in the index set through DQN, and selecting the index corresponding to the maximum value among the selection values of all indexes in the index set of each target as the optimal index of the target, the method for obtaining the selection value includes the following expression: In the formula, C represents the selection value, R represents the expected query response speed, V represents the query speed for querying through the index selected for the query pattern, and S represents the storage space consumed when querying using the selected index.
7. A DQN index recommendation system based on multi-objective reinforcement learning, characterized in that, Including: An index set generation module, configured to configure an initial index set for each target among multiple targets, obtain a feasible new index through reinforcement learning on the initial index set of the target, and add the new index to the initial index set to generate the index set of the target; The target optimal index selection module is used to obtain the user's query pattern, obtain the selection values when each index in the index set is selected for the query pattern through DQN, and select the index corresponding to the maximum value of the selection values of all indexes in the index set of each target as the optimal index of the target; The query pattern optimal index acquisition module is used to obtain the optimal indexes of all the targets in the multi-targets, and obtain the optimal index with the largest selection value among all the optimal indexes as the optimal index of the query pattern; The index recommendation module is used to recommend the optimal index of the query pattern to the user.
8. A DQN index recommendation system based on multi-objective reinforcement learning according to claim 7, characterized in that, It further includes: The optimal index recommendation model construction module is used to construct an optimal index recommendation model based on the query pattern data input by the user and the optimal index data of the query pattern, and recommend the optimal index for the user through the optimal index recommendation model.
9. A DQN index recommendation system based on multi-objective reinforcement learning according to claim 7, characterized in that, It further includes: The current index change module sorts the selection values of the indexes selected for the query pattern in descending order. When the storage space size required for the current index selected for the query pattern exceeds the preset threshold, the current index is changed to an index with a selection value smaller than the selection value of the current index, and the index with a selection value smaller than the selection value of the current index is recommended to the user who inputs the query pattern.
10. A DQN index recommendation system based on multi-objective reinforcement learning according to claim 7, characterized in that, It further includes: The query response speed test module is used to test the query response speed of the optimal index recommended for the query pattern through the query pattern test data. When the tested query response speed is lower than the preset speed, the reinforcement learning of the initial index set of the target is strengthened to increase the number of feasible new indexes.