Query Recommendation Method, Apparatus, Electronic Device, Medium and Program Product
Through the query recommendation method based on the reinforcement learning model, the optimization method of predicting parameters including attention value is solved, and the problem of difficult to capture user behavior and query relationships in the prior art is achieved, and a simple, intelligent and accurate query recommendation effect is achieved.
Patent Information
- Application Number
- CN202210844752.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-18
AI Technical Summary
In the prior art, query recommendation methods are difficult to effectively capture the potential relationship between user behavior and query, and the deep neural network-based method is poor in context when processing sparse data.
The query recommendation method based on the reinforcement learning model is adopted. By obtaining the query data and inputting it into the reinforcement learning model, the optimization method of predicting parameters including attention value is used to calculate the query recommendation result.
Simple, intelligent and accurate query recommendations are implemented, avoiding the dependency between users and queries manually, and no need to perceive the context of query and recommendations.
Smart Images

Figure CN115098783B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly, to a query recommendation method, apparatus, electronic device, computer-readable storage medium, and computer program product based on a reinforcement learning model. Background Art
[0002] Search engines capture users' query intentions through multiple historical information in user query logs. Query recommendation is a very important function of search engines, which can provide users with a series of query recommendations to help users continue the search session. Banks are no exception. For example, when a user searches for financial products, such as funds, a good query recommendation system can reduce the cognitive difficulties during the user's query process and bring a high-quality experience to the user. Commonly used query recommendation methods in the prior art include methods based on statistical learning and methods based on deep neural networks.
[0003] The methods based on statistical learning mainly rely on corresponding features, such as the number of user clicks, stay time, query frequency, and query time, etc., and use sorting algorithms to give query recommendation results. Commonly used methods based on deep neural networks mainly process user query logs through recurrent neural networks (RNNs), and then recommend content that users are interested in. Summary of the Invention
[0004] In view of this, the present disclosure provides a simple, intelligent, and accurate query recommendation method, apparatus, electronic device, computer-readable storage medium, and computer program product based on a reinforcement learning model.
[0005] One aspect of the present disclosure provides a query recommendation method based on a reinforcement learning model, including: in response to a query request, obtaining query data; and using the query data as an input of the reinforcement learning model, and obtaining a query recommendation result according to the prediction parameters of the reinforcement learning model, where the prediction parameters include attention values, and the prediction parameters are obtained by performing a parameter optimization method.
[0006] According to the query recommendation method based on the reinforcement learning model of the embodiments of the present disclosure, a query recommendation result is calculated in the reinforcement learning model through prediction parameters that can reflect the user's attention value. Therefore, the present disclosure does not require artificial construction of the dependency relationship between users and queries, nor does it require perception of the context of queries and recommendations, making the recommended queries simple, intelligent, and accurate.
[0007] In some embodiments, the parameter optimization method includes: obtaining historical query data; determining a reward value of the reinforcement learning model according to the historical query data and training parameters; and optimizing the training parameters according to the reward value to obtain prediction parameters.
[0008] In some embodiments, the historical query data includes first historical query data and second historical query data. Determining the reward value of the reinforcement learning model according to the historical query data and training parameters includes: determining a historical prediction result according to the first historical query data and training parameters; determining a historical true result according to the second historical query data; and calculating the cosine similarity between the historical prediction result and the historical true result to obtain the reward value.
[0009] In some embodiments, the first historical query data includes first text information and first time information, and the training parameters include a training attention value, a training weight, and a training bias. Determining a historical prediction result according to the first historical query data and training parameters includes: determining a historical query encoding according to the first text information and first time information; determining a historical attention encoding according to the first text information and the training attention value; determining a historical recommendation encoding according to the historical query encoding and the historical attention encoding; and determining a historical prediction result according to the historical recommendation encoding, the training weight, and the training bias.
[0010] In some embodiments, the second historical query data includes second text information and second time information. Determining a historical true result according to the second historical query data includes: determining a second text vector according to the second text information; determining a second time vector according to the second time information; and concatenating the second text vector and the second time vector to obtain the historical true result.
[0011] In some embodiments, optimizing the training parameters according to the reward value to obtain prediction parameters includes: creating an objective function, where the objective function includes an expected reward value, the reward value, and the training parameters; comparing the value of the objective function with an optimization threshold; and optimizing the training parameters according to the value of the objective function that satisfies the optimization threshold.
[0012] In some embodiments, the query data includes text information and time information, and the prediction parameters further include a weight and a bias. Obtaining a query recommendation result according to the prediction parameters of the reinforcement learning model includes: determining a query encoding according to the text information and time information; determining an attention encoding according to the text information and the attention value; determining a recommendation encoding according to the query encoding and the attention encoding; and determining a query recommendation result according to the recommendation encoding, the weight, and the bias.
[0013] Another aspect of the present disclosure provides a query recommendation device based on a reinforcement learning model, including: an acquisition module configured to acquire query data in response to a query request; and a determination module configured to use the query data as an input to the reinforcement learning model and obtain a query recommendation result according to prediction parameters of the reinforcement learning model, where the prediction parameters include attention values and are obtained by performing a parameter optimization method.
[0014] Another aspect of the present disclosure provides an electronic device, including one or more processors and one or more memories, where the memory is configured to store executable instructions that, when executed by the processor, implement the method as described above.
[0015] Another aspect of the present disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the method as described above.
[0016] Another aspect of the present disclosure provides a computer program product including a computer program, where the computer program includes computer-executable instructions that, when executed, are used to implement the method as described above. Description of the Drawings
[0017] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0018] Figure 1 Schematically shows an exemplary system architecture to which the method and device according to the embodiments of the present disclosure can be applied;
[0019] Figure 2 Schematically shows a flowchart of a query recommendation method based on a reinforcement learning model according to an embodiment of the present disclosure;
[0020] Figure 3 Schematically shows a flowchart of obtaining a query recommendation result according to prediction parameters of a reinforcement learning model according to an embodiment of the present disclosure;
[0021] Figure 4 Schematically shows a flowchart of a parameter optimization method according to an embodiment of the present disclosure;
[0022] Figure 5 Schematically shows a flowchart of determining a reward value of a reinforcement learning model according to historical query data and training parameters according to an embodiment of the present disclosure;
[0023] Figure 6 Schematically shows a flowchart of determining a historical prediction result according to first historical query data and training parameters according to an embodiment of the present disclosure;
[0024] Figure 7 Schematically shows a flowchart for determining a historical true result according to second historical query data according to an embodiment of the present disclosure;
[0025] Figure 8 Schematically shows a flowchart for optimizing training parameters according to a reward value to obtain prediction parameters according to an embodiment of the present disclosure;
[0026] Figure 9 Schematically shows a structural block diagram of a query recommendation device based on a reinforcement learning model according to an embodiment of the present disclosure;
[0027] Figure 10 Schematically shows a structural block diagram of a determination module according to an embodiment of the present disclosure;
[0028] Figure 11 Schematically shows a structural block diagram of a parameter optimization module according to an embodiment of the present disclosure;
[0029] Figure 12 Schematically shows a structural block diagram of a fifth determination unit according to an embodiment of the present disclosure;
[0030] Figure 13 Schematically shows a structural block diagram of a first determination element according to an embodiment of the present disclosure;
[0031] Figure 14 Schematically shows a structural block diagram of a second determination element according to an embodiment of the present disclosure;
[0032] Figure 15 Schematically shows a structural block diagram of a sixth determination unit according to an embodiment of the present disclosure;
[0033] Figure 16 Schematically shows a block diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0034] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0035] In the technical solutions of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated. In the technical solutions of the present disclosure, the processing of data such as acquisition, collection, storage, use, processing, transmission, provision, disclosure, and application all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.
[0036] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0037] In the case of using expressions such as "at least one of A, B, or C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, or C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). The terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features.
[0038] The search engine captures the user's query intention through multiple historical information in the user query log. Query recommendation is a very important function of the search engine, which can provide a series of query recommendations for the user to help the user continue the search session. Banks are no exception. For example, when a user searches for wealth management products, such as funds, a good query recommendation system can reduce the cognitive difficulties in the user's query process and bring a high-quality experience to the user. The commonly used query recommendation methods in the prior art include methods based on statistical learning and methods based on deep neural networks.
[0039] The method based on statistical learning mainly relies on corresponding features, such as the number of user clicks, stay time, query frequency, and query time, etc., and uses a sorting algorithm to give the query recommendation result. In the query recommendation method based on statistical learning, since the dependence relationship between the user and the query is artificially constructed, it is difficult to capture the potential relationship between the user behavior and the query. The commonly used method based on deep neural networks mainly processes the user query log through a recurrent neural network (RNN) to recommend the content that the user is interested in. However, the data in the query log is often relatively sparse, and the query recommendation context awareness predicted by the RNN is poor, and the obtained result is not ideal. At the same time, the time information in the user query log is also ignored.
[0040] Embodiments of the present disclosure provide a query recommendation method, apparatus, electronic device, computer-readable storage medium, and computer program product based on a reinforcement learning model. The query recommendation method based on the reinforcement learning model includes: in response to a query request, obtaining query data; and using the query data as an input to the reinforcement learning model, and obtaining a query recommendation result according to the prediction parameters of the reinforcement learning model, where the prediction parameters include attention values, and the prediction parameters are obtained by performing a parameter optimization method.
[0041] It should be noted that the query recommendation method, apparatus, electronic device, computer-readable storage medium, and computer program product based on the reinforcement learning model of the present disclosure can be used in the field of artificial intelligence technology, and can also be used in any field other than the artificial intelligence technology field, such as the financial field. The field of the present disclosure is not limited here.
[0042] Figure 1 Exemplary system architecture 100 to which the query recommendation method, apparatus, electronic device, computer-readable storage medium, and computer program product based on the reinforcement learning model according to embodiments of the present disclosure can be applied is schematically shown. It should be noted that Figure 1 The shown is only an example of the system architecture to which embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.
[0043] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0044] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0045] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0046] Server 105 may be a server that provides various services. For example, it can be a back-end management server (only for example) that supports websites browsed by users using terminal devices 101, 102, and 103. The back-end management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0047] It should be noted that the query recommendation method based on the reinforcement learning model provided by the embodiments of the present disclosure can generally be executed by server 105. Correspondingly, the query recommendation device based on the reinforcement learning model provided by the embodiments of the present disclosure can generally be set in server 105. The query recommendation method based on the reinforcement learning model provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the query recommendation device based on the reinforcement learning model provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.
[0048] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0049] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figures 2 to 8 The following will describe in detail the query recommendation method based on the reinforcement learning model of the embodiments of the present disclosure through
[0050] Figure 2 FIG. schematically shows a flowchart of the query recommendation method based on the reinforcement learning model according to the embodiments of the present disclosure.
[0051] As Figure 2 shown, the query recommendation method based on the reinforcement learning model of this embodiment includes operation S210 and operation S220.
[0052] In operation S210, in response to a query request, query data is obtained. Among them, the user can input search content to submit a query request, so the query data can be obtained from the search content input by the user.
[0053] In operation S220, the query data is used as the input of the reinforcement learning model, and according to the prediction parameters of the reinforcement learning model, a query recommendation result is obtained. Among them, the prediction parameters include attention values, and the prediction parameters are obtained by performing the parameter optimization method of operation S300.
[0054] As a possible implementation, the query data includes text information and time information, and the prediction parameters further include weights and biases. For example, Figure 3 As shown, operation S220 obtains a query recommendation result according to the prediction parameters of the reinforcement learning model, including operations S221 to S224.
[0055] In operation S221, a query encoding is determined according to the text information and the time information.
[0056] As a specific example, the text information q i can be vector-converted through formula (1) to obtain a text vector
[0057]
[0058] The time information t i can be vector-converted through formula (2) to obtain a time vector
[0059]
[0060] The text vector and the time vector can be concatenated through formula (3) to obtain a concatenated vector
[0061]
[0062] When there is 1 piece of query data, the concatenated vector converted from 1 piece of query data can be converted into a query encoding through formula (4).
[0063]
[0064] When there are m pieces of query data, where m is an integer greater than 1, the m concatenated vectors respectively converted from m pieces of query data can be converted into a query encoding through formula (5).
[0065]
[0066] In operation S222, an attention encoding is determined according to the text information and the attention value. The attention value is the attention value α obtained according to the execution parameter optimization method i , and the attention value reflects the user's attention to the key information in the text information. The key information can be understood as all or part of the word segments extracted after word segmentation of the text information. In some specific examples, when there are m pieces of query data, where m is an integer greater than or equal to 1, the attention encoding can be obtained through formula (6)
[0067]
[0068] In operation S223, a recommendation code is determined based on the query code and the attention code. In some specific examples, the recommendation code S can be calculated by formula (7).
[0069]
[0070] Where β is a set constant.
[0071] In operation S224, a query recommendation result is determined based on the recommendation code, the weight, and the bias. The weight and the bias are the weight W and the bias b obtained by the execution parameter optimization method. In some specific examples, the query recommendation result Q can be calculated by formula (8).
[0072]
[0073] By operations S221 to S224, it is convenient to obtain the query recommendation result according to the prediction parameters of the reinforcement learning model.
[0074] According to the query recommendation method based on the reinforcement learning model of the present disclosure, the query recommendation result is calculated in the reinforcement learning model through the prediction parameters that can reflect the user's attention value. Therefore, the present disclosure does not require artificially constructing the dependency relationship between the user and the query, nor does it require perceiving the context of the query and the recommendation, making the recommended query simple, intelligent, and accurate.
[0075] Figure 4 A flowchart of the parameter optimization method according to an embodiment of the present disclosure is schematically shown.
[0076] The parameter optimization method of operation S300 includes operations S310 to S330.
[0077] In operation S310, historical query data is obtained.
[0078] In operation S320, the reward value of the reinforcement learning model is determined according to the historical query data and the training parameters.
[0079] As an implementable manner, as Figure 5 shown, the historical query data includes first historical query data and second historical query data. Operation S320 determines the reward value of the reinforcement learning model according to the historical query data and the training parameters, including operations S321 to S323.
[0080] In operation S321, the historical prediction result is determined according to the first historical query data and the training parameters.
[0081] In some specific examples, the first historical query data includes first text information and first time information, and the training parameters include training attention values, training weights, and training bias amounts. As shown in Figure 6 Figure, operation S321 determines a historical prediction result according to the first historical query data and the training parameters, including operations S3211 to S3214.
[0082] In operation S3211, a historical query encoding is determined according to the first text information and the first time information.
[0083] As a specific example, the first text information q can be vector-converted through formula (9) i历史 to obtain a first text vector
[0084]
[0085] The first time information t can be vector-converted through formula (10) i历史 to obtain a first time vector
[0086]
[0087] The first text vector and the first time vector can be concatenated through formula (11) to obtain a concatenated vector
[0088]
[0089] When there is 1 piece of first historical query data, the concatenated vector converted from 1 piece of first historical query data can be converted into a historical query encoding through formula (12).
[0090]
[0091] When there are m pieces of first historical query data, where m is an integer greater than 1, the m concatenated vectors respectively converted from m pieces of first historical query data can be converted into a historical query encoding through formula (13).
[0092]
[0093] In operation S3212, a historical attention encoding is determined according to the first text information and the training attention value, where the training attention value is the attention value α that is continuously optimized and updated according to the execution parameter optimization method i训练 , and the training attention value needs to be continuously optimized and updated according to the user's attention to the key information in the text information. In some specific examples, when there are m pieces of first historical query data, where m is an integer greater than or equal to 1, the attention encoding can be obtained through formula (14)
[0094]
[0095] In operation S3213, based on the historical query encoding and the historical attention encoding, determine the historical recommendation encoding. In some specific examples, the historical recommendation encoding can be calculated through formula (15).
[0096]
[0097] Among them, β is a set constant.
[0098] In operation S3214, based on the historical recommendation encoding, the training weight, and the training bias, determine the historical prediction result. Among them, the training weight and the training bias are the continuously optimized and updated weight W 训练 and bias b 训练 , in some specific examples, the historical prediction result Q can be calculated through formula (16). 历史 .
[0099]
[0100] Through operations S3211 to S3214, it is convenient to determine the historical prediction result according to the first historical query data and the training parameters.
[0101] In operation S322, based on the second historical query data, determine the historical true result.
[0102] As a possible implementation method, the second historical query data includes the second text information and the second time information. It should be noted that the second historical query data is the true data corresponding to the historical prediction result obtained through operations S3211 to S3214. For example, if the historical prediction result is the query recommendation result predicted for the 7th day based on the historical data of the previous 6 days, then the second historical query data is the true query data for the 7th day. As Figure 7 shown, operation S322 determines the historical true result according to the second historical query data, including operations S3221 to S3223.
[0103] In operation S3221, determine the second text vector according to the second text information.
[0104] In operation S3222, determine the second time vector according to the second time information.
[0105] In operation S3223, splice the second text vector and the second time vector to obtain the historical true result.
[0106] As a specific example, the second text information q can be j真实 vector-converted to obtain a second text vector
[0107]
[0108] The second time information t can be i真实 vector-converted to obtain a second time vector
[0109]
[0110] By concatenating the second text vector and the second time vector through formula (19), a concatenated vector can be obtained
[0111]
[0112] Through operations S3221 to S3223, it is convenient to implement determining the historical true result according to the second historical query data.
[0113] In operation S323, calculate the cosine similarity between the historical prediction result and the historical true result to obtain a reward value. For example, the reward value R can be calculated through formula (20).
[0114]
[0115] Through operations S321 to S323, it is convenient to implement determining the reward value of the reinforcement learning model according to the historical query data and training parameters.
[0116] In operation S330, optimize the training parameters according to the reward value to obtain prediction parameters.
[0117] As an implementable manner, as Figure 8 shown, operation S330 optimizes the training parameters according to the reward value to obtain prediction parameters, including operations S331 to S333.
[0118] In operation S331, create an objective function, which includes an expected reward value, a reward value, and training parameters. For example, the objective function J(α i训练 , W 训练 , b 训练 ) can be as shown in formula (21).
[0119] J(α i训练 , W 训练 , b 训练 ) = E[R|(α i训练 , W 训练 , b训练 )] (21)
[0120] Among them, E(X) represents the expected reward value, where X is R|(α i训练 , W 训练 , b 训练 ), so the value of the objective function can be obtained.
[0121] In operation S332, compare the value of the objective function with the optimization threshold.
[0122] In operation S333, according to the value of the objective function that meets the optimization threshold, optimize the training parameters. For example, when the value of the objective function is greater than or equal to the optimization threshold, α i训练 , W 训练 and b 训练 in the objective function are used as the prediction parameters of the reinforcement learning model. Thus, through operations S331 to S333, it is convenient to optimize the training parameters according to the reward value and obtain the prediction parameters.
[0123] Based on the above query recommendation method based on the reinforcement learning model, the present disclosure also provides a query recommendation device 10 based on the reinforcement learning model. The following will be combined with Figures 9 - 15 to describe the query recommendation device 10 based on the reinforcement learning model in detail.
[0124] Figure 9 Schematically shows a structural block diagram of a query recommendation device 10 based on an embodiment of the present disclosure.
[0125] The query recommendation device 10 based on the reinforcement learning model includes an acquisition module 1, a determination module 2, and a parameter optimization module 3.
[0126] The acquisition module 1 is used to execute operation S210: in response to a query request, acquire query data.
[0127] The determination module 2 is used to execute operation S220: take the query data as the input of the reinforcement learning model, and obtain a query recommendation result according to the prediction parameters of the reinforcement learning model, where the prediction parameters include attention values, and the prediction parameters are obtained by the parameter optimization module 3 executing the parameter optimization method of operation S300.
[0128] Figure 10 Schematically shows a structural block diagram of the determination module 2 according to an embodiment of the present disclosure. The query data includes text information and time information, the prediction parameters further include weights and bias amounts, and the determination module 2 includes a first determination unit 21, a second determination unit 22, a third determination unit 23, and a fourth determination unit 24.
[0129] The first determination unit 21 determines a query code according to the text information and the time information.
[0130] The second determination unit 22 determines an attention code according to the text information and the attention value.
[0131] The third determination unit 23 determines a recommendation code according to the query code and the attention code.
[0132] The fourth determination unit 24 determines a query recommendation result according to the recommendation code, the weight, and the bias.
[0133] Figure 11 The structural block diagram of the parameter optimization module 3 according to an embodiment of the present disclosure is schematically shown. The parameter optimization module 3 includes an acquisition unit 31, a fifth determination unit 32, and a sixth determination unit 33.
[0134] The acquisition unit 31 acquires historical query data.
[0135] The fifth determination unit 32 determines the reward value of the reinforcement learning model according to the historical query data and the training parameters.
[0136] The sixth determination unit 33 optimizes the training parameters according to the reward value to obtain prediction parameters.
[0137] Figure 12 The structural block diagram of the fifth determination unit 32 according to an embodiment of the present disclosure is schematically shown. The historical query data includes first historical query data and second historical query data, and the fifth determination unit 32 includes a first determination element 321, a second determination element 322, and a third determination element 323.
[0138] The first determination element 321 determines a historical prediction result according to the first historical query data and the training parameters.
[0139] The second determination element 322 determines a historical true result according to the second historical query data.
[0140] The third determination element 323 calculates the cosine similarity between the historical prediction result and the historical true result to obtain the reward value.
[0141] Figure 13A structural block diagram of a first determination element 321 according to an embodiment of the present disclosure is schematically shown. The first historical query data includes first text information and first time information, the training parameters include training attention values, training weights, and training bias amounts, and the first determination element 321 includes a first determination member 3211, a second determination member 3212, a third determination member 3213, and a fourth determination member 3214.
[0142] The first determination member 3211, the first determination member 3211 determines a historical query encoding according to the first text information and the first time information.
[0143] The second determination member 3212, the second determination member 3212 determines a historical attention encoding according to the first text information and the training attention value.
[0144] The third determination member 3213, the third determination member 3213 determines a historical recommendation encoding according to the historical query encoding and the historical attention encoding.
[0145] The fourth determination member 3214, the fourth determination member 3214 determines a historical prediction result according to the historical recommendation encoding, the training weight, and the training bias amount.
[0146] Figure 14 A structural block diagram of a second determination element 322 according to an embodiment of the present disclosure is schematically shown. The second historical query data includes second text information and second time information, and the second determination element 322 includes a fifth determination member 3221, a sixth determination member 3222, and a splicing member 3223.
[0147] The fifth determination member 3221, the fifth determination member 3221 determines a second text vector according to the second text information.
[0148] The sixth determination member 3222, the sixth determination member 3222 determines a second time vector according to the second time information.
[0149] The splicing member 3223, the splicing member 3223 splices the second text vector and the second time vector to obtain a historical true result.
[0150] Figure 15 A structural block diagram of a sixth determination unit 33 according to an embodiment of the present disclosure is schematically shown. The sixth determination unit 33 includes a creation element 331, a comparison element 332, and an optimization element 333.
[0151] The creation element 331, the creation element 331 creates an objective function, and the objective function includes an expected reward value, a reward value, and training parameters.
[0152] The comparison element 332, the comparison element 332 compares the value of the objective function with an optimization threshold.
[0153] Optimization component 333, which optimizes training parameters according to the value of the objective function that meets the optimization threshold.
[0154] According to the query recommendation device 10 based on the reinforcement learning model of the embodiments of the present disclosure, a query recommendation result is calculated in the reinforcement learning model through prediction parameters that can reflect the user attention value. Therefore, the present disclosure does not require artificially constructing the dependency relationship between the user and the query, nor does it require perceiving the context of the query and the recommendation, making the recommended query simple, intelligent, and accurate.
[0155] In addition, according to the embodiments of the present disclosure, any multiple of the acquisition module 1, the determination module 2, and the parameter optimization module 3 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module.
[0156] According to the embodiments of the present disclosure, at least one of the acquisition module 1, the determination module 2, and the parameter optimization module 3 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them.
[0157] Or, at least one of the acquisition module 1, the determination module 2, and the parameter optimization module 3 can be at least partially implemented as a computer program module, which can execute the corresponding functions when the computer program module is run.
[0158] Figure 16 A block diagram of an electronic device suitable for implementing the above method according to the embodiments of the present disclosure is schematically shown.
[0159] As Figure 16As shown, an electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. The processor 901 can include, for example, a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 901 can also include on-board memory for caching purposes. The processor 901 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0160] In the RAM 903, various programs and data required for the operation of the electronic device 900 are stored. The processor 901, ROM 902, and RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 902 and / or RAM 903. It should be noted that the program can also be stored in one or more memories other than the ROM 902 and RAM 903. The processor 901 can also perform various operations of the method flow according to an embodiment of the present disclosure by executing the programs stored in the one or more memories.
[0161] According to an embodiment of the present disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, and the input / output (I / O) interface 905 is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage section 908 as needed.
[0162] The present disclosure also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiment; or may exist separately without being assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to an embodiment of the present disclosure is implemented.
[0163] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903.
[0164] An embodiment of the present disclosure further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the method of the embodiment of the present disclosure.
[0165] When the computer program is executed by the processor 901, it executes the above-described functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described system, apparatus, module, unit, etc. may be implemented by computer program modules.
[0166] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium and downloaded and installed through the communication part 909, and / or installed from the removable medium 911. The program code included in the computer program may be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0167] In such an embodiment, the computer program may be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, it executes the above-described functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. may be implemented by computer program modules.
[0168] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0170] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0171] The embodiments of the present disclosure have been described above. However, these embodiments are merely for illustrative purposes and not for limiting the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present disclosure.
Claims
1. A query recommendation method based on a reinforcement learning model, characterized in that including: in response to a query request, obtaining query data, where the query data includes text information and time information; and using the query data as an input to a reinforcement learning model, and obtaining a query recommendation result according to prediction parameters of the reinforcement learning model, where the prediction parameters include attention values, and the prediction parameters are obtained by performing a parameter optimization method; wherein the prediction parameters further include weights and bias amounts, and obtaining the query recommendation result according to the prediction parameters of the reinforcement learning model includes: determining a query encoding according to the text information and the time information; determining an attention encoding according to the text information and the attention values; determining a recommendation encoding according to the query encoding and the attention encoding; and determining the query recommendation result according to the recommendation encoding, the weights, and the bias amounts.
2. The method according to claim 1, wherein The parameter optimization method includes: obtaining historical query data; determining a reward value of the reinforcement learning model according to the historical query data and training parameters; and optimizing the training parameters according to the reward value to obtain prediction parameters.
3. The method according to claim 2, characterized in that, The historical query data includes first historical query data and second historical query data, and determining the reward value of the reinforcement learning model according to the historical query data and training parameters includes: determining a historical prediction result according to the first historical query data and training parameters; determining a historical true result according to the second historical query data; and calculating a cosine similarity between the historical prediction result and the historical true result to obtain the reward value.
4. The method according to claim 3, wherein The first historical query data includes first text information and first time information, the training parameters include training attention values, training weights, and training bias amounts, and determining the historical prediction result according to the first historical query data and training parameters includes: determining a historical query encoding according to the first text information and the first time information; determining a historical attention encoding according to the first text information and the training attention values; determining a historical recommendation encoding according to the historical query encoding and the historical attention encoding; and determining the historical prediction result according to the historical recommendation encoding, the training weights, and the training bias amounts.
5. The method according to claim 3, characterized in that, The second historical query data includes second text information and second time information, and determining the historical true result according to the second historical query data includes: determining a second text vector according to the second text information; determining a second time vector according to the second time information; and concatenating the second text vector and the second time vector to obtain the historical true result.
6. The method according to claim 2, wherein Optimizing the training parameters according to the reward value to obtain prediction parameters includes: creating an objective function, where the objective function includes an expected reward value, the reward value, and the training parameters; comparing a value of the objective function with an optimization threshold; and optimizing the training parameters according to the value of the objective function that satisfies the optimization threshold.
7. A query recommendation device based on a reinforcement learning model, characterized in that, including: an obtaining module, configured to perform, in response to a query request, obtaining query data, where the query data includes text information and time information; and A determination module, which is configured to use the query data as an input to a reinforcement learning model, and obtain a query recommendation result according to prediction parameters of the reinforcement learning model, where the prediction parameters include attention values, and the prediction parameters are obtained by performing a parameter optimization method; Wherein, the prediction parameters further include weights and bias amounts, and obtaining a query recommendation result according to the prediction parameters of the reinforcement learning model includes: Determining a query encoding according to the text information and time information; Determining an attention encoding according to the text information and the attention values; Determining a recommendation encoding according to the query encoding and the attention encoding; and Determining a query recommendation result according to the recommendation encoding, the weights, and the bias amounts.
8. An electronic device, characterized in that, Comprising: One or more processors; One or more memories for storing executable instructions, which, when executed by the processor, implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, An executable instruction is stored on the storage medium, and when the instruction is executed by the processor, it implements the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, Comprising a computer program, the computer program comprising one or more executable instructions, which, when executed by the processor, implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Rolling optimization-based urban-level global traffic signal recommendation method and system
CN110533932A
Query recommendation method and system based on improved VHRED and reinforcement learning
CN111274359A