Data Processing Method and Apparatus, Electronic Device, and Storage Medium
Through MCTS, the search tree is constructed and the optimization sorting algorithm is integrated into the value and cost estimate model, which solves the problem that the sorting algorithm in the existing technology is affected by push demand, and the expected effect of the click-through rate and conversion rate of recommended products is achieved, and the accuracy and fairness of recommendations are improved.
Patent Information
- Application Number
- CN202111355773.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-11-16
AI Technical Summary
When facing new push demands, existing sorting algorithms can easily affect the exposure sorting of products, resulting in the clicks and conversions of recommended products not reaching the expected results, and even leading to a decline in recommendation fairness and overall click-through rate.
MCTS (Monte Carlo Tree Search) is used to explore recalled products, integrate strategies such as product value, repeated recommendation punishment and popular punishment, build a search tree, comprehensively considering click-through rate, conversion rate and user step size, and optimize the sorting results through the value estimate model and the cost estimate model.
Ensure that the click-through rate and conversion rate of the target data pushed to the target object can achieve the expected results, improve the accuracy and fairness of recommended products, and avoid the interference of new push requirements on sorting.
Smart Images

Figure CN114036388B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to the fields of artificial intelligence, reinforcement learning, and intelligent recommendation. Specifically, a data processing method, an apparatus, an electronic device, and a storage medium are provided. Background Art
[0002] In an intelligent recommendation scenario, it is often necessary to process recalled data through a sorting algorithm. However, when facing new push requirements, the current commonly used sorting algorithms will directly affect the exposure sorting of products, resulting in the clicks and conversions of the recommended products not meeting the expected effects, and even leading to an overall decline in recommendation fairness and click-through rate. Summary of the Invention
[0003] The present disclosure provides a data processing method, an apparatus, an electronic device, and a storage medium.
[0004] According to a first aspect of the present disclosure, a data processing method is provided, including: obtaining a recalled data set; constructing a search tree corresponding to the recalled data set, where the search tree includes: a root node and a plurality of data nodes at different levels, each data node is used to represent the recalled data in the recalled data set, and each data node is used to store the push value and the search times of the corresponding recalled data, and the push value is used to represent the value of the feedback result received after the corresponding recalled data is pushed to a target object; determining target data in the recalled data set based on the push value of each recalled data in the recalled data set, where the target data is the data pushed to the target object.
[0005] According to a second aspect of the present disclosure, a data processing apparatus is provided, including: an obtaining module, configured to obtain a recalled data set; a constructing module, configured to construct a search tree corresponding to the recalled data set, where the search tree includes: a root node and a plurality of data nodes at different levels, each data node is used to represent the recalled data in the recalled data set, and each data node is used to store the push value and the search times of the corresponding recalled data, and the push value is used to represent the value of the feedback result received after the corresponding recalled data is pushed to a target object; a decision module, configured to determine target data in the recalled data set based on the push value of each recalled data in the recalled data set, where the target data is the data pushed to the target object.
[0006] According to a third aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.
[0007] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method described above.
[0008] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method described above.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0011] Figure 1 is a flowchart of a product recommendation process in the related art;
[0012] Figure 2 is a flowchart of a product recommendation process according to the present disclosure;
[0013] Figure 3 is a flowchart of a data processing method according to the present disclosure;
[0014] Figure 4 is a schematic diagram of an optional MDP tree according to the present disclosure;
[0015] Figure 5 is a schematic diagram of an optional user step length according to the present disclosure;
[0016] Figure 6 is a product recommendation scenario diagram that can implement the embodiments of the present disclosure;
[0017] Figure 7 is a schematic diagram of a data processing apparatus according to the present disclosure;
[0018] Figure 8 is a block diagram of an electronic device for implementing the data processing method of the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0020] AsFigure 1 As shown in the figure, the current product recommendation process mainly includes: using different recall algorithms (including but not limited to: collaborative filtering, vectorized recall, category recall, label recall, new product recall and popularity recall, etc.) to recall the candidate products in the candidate product library to obtain a set of recalled products; then estimating the click rate and conversion rate of the recalled products through the click rate prediction model and the conversion rate prediction model, and sorting the recalled products according to the click rate and conversion rate to obtain sorted recalled products; further screening the sorted recalled products according to supplementary strategies such as push intervention rules, diversity strategies and repeated recommendation strategies to obtain a final list of recommended products.
[0021] At present, the relevant technology provides a variety of sorting algorithms to realize the sorting process of recalled products: the first is traditional collaborative filtering, LR (Logistic Regression) + GBDT (Gradient Boosting Decision Tree), FM (Factorization Machines), etc.; the second is the method based on deep learning, such as Wide&Deep, DeepFM, etc.; the third is the deep reinforcement learning model, for example, Policy Gradient, DQN (Deep Q-Learning), Actor-Critic.
[0022] However, when the above sorting algorithm faces new push demands, the new push demands will directly affect the exposure ranking of the products, resulting in the clicks and conversions of the recommended products failing to achieve the expected results, and even leading to an overall decline in the fairness of the recommendations and the click-through rate.
[0023] For recommendation scenarios that take into account the click and conversion value of products, many recommendation methods use static grouping to uniformly assign certain scores to products in different groups at one time or on a regular basis to distinguish their value. However, this method fails to take into account the dynamic characteristics of product value. The value of a product may change in real time with the recommendation strategy and user feedback (such as exposure, click, and conversion). For example, when the daily exposure and click conversion of a product have met the push requirements, other unexposed products or products that can bring higher click conversion values should be recommended to achieve recommendation fairness.
[0024] In order to solve the above problems, the present disclosure improves the sorting layer in the related art, such as Figure 2 As shown in Figure 2, MCTS (Monte Carlo Tree Search) is mainly used to explore recalled products. The exploration process incorporates product value (such as Figure 2The value prediction model in), repeated recommendation penalty (such as Figure 2 The cost prediction model in), popularity penalty and other strategies, and comprehensively consider objectives such as click-through rate, conversion rate, and user step length for recommendation. Further, according to the diversity strategy, the recalled items after sorting are screened to obtain the final recommended item list.
[0025] According to an embodiment of the present disclosure, the present disclosure provides a data processing method.
[0026] Figure 3 Is a flowchart of the data processing method according to an embodiment of the present disclosure, as Figure 3 shown, the method includes the following steps:
[0027] Step S302, obtain a recalled data set.
[0028] The recalled data set in the above step may be a set of recalled data determined by a recall algorithm during the data recommendation process. The recalled data here may be candidate data that matches the user successfully in the candidate data set. In different data recommendation scenarios, the types of recalled data are different. For example, for the Figure 2 shown product recommendation scenario, the recalled data set may be a set of recalled products determined by a recall algorithm.
[0029] Step S304, construct a search tree corresponding to the recalled data set. Among them, the search tree includes: a root node and multiple data nodes at different levels. Each data node is used to represent the recalled data in the recalled data set. Each data node is used to store the push value and search times of the corresponding recalled data. The push value is used to represent the value of the feedback result received after the corresponding recalled data is pushed to the target object.
[0030] The search tree in the above step may be a tree constructed by simulating the push process of the recalled data within a limited time. The search tree may include a root node and multiple data nodes at different levels. The root node represents a request to push data to the user, and each data node represents a recalled data. During the simulation process, node selection can be based on the push value and search times of the nodes, and the estimated click-through rate and conversion rate are used as the state transition probabilities when expanding the nodes. Simulate pushing each recalled data to the target object (i.e., the user), and the feedback result of the target object for the pushed recalled data. Here, the feedback result may be that the user clicks on the pushed recalled data, consults the pushed recalled data, leaves, etc., which can be determined according to the specific data push scenario.
[0031] To ensure that the click-through rate and conversion rate of the target data pushed to the target object can achieve the expected results, a value prediction model can be constructed to predict the rewards corresponding to different feedback results. Among them, for the user's "leave" behavior, the negative of the predicted reward can be used as the cost; for the user's "click" and "consult" behaviors, the click and conversion values can be predicted and used as the rewards. At the same time, to ensure recommendation fairness, certain penalties can be imposed on the recalled data with more exposures, and a higher value can be given to some potential recalled data. It should be noted that this higher value will not directly expose the recalled data, but only push the recalled data in a short period of time. However, if the click-through rate / conversion rate of the recalled data does not increase, the probability of the recalled data being selected will decrease, and the recall data will stop being pushed continuously.
[0032] In addition, considering the impact of the repeated recommendation penalty strategy on sorting, for the recalled data repeatedly pushed within a short period of time, a cost prediction module can be constructed to predict the cost of each repeated push according to the target object characteristics, background information, etc. The background information here can refer to the time interval between the recalled data and the current exhibition, the historical exhibition position, the page, the user's preference or acceptance of the repeated recommended data, etc.
[0033] In some alternative embodiments, MCTS can be used to construct an MDP (Markov Decision Process) tree. Through each simulation, the evaluated state values are stored on the nodes of the tree. These state values can be accumulated through the loop iteration of the four steps of selection, expansion, simulation, and backpropagation. Among them, each data node of the tree will store the following two state values V(s): the push value and the search times. The push value here is jointly determined based on the reward and the cost.
[0034] The expressions of each node of the MDP tree, the state transition process, and the parameters required in the actual operation are as follows:
[0035] State: s t It can be represented by user preferences, user requests, and system status;
[0036] Action: a can be that the system selects a recalled data from the candidate recommendation list and recommends it to the user;
[0037] Transitions: P a(s|s′) can be the state transition probability, and successor states are obtained through user feedback. The probability equation of Transition is generally equivalent to the probability of user behavior, and p is estimated through a probability network (including: click-through rate prediction model, conversion rate prediction model);
[0038] Reward: r(s t ,a,s t+1 ) can be a measurable indicator of the satisfaction of the recommendation accuracy and exposure conversion after the user takes an action. This indicator is given by a value prediction model, which can combine the exposure, click, and conversion of the recall data, and will also consider the recommendation fairness and the value brought by clicks and conversions to estimate the reward;
[0039] Cost: c(s t ,a,s t+1 ) can be the cost of enabling the user to take a certain action, such as repeated recommendations in the short term, etc., and the cost can be estimated according to the time interval of repeated occurrences, page type, location, and the user's acceptance of repeated recommendations;
[0040] Discount rate: γ is used to measure the contribution rate of the long-term reward to the current value. Generally, it is considered that the greater the speculation depth of the behavior, the higher the uncertainty, and the contribution to the current will decrease.
[0041] Step S306, based on the push value of each recall data in the recall data set, determine the target data in the recall data set, where the target data is the data pushed to the target object.
[0042] In this embodiment, the push value of each recall data reflects the value of clicks and conversions of the recall data. In order to achieve the purpose of maximizing the overall value, all recall data can be sorted according to the push value, and multiple recall data with the top rankings can be selected as the target data, and then the target data is pushed to the target user according to a certain push rule. The push rule here can be set according to different application scenarios and push requirements, and the present disclosure does not make specific limitations in this regard.
[0043] Through the above solution, after obtaining the recall data set, a search tree corresponding to the recall data set can be constructed, and then based on the push value of each recall data in the search tree, the target data to be pushed to the target object can be determined, achieving the purpose of sorting the recall data. It is easy to notice that the target data pushed to the target object is determined based on the push value of each recall data, and the push value is used to characterize the value of the feedback result received after the corresponding recall data is pushed to the target object. Therefore, the push value can reflect the click and conversion effects of the recall data, ensuring that the clicks and conversions of the target data pushed to the target object can achieve the expected effects. In addition, the push value of the recall data is determined during the process of constructing the search tree. Therefore, the push values of different recall data can be adjusted according to actual push requirements, realizing the adjustment of the sorting results of all recall data, so that the newly pushed intervention rules will not affect the push values of different recall data, thereby ensuring that the target data pushed to the target object meets the push requirements, and further solving the technical problem that the sorting algorithm in the related art is easily interfered by different push requirements, and the clicks and conversions of the recommended products cannot achieve the expected effects.
[0044] Optionally, constructing a search tree corresponding to the recall data set includes: Step A, determining the target node to be searched, where the target node is used to represent the root node or the recall data in the recall data set; Step B, expanding new child nodes under the target node, and using the value prediction model to determine the rewards of the new child nodes; Step C, simulating the target child nodes in the new child nodes, and determining the push value of the last child node at the end of the simulation; Step D, based on the push value of the last child node, performing backward iteration on the search tree, updating the push values of the target node and each layer of child nodes under the target node, and updating the search times of the target node; repeating Steps A to D until the exploration time of the search tree reaches the preset exploration time, or the exploration depth reaches the preset exploration depth.
[0045] The preset exploration time in the above steps can be the simulation construction time of the search tree, and the preset exploration depth can be the simulation exploration depth of the search tree, which can be set according to actual needs, and the present disclosure does not make specific limitations thereto.
[0046] In the embodiments of the present disclosure, the search tree needs to start exploring from the root node. After the exploration of the root node is completed, the exploration of other data nodes can continue. Therefore, at the beginning of each loop iteration, the selection step (i.e., step A above) is first executed to determine the target node to be explored this time; then the expansion step (i.e., step B above) is executed to expand the target node through the action, and all feedback results generated by expanding this action are expanded. Among them, a new child node can be created for the new feedback result, and at the same time, the value evaluation model can be used to evaluate this feedback result to obtain the reward of this child node; further, the simulation step (i.e., step C above) is executed, and a child node can be randomly selected for simulation until the simulation ends (the simulation end condition can be set as needed, which can be reaching the set simulation time or reaching the set simulation depth). At this time, the push value of the simulation termination state (including reward, cost, etc.) can be given, that is, the push value of the last child node is given; finally, the backpropagation step (i.e., step D above) is executed to backpropagate the push value after simulation upward, and the push value of each layer of child nodes is updated recursively, and at the same time, the search times of the target node are updated.
[0047] In some alternative embodiments, for the Figure 4 MDP tree shown in the figure, during the first loop iteration, the root node can be selected as the target node, and then an action is selected (assuming that Item 1 is selected for pushing), as shown in the solid circle in Figure 4 ; then, the child nodes corresponding to the feedback results (including click, leave, and conversion) after pushing Item 1 can be expanded according to the state transition probability, which respectively correspond to the three ellipses shown in the upper left part of Figure 4 , and the reward corresponding to each child node is evaluated using the value evaluation model. Assume that the rewards corresponding to the three child nodes are 3, 0, and 100 respectively; a child node is randomly selected for simulation. Assume that the child node corresponding to "click" is selected, as shown by the solid line ellipse in Figure 4 , and the unselected child nodes are shown by the dashed line ellipse in Figure 4 . The simulation process can be to select a result after this node according to the state transition probability (assuming that Item k is selected for pushing), and then expand the child nodes after pushing Item k according to the state transition probability, including click, leave, and conversion. At this time, the simulation ends, and the push value of the last child node can be directly obtained according to the reward estimated by the value prediction module; finally, the push value of the last child node is backpropagated according to the simulation depth, so that the push value of each layer of child nodes corresponding to the parent node can be updated recursively, and at the same time, the search times of the target node are incremented by 1. It should be noted that since the child nodes after pushing Item k have not been simulated before, the rewards of the three child nodes are all 0.
[0048] It should be noted that the push value of the upper-level child node can be obtained by taking the maximum value of the total value of the lower-level child nodes (i.e., the formaldehyde sum of the push value and the reward), but it is not limited to this.
[0049] Through the above solution, by repeatedly performing steps such as selection, expansion, simulation, and backpropagation, the push value of each recalled data can be continuously updated in the loop iteration, so as to ensure that the push value not only meets the push requirements but also can truly reflect the click and conversion value of the recalled data, achieving the effect of improving the accuracy of the push value and the accuracy of data push.
[0050] Optionally, determining the target node to be searched includes: starting from the root node and traversing to determine whether there is an unexpanded node; in response to the existence of an unexpanded node, determining the unexpanded node as the target node; in response to the non-existence of an unexpanded node, determining the target node based on the push value and search times of each data node.
[0051] The unexpanded node in the above embodiment may refer to a node that has an action that has not been simulated or a child node that has not been simulated.
[0052] In the embodiments of the present disclosure, the construction process of the search tree is to start expanding the next node after the expansion of one node is completed. Therefore, after each iteration starts, first determine whether there is an unexpanded node. If there is, the node can be directly used as the target node and the subsequent expansion, simulation, and backpropagation steps can be executed; if not, the selection of the target node can be achieved through UCT (Upper Confidence Bound Apply to Tree), and the calculation formula of the target node is as follows:
[0053]
[0054] where Q(s,a) is the push value of state s, N(s) is the number of searches, N(s,a) is the number of times action a is executed in state s, and Q(s,a) encourages "exploitation", encourages "exploration".
[0055] In some alternative embodiments, for Figure 4For the MDP tree shown, during the first loop iteration, the root node can be directly determined as the target node. At this time, both the push value and the search count of the root node are 0. During the second loop iteration, since the root node is not fully expanded, the root node can still be determined as the target node. At this time, the push value and the search count of the root node are no longer 0 and have been updated in the previous loop iteration. During the third loop iteration, since the root node has been fully expanded, traversal can continue to determine the node corresponding to Item 1 as the target node.
[0056] Through the above solution, by determining whether there are unexpanded nodes, a determination result is obtained, and different methods are used to determine the target object for different determination results, thereby achieving the effect of improving the determination efficiency and accuracy of the target node.
[0057] Optionally, expanding new child nodes under the target node includes: obtaining the state transition probability corresponding to the target node; determining the target execution operation corresponding to the target node based on the state transition probability; creating new child nodes under the target node based on the target execution operation.
[0058] The state transition probability in the above steps can be the probability determined based on the click-through rate estimated by the click-through rate estimation model and the conversion rate estimated by the conversion rate estimation model, and is used to judge the possible operations that the user may perform on the recalled data; the target execution operation can be the execution operation corresponding to the maximum probability in the state transition probability.
[0059] In some alternative embodiments, for the Figure 4 For the MDP tree shown, during the first loop iteration, after determining the target node, an unexpanded action can be selected, that is, push Item 1, and then the target execution operations that the user may perform are determined according to the state transition probability, which are click, leave, and consult respectively, and three child nodes are created at the next level of the target node. At this time, both the push value and the search count of the three child nodes are 0. During the second loop iteration, after determining the target node, an unexpanded action can be selected, that is, push Item2, and then the target execution operations that the user may perform are determined according to the state transition probability, which are click, leave, and consult respectively, and three child nodes are created at the next level of the target node. At this time, both the push value and the search count of the three child nodes are 0.
[0060] Through the above solution, the purpose of expanding child nodes is achieved through the state transition probability, ensuring that the expanded child nodes can truly reflect the click and conversion value of the recalled data, achieving the effect of improving the accuracy of the push value and the accuracy of data push.
[0061] Optionally, simulating the target child node among the new child nodes to determine the push value corresponding to the last child node at the end of the simulation includes: determining the target child node based on the probability corresponding to the new child node; simulating the target child node; determining the end of the simulation when the simulation time reaches the preset simulation time or the simulation depth reaches the preset simulation depth; and determining the push value of the last child node.
[0062] The preset simulation time in the above steps can be the time for simulating the child node set in advance, which can be set according to actual needs, and the present disclosure does not make specific limitations thereto. The preset simulation depth in the above steps can be the depth for simulating the child node set in advance, and the preset simulation depth here can be determined according to user habits or the average step length. In most data push scenarios, the user step length is often small. As Figure 5 shown, most of the user step lengths are concentrated below 5 steps. Therefore, the exploration depth can be determined by slightly increasing the user's historical step length. It should be noted that for new users, the user's historical step length can be estimated based on the average step length of all users on the website, or directly use the average historical step length of similar users as the user's historical step length.
[0063] In the embodiments of the present disclosure, after expanding the child nodes, a child node can be selected and simulated according to MDP until the set simulation time or simulation depth is reached. Then, the reward of the last child node can be estimated according to the value prediction model, and then the push value of the child node can be updated based on the estimated reward. If the recalled data is pushed repeatedly, the cost of the last node can also be estimated through the cost prediction model, and then the push value of the child node can be updated based on the estimated reward and cost. For child nodes at different levels, different discount rates corresponding to different levels can be set, and the deeper the level, the lower the discount rate, so as to update the corresponding push value by obtaining the product of the reward and the discount rate. For example, the discount rate can be 0.9, but it is not limited thereto.
[0064] In some alternative embodiments, for the MDP tree as Figure 4 shown, in the first loop iteration, after creating three child nodes, a child node (i.e., clicking on the corresponding child node) can be selected for simulation, and then according to the state transition probability, Item k is selected to create three child nodes at the next level, and the corresponding child node is selected to be clicked. At this time, it is determined that the exploration depth is reached. Therefore, the push value of the corresponding child node clicked can be determined. Assuming that the reward estimated for the child node is 10 and the simulation depth is 3, then the push value V(y) of the node is updated as V(y)=γ 3 ×reward = 0.9 3× 10 = 7, where γ represents the discount rate. The deeper the exploration depth, the lower the discount rate. During the second loop iteration, after creating three child nodes, one child node can be selected (i.e., click on the corresponding child node) for simulation, and then Itemk’ is selected according to the state transition probability. Three child nodes of the next level are created, and the corresponding child node is selected and clicked. At this time, it is determined that the exploration depth is reached. Therefore, the push value of the corresponding child node can be determined. Assuming that the estimated reward of this child node is 100 and the simulation depth is 3, then update the push value V(n) of this node = γ 3 × reward = 0.9 3 × 100 = 70.
[0065] Through the above solution, the target child node is determined based on the probability corresponding to the new child node, the target child node is simulated, it is determined whether the simulation ends through the preset simulation time or preset simulation depth, and by determining the push value of the last child node, the push value is adjusted in real time to improve the push accuracy.
[0066] Optionally, before determining the push value of the last child node, it further includes: determining whether the associated data corresponding to the last child node is duplicate data repeatedly pushed within the target time period; in response to the associated data being duplicate data, processing the last child node using the cost estimation model and the value estimation model to obtain the push value of the last child node; in response to the associated data not being duplicate data, processing the last child node using the value estimation model to obtain the push value of the last child node.
[0067] The target time period in the above steps can be a preset short time period. For example, it can be the entire construction process of the search tree or a historical time period before the construction of the search tree, but it is not limited to this.
[0068] In the embodiments of the present disclosure, for the recall data repeatedly pushed within a short time, in order to avoid a poor experience brought to users by repeated pushing, it can first be determined whether the recall data corresponding to the last child node is duplicate data. If so, it is necessary to combine the estimation results of the cost estimation model and the value estimation model to determine the push value; if not, only the estimation result of the value estimation model is needed to determine the push value.
[0069] Through the above solution, different determination processes are given for duplicate data and non-duplicate data, achieving the effect of improving the determination accuracy of the push value and further improving the accuracy of data pushing.
[0070] Optionally, based on the push value of the last child node, perform a reverse iteration on the search tree to update the push values of the target node and each layer of child nodes below the target node, including: Step a, obtain the total value of the expanded node based on the push value, reward, and state transition probability of at least one child node located in the current layer; Step b, update the push value of the parent node located in the upper layer based on the maximum value of the total value of the expanded node; repeat Steps a to b until the update of the push value of the root node is completed.
[0071] It should be noted that the update here can be to update the current push value of the parent node to the maximum value of the total value of the expanded node, or to superimpose the current push value of the parent node and the maximum value of the total value of the expanded node. In the embodiments of the present disclosure, taking the example of updating the current push value of the parent node to the maximum value of the total value of the expanded node for illustration.
[0072] During the backpropagation process, there will be a discount for the push value of each child node. Therefore, the product of the push value and the discount rate can be obtained, then accumulated with the reward of the child node, and finally multiplied by the state transition probability of the child node to obtain the final total value. In some optional embodiments, for the Figure 4 MDP tree shown as follows, during the first loop iteration, the push value V(t) of node t can be updated according to the following formula: V(t) = max(P(click|t) × [r(y) + γV(y)] + P(leave|t) × [r(y′) + γV(y′)] + P(convert|t) × [r(y″) + γV(y″)]) = max(0.1 × (0 + 0.9 × 7) + 0.79 × (0 + 0) + 0.01 × (0 + 0)) = 0.63. Then, the push value V(s) of node s can be updated according to the following formula: V(s) = max a∈{1,2,…,k} (P a (t|s) × [r(t,a,s′) + γV(s′)]) = max(0.1 × (3 + 0.9 × 0.63) + 0.89 × (0 + 0) + 0.01 × (100 + 0)) - action:Item1 = 1.356. At this time, the first loop iteration process ends, the search count N of the root node is updated to 1, and the push value value is updated to 1.356. During the second loop iteration, the push value V(m) of node m can be updated according to the following formula: V(m) = max(P(click|n) × [r(n) + γV(n)] + P(leave|n) × [r(n′) + γV(n′)] + P(convert|n) × [r(n″) + γV(n″)]) = max(0.1 × (0 + 0.9 × 70) + 0.78 × (0 + 0) + 0.02 × (0 + 0)) = 6.3. Then, the push value V(s) of node s can be updated according to the following formula: V(s) = max a∈{1,2,…,k} (Pa (t|s)×[r(t,a,s′)+γV(s′)]) = max(0.1×(3 + 0.9×0.63)+0.89×(0 + 0)+0.01×(100 + 0)-action:Item1, 0.15×(5 + 0.9×6.3)+0.8×(0 + 0)+0.05×(80 + 0)-action:Item2) = max(1.356, 4.8505) = 4.8505. At this time, the second loop iteration process ends, the search times N of the root node is updated to 2, and the push value value is updated to 4.8505.
[0073] It should be noted that in the above case, the sorting method of Item2 > Item1 can be selected to give the final result.
[0074] Through the above solution, the total value is determined by the push value, reward, and state transition probability, and then the push value of each node is updated by backpropagation, so as to accurately determine the push value of each node, improve the determination accuracy of the target data, and thus achieve the effect of improving the push accuracy.
[0075] Next, in combination with Figure 4 and Figure 6 Take the product recommendation scenario as an example to describe a preferred embodiment of the present disclosure in detail. First, a user request is received. Here, the user request can be different requests for different scenarios. For example, in the search scenario, the user request can be the search text entered by the user; in the list page recommendation scenario, the user request can be the search text entered by the user; in the detail page recommendation scenario, the user request can be the behavior information of the user clicking on the current product detail page. In addition, search results and the user's historical browsing and clicking behaviors can be used as background information.
[0076] The specific process of this solution is as follows: As Figure 6 shown, the user can retrieve the information to be queried in the search box. Suppose the user searches for "spicy hot pot franchise" in the search box. It can be determined that "spicy hot pot franchise" is the user request. After receiving this request, the process of recommending result decision can be started. The recall data can be sorted in combination with the MCTS tree, and recommendation results such as "Recommendation Result 1", "Recommendation Result 2", etc. can be given below the retrieval results. As Figure 4As shown by the solid circles in the figure. Then, through simulation, it is assumed that if Item1 is recommended (such as Recommendation Result 1), the user may click on the product to enter the detail page, may directly click the "Consult" button to leave a lead conversion, or may just browse and then leave. Assume that the user clicks on the product in the recommendation result and enters the detail page. At this time, the recommendation result is displayed again on the detail page. Further assume that the user clicks on data such as "Join Diary", "Content Temperament", "Certification Rating", "Recommendation", etc., and a series of candidate lists (Item a - Item k) can be given for exploration. Observe whether these products will bring sufficient subsequent value after being recommended as Item1 on the list page and assuming the user clicks on them.
[0077] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0078] According to the embodiments of the present disclosure, the present disclosure provides a data processing device. This device is used to implement the above embodiments and preferred real-time methods, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0079] Figure 7 is a schematic diagram of the data processing device according to the present disclosure, as Figure 7 shown, the device includes: an acquisition module 72, configured to acquire a recall data set; a construction module 74, configured to construct a search tree corresponding to the recall data set, where the search tree includes: a root node and a plurality of data nodes at different levels, each data node is used to represent the recall data in the recall data set, and each data node is used to store the push value and search times of the corresponding recall data, and the push value is used to represent the value of the feedback result received after the corresponding recall data is pushed to the target object; a decision module 76, configured to determine the target data in the recall data set based on the push value of each recall data in the recall data set, where the target data is the data pushed to the target object.
[0080] Optionally, the building block includes: a first determination unit configured to determine a target node to be searched, where the target node is used to represent the root node or the recalled data in the recall data set; an expansion unit configured to expand new child nodes under the target node and use a value estimation model to determine the rewards of the new child nodes; a simulation unit configured to simulate a target child node in the new child nodes and determine the push value of the last child node at the end of the simulation; a second determination unit configured to perform a backward iteration on the search tree based on the push value of the last child node to determine the push values of the target node and each layer of child nodes under the target node; and an execution unit configured to repeatedly execute the functions of the determination unit, the expansion unit, the simulation unit, and the second determination unit until the exploration time of the search tree reaches a preset exploration time or the exploration depth reaches a preset exploration depth.
[0081] Optionally, the first determination unit includes: a traversal subunit configured to traverse from the root node to determine whether there is an unexpanded node; a first node determination subunit configured to, in response to the existence of an unexpanded node, determine the unexpanded node as the target node; and a second node determination subunit configured to, in response to the non-existence of an unexpanded node, determine the target node based on the push value and the search times of each data node.
[0082] Optionally, the expansion unit includes: a probability acquisition subunit configured to acquire the state transition probability corresponding to the target node; an operation determination subunit configured to determine the target execution operation corresponding to the target node based on the state transition probability; and a creation subunit configured to create a new child node under the target node based on the target execution operation.
[0083] Optionally, the simulation unit includes: a probability determination subunit configured to determine the target child node based on the probability corresponding to the new child node; a simulation subunit configured to simulate the target child node; a simulation determination subunit configured to determine the end of the simulation when the simulation time reaches a preset simulation time or the simulation depth reaches a preset simulation depth; and a value determination subunit configured to determine the push value corresponding to the last child node.
[0084] Optionally, the simulation unit further includes: a data determination subunit configured to determine whether the associated data corresponding to the last child node is duplicate data repeatedly pushed within a target time period; a first processing subunit configured to, in response to the associated data being duplicate data, process the last child node using a cost estimation model and a value estimation model to obtain the push value of the last child node; and a second processing subunit configured to, in response to the associated data not being duplicate data, process the last child node using the value estimation model to obtain the push value of the last child node.
[0085] Optionally, the second determination unit includes: a value acquisition subunit, configured to obtain the total value of the expansion node based on the push value, reward, and state transition probability of at least one child node located in the current layer; a value update subunit, configured to update the push value of the parent node located in the upper layer based on the maximum value of the total value of the expansion node, where the target child node is a child node associated with at least one child node; and an execution subunit, configured to repeatedly execute the functions of the value acquisition subunit and the value update subunit until the push value update of the root node is completed.
[0086] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0087] Figure 8 A schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0088] As Figure 8 shown, the device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0089] A plurality of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0090] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the data processing method. For example, in some embodiments, the data processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the data processing method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the data processing method in any other suitable way (e.g., by means of firmware).
[0091] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0092] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0093] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0094] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0095] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0096] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0097] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in the present disclosure can be achieved, and no limitation is imposed herein.
[0098] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A data processing method, comprising: Obtaining a recall data set; Constructing a search tree corresponding to the recall data set, wherein the search tree includes: a root node and a plurality of data nodes at different levels, each data node is used to represent the recall data in the recall data set, and each data node is used to store the push value and search times of the corresponding recall data, and the push value is used to represent the value of the feedback result received after the corresponding recall data is pushed to the target object; Based on the push value of each recall data in the recall data set, determining the target data in the recall data set, wherein the target data is the data pushed to the target object; Wherein, constructing the search tree corresponding to the recall data set includes: Step A, determining a target node to be searched; Step B, expanding new child nodes under the target node, and using a value prediction model to determine the reward of the new child nodes; Step C, simulating the target child nodes in the new child nodes, and determining the push value of the last child node at the end of the simulation; Step D, based on the push value of the last child node, performing reverse iteration on the search tree, updating the push value of the target node and each layer of child nodes under the target node, and updating the search times of the target node; repeating Steps A to D until the exploration time of the search tree reaches a preset exploration time or the exploration depth reaches a preset exploration depth.
2. The method according to claim 1, wherein determining the target node to be searched includes: Traversing from the root node to determine whether there is an unexpanded node; In response to the existence of the unexpanded node, determining the unexpanded node as the target node; In response to the non-existence of the unexpanded node, determining the target node based on the push value and search times of each data node.
3. The method according to claim 1, wherein expanding new child nodes under the target node includes: Obtaining the state transition probability corresponding to the target node; Based on the state transition probability, determining the target execution operation corresponding to the target node; Based on the target execution operation, creating the new child nodes under the target node.
4. The method according to claim 1, wherein simulating the target child nodes in the new child nodes and determining the push value corresponding to the last child node at the end of the simulation includes: Determining the target child nodes based on the probability corresponding to the new child nodes; Simulating the target child nodes; In the case where the simulation time reaches a preset simulation time or the simulation depth reaches a preset simulation depth, determining the end of the simulation; Determining the push value corresponding to the last child node.
5. The method according to claim 4, further comprising, before determining the push value corresponding to the last child node: Determining whether the associated data corresponding to the last child node is duplicate data repeatedly pushed within a target time period; In response to the associated data being the duplicate data, process the last child node by using a cost estimation model and a value estimation model to obtain the push value of the last child node; In response to the associated data not being the duplicate data, process the last child node by using the value estimation model to obtain the push value of the last child node.
6. The method according to claim 1, based on the push value of the last child node, performing a reverse iteration on the search tree to update the push values of the target node and each layer of child nodes below the target node, including: Step a, obtaining the total value of the expansion node based on the push value, reward, and state transition probability of at least one child node located in the current layer; Step b, obtaining the maximum value of the total value of the expansion node to obtain the push value of the parent node located in the upper layer; Repeat steps a to b until the update of the push value of the root node is completed.
7. A data processing device, comprising: An acquisition module, configured to acquire a recall data set; A construction module, configured to construct a search tree corresponding to the recall data set, where the search tree includes: a root node and a plurality of data nodes at different levels, each data node is used to represent the recall data in the recall data set, and each data node is used to store the push value and search times of the corresponding recall data, and the push value is used to represent the value of the feedback result received after the corresponding recall data is pushed to the target object; A decision module, configured to determine the target data in the recall data set based on the push value of each recall data in the recall data set, where the target data is the data pushed to the target object; Wherein, the construction module includes: a first determination unit, configured to determine a target node to be searched, where the target node is used to represent the root node or the recall data in the recall data set; an expansion unit, configured to expand new child nodes below the target node and use a value estimation model to determine the reward of the new child nodes; a simulation unit, configured to simulate the target child node in the new child nodes to determine the push value of the last child node at the end of the simulation; a second determination unit, configured to perform a reverse iteration on the search tree based on the push value of the last child node to determine the push values of the target node and each layer of child nodes below the target node; an execution unit, configured to repeat the functions of the determination unit, the expansion unit, the simulation unit, and the second determination unit until the exploration time of the search tree reaches a preset exploration time or the exploration depth reaches a preset exploration depth.
8. The device according to claim 7, wherein the first determination unit includes: A traversal sub-unit, configured to start traversing from the root node to determine whether there is an unexpanded node; A first node determination sub-unit, configured to, in response to the existence of the unexpanded node, determine the unexpanded node as the target node; A second node determination subunit, configured to determine the target node based on the push value and the search times of each data node in response to the absence of the unexpanded node.
9. The apparatus according to claim 7, wherein the expansion unit comprises: A probability acquisition subunit, configured to acquire the state transition probability corresponding to the target node; An operation determination subunit, configured to determine the target execution operation corresponding to the target node based on the state transition probability; A creation subunit, configured to create the new child node under the target node based on the target execution operation.
10. The apparatus according to claim 7, wherein the simulation unit comprises: A probability determination subunit, configured to determine the target child node based on the probability corresponding to the new child node; A simulation subunit, configured to simulate the target child node; A simulation determination subunit, configured to determine the end of the simulation when the simulation time reaches a preset simulation time or the simulation depth reaches a preset simulation depth; A value determination subunit, configured to determine the push value corresponding to the last child node.
11. The apparatus according to claim 10, wherein the simulation unit further comprises: A data determination subunit, configured to determine whether the associated data corresponding to the last child node is duplicate data repeatedly pushed within a target time period; A first processing subunit, configured to process the last child node by using a cost estimation model and a value estimation model to obtain the push value of the last child node in response to the associated data being the duplicate data; A second processing subunit, configured to process the last child node by using the value estimation model to obtain the push value of the last child node in response to the associated data not being the duplicate data.
12. The apparatus according to claim 7, wherein the second determination unit comprises: A value acquisition subunit, configured to obtain the total value of the expansion node based on the push value, the reward, and the state transition probability of at least one child node located in the current layer; A value update subunit, configured to update the push value of the parent node located in the upper layer based on the maximum value of the total value of the expansion node; An execution subunit, configured to repeatedly execute the functions of the value acquisition subunit and the value update subunit until the push value update of the root node is completed.
13. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.
15. A computer program product, comprising a computer program, where the computer program implements the method according to any one of claims 1-6 when being executed by a processor.
Citation Information
Patent Citations
Information processing method and device thereof
CN104699693A
Artificial intelligence-based news recalling method and device, equipment and storage medium
CN107391549A