Content pushing method and device, computer equipment, readable storage medium and program product
By disassembling the content push task into multiple subtasks, the number of contents generated by each subtask is smaller than the total number of pushes, the problem of excessive search space in the prior art is solved, and the content push effect is significantly improved.
Patent Information
- Application Number
- CN202510146486.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-27
AI Technical Summary
When pushing content in the prior art, if the length of the list required for the push request is long, the number of candidate lists will increase sharply, the search space will be too large, and it will be difficult to find the optimal combination method, and the content push effect will be poor.
In response to the push request, a collection of candidate content related to the push request is determined, and the generated total task is broken down into multiple subtasks. The number of contents of the target sublist generated by each subtask is smaller than the total number of pushes until the stop condition is reached.
This greatly reduces the search space, allowing more adaptable target sublists to be filtered out, significantly improving the content push effect.
Smart Images

Figure CN120050326A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a content push method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the development of computer technology and Internet technology, more and more content is pushed through the network. Due to the increasing diversity and personalization of users' content needs, personalized content push for users is usually required.
[0003] In the related art, based on a user's push request, after passing through a recall stage, a rough ranking stage, and a fine ranking stage in sequence, a large number of finely ranked contents are obtained. Then, based on the large number of finely ranked contents, multiple candidate lists are freely combined to form a search space, and then a list that meets the push requirements is searched from the search space for exposure.
[0004] However, in the related art, if the length of the list required by the push request is long, the number of combined candidate lists will increase sharply, thus greatly increasing the search space. Furthermore, it is difficult to find an optimal combination method for exposure in the huge search space, and the content push effect is not good. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a content push method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the content push effect.
[0006] In a first aspect, this application provides a content push method, including:
[0007] Responding to a push request, determining a set of candidate contents related to the push request;
[0008] Determining the total number of pushes that match the push request;
[0009] Performing multiple subtasks based on the set of candidate contents, and pushing the target sub-lists obtained each time the subtask is executed until the stop condition is reached and then stopping; the number of contents in the target sub-list is less than the total number of pushes;
[0010] Wherein, the execution steps of any subtask include: searching for multiple candidate sub-lists based on the un-pushed contents in the set of candidate contents, and screening out the target sub-list from the multiple candidate sub-lists.
[0011] In a second aspect, this application further provides a content push apparatus, including:
[0012] A set determination module, configured to determine a candidate content set related to the push request in response to the push request;
[0013] A quantity determination module, configured to determine the total number of pushes that match the push request;
[0014] A list push module, configured to perform multiple subtasks based on the candidate content set, and push the target sub-lists obtained by each execution of the subtask until the stop condition is reached and stop; the number of contents in the target sub-list is less than the total number of pushes; wherein, the execution steps of any one subtask include: searching for multiple candidate sub-lists based on the un-pushed contents in the candidate content set, and screening out the target sub-list from the multiple candidate sub-lists.
[0015] In a third aspect, the present application further provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above content push method are implemented.
[0016] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above content push method are implemented.
[0017] In a fifth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above content push method are implemented.
[0018] The above content push method, device, computer device, computer-readable storage medium and computer program product determine a candidate content set related to the push request in response to the push request; determine the total number of pushes that match the push request; perform multiple subtasks based on the candidate content set, and push the target sub-lists obtained by each execution of the subtask until the stop condition is reached and stop; the number of contents in the target sub-list is less than the total number of pushes. That is to say, the total task of generating the total number of pushes is decomposed into multiple serial subtasks, and the growth length of the target sub-list generated by each subtask is less than the growth length of the list corresponding to the total task. Among them, the execution steps of any one subtask include: searching for multiple candidate sub-lists based on the un-pushed contents in the candidate content set, and screening out the target sub-list from the multiple candidate sub-lists. That is to say, for each subtask, a plurality of candidate sub-lists with a smaller growth length are selected from the selection range less than or equal to the candidate content set, and there is no need to select a list with a larger growth length from the total number of pushes, which greatly reduces the search space and it is very easy to screen out a more suitable target sub-list from this search space, thereby improving the content push effect. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0020] Figure 1 It is an application environment diagram of the content push method in an embodiment;
[0021] Figure 2 It is a schematic flowchart of the content push method in an embodiment;
[0022] Figure 3 It is a schematic flowchart of the first future revenue value determination step in an embodiment;
[0023] Figure 4 It is a schematic diagram of the candidate sub - list determination step in an embodiment;
[0024] Figure 5 It is a schematic diagram of the process of pushing content in an embodiment;
[0025] Figure 6 It is a schematic flowchart of the second future revenue prediction model update step in an embodiment;
[0026] Figure 7 It is a structural block diagram of the content push device in an embodiment;
[0027] Figure 8 It is an internal structure diagram of a computer device in an embodiment. Specific embodiments
[0028] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further details the present application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0029] Before introducing the present application, the process of generating the target sub - list corresponding to the subtasks involved in the present application is first described. The embodiments of the present application regard the generation of a target sub - list once as a Markov decision process. Among them:
[0030] : State set, a specific state , where each state represents the state of selecting a certain target sub - list for exposure and being viewed. The user characteristics of the viewers, the characteristics of the un - pushed content in the candidate content set (which can be regarded as the context characteristics in this state), and the characteristics of the pushed content before the target sub - list (which can be regarded as the context characteristics in this state) can be counted. Based on the user characteristics, the corresponding context characteristics before and after, the feature combination in this state is determined;
[0031] : The set of actions. In the embodiments of the present application, it can be understood as the set of multiple candidate sub - lists composed of un - pushed content. Among them, a specific action , in the embodiments of the present application, represents a certain candidate sub - list in A.
[0032] : Immediate benefit, represents the immediate benefit after taking the action in the state . In the embodiments of the present application, it can be understood that in the state s, after selecting a certain candidate sub - list for exposure, the immediate benefit brought by this candidate sub - list;
[0033] : State transition function, represents the probability that after taking the action in the state , the state transfers to .
[0034] , which can be understood as the push strategy for generating the target sub - list, that is, the strategy for selecting the target sub - list from multiple candidate sub - lists.
[0035] , which is also V mentioned later: represents the total benefit under the push strategy and the initial state . This total benefit is obtained by integrating the immediate benefit and the future benefit. The future benefit refers to the benefit that the un - pushed content in the candidate content set in a certain state can bring. The specific formula is as follows:
[0036]
[0037] Among them, is the new state. is the coefficient. refers to the push strategy selected in the initial state . refers to the total benefit under the push strategy and the initial state .
[0038] In the process of personalized content push, the push system will successively perform recall, rough ranking, fine ranking, and re-ranking. Among them, the main purpose of re-ranking is to refine the sorting of the content obtained from fine ranking to generate a more accurate push list, so as to improve the quality of push results and user satisfaction.
[0039] Among them, for the push presented in a two-column form, re-ranking can be completed based on the point level, that is, jointly perform point-level re-estimation of the candidate set and the above-text perception. For example, when predicting the (x + 1)-th push content, it is necessary to predict based on the first x generated push contents by calling a model once. That is to say, to generate a push list with a growth length (number of contents) of M, the model needs to be called multiple times, which will result in a relatively large time delay, and the maximization of the target benefit is not considered throughout the process.
[0040] Therefore, in related technologies, re-ranking is usually completed in a list-level manner. In this process, it is necessary to first generate multiple candidate lists based on the content in the candidate set, and then estimate the final push list for exposure from the multiple candidate lists. However, if the growth length of the push list required by the push request is relatively long, such as M, and the candidate set contains N candidate contents, where N is much larger than M. At this time, first selecting M from N There are combinations, and each combination contains M elements (contents). For each combination, there are permutation ways (that is, the factorial of M); thus, there are a total of permutation selections. At this time, the search space is That is, there are candidate lists.
[0041] Since there are many candidate contents obtained from fine ranking (the value of N is very large), it can be understood that there is a problem of combinatorial explosion, resulting in a large search space, and it is difficult to quickly query a suitable push list in the huge search space, and the content push effect is not good.
[0042] In the embodiment of the present application, by responding to a push request, a candidate content set related to the push request is determined; the total number of pushes matching the push request is determined; multiple subtasks are executed based on the candidate content set, and the target sub-lists obtained by each execution of the subtask are pushed until the stop condition is reached and then stopped; the number of contents in the target sub-list is less than the total number of pushes. That is to say, the total task of generating the total number of pushes is decomposed into multiple serial subtasks, and the growth length of the target sub-list generated by each subtask is less than the growth length of the list corresponding to the total task. Among them, the execution steps of any subtask include: searching for multiple candidate sub-lists based on the un-pushed contents in the candidate content set, and screening out the target sub-list from the multiple candidate sub-lists. That is to say, for each subtask, multiple candidate sub-lists with a smaller growth length are selected from the selection range less than or equal to the candidate content set, without selecting a list with a larger growth length from the total number of pushes, greatly reducing the search space, and it is very easy to screen out a more suitable target sub-list from this search space, thereby improving the content push effect.
[0043] An example is used to illustrate the search space of the subtask in the embodiment of the present application. For the (i + 1)-th subtask, the number of contents in the target sub-list is K, K is less than M, and the number of un-pushed contents in the candidate content set is N - i×K. At this time, first select K from N - i×K, and there are combinations. Since each combination contains K contents, therefore, arranging the contents in each combination, there are arrangement methods. Therefore, there are candidate sub-lists. At this time, the search space is . Comparing with the search space in the related technology above, obviously, the search space in the embodiment of the present application is smaller. In this way, it is easier to screen out a more suitable target sub-list to significantly improve the content push effect.
[0044] The content push method provided by the embodiment of the present application can be applied to the application environment as shown in Figure 1 . Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers.
[0045] The terminal 102 initiates a push request to the server 104. The server 104, in response to the push request, determines a set of candidate content related to the push request; the server 104 determines the total number of pushes that match the push request; the server 104 performs multiple subtasks based on the set of candidate content and pushes the target sub-list obtained from each execution of the subtask to the terminal 102 for the terminal 102 to expose the target sub-list until the stop condition is reached and then stops; the number of contents in the target sub-list is less than the total number of pushes; wherein, the execution steps of any one subtask include: the server 104 searches for multiple candidate sub-lists based on the un-pushed content in the set of candidate content and filters out the target sub-list from the multiple candidate sub-lists.
[0046] Among them, the terminal 102 is a content receiving terminal. It can be understood that the account logged in on the terminal 102 can be understood as the account of the content recipient. The terminal 102 is used to receive the target sub-list. The terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0047] In an exemplary embodiment, as Figure 2 shown, a content push method is provided. Taking the method applied to the Figure 1 server 104 as an example for description, it includes the following steps S202 to step S206. Among them:
[0048] Step S202, in response to the push request, determine a set of candidate content related to the push request.
[0049] Among them, the push request refers to a request for content push. The content includes but is not limited to videos, notes, items, etc. Exemplarily, the terminal generates a push request in response to a sliding operation on the content display interface and sends it to the server. Exemplarily, the terminal generates a search term in response to an input operation on the search box of the content display interface, generates a push request based on the search term, and sends it to the server. Exemplarily, the terminal automatically generates a trigger request when detecting an open operation on the target application; Exemplarily, the terminal automatically triggers and generates a push request every preset period and sends it to the server.
[0050] The candidate content set is a set filtered from the content pool based on the push request. The candidate content set contains a preset number of candidate contents, for example, it contains N candidate contents.
[0051] Optionally, in response to the push request, the server sequentially performs recall, rough ranking, and fine ranking based on the content provided in the content pool to obtain a candidate content set containing multiple candidate contents. Among them, the candidate content set can be regarded as a set of fine-ranked contents obtained through fine ranking.
[0052] Optionally, in response to the push request, the server obtains the search term from the push request, filters out the candidate contents with a high degree of association with the search term from the content pool, and takes the set of the filtered candidate contents as the candidate content set.
[0053] Step S204, determine the total number of pushes that match the push request.
[0054] Among them, the total number of pushes that match the push request refers to the total number M of contents that need to be pushed in theory each time the push request is responded to.
[0055] It should be noted that the total number N of contents in the candidate content set is greater than the total number of pushes M.
[0056] Step S206, perform multiple subtasks based on the candidate content set, and push the target sub-lists obtained each time the subtask is executed until the stop condition is reached and stop; the number of contents in the target sub-list is less than the total number of pushes; among them, the execution steps of any subtask include: searching for multiple candidate sub-lists based on the un-pushed contents in the candidate content set, and filtering out the target sub-list from the multiple candidate sub-lists.
[0057] Among them, each subtask is used to generate a target sub-list, and the number of contents in each target sub-list is the same; of course, the number of contents in each target sub-list can also be different. For example, according to the feedback information of the current target sub-list, adjust the number of contents in the next target sub-list. For example, for the current target sub-list, if the number of times the content recipient favorites, clicks, or comments is higher than the upper limit, then increase the number of contents in the next target sub-list. If it is higher than the lower limit and lower than the upper limit, then keep the number of contents in the current target sub-list and the next target sub-list the same. If it is lower than the lower limit, then reduce the number of contents in the next target sub-list.
[0058] The number of contents in each target sub-list must be less than the total number of pushes.
[0059] The stop condition refers to the condition for stopping the execution of subtasks. For example, the stop condition can be that the number of times of executing the subtask meets a preset number; the stop condition can also be that the current push quantity of the push request reaches the total push quantity; the stop condition can also be that the content that has not been pushed in the candidate content set cannot meet the benefit requirement.
[0060] The candidate sub-list is the sub-list to be screened, and the number of contents in the candidate sub-list is equal to the number of contents in the target sub-list.
[0061] Optionally, the server determines the number of contents in the target sub-list, and determines the preset number of times of executing the subtask according to the number of contents and the total push quantity. The server executes the subtask multiple times based on the candidate content set, and pushes the target sub-list obtained each time the subtask is executed until the number of times of the executed subtasks meets the preset number of times. For example, if the number of contents is K and the total push quantity is M, calculate the quotient obtained by dividing M by K, and determine the preset number of times based on the quotient.
[0062] Optionally, after the current subtask is executed, it is statistically analyzed whether the current push quantity of the push request reaches the total push quantity. If so, stop. If not, execute the next subtask.
[0063] Optionally, before the current subtask is executed, it is judged whether the content that has not been pushed in the candidate content set cannot meet the benefit requirement. If it meets, execute the current subtask. If it does not meet, stop.
[0064] In any of the above three optional examples, the execution steps of any subtask include: searching for multiple candidate sub-lists based on the content that has not been pushed in the candidate content set, and screening out the target sub-list from the multiple candidate sub-lists.
[0065] It should be noted that in the related art, a single push request will generate a target list containing the total push quantity M at one time. However, for the content receiver, although it can quickly scroll through when browsing the double column, it cannot view M contents in a short time, and statistical analysis shows that the maximum scroll depth of most content receivers is less than M. Therefore, the one-time output of M contents in the related art is likely to cause waste of computing power. In this application, it is pushed in multiple times, and the growth length of the target sub-list pushed each time is less than M. In this way, it is not only convenient for browsing, but also can avoid waste of computing power.
[0066] In the above content pushing method, in response to a pushing request, a candidate content set related to the pushing request is determined; the total number of pushes matching the pushing request is determined; multiple subtasks are executed based on the candidate content set, and the target sub - lists obtained from each execution of the subtask are pushed until the stop condition is met and then stopped; the number of contents in the target sub - list is less than the total number of pushes. That is to say, the total task of generating the total number of pushes is decomposed into multiple serial subtasks, and the growth length of the target sub - list generated by each subtask is less than the growth length of the list corresponding to the total task. Among them, the execution steps of any subtask include: searching for multiple candidate sub - lists based on the un - pushed contents in the candidate content set, and screening out the target sub - list from the multiple candidate sub - lists. That is to say, for each subtask, multiple candidate sub - lists with a smaller growth length are selected from the selection range less than or equal to the candidate content set, without having to select a list with a larger growth length from the total number of pushes, greatly reducing the search space, and it is very easy to screen out a more suitable target sub - list from this search space, thereby improving the content pushing effect.
[0067] In some embodiments, the stop condition includes at least one of the following: the first future benefit value determined by the un - pushed contents in the candidate content set does not reach the preset benefit threshold; the total number of contents in the pushed target sub - lists reaches the total number of pushes.
[0068] Among them, the first future benefit value refers to the value of the future benefit that the un - pushed contents in the candidate content set can bring. The first future benefit value determined by the un - pushed contents in the candidate content set does not reach the preset benefit threshold can be understood as that the un - pushed contents in the candidate content set mentioned above cannot meet the benefit requirements.
[0069] The pushed target sub - list refers to the target sub - list that has been pushed. For example, in response to a pushing request, if the x - th subtask has been completed currently, the pushed target sub - lists are respectively the target sub - list 1 generated by the first subtask, ……, the target sub - list x generated by the x - th subtask.
[0070] The total number of contents in the pushed target sub - lists reaching the total number of pushes can also be understood as that the current number of pushes for the pushing request mentioned above reaches the total number of pushes. For example, in response to a pushing request, if the x - th subtask has been completed currently and x target sub - lists have been pushed, if the growth length of each target sub - list is K, then the current number of pushes for the pushing request is the product of K and x.
[0071] The above mentioned how to perform content pushing based on a single stop condition. Next, the content pushing process of the embodiments of this application by combining these two stop conditions is introduced:
[0072] Step 1: For each subtask, before the execution of the subtask, the server determines whether the first future benefit value determined by the un-pushed content in the candidate content set reaches the preset benefit threshold based on the un-pushed content in the candidate content set.
[0073] Step 2: If it does not reach, stop pushing content for this push request. If it reaches, start executing the subtask. During the execution of the subtask, the server searches for multiple candidate sub-lists based on the un-pushed content in the candidate content set, filters out the target sub-list from the multiple candidate sub-lists, and returns the target sub-list to the terminal for pushing.
[0074] Step 3: The server determines whether the total number of contents in the pushed target sub-list reaches the total push number. If it reaches, stop pushing content for this push request. If it does not reach, the server takes each content in the target sub-list as the pushed content, updates the un-pushed content in the candidate content set based on the target sub-list, returns to judge whether the stop condition is reached based on the un-pushed content in the candidate content set, and continues to execute until the total number of contents in the pushed target sub-list reaches the total push number.
[0075] Exemplarily, determining whether to start a subtask can also be: The server determines the resource consumption generated by estimating the generation of the subtask. When the first future benefit value reaches the preset benefit threshold and the resource consumption is less than the consumption threshold, return to the step of starting to execute this subtask and continue to execute.
[0076] When the first future benefit value does not reach the preset benefit threshold and the resource consumption is greater than or equal to the consumption threshold, stop pushing content for this push request.
[0077] In some scenarios, for the case where the first future benefit value reaches the preset benefit threshold and the resource consumption is greater than or equal to the consumption threshold, the subtask can be executed. Of course, it can also not be executed, which is determined according to actual requirements.
[0078] Among them, the resource consumption can be the time taken to generate the subtask, or the CPU (Central Processing Unit) occupied, etc.
[0079] Exemplarily, taking the i-th subtask as an example, before executing the i-th subtask, it is determined that there are y un-pushed contents in the candidate content set (y = N - (i - 1) × k), and it is estimated whether the first future benefit value brought by the y un-pushed contents reaches the preset benefit threshold. If it does not reach, stop starting the i-th subtask and stop responding to this push request.
[0080] If it is reached, the i-th sub-task is started. Specifically, the server searches C candidate sub-lists from y un-pushed contents, the growth length of each candidate sub-list is k, and the target sub-list is screened out from multiple candidate sub-lists, and the target sub-list is returned to the terminal for pushing.
[0081] The server counts the total number i×k of the contents in the pushed target sub-list, and judges whether the total number i×k is greater than or equal to the total pushing number N. If it is reached, the pushing is stopped. If not, the un-pushed contents are updated to obtain z un-pushed contents (z = N - i×k), and based on the un-pushed contents in the candidate content set, it is judged whether the stop condition is reached and the steps are continued to be executed until the total number of the contents in the pushed target sub-list reaches the total pushing number.
[0082] Exemplarily, after the server determines that the browsing depth of the previous target sub-list reaches the preset depth, the step of judging whether the first future revenue value reaches the preset revenue threshold is continued to be executed.
[0083] Among them, the previous target sub-list is generated by the previous sub-task.
[0084] For example, the preset depth can be the growth length of the target sub-list. At this time, after the content receiver browses the previous target sub-list, the step of judging whether the first future revenue value reaches the preset revenue threshold is continued to be executed. Another example is that the preset depth can be less than the growth length of the target sub-list. For example, after the content receiver browses 80% of the contents of the previous target sub-list, the step of judging whether the first future revenue value reaches the preset revenue threshold is continued to be executed.
[0085] In this embodiment, by setting the stop condition as the first future revenue value determined by the un-pushed contents in the candidate content set and not reaching the preset revenue threshold, it can effectively estimate whether the un-pushed contents can generate a target sub-list with high revenue, effectively avoid the generation of sub-tasks with low revenue, avoid waste of computing power, and save computing power. By setting the stop condition as the total number of the contents in the pushed target sub-list reaching the total pushing number, it is ensured that the pushing requirements of the push request can be met, and the user experience is improved.
[0086] In some embodiments, the step of determining the first future revenue value includes: based on the characteristics of the un-pushed contents in the candidate content set, calling the updated first future revenue prediction model to obtain the first future revenue value.
[0087] Among them, the updated first future revenue prediction model is used to predict future revenue. The updated first future revenue prediction model is a neural network model.
[0088] Optionally, the server determines the content that has not been pushed in the candidate content set, and inputs the features of the content that has not been pushed into the updated first future revenue prediction model to output the first future revenue value.
[0089] Exemplarily, when it is necessary to determine whether to start a subtask in combination with resource consumption, the updated first future revenue prediction model can also determine resource consumption. That is, the features of the content that has not been pushed are input into the updated first future revenue prediction model to output the first future revenue value and resource consumption.
[0090] In this embodiment, based on the features of the content that has not been pushed in the candidate content set, through the updated first future revenue prediction model, the first future revenue value that the content that has not been pushed can bring can be predicted more quickly and timely.
[0091] In some embodiments, as Figure 3 shown, it is a schematic flowchart of the steps for determining the first future revenue value in an embodiment. Based on the features of the content that has not been pushed in the candidate content set, the updated first future revenue prediction model is called to obtain the first future revenue value, including:
[0092] Step S302, using the features of the pushed content as the above - text features in the current state, where the current state refers to the state where the latest pushed target sub - list is browsed.
[0093] Among them, the latest pushed target sub - list can be pushed in response to a push request or in response to the previous push request.
[0094] For example, before the first subtask (the first subtask required to be executed in response to a push request) is not started, the latest pushed target sub - list is the sub - list pushed last in response to the previous push request.
[0095] Before the non - first subtask (the non - first subtask required to be executed in response to a push request) is not started, the latest pushed target sub - list is the target sub - list obtained from the previous subtask in response to this push request.
[0096] Among them, the pushed content includes at least the content of the latest pushed target sub - list. For example, it only includes the content of the latest pushed target sub - list, or for example, it includes the content of the list before the latest pushed target sub - list and the content of the latest pushed target sub - list.
[0097] Among them, the first future revenue value is the future revenue evaluated based on the content that has not been pushed. The future revenue is long - term revenue, which can also be understood as a long - term and potential effect or return. It can reflect the user's satisfaction with the push application platform. The higher the future revenue value, the better the satisfaction, and it can better improve the user's dependence and stickiness to the push application platform.
[0098] Optionally, for each subtask, the server determines the latest pushed target sublist, uses the latest pushed target sublist as the pushed content, and uses the features of the pushed content as the context features in the current state.
[0099] Exemplarily, for the i-th subtask, if i equals 1, the server uses the features of each content in the sublist pushed last in response to the previous push request as the context features in the current state.
[0100] If i is greater than 1, the server uses the features of each content in the target sublist pushed for the (i - 1)-th time in response to this push request as the context features in the current state.
[0101] Step S304: Use the features of the unpushed content in the candidate content set as the context features in the current state, and obtain the object features of the content recipient in the current state.
[0102] Optionally, the server determines the unpushed content in the candidate content based on the content in the target sublist obtained from the completed subtasks. The server uses the unpushed content as the context features in the current state. The server obtains the object features of the content recipient.
[0103] Exemplarily, for the i-th subtask, determine the ((i - 1) × k) contents obtained from the previous i - 1 subtasks and remove them from the candidate content set, and use the features of each content in the set after removal as the context features in the current state.
[0104] Step S306: Based on the context features, context features, and object features, construct a feature combination in the current state, and input the feature combination into the updated first future revenue prediction model to obtain a first future revenue value.
[0105] In this embodiment, by using the features of the pushed content and the features of the unpushed content in the candidate content set as the context features in the current state, based on the object features of the content recipient in the current state and the context features, through the updated first future revenue prediction model, it is possible to comprehensively consider the details of the context, context, and object in three dimensions, and more accurately and comprehensively predict the future revenue of the current state, ensuring the accuracy of the future revenue prediction.
[0106] In some embodiments, multiple candidate sublists are searched based on the unpushed content in the candidate content set, including: generating a search tree based on the unpushed content in the candidate content set and the interaction data generated by the pushed target sublist, and determining multiple candidate sublists based on the search tree.
[0107] Among them, the interaction data refers to the data feedback by the content receiver, including but not limited to the number of likes, the number of forwards, the number of comments, the browsing duration, etc.
[0108] A search tree refers to a tree-like data structure constructed during the execution of the beam search algorithm, used to represent and search the possible solution space.
[0109] Optionally, for the first layer, use the content in the pushed target sublist as the above-content of the first layer. Based on the content information of the above-content and the content information of the unpushed content, call the scoring model to score the unpushed content, and use the preset number of unpushed content with the top scores as the leaf nodes of the first layer. Each leaf node of the first layer corresponds to a candidate sublist, and each leaf node of the first layer is the content in the first push position of the corresponding candidate sublist;
[0110] For the current layer that is not the first layer, obtain each leaf node of the previous layer. For each leaf node of the previous layer, determine the leaf node of the first layer from which each leaf node of the previous layer originates, and use the leaf node of the first layer, the leaf node of the previous layer, and each intermediate layer leaf node between the leaf node of the first layer and the leaf node of the previous layer as the above-content of the corresponding candidate sublist in the current layer.
[0111] Use the other content in the candidate content set except the above-content as the content to be screened corresponding to the candidate sublist in the current layer, and based on the content information of the above-content and the content information of the content to be screened, call the scoring model to score each content to be screened respectively. Use the content to be screened with the highest score as the leaf node of the current layer pointed to by the leaf node of the previous layer.
[0112] After determining the leaf node of the last layer, stop generating the search tree.
[0113] Among them, the content information of the above-content includes at least the features of the content, the arrangement order of each content in the above-content, etc. The content information of other content includes at least the features of other content, and the content information of the unpushed content includes at least the features of the unpushed content.
[0114] Optionally, after determining the search tree, the server traverses the search tree layer by layer, determines the path from each leaf node of the first layer to the corresponding leaf node of the last layer. For each path, determine the order of each leaf node in the direction from the leaf node of the first layer to the corresponding leaf node of the last layer in this path, and arrange each leaf node in this path in the order corresponding to this path to generate the candidate sublist corresponding to this path.
[0115] Exemplarily, as Figure 4 shown, it is a schematic diagram of the steps for determining the candidate sublist in an embodiment. For Figure 4In the stage of generating the candidate sub - list, the specific steps are as follows:
[0116] Step 4.1: The server determines the beam width of beam search and the number of layers of beam search. The beam width is used to determine the number of candidate sub - lists, that is, to determine the preset quantity in the first layer. For example Figure 4 If the beam width is 2, then finally 2 candidate sub - lists are determined.
[0117] Step 4.2: If the content that has not been pushed includes Content1 (Content 1), Content2 (Content 2), Content3 (Content 3) and Content4 (Content 4).
[0118] Starting from the first layer, take the content in the pushed target sub - list as the above - text content of the first layer. Based on the content information of the above - text content and the content information of the content that has not been pushed, call the scoring model to score the content that has not been pushed, and select the top 2 pieces of content, which are Content2 and Content4 respectively. These two pieces of content are used as the leaf nodes of the first layer respectively, and each leaf node of the first layer is used as the content of the first push position in the corresponding candidate sub - list. For example, the content of the first push position of candidate sub - list 1 is Content2, and the content of the first push position of candidate sub - list 2 is Content4.
[0119] For the second layer, Content2 is the above - text content of candidate sub - list 1 in the second layer, and Content1, Content3 and Content4 are the content to be screened corresponding to candidate sub - list 1 in the second layer. Thus, Content1 is screened out from the content to be screened as the leaf node of the second layer.
[0120] Similarly, Content4 is the above - text content of candidate sub - list 2 in the second layer, and Content1, Content2 and Content3 are the content to be screened corresponding to candidate sub - list 2 in the second layer. Thus, Content1 is screened out from the content to be screened as the leaf node of the second layer.
[0121] At this time, Path 1: Content2 points to Content1, then the candidate sub - list 1 of Path 1 is [Content2, Content1], with Content2 pushed first and Content1 pushed later.
[0122] Path 2: Content4 points to Content1, then the candidate sub - list 2 of Path 2 is [Content4, Content2], with Content4 pushed first and Content2 pushed later.
[0123] In this embodiment, based on the interaction data generated from the un-pushed content in the candidate content set and the target sub-list of the pushed content, the beam search algorithm is used to adopt a breadth-first strategy to generate a search tree. In this way, based on the search tree, multiple ordered sequences can be generated more quickly, that is, multiple candidate sub-lists are generated, ensuring the efficiency of the candidate sub-lists.
[0124] In some embodiments, screening the target sub-list from multiple candidate sub-lists includes: for each candidate sub-list, predicting the immediate benefit value generated after pushing the candidate sub-list; for each candidate sub-list, predicting the second future benefit value determined by the remaining content in the candidate content set after pushing the candidate sub-list; fusing the immediate benefit value and the second future benefit value to obtain the total benefit value corresponding to the candidate sub-list; selecting the candidate sub-list with the largest total benefit value as the target sub-list.
[0125] The immediate benefit value refers to the value of the immediate benefit, which reflects the direct and short-term effects brought by the pushed list or content after the pushed list or content is sent. For example, if the user immediately generates interaction behaviors (including but not limited to clicking, forwarding, favoriting, purchasing) after the pushed list or content is sent, the higher the immediate benefit value.
[0126] The second future benefit value is the future benefit determined based on the remaining content. For example, the un-pushed content includes Content1 to Content100. After determining multiple candidate sub-lists, for candidate sub-list 1 (including Content1 to Content10), the remaining content refers to Content11 to Content100, that is, the future benefit that these 90 contents can generate.
[0127] Optionally, for each candidate sub-list, the server calls the updated immediate benefit prediction model, and based on the characteristics of the content in the pushed candidate sub-list, determines the immediate benefit value. After the server determines to push the candidate sub-list, for the remaining content in the candidate content set, based on the characteristics of the remaining content and the characteristics of the content in the pushed candidate sub-list, the updated second future benefit prediction model is called to predict the second future benefit value determined by the remaining content in the candidate content set after pushing the candidate sub-list.
[0128] The server obtains the weights of the immediate benefit value and the second future benefit value respectively, and performs weighted summation on the immediate benefit value and the second future benefit value according to the weights to obtain the total benefit value of the candidate sub-list.
[0129] The server selects the candidate sub-list with the largest total benefit value as the target sub-list.
[0130] Exemplarily, continuing with Figure 4Taking this as an example, after determining candidate sub - list 1 and candidate sub - list 2, candidate sub - list evaluation is carried out. Specifically, based on the updated immediate revenue prediction model, immediate revenue evaluation is performed to obtain the immediate revenue value. Based on the updated second future revenue prediction model, future revenue evaluation is performed to obtain the second future revenue value. In this way, the total revenue value of each of candidate sub - list 1 and candidate sub - list 2 can be determined, and the candidate sub - list 1 with the highest total revenue value: [Content2, Content1] is selected as the target sub - list of the subtask.
[0131] In this embodiment, for each candidate sub - list, by comprehensively considering the immediate revenue value and the second future revenue value corresponding to the candidate sub - list, the total revenue of the candidate sub - list is obtained, which can effectively avoid falling into a local optimal solution. Thus, it can ensure that the push of the target sub - list can maximize the overall revenue. While performing accurate push for users, it can also take into account the benefits of the push service.
[0132] In some embodiments, predicting the immediate revenue value generated after pushing a candidate sub - list includes: taking the list information of the candidate sub - list as the policy information of the next push policy selected in the current state, and based on the feature set and policy information in the current state, calling the updated immediate revenue prediction model to predict the immediate revenue value generated after pushing the candidate sub - list.
[0133] Among them, the push policy refers to the policy for selecting the target sub - list for pushing, and the policy information includes the list information of the candidate sub - list selected by the push policy. The list information includes, but is not limited to, the content identifiers of each content in the selected candidate sub - list, the arrangement order of each content (i.e., the placement order in this candidate sub - list), and the features of each content.
[0134] Optionally, after the server determines the feature set in the current state, for each candidate sub - list among multiple candidate sub - lists, the server takes this candidate sub - list as the next push policy selected in the current state, and then takes the list information of the candidate sub - list as the policy information of the next push policy.
[0135] The server splices this policy information and the feature set in the current state to obtain the spliced information, and inputs the spliced information into the updated immediate revenue prediction model to predict the immediate revenue that can be generated after pushing this candidate sub - list, so as to obtain the immediate revenue value corresponding to this candidate sub - list.
[0136] Exemplarily, when executing the 3rd subtask, the target sub - list 2 (the latest pushed target sub - list) determined by the 2nd subtask is being pushed, and the current state s refers to the state after target sub - list 2 has been browsed.
[0137] After the server determines the feature set in the current state using the foregoing feature set determination process, the server obtains P candidate sub-lists, that is, obtains P push strategies, and obtains the strategy set {π}. For the push strategies in the set (j ranges from 1 to P), it means pushing the candidate sub-list j. At this time, the push strategy in the current state s is denoted as . The feature set in the current state s and 's strategy information are used as the input combination and input into the updated immediate revenue prediction model to output the corresponding immediate revenue value .
[0138] It can be understood that the other P - 1 candidate sub-lists are all processed in a similar manner to obtain the corresponding immediate revenue values. Thus, subsequently, based on the P immediate revenue values and P second future revenue values, a candidate sub-list with the maximum total revenue can be selected from the P candidate sub-lists. For example, candidate sub-list 1. At this time, the candidate sub-list is used as the target sub-list corresponding to the third sub-task.
[0139] In this embodiment, the list information of the candidate sub-list is used as the strategy information of the next push strategy selected in the current state. In this way, based on the feature set and strategy information in the current state, by invoking the updated immediate revenue prediction model, the immediate revenue value generated after pushing the candidate sub-list can be accurately predicted, presenting the short-term revenue situation that the candidate sub-list can bring in an intuitive and accurate form.
[0140] In some embodiments, after predicting the push of the candidate sub-list, the second future revenue value determined by the remaining content in the candidate content set includes: regarding the features of the pushed content and the features of the content in the candidate sub-list as the foregoing features in the next state, where the next state refers to the state where the candidate sub-list is viewed; regarding the features of the remaining content in the candidate content set as the following features in the next state, and obtaining the object features of the content recipient in the next state; determining the feature set in the next state based on the foregoing features, the following features, and the object features; and based on the feature set in the next state, invoking the updated second future revenue prediction model to predict the second future revenue value determined by the remaining content in the candidate content set after pushing the candidate sub-list.
[0141] Among them, as mentioned above for the remaining content, it can be known that the remaining content is obtained by removing the content in the candidate sub-list from the unpushed content, so as to estimate the future revenue brought by the remaining content after pushing the candidate sub-list. The content for future estimation of the second future revenue value is different from that of the first future revenue value mentioned above, so the two are not equivalent.
[0142] Among them, the second future revenue prediction model is used to predict future revenues and is a model constructed based on a neural network. The first future revenue prediction model may be the same as or different from the second future revenue prediction model. Exemplarily, when the first future revenue prediction model is not used to output resource consumption, it is the same as the second future revenue prediction model; when the first future revenue prediction model is also used to output resource consumption, it is different from the second future revenue prediction model.
[0143] Optionally, for each candidate sub-list, the server uses the features of the content that has been pushed and the content in this candidate sub-list as the context features in the next state. The server obtains the content that has not been pushed for the subtask, removes the content in this candidate sub-list from the content that has not been pushed to obtain the remaining content, and the server uses the features of the remaining content as the context features in the next state. The server uses the object features in the current state as the object features in the next state.
[0144] The server constructs a feature set in the next state based on the context features, context features, and object features; inputs the feature set in the next state into the updated second future revenue prediction model, predicts the second future revenue value determined by the remaining content in the candidate content set after pushing the candidate sub-list, and outputs it.
[0145] Exemplarily, when the current state is s and the next state is , after calling the updated second future revenue prediction model, the second future revenue value is output.
[0146] Exemplarily, as Figure 5 shows, it is a schematic diagram of the process of pushing content in an embodiment. Referring to Figure 5 , among them, the user, as the content recipient, logs in to the target application with the target account, and the target application is an application with content push function. The terminal logged in by the target account is the content receiving terminal. The whole process involves the interaction between two end-sides: the server and the content receiving end.
[0147] The following will combine the foregoing to understand the entire technical concept of content pushing:
[0148] Step 5.1: The content terminal sends a push request to the server. In response to the push request, the server performs recall, rough ranking, and fine ranking in sequence based on the content pool to perform content screening, obtains a candidate content set including the content that has not been pushed, and performs multiple subtasks based on the content candidate set.
[0149] Step 5.2: For each subtask, before executing the subtask, the server conducts an intelligent computing power evaluation phase, that is, based on the characteristics of the content that has not been pushed, it determines whether the future revenue evaluation passes. If it passes, then execute Step 5.3 below; if it does not pass, then execute Step 5.5 below.
[0150] Specifically, the server takes the characteristics of the content that has been pushed as the above - text characteristics in the current state. The current state refers to the state where the latest pushed target sub - list has been viewed. It takes the characteristics of the content that has not been pushed in the candidate content set as the below - text characteristics in the current state, and obtains the object characteristics of the content recipient in the current state. Based on the above - text characteristics, below - text characteristics, and object characteristics, it constructs a feature combination in the current state, and inputs the feature combination into the updated first future revenue prediction model to obtain the first future revenue value.
[0151] The server determines whether the first future revenue value reaches the preset revenue threshold. If it reaches, it determines that the future revenue evaluation passes; if it does not reach, it determines that the future revenue evaluation does not pass.
[0152] Step 5.3: The server conducts the subtask execution phase, that is, it starts the subtask to generate a target subtask corresponding to the subtask. That is, based on the content that has not been pushed, it searches for multiple candidate sub - lists, and selects the candidate sub - list with the largest total revenue from the multiple candidate sub - lists and determines it as the target subtask.
[0153] Specifically, the server obtains the interaction data (real - time feedback information) about the pushed target sub - list sent by the content receiving terminal. Based on the content that has not been pushed and the interaction information, it uses beam search to generate a search tree, and determines multiple candidate sub - lists corresponding to the subtask according to the search tree. The determination steps of the search beam are detailed above.
[0154] For each candidate sub - list, the server takes the list information of the candidate sub - list as the policy information of the next push policy selected in the current state, and based on the feature set and policy information in the current state, calls the updated immediate revenue prediction model to predict the immediate revenue value generated after pushing the candidate sub - list.
[0155] The server takes the characteristics of the content that has been pushed and the characteristics of the content in the candidate sub - list as the above - text characteristics in the next state. The next state refers to the state where the candidate sub - list has been viewed; it takes the characteristics of the remaining content in the candidate content set as the below - text characteristics in the next state, and obtains the object characteristics of the content recipient in the next state; based on the above - text characteristics, below - text characteristics, and object characteristics, it determines the feature set in the next state; based on the feature set in the next state, it calls the updated second future revenue prediction model to predict the second future revenue value determined by the remaining content in the candidate content set after pushing the candidate sub - list.
[0156] The server performs weighted summation according to the weights of the immediate revenue value and the second future revenue value respectively, obtains the total revenue value corresponding to the candidate sub-list, determines the maximum total revenue value from the total revenue values corresponding to multiple candidate sub-lists, and uses the candidate sub-list corresponding to the maximum total revenue value as the target sub-list corresponding to the sub-task, and pushes the target sub-list to the content receiving terminal where the user is located.
[0157] Step 5.4: The server determines whether the total number of contents in the pushed target sub-list has reached the total push quantity. If it has reached, the server stops pushing contents for this push request. If it has not reached, the server takes each content in the target sub-list as the pushed content, updates the unpushed contents in the candidate content set based on the target sub-list, returns to judge whether the stop condition is reached based on the unpushed contents in the candidate content set and continues to execute until the total number of contents in the pushed target sub-list reaches the total push quantity.
[0158] Step 5.5: End the response to this push request.
[0159] In this embodiment, to predict the future revenue after pushing the candidate sub-list, assume that the candidate sub-list is the latest pushed sub-list and the corresponding state is the next state. In this way, based on the above-text features, below-text features, and object features in the next state, by calling the updated second future revenue prediction model, comprehensively considering the details of the three dimensions of context and object, the future revenue of the next state is reasonably and effectively predicted. Thus, the total revenue values brought by pushing each candidate sub-list can be predicted, and the target sub-list that can bring considerable revenue can be selected to achieve the maximization of the overall revenue. While performing precise push for users, the benefits of the push service can also be taken into account.
[0160] In some embodiments, the step of updating the immediate revenue prediction model includes: obtaining an immediate revenue sample, where the immediate revenue sample includes a sample feature set in the first current sample state and sample strategy information of the next push strategy selected in the first current sample state; obtaining a first label corresponding to the immediate revenue sample, where the first label is the true immediate revenue value after pushing the first to-be-pushed sample sub-list targeted by the next push strategy online; and updating the model of the first immediate revenue prediction model to be updated based on the immediate revenue sample and the first label to obtain the updated immediate revenue prediction model.
[0161] Exemplarily, every preset period, or based on actual business requirements, the currently used immediate revenue prediction model is used as the immediate revenue prediction model to be updated, and the process of updating the model of the immediate revenue prediction model to be updated is started.
[0162] The server obtains an instant revenue sample, which includes a set of sample features in the first current sample state and sample strategy information of the next push strategy selected in the first current sample state. Herein, the first current sample state refers to the state in which the first sample sub - list of the current latest push is viewed. The next push strategy refers to the strategy of pushing an arbitrarily selected first sample sub - list to be pushed, and the strategy information refers to the list information of this first sample sub - list to be pushed.
[0163] After the server counts the instant revenue brought by the first sample sub - list to be pushed after it is pushed online, it obtains the true instant revenue value.
[0164] The server inputs the instant revenue sample into the instant revenue prediction model to be updated, outputs the predicted instant revenue value, and updates the instant revenue prediction model to be updated according to the difference between the predicted instant revenue value and the true instant revenue value, so as to obtain an updated instant revenue prediction model.
[0165] In this embodiment, by returning a randomly selected push strategy to be allowed online and counting the instant revenue brought after the push strategy is pushed online to obtain the true instant revenue value, thus, the true instant revenue value can be used as the first label to update the instant revenue prediction model timely and accurately.
[0166] In some embodiments, as Figure 6 shown, it is a schematic flowchart of the update steps of the second future revenue prediction model in an embodiment. The update steps of the second future revenue prediction model include:
[0167] Step S602: Obtain a future revenue sample, which includes a set of sample features in the second current sample state.
[0168] Herein, the second current sample state refers to the state in which the second sample sub - list of the current latest push is viewed.
[0169] Step S604: Obtain the true instant revenue value after the second sample sub - list to be pushed, which the next push strategy targets, is pushed online, and obtain a set of sample features in the next sample state. The next sample state is the state in which the second sample sub - list to be pushed is viewed.
[0170] Step S606: Based on the set of sample features in the second current sample state, call the second future revenue prediction model to be updated and predict the current predicted future revenue value in the second current sample state.
[0171] Step S608: Based on the sample feature set in the next sample state, call the second future revenue prediction model to be updated to predict the next predicted future revenue value in the next sample state. Based on the true immediate revenue value, the current predicted future revenue value, and the next predicted future revenue value, determine the current true future revenue value in the second current sample state. The current true future revenue value serves as the second label corresponding to the future revenue sample.
[0172] Exemplarily, let the second current sample state be , and the current predicted future revenue value in the second current sample state ;
[0173] The selected second sub-list of samples to be pushed is , and the true immediate revenue value obtained after going online is , the next sample state is , and the next predicted future revenue value in the next sample state .
[0174] Obtain and their respective weights, and according to the weights, perform weighted summation on and to obtain a sum value. Calculate the difference between this sum value and to obtain an error value.
[0175] Obtain an error coefficient, and superimpose the product of the error coefficient and the error value with to obtain the current true future revenue value.
[0176] It should be noted that the future revenue in the current state can be obtained by fusing the immediate revenue corresponding to this push strategy and the future revenue in the next state (the state obtained by executing this push strategy).
[0177] Thus, the above sum value reflects the future revenue of the second current sample state, and the subsequent obtained difference can obtain the theoretical error value. In this way, based on the product of this error value and the pre-determined error coefficient, the actual error value after going online is determined. Thus, based on this actual error value and the current predicted future revenue value predicted by the second future revenue prediction model to be updated in the second current sample state, the actual future revenue in the second current sample state can be determined.
[0178] Step S610: Update the second future revenue prediction model based on the future revenue samples and the second label to obtain an updated second future revenue prediction model.
[0179] Exemplarily, the difference between the current predicted future revenue value in the second current sample state and the current actual future revenue value in the second current sample state is used to update the second future revenue prediction model, resulting in an updated second future revenue prediction model. Exemplarily, this difference can be the product of the error coefficient and the error value mentioned above.
[0180] It should be noted that if the first future revenue prediction model is equal to the second future revenue prediction model, then the updated second future revenue prediction model can be directly used as the updated second future revenue prediction model.
[0181] If they are not equal, the update execution steps for the first future revenue prediction model are as follows:
[0182] Obtain another future revenue sample, which includes a sample feature set in the third current sample state; obtain the actual immediate revenue value after the online push of the third sub-list of samples to be pushed for the next push strategy, and obtain the sample feature set in the next sample state, where the next sample state is the state in which the third sub-list of samples to be pushed is viewed; based on the sample feature set in the third current sample state, call the first future revenue prediction model to be updated and output the current predicted future revenue value and the corresponding predicted resource consumption in the third current state; based on the sample feature set in the next sample state, call the first future revenue prediction model to be updated and output the next predicted future revenue value and the corresponding predicted resource consumption in the next sample state. Based on the actual immediate revenue value, the current predicted future revenue value, and the next predicted future revenue value, determine the current actual future revenue value in the third current sample state, and use the current actual future revenue value as the third label corresponding to the future revenue sample. Statistically calculate the actual resource consumption when generating the third sub-list of samples to be pushed, and use the actual resource consumption as the fourth label; based on the future revenue sample, the third label, and the fourth label, update the first future revenue prediction model to obtain an updated first future revenue prediction model.
[0183] In this embodiment, by using the actual immediate revenue value actually determined after the online push of the second sub-list of samples to be pushed, combined with the predicted future revenue values corresponding to the second current sample state and the next sample state respectively, the actual future revenue value in the second current sample state can be accurately obtained to serve as the corresponding second label. Thus, the second future revenue prediction model can be updated in a timely and accurate manner based on the second label.
[0184] The present application also provides an application scenario, which applies the above content push method. The application of the content push method in this application scenario is as follows: In the note push scenario (where a note is a diversified content carrier created and published by a creator, which can integrate at least one form of expression such as text, pictures, and videos), in response to a push request, a candidate note set related to the push request is determined; the total number of pushes that match the push request is determined; multiple subtasks are executed based on the candidate note set, and the target sub-lists obtained from each execution of the subtask are pushed until the stop condition is reached and the process stops; the number of notes in the target sub-list is less than the total number of pushes; wherein, the execution steps of any one subtask include: searching for multiple candidate sub-lists based on the notes in the candidate note set that have not been pushed, and screening out the target sub-list from the multiple candidate sub-lists.
[0185] Of course, it is not limited to this. The content push method provided by the present application can also be applied in other scenarios. For example, in the audio-visual push scenario, the corresponding content is audio-visual; for another example, in the commodity interaction scenario, the corresponding content is commodities; for another example, in the advertisement push scenario, the corresponding content is advertisements.
[0186] In a specific embodiment, it is described in combination with the following code example:
[0187] 1: for each subtask i do / / For subtask i
[0188] 2: s ← current state / / Current state s
[0189] 3: if <preset revenue threshold then / / Intelligent computing power evaluation
[0190] 4: Initiate a new push request
[0191] 5: return
[0192] 6: end if
[0193] 7: ← null / / Initialize the best push strategy is null (empty)
[0194] 8: ← -inf / / Initialize the best revenue is -inf (regarded as the initial value)
[0195] 9: {π} ← beam search / / Obtain a set containing multiple candidate sub-lists (push strategy set)
[0196] 10: for ∈{π} do / / Select the j-th candidate list from multiple candidate sub-lists
[0197] 11: ← Adopt The new state after / / Select The next state after
[0198] 12: / / Determine the total revenue after selection The total revenue after
[0199] 13: if V > then / / Push policy iteration
[0200] 14: ← V / / Update V to the best revenue
[0201] 15: ← / / The Update to the best push policy
[0202] 16: end if
[0203] 17: end for
[0204] 18: Randomly select a π with probability , and with probability 1 - select , and return the result Lw corresponding to the selected push policy to the online
[0205] 19: Obtain the next state corresponding to Lw , and the online immediate revenue
[0206] 20: / / Value iteration, use the online immediate revenue as the true immediate revenue value brought by pushing Lw in the current sample state
[0207] 21: / / Use the updated as the second label of the second future revenue prediction model
[0208] 22: end for
[0209] Specifically: Step 1: Before executing the target sub-list of the i-th sub-task, first input the feature set corresponding to the current state s into the first future revenue prediction model , and obtain the corresponding future revenue , if If it is less than the preset future threshold, stop pushing the current push request. If the future revenue is greater than or equal to the preset future threshold, start generating the target sub-list of the i-th sub-task. (The corresponding code is the code from line 1 to line 6, and the specific implementation process refers to the previous text).
[0210] Step 2: When starting to execute the i-th sub-task, initialize the best push strategy and the best revenue, and initialize the best push strategy to be null, and the best revenue value to be -inf; then, use the beam search method mentioned above to obtain a set {π} containing multiple candidate sub-lists.
[0211] Step 3: Calculate the total revenue corresponding to each candidate sub-list in {π} in turn. Taking the j-th candidate sub-list as an example, take as the next push strategy, take the state of browsing the j-th candidate sub-list as the next state (new state), and based on the feature set in the next state, call the second future revenue prediction model to determine the second future revenue value ( ).
[0212] Input the feature set in the current state and the strategy information of the next push strategy selected in the current state s into the updated immediate revenue prediction model to obtain the immediate revenue value . Let the weights of the immediate revenue value and the second future revenue value be 1 and respectively, and obtain the total revenue value V through weighted summation.
[0213] If the total revenue value V is greater than the current best revenue value , update the total revenue value V to the current best revenue value , and take the corresponding as the best push strategy.
[0214] Continue to estimate the total revenue of the j+1 candidate sub-list until the total revenue value of the last candidate sub-list is estimated. Take the candidate sub-list corresponding to the best push strategy obtained for all candidate sub-lists as the target sub-list corresponding to the i-th sub-task.
[0215] The above steps 2-3 correspond to the code from line 7 to line 17 above, that is, the process of determining the target sub-list corresponding to the previous text.
[0216] Step 4: Update the above-mentioned models at preset time intervals. The following describes the process of jointly updating the above-mentioned immediate revenue prediction model and the second future revenue prediction model:
[0217] Select a π randomly with the probability of , and select with the probability of 1 - . Construct an updated sample from the candidate sub - list Lw corresponding to the push strategy w randomly selected from the push strategy set π. The updated sample includes the sample feature set under the current sample state and the sample strategy information of the selected push strategy w under the current sample state. Among them, the current sample state is the viewed state of the pushed target sub - list.
[0218] Then, after putting the push strategy w online, obtain the online immediate revenue , which is used as the first label for updating the immediate revenue prediction model, and obtain the sample feature set under the next sample state . The next sample state is the viewed state of the candidate sub - list Lw.
[0219] Input the sample feature set under the current sample state and the sample strategy information of the selected push strategy w under the current sample state into the immediate revenue update model to obtain the predicted immediate revenue value.
[0220] Based on the sample feature set under the current sample state , call the second future revenue prediction model to be updated to predict the current predicted future revenue value under the current state.
[0221] Based on the sample feature set under the next sample state , call the second future revenue prediction model to be updated to predict the next predicted future revenue value under the next sample state. Based on the true immediate revenue value R, the current predicted future revenue value and the next predicted future revenue value , determine the current true future revenue value under the current sample state . The current true future revenue value is used as the second label for updating the second future revenue prediction model.
[0222] Update the immediate revenue prediction model according to the difference between the first label and the predicted immediate revenue value to obtain the updated immediate revenue prediction model.
[0223] Update the second future revenue prediction model according to the difference between the second label and the current predicted future revenue value to obtain the updated second future revenue prediction model.
[0224] Among them, the update of the first future revenue prediction model can be found in the previous text.
[0225] Step 4 corresponds to lines 18 to 22 of the above code.
[0226] In this embodiment, by decomposing the total task of generating the total number of pushes into multiple subtasks and respectively determining the target sub - list of each subtask. In this way, for each subtask, a plurality of candidate sub - lists with smaller growth lengths are selected from the selection range less than or equal to the candidate content set, rather than selecting a list with a larger growth length from the total number of pushes, which greatly reduces the search space and makes it very easy to screen out a more suitable target sub - list from this search space. Thus, the content push effect is improved. And when screening out a target sub - list from multiple search spaces, it combines the total revenue of each candidate sub - list to maximize the total revenue of the target sub - list, rather than maximizing the immediate revenue, effectively avoiding falling into a local optimal solution during rearrangement. In addition, by predicting in advance whether the future revenue that the un - pushed content can bring reaches the preset revenue threshold before executing each subtask, it is possible to reduce low - efficiency rearrangement and effectively save the push computing power.
[0227] It should be understood that although the steps in the flowcharts involved in the above - mentioned embodiments are shown in sequence according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above - mentioned embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps in other steps.
[0228] Based on the same inventive concept, the embodiments of the present application also provide a content push device for implementing the content push method involved above. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the content push device provided below can refer to the limitations on the content push method in the above text, and will not be repeated here.
[0229] In an exemplary embodiment, as Figure 7 shown, a content push device 700 is provided, including: a set determination module 702, a quantity determination module 704, and a list push module 706, where:
[0230] A set determination module 702, configured to determine a set of candidate content related to a push request in response to the push request;
[0231] A quantity determination module 704, configured to determine the total number of pushes that match the push request;
[0232] A list push module 706, configured to perform multiple subtasks based on the set of candidate content, and push the target sub - lists obtained from each execution of the subtask until the stop condition is reached and then stop; the number of contents in the target sub - list is less than the total number of pushes; wherein, the execution steps of any one subtask include: searching for multiple candidate sub - lists based on the un - pushed content in the set of candidate content, and screening out the target sub - list from the multiple candidate sub - lists.
[0233] In some embodiments, the stop condition includes at least one of the following: the first future benefit value determined by the un - pushed content in the set of candidate content does not reach the preset benefit threshold; the total number of contents in the pushed target sub - list reaches the total number of pushes.
[0234] In some embodiments, the apparatus further includes a benefit determination module, configured to call the updated first future benefit prediction model based on the characteristics of the un - pushed content in the set of candidate content to obtain the first future benefit value.
[0235] In some embodiments, the benefit determination module is configured to use the characteristics of the pushed content as the above - text characteristics in the current state, where the current state refers to the state where the latest pushed target sub - list is viewed; use the characteristics of the un - pushed content in the set of candidate content as the below - text characteristics in the current state, obtain the object characteristics of the content recipient in the current state; construct a feature combination in the current state based on the above - text characteristics, below - text characteristics, and object characteristics, and input the feature combination into the updated first future benefit prediction model to obtain the first future benefit value.
[0236] In some embodiments, the list push module 706 is further configured to generate a search tree based on the un - pushed content in the set of candidate content and the interaction data generated by the pushed target sub - lists, and determine multiple candidate sub - lists based on the search tree.
[0237] In some embodiments, the list push module 706 is further configured to, for each candidate sub - list, predict the immediate benefit value generated after pushing the candidate sub - list; for each candidate sub - list, predict the second future benefit value determined by the remaining content in the set of candidate content after pushing the candidate sub - list; fuse the immediate benefit value and the second future benefit value to obtain the total benefit value corresponding to the candidate sub - list; select the candidate sub - list with the largest total benefit value as the target sub - list.
[0238] In some embodiments, the list push module 706 is further configured to use the list information of the candidate sub - list as the policy information of the next push policy selected in the current state, and based on the feature set and the policy information in the current state, call the updated immediate revenue prediction model to predict the immediate revenue value generated after pushing the candidate sub - list.
[0239] In some embodiments, the device further includes a model update module. The model update module is configured to obtain an immediate revenue sample, where the immediate revenue sample includes a sample feature set in a first current sample state and sample policy information of the next push policy selected in the first current sample state; obtain a first label corresponding to the immediate revenue sample, where the first label is the true immediate revenue value after pushing the first sub - list of samples to be pushed targeted by the next push policy online; and update the first immediate revenue prediction model to be updated based on the immediate revenue sample and the first label to obtain an updated immediate revenue prediction model.
[0240] In some embodiments, the list push module 706 is further configured to use the features of the pushed content and the features of the content in the candidate sub - list as the above - text features in the next state, where the next state refers to the state where the candidate sub - list is browsed; use the features of the remaining content in the candidate content set as the below - text features in the next state, and obtain the object features of the content recipient in the next state; determine the feature set in the next state based on the above - text features, below - text features, and object features; and based on the feature set in the next state, call the updated second future revenue prediction model to predict the second future revenue value determined by the remaining content in the candidate content set after pushing the candidate sub - list.
[0241] In some embodiments, the model update module is configured to obtain a future revenue sample, where the future revenue sample includes a sample feature set in a second current sample state; obtain the true immediate revenue value after pushing the second sub - list of samples to be pushed targeted by the next push policy online, and obtain the sample feature set in the next sample state, where the next sample state is the state where the second sub - list of samples to be pushed is browsed; based on the sample feature set in the second current sample state, call the second future revenue prediction model to be updated to predict the current predicted future revenue value in the second current sample state; based on the sample feature set in the next sample state, call the second future revenue prediction model to be updated to predict the next predicted future revenue value in the next sample state, and determine the current true future revenue value in the second current sample state based on the true immediate revenue value, the current predicted future revenue value, and the next predicted future revenue value, where the current true future revenue value is used as the second label corresponding to the future revenue sample; and update the second future revenue prediction model based on the future revenue sample and the second label to obtain an updated second future revenue prediction model.
[0242] Each module in the above content pushing device can be implemented in whole or in part by software, hardware, or a combination thereof. Each of the above modules can be embedded in the processor of a computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0243] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a content pushing method.
[0244] Those skilled in the art can understand that Figure 8 the structure shown in
[0245] is only a block diagram of a part of the structure related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0246] In an embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0246] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0247] In an embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0248] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0249] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.
[0250] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0251] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A content push method, characterized in that: The method comprises: In response to a push request, determining a set of candidate content related to the push request; Determining the total number of pushes that match the push request; Executing multiple subtasks based on the candidate content set, and pushing the target sublist obtained by executing the subtask each time until a stop condition is met; the number of contents in the target sublist is less than the total number of contents pushed; The execution steps of any one subtask include: searching for a plurality of candidate sublists based on the unpushed content in the candidate content set, and selecting a target sublist from the plurality of candidate sublists.
2. The method according to claim 1, characterized in that The stop condition includes at least one of the following: The first future revenue value determined by the unpushed content in the candidate content set does not reach a preset revenue threshold; The total number of contents in the pushed target sub-list reaches the total number of pushes.
3. The method according to claim 2, characterized in that The step of determining the first future benefit value comprises: Based on the features of the unpushed content in the candidate content set, the updated first future revenue estimation model is called to obtain a first future revenue value.
4. The method according to claim 3, characterized in that The step of calling the updated first future revenue estimation model based on the features of the unpushed content in the candidate content set to obtain the first future revenue value includes: The characteristics of the pushed content are used as the above characteristics in the current state, where the current state refers to the state in which the most recently pushed target sublist is browsed; Using the features of the unpushed content in the candidate content set as the context features in the current state, and acquiring the object features of the content receiver in the current state; Based on the above features, the below features, and the object features, a feature combination in the current state is constructed, and the feature combination is input into an updated first future benefit estimation model to obtain the first future benefit value.
5. The method according to claim 1, characterized in that The step of searching for multiple candidate sublists based on the unpushed content in the candidate content set includes: A search tree is generated based on the unpushed content in the candidate content set and the interaction data generated by the pushed target sublist, and a plurality of candidate sublists are determined based on the search tree.
6. The method according to claim 1, characterized in that The step of selecting a target sublist from the plurality of candidate sublists includes: For each candidate sublist, predict the immediate benefit value generated after pushing the candidate sublist; For each candidate sub-list, predicting a second future revenue value determined by the remaining contents in the candidate content set after the candidate sub-list is pushed; Merging the instant benefit value and the second future benefit value to obtain a total benefit value corresponding to the candidate sublist; Select the candidate sublist with the largest total benefit value as the target sublist.
7. The method according to claim 6, characterized in that The instant benefit value generated after the prediction and pushing of the candidate sublist includes: The list information of the candidate sublist is used as the strategy information of the next push strategy selected in the current state, and based on the feature set in the current state and the strategy information, the updated instant benefit estimation model is called to predict the instant benefit value generated after pushing the candidate sublist.
8. The method according to claim 7, characterized in that The instant revenue estimation model updating step comprises: Acquire an instant profit sample, where the instant profit sample includes a sample feature set under a first current sample state and sample strategy information of a next push strategy selected under the first current sample state; Acquire a first label corresponding to the instant revenue sample, where the first label is a real instant revenue value after the first to-be-pushed sample sublist targeted by the next push strategy is pushed online; Based on the instant profit sample and the first label, the first instant profit estimation model to be updated is updated to obtain an updated instant profit estimation model.
9. The method according to claim 6, characterized in that The predicting of a second future revenue value determined by the remaining content in the candidate content set after the candidate sublist is pushed includes: The features of the pushed content and the features of the content in the candidate sublist are used as the above-context features in the next state, where the next state refers to the state in which the candidate sublist is browsed; Using the features of the remaining content in the candidate content set as the context features in the next state, and obtaining the object features of the content receiver in the next state; Determine a feature set in a next state based on the preceding feature, the following feature, and the object feature; Based on the feature set in the next state, the updated second future revenue prediction model is called to predict the second future revenue value determined by the remaining content in the candidate content set after the candidate sub-list is pushed.
10. The method according to claim 9, characterized in that The second future earnings estimation model updating step comprises: Acquire a future income sample, wherein the future income sample includes a sample feature set in a second current sample state; Obtaining the real instant revenue value of the second sublist of samples to be pushed targeted by the next push strategy after being pushed online, and obtaining a sample feature set in the next sample state, where the next sample state is a state in which the second sublist of samples to be pushed is browsed; Based on the sample feature set under the second current sample state, calling the second future profit estimation model to be updated to predict the current predicted future profit value under the second current sample state; Based on the sample feature set under the next sample state, calling the second future profit estimation model to be updated, predicting the next predicted future profit value under the next sample state, and determining the current real future profit value under the second current sample state based on the real instant profit value, the current predicted future profit value, and the next predicted future profit value, and the current real future profit value is used as the second label corresponding to the future profit sample; Based on the future profit sample and the second label, the second future profit estimation model is updated to obtain an updated second future profit estimation model.
11. A content push device, characterized in that: The device comprises: A set determination module, configured to determine, in response to a push request, a set of candidate content related to the push request; A quantity determination module, used to determine the total number of pushes that match the push request; A list push module is used to execute multiple subtasks based on the candidate content set, and push the target sublist obtained by executing the subtask each time until the stop condition is reached; the number of contents in the target sublist is less than the total number of pushes; wherein the execution steps of any subtask include: searching for multiple candidate sublists based on the content that has not been pushed in the candidate content set, and filtering out the target sublist from the multiple candidate sublists.
12. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Cited By
Content pushing method and apparatus, computer device, readable storage medium, and program product
WO2026166396A1