Method, device and equipment for obtaining push strategy and storage medium

By acquiring historical feedback data and state transition matrices from target users, and combining them with preset short-term and long-term values, a binary search method is used to determine the optimal push strategy. This solves the problem of poor information push performance in e-commerce scenarios and achieves more efficient user value assessment and push strategy selection.

CN114756756BActive Publication Date: 2026-04-10HANGZHOU ALIBABA INT INTERNET IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU ALIBABA INT INTERNET IND CO LTD
Filing Date
2022-04-20
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, the effectiveness of information push in e-commerce scenarios is poor, mainly because it focuses on analyzing the short-term value of users (such as click-through rates) and fails to effectively consider the long-term value of users, resulting in ineffective push strategies.

Method used

By acquiring historical feedback data from target users for different push strategies, the state transition matrix and preset short-term value are determined. Combined with the benchmark long-term value, the target push strategy with the maximum short-term and long-term value is selected from the set of push strategies. The benchmark long-term value is quickly solved using the bisection method to determine the optimal push strategy.

Benefits of technology

It improves the reliability of user value assessment and the effectiveness of push strategies, ensuring optimal push results in both the short and long term.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114756756B_ABST
    Figure CN114756756B_ABST
Patent Text Reader

Abstract

The application provides an acquisition method and device of a push strategy, equipment and a storage medium. The acquisition method of the push strategy comprises the following steps: obtaining historical feedback data of a target user for different push strategies; firstly, determining a state transition matrix of the target user according to the historical feedback data, wherein the state transition matrix is used for indicating the probability value of state change of the target user under different push strategies; and secondly, determining a target push strategy with the maximum long-term and short-term value from a push strategy set according to the state transition matrix of the target user, the preset short-term value corresponding to the state change under different push strategies and the benchmark long-term value. Compared with the prior art, the reliability of the user value evaluation is higher and the selected push strategy is better because the value of the user in the short term and the long term is considered simultaneously.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a push strategy acquisition method and device, equipment and a storage medium. BACKGROUND

[0002] Life time value (LTV) refers to the cumulative value that a user can bring to a target for a lifetime in the case that the user may exist for a long time. Taking an e-commerce scenario as an example, a platform pushes information to a user, the user may browse the pushed information, may set to close the pushed information, or may not do anything (i.e. ignore the pushed information). Therefore, in the e-commerce scenario, the life time value of the user in the pushed information can be the value that the user will contribute for a lifetime, rather than the opening rate at the moment. At present, information pushing is mainly performed by analyzing the short-term value (such as user click rate and the like) of the user, and the pushing effect is not good. SUMMARY

[0003] Embodiments of the present application provide a push strategy acquisition method, device, equipment and storage medium to improve the pushing effect.

[0004] A first aspect of embodiments of the present application provides a push strategy acquisition method, comprising:

[0005] obtaining historical feedback data of a target user for different push strategies, the historical feedback data being used to indicate behavior statistical data of the target user for push information under different push strategies;

[0006] determining a state transition matrix of the target user according to the historical feedback data, the state transition matrix being used to indicate a probability value of state change of the target user under different push strategies;

[0007] obtaining a preset short-term value corresponding to the state change under different push strategies;

[0008] determining a target push strategy with maximum long-short-term value from a push strategy set according to the state transition matrix of the target user, the preset short-term value and a benchmark long-term value.

[0009] In an optional embodiment of the first aspect of the present application, the long-short-term value comprises a user long-term value of user state retention to the next period and a user short-term value of the current period, and the user long-term value is determined according to the product of the probability value of user state retention to the next period and the benchmark long-term value.

[0010] In an optional embodiment of the first aspect of the present application, the determining, according to the state transition matrix of the target user, the preset short-term value, and the reference long-term value, a target push strategy in which a long-short-term value is maximum from the set of push strategies comprises:

[0011] selecting a middle value from the preset interval of the long-term value as an initial value of the long-term value;

[0012] determining the reference long-term value from the preset interval of the long-term value based on the initial value of the long-term value by using a dichotomy method;

[0013] determining, as the target push strategy, a push strategy in which a long-short-term value is maximum based on the reference long-term value, the state transition matrix of the target user, and the preset short-term value.

[0014] In an optional embodiment of the first aspect of the present application, the determining, according to the state transition matrix of the target user, the preset short-term value, and the reference long-term value, a target push strategy in which a long-short-term value is maximum from the set of push strategies comprises:

[0015] determining, from the set of push strategies, a first push strategy in which a long-short-term value is maximum based on the initial value of the long-term value, the state transition matrix of the target user, and the preset short-term value;

[0016] if the long-short-term value corresponding to the first push strategy is not equal to the initial value of the long-term value, shortening the preset interval of the long-term value;

[0017] selecting a middle value from the shortened preset interval of the long-term value as a new long-term value;

[0018] determining, from the set of push strategies, a second push strategy in which a long-short-term value is maximum based on the new long-term value, the state transition matrix of the target user, and the preset short-term value;

[0019] if the long-short-term value corresponding to the second push strategy is equal to the new long-term value, determining the new long-term value as the reference long-term value.

[0020] In an optional embodiment of the first aspect of the present application, the determining, according to the state transition matrix of the target user, the preset short-term value, and the reference long-term value, a target push strategy in which a long-short-term value is maximum from the set of push strategies comprises:

[0021] determining, as the target push strategy, the second push strategy in which a long-short-term value is maximum based on the new long-term value, the state transition matrix of the target user, and the preset short-term value.

[0022] In an optional embodiment of the first aspect of the present application, if the long-short term value corresponding to the first push strategy is not equal to the initial value of the long term value, the preset interval of the long term value is shortened, comprising:

[0023] If the long-short term value corresponding to the first push strategy is greater than the initial value of the long term value, the left end point value of the preset interval of the long term value is replaced by the initial value; or

[0024] If the long-short term value corresponding to the first push strategy is less than the initial value of the long term value, the right end point value of the preset interval of the long term value is replaced by the initial value.

[0025] In an optional embodiment of the first aspect of the present application, the state transition matrix of the target user is determined according to the historical feedback data, comprising:

[0026] A first statistical value of the target user having a click operation on a third push strategy within a preset time period is obtained, and a probability value of the target user having a state change under the third push strategy is determined according to the first statistical value; or

[0027] A second statistical value of the target user having an unsubscribe operation on the third push strategy within a preset time period is obtained, and a probability value of the target user having a state change under the third push strategy is determined according to the second statistical value; or

[0028] A third statistical value of the target user having no operation on the third push strategy within a preset time period is obtained, and a probability value of the target user having no state change under the third push strategy is determined according to the third statistical value.

[0029] The third push strategy is any push strategy in the push strategy set.

[0030] In an optional embodiment of the first aspect of the present application, the target push strategy comprises at least one of the following:

[0031] Sending push information at a specified time within a preset time period;

[0032] The number of times of sending push information within a preset time period;

[0033] The information type of sending push information within a preset time period;

[0034] Sending push information containing specified content within a preset time period.

[0035] The second aspect of the embodiment of the present application provides a push strategy acquisition device, comprising:

[0036] The acquisition module is configured to acquire historical feedback data of a target user for different push strategies, the historical feedback data being used to indicate behavior statistics of the target user for push information of different push strategies.

[0037] The processing module is configured to determine a state transition matrix of the target user according to the historical feedback data, the state transition matrix being used to indicate probability values of state changes of the target user under different push strategies.

[0038] The acquisition module is configured to acquire preset short-term values corresponding to state changes under different push strategies.

[0039] The processing module is configured to determine a target push strategy with maximum long-short-term values from a push strategy set according to the state transition matrix of the target user, the preset short-term values, and a benchmark long-term value.

[0040] A third aspect of the embodiments of the present application provides an electronic device, comprising a memory, a processor, and a computer program; the computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of the first aspect of the present application.

[0041] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method according to any one of the first aspect of the present application.

[0042] A fifth aspect of the embodiments of the present application provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the method according to any one of the first aspect of the present application.

[0043] The embodiments of the present application provide a push strategy acquisition method and device, equipment, and a storage medium. The method comprises the following steps: acquiring historical feedback data of a target user for different push strategies; determining a state transition matrix of the target user according to the historical feedback data, the state transition matrix being used to indicate probability values of state changes of the target user under different push strategies; and determining a target push strategy with maximum long-short-term values from a push strategy set according to the state transition matrix of the target user, preset short-term values corresponding to state changes under different push strategies, and a benchmark long-term value. Compared with the prior art, the reliability of user value evaluation is higher, and the effect of the selected push strategy is better, because the values of the user in the short term and the long term are considered simultaneously. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1 The state transition matrix provided by the embodiments of the present application Figure 1 ;

[0045] Figure 2 State transition provided for the embodiment of the present application Figure 2 ;

[0046] Figure 3 State transition provided for the embodiment of the present application Figure 3 ;

[0047] Figure 4 Application scenario diagram of the method for acquiring the push strategy provided for the embodiment of the present application

[0048] Figure 5 Flow diagram of the method for acquiring the push strategy provided for the embodiment of the present application Figure 1 ;

[0049] Figure 6 Flow diagram of the method for acquiring the push strategy provided for the embodiment of the present application Figure 2 ;

[0050] Figure 7 Schematic diagram of the method for acquiring the push strategy provided for the embodiment of the present application

[0051] Figure 8 State transition provided for the embodiment of the present application Figure 4 ;

[0052] Figure 9 Structure diagram of the device for acquiring the push strategy provided for the embodiment of the present application

[0053] Figure 10 Hardware structure diagram of the electronic device provided for the embodiment of the present application. DETAILED DESCRIPTION

[0054] To make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the accompanying drawings for the embodiments of the present application to clearly and completely describe the technical solutions of the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0055] The terms “first”, “second”, and the like in the specification, claims, and above-mentioned drawings of the embodiments of the present application are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0056] It should be understood that the terms “comprising” and “having” as used in this application, and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.

[0057] In the description of the embodiments of this application, the term "correspondence" may indicate that there is a direct or indirect correspondence between two things, or that there is an association between two things, or that there is a relationship of instruction and being instructed, configuration and being configured, etc.

[0058] Before introducing the technical solution of this application, the theoretical knowledge involved in the embodiments of this application will be explained first.

[0059] In the method described in the embodiments of this application, although the state value is defined as a single value, it must be split into two target values ​​in the long-term and short-term directions during the solution process. The following is a macroscopic description of the problem using the single-value version.

[0060] The Markov Decision Process (MDP) quadruple (S, A, P, R) consists of the state (S), action (A), transition probability (P), and reward (R).

[0061] Where, state s∈S: represents a stage in the process, fully encompassing information that influences subsequent decisions; in other words, any trajectory leading to this state is equivalent in subsequent decisions. Action a∈A: represents the behavior that can be chosen when facing a decision, such as in push notifications, the action set is send or not send; in item sorting problems, the action set is a candidate set of items, etc. Transition probability. : Indicates that after taking action a in state s, the state transitions to state s. The probability, reward : Indicates the corresponding instantaneous reward, such as the reward for clicking on item 'a' in the item recommendations.

[0062] Single cycle: This refers to a single, independent experimental process, such as daily or monthly. The following text will use daily as an example of a single cycle.

[0063] intermediate state The state value f(s) i ): Indicates that the state is reached at time t. Then, the expected cumulative return generated in the future under the optimal strategy, where t is implied in In MDP research, this form is also written as... This indicates that in the optimal strategy From The expected return of the trajectory obtained from the start will be uniformly represented by f in this application.

[0064] Initial state (denoted as s) init ): The initial state at the start of a single experiment.

[0065] Lost state (denoted as s) loss ): The state of a user permanently churned, with f(s) loss )=0.

[0066] Survival absorption state (denoted as s) surv ): The final state of a user's survival in a single experiment, under the memoryless setting, has f(s) surv )=f(s init ).

[0067] State Transition (within a single period, only the number of push notifications is selected, a special case for push messages): Given the above state definitions, assuming only the number of push notifications is selected within a single period, and the server wants to send a maximum of 3 push notifications to attract users, how do the above states relate, given the rewards for sending 1 to 3 push notifications and the probability of churn? This can be addressed through… Figure 1 describe.

[0068] Figure 1 State transitions provided in the embodiments of this application Figure 1 .like Figure 1 As shown, this example only has 3 intermediate states, representing the cases where 1, 2, and 3 messages have been sent and the message has not been closed. Figure 2 Arrows 1 and 2 appear in pairs, representing two possible state transitions resulting from sending a message (a decision-making action): user churn or not churn, each generating the expected number of clicks; arrow 3 represents the state transition where no more messages are sent that day (another decision-making action), and the user remains in the next cycle to generate long-term benefits. It can be observed that, ignoring long-term benefits, the above modeling is a special case of the stochastic shortest path (SSP) problem.

[0069] State transition (within a single cycle, selecting the number of messages and time, a special case for push notifications): This method extends the above approach by incorporating time (here, time represents the decision-making action) into the state. Assuming the sending time of the first push notification < the sending time of the second push notification < the sending time of the third push notification, the relationships between these states can be established through... Figure 2 describe.

[0070] Figure 2 State transitions provided in the embodiments of this application Figure 2 .like Figure 3As shown in the example, the state in row t and column m represents the state where m push notifications were sent within the previous t time units. At each state, the user can choose to send or not send at the current time. Not sending will transition to (t+1, m) along arrow 4, while sending has a probability of transitioning to a non-churned state along arrow 2, or a probability of transitioning to a churned state along arrow 1. Similarly, without considering long-term benefits, the transition graph forms a directed acyclic graph (DAG), allowing the optimal strategy to be deduced by working backward from its topological order.

[0071] State transition (within the lifetime, general form): Let the directed graph G=(V,E) be the state transition graph:

[0072]

[0073]

[0074] in, and It is a state transition due to successful decision-making. This is a shift in user churn caused by decision-making. This refers to the transfer that ultimately leads to the user's survival. Furthermore, each transfer is associated with a transfer probability P and a short-term value (short-term reward / short-term benefit) R, satisfying the following conditions:

[0075] ,

[0076] Figure 3 State transitions provided in the embodiments of this application Figure 1 .like Figure 2 As shown, this example uses an infinite time frame, including day 1, day 2, ..., with each day's state transition being the same. Figure 3 or Figure 3 However, the initial state is the same every day, such as the initial state of day 1 being the same as the initial state of day 2. Unlike regular dynamic programming, this example cannot be easily reduced to solving an MDP problem.

[0077] Memoryless Repeated Experiments: In models that require considering the continued revenue generated by user survival, the memoryless condition is a highly efficient assumption that simplifies analysis. A well-known model type is Buy Till You Die, which assumes that if a user remains in the next cycle, the value they generate is exactly the same as the cumulative return they could have generated at the start of the current cycle. For example, Figure 3The initial state of the first day and the initial state of the second day are completely identical in the state value, and are also completely identical with the state value of the initial state of the third day. Therefore, the no-memory repeated experiment can be understood as continuously launching the optimal strategy for an active user who does not remember anything every day until the user churns.

[0078] User lifetime value: assumptions denotes the state value of the initial state of the i-th day. Under the no-memory condition, there is , and is equivalent to .

[0079] The algorithm goal is to optimize , The solution of the optimal strategy can be written as follows:

[0080]

[0081] In the formula, is the expected function, denotes the short-term revenue of the i-th day, and since Figure 4 The no-memory repeated experiment shown in formula (1) is infinite, and direct optimization cannot be performed here.

[0082] It should be noted that the technical solutions of the embodiments of the present application are applicable to Figure 4 the state transition diagram shown in formula (2).

[0083] Figure 4 An application scenario diagram of the push strategy acquisition method provided by the embodiments of the present application. As shown in Figure 5 , the application scenario of the present embodiment includes a terminal device 401 and a server 402. A user accesses the server 402 through an application software, a program or a browser in the terminal device 401 to obtain various information. The server 402 can push information to the terminal device 401. The user can click to view the information on the terminal device 401, can set not to receive the information, or can do nothing. The server 402 determines whether to continue pushing (maintain the existing push strategy) or adjust the push strategy by collecting the response of the user to the pushed information.

[0084] In one possible scenario, the server is an e-commerce server, and the user accesses the e-commerce server through the terminal device to obtain various item information. The e-commerce server can push item information of interest to the user, such as item links and item activity information, according to historical behavior data of the user, and the user can select to view the item information details. The e-commerce server determines whether to continue pushing or adjust the push strategy by collecting the response of the user to the pushed item information.

[0085] In a possible scenario, the server is a news server, and a user accesses the news server through a terminal device to obtain various hot news. The news server can push hot news that the user is interested in according to historical browsing data of the user, and the user can select to view details of the news. The news server determines whether to continue pushing or adjusts the pushing strategy by collecting responses of the user to the pushed hot news.

[0086] It should be noted that the above two scenarios are only examples, and the technical solutions provided in this application can also be applied to other possible scenarios as long as the scenario involves pushing information and user value evaluation.

[0087] Based on the above application scenarios, in the related art, the server mainly determines the pushing strategy by short-term value evaluation of the user, for example, by determining the pushing strategy through short-term click rate of the user. At present, there is little research on the long-term value of the user, and there is no complete solution.

[0088] To solve the above problems, an embodiment of the present application provides a method for obtaining a pushing strategy. First, the base long-term value is solved according to behavior data of the user for different pushing strategies and the preset short-term value of the user state change under different pushing strategies. The solving process considers the combination of long-term and short-term values of the user. When the base long-term value is determined, the pushing strategy corresponding to the maximum long-term and short-term values can be obtained, and then information is pushed based on the pushing strategy. The above process can be referred to in Figure 1 The main invention points of the present case are:

[0089] 1) The problem that the long-term value of the user is difficult to quantify is solved by newly defining the long-term and short-term value of the user.

[0090] The long-term and short-term value includes the short-term value of the user in a single cycle (i.e. in a preset period, for example, a week, a month, a quarter or a year, etc.) and the long-term value of the user. The long-term value of the user can be represented by the product of the single-cycle user survival probability and the base long-term value.

[0091] 2) A dynamic solution for optimizing the long-term and short-term value of the user under the condition of no memory repeated experiments is introduced. The state of the user under the condition of no memory repeated experiments is independent of the previous behavior, which can be used in the analysis scenario of the long-term value of the user. The main idea of the scheme is as follows:

[0092] The preset interval of the long-term value is bisected, and it is judged whether the long-term value of the user under the optimal push strategy (the long-term value of the user under the push strategy is maximum) at the benchmark long-term value (the middle value of the preset interval) is equal to the benchmark long-term value. If not, it indicates that the set benchmark long-term value is too large or too small. The benchmark long-term value is re-determined in the left half interval or the right half interval of the preset interval, and the above judgment process is repeatedly executed until the appropriate benchmark long-term value is selected. The optimal push strategy determined at the benchmark long-term value is the target push strategy. The above process uses the bisection method to solve the benchmark long-term value, which can quickly solve the appropriate benchmark long-term value, helps to determine the optimal push strategy, and further improves the push effect.

[0093] The technical solutions provided by the embodiments of the present application will be described in detail below through specific embodiments. It should be noted that the technical solutions provided by the embodiments of the present application can include part or all of the following contents, and the following several specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in detail in some embodiments.

[0094] Figure 4 Flowchart of the push strategy acquisition method provided by the embodiments of the present application Figure 5 The method of the present embodiment can be applied to Figure 6 The server shown in FIG. 1. As shown in Figure 2 The push strategy acquisition method of the present embodiment includes the following steps:

[0095] Step 501, obtaining historical feedback data of a target user for different push strategies, the historical feedback data being used to indicate behavior statistical data of the target user for push information of different push strategies.

[0096] In the present embodiment, the server obtains the historical feedback data of the target user for different push strategies from a terminal device to which the target user belongs. The push strategy includes push content, push style, push frequency, push time, etc., and the present application does not make any limitation on this.

[0097] The push style includes the display mode of the push content, such as link, picture, video, voice, etc. The push frequency refers to the number of times of push information in a preset period, such as only one piece of information is pushed per week. The push time refers to the time of push information in a preset period, such as information is pushed at 19 o'clock every night, or information is pushed in the middle of every month.

[0098] The preset operation of the target user includes a click-to-view operation, a click-to-share operation, a click-to-favorite operation, a click-to-buy operation, a push setting operation, etc. The push setting operation includes setting not to receive push and customizing push setting.

[0099] It should be understood that the response of the target user is usually different for different push strategies, and therefore the push effect can be improved by adjusting the push strategy. For example, pushing information to the same user at different time periods, such as the morning period and the evening period, can be regarded as different push strategies, and the time at which the user clicks the pushed information can reflect the browsing habits of the user, such as the user only having browsing behavior in the evening period.

[0100] It should also be understood that for the same push strategy, the response of the target user can change over time or with changes in the environment, and therefore the historical feedback data of the target user for different push strategies can be periodically obtained, and the following steps of the embodiment are performed based on the latest historical feedback data to analyze whether the user has a state change (i.e., to analyze the user behavior), adjust the push strategy, and improve the push effect.

[0101] Step 502, determining a state transition matrix of the target user according to the historical feedback data, the state transition matrix being used to indicate a probability value of a state change of the target user under different push strategies.

[0102] In the embodiment, the state transition matrix of the target user can be represented as , wherein , , , represents a state set of the user, represents a push strategy set. refers to a current state of the target user being , and the state transition being after taking the push strategy , wherein .

[0103] It should be understood that the latest state of the user can be determined according to the response of the user to the pushed information. The state of the target user includes but is not limited to a survival state, a loss state (a push subscription state), an active state, and an accepted push state, and the state of the target user can be refined according to actual needs, and the embodiments of the present application do not make any limitation thereto.

[0104] Step 503, obtaining a preset short-term value corresponding to a state change under different push strategies.

[0105] In combination with the state transition matrix in step 502, , if the target user clicks to view the pushed information corresponding to the push strategy , the corresponding user short-term value can be obtained, which can be represented as . The short-term value can be preset for the state change under different push strategies, and a preset short-term value set is obtained, which can be represented as .

[0106] It should be noted that according to a series of operation processes of the target user for the push information, such as clicking to view, clicking to view sharing, clicking to view collection, clicking to view forwarding, etc., different short-term values can be preset. For example, the user clicks to view the item recommendation information, and the obtained short-term value is 0.2; after the user clicks to view the item recommendation information, the user also shares it with friends, and the obtained short-term value is 0.2+0.1=0.3; after the user clicks to view the item recommendation information, the user also forwards it to the friend circle, and the obtained short-term value is 0.2+0.3=0.5.

[0107] Step 504, determining the target push strategy with the maximum long-short-term value from the push strategy set according to the state transition matrix of the target user, the preset short-term value, and the benchmark long-term value.

[0108] In this embodiment, the short-term value of the target user can be obtained according to the state transition matrix of the target user and the preset short-term value. However, only considering the short-term value of the user for selecting the push strategy, the push effect may be short-term, and therefore the factors in the long-term value direction of the user should also be considered. Therefore, in this embodiment, when determining the target push strategy, the conditions of the user in the long-term value and the short-term value are considered comprehensively, and the push strategy with the maximum long-short-term value is selected from the push strategy set as the target push strategy in combination with the benchmark long-term value, which can further improve the push effect and maintain or enhance the push effect in the time dimension.

[0109] Optionally, the long-short-term value of the embodiment includes a user long-term value of a user state remaining to a next period and a user short-term value of a current period. The user long-term value is determined according to a product of a probability value of the user state remaining to the next period and the benchmark long-term value. The period here can correspond to the preset time period in the above, including, for example, every day, every week, every month, every year, etc. For example, the probability value of the user state remaining to the next period is denoted as p, the benchmark long-term value is denoted as g, and the user short-term value of the current period is denoted as r, and then the long-short-term value of the user can be represented as p×g+r.

[0110] Optionally, the target push strategy of the embodiment includes at least one of the following: sending push information at a specified time within a preset time period; the number of times of sending push information within a preset time period; the type of information sent within a preset time period; and sending push information containing specified content within a preset time period.

[0111] For example, in an e-commerce scenario, the target push strategy can be to push a push link containing, for example, electronic products that the user is interested in once a week. In a news information scenario, the target push strategy can be to push real-time hot news of the day at 12 noon and 9 pm every day.

[0112] The method for obtaining a push strategy provided in the embodiments of the present application obtains historical feedback data of a target user for different push strategies, first determines a state transition matrix of the target user according to the historical feedback data, and uses the state transition matrix to indicate a probability value of state change of the target user under different push strategies. Then, according to the state transition matrix of the target user, a preset short-term value corresponding to state change under different push strategies, and a benchmark long-term value, a target push strategy with maximum long-short-term value is determined from a push strategy set. Compared with the prior art, since the value of the user in the short term and the long term is considered at the same time, the reliability of the evaluation of the value of the user is higher, and the selected push strategy is better.

[0113] On the basis of the above embodiments, the scheme of how to determine a target push strategy from a push strategy set is described in detail below through one specific embodiment.

[0114] Figure 4 Flowchart of the method for obtaining a push strategy provided in the embodiments of the present application Figure 6 The method of the present embodiment can also be applied to Figure 7 the server shown in FIG. 1. As shown in FIG. 1, the method for obtaining a push strategy of the present embodiment comprises the following steps: Figure 7

[0115] Step 601: Obtain historical feedback data of a target user for different push strategies, and the historical feedback data is used to indicate behavior statistical data of the target user for push information under different push strategies.

[0116] In one optional embodiment of the present embodiment, a first statistical value of the target user having a click operation for a third push strategy in a preset time period is obtained, and a probability value of the target user having a state change under the third push strategy is determined according to the first statistical value. The third push strategy in this step is any push strategy in the push strategy set.

[0117] Specifically, the server sends push information to a terminal device to which the user belongs based on the third push strategy, and obtains a probability value of the user having a state change (state change caused by a click behavior) under the third push strategy by taking the ratio of the total number of times that the user clicks the push information to the total number of times that the push information is sent in a preset time period, for example, one week. For example, it is assumed that a strategy with a fixed total number of times of sending in one week is adopted, and 3 pieces of push information are sent to the user in one week. If the user only clicks and views 2 pieces of push information, then the probability value of the user having a click behavior under the push strategy is 2 / 3.

[0118] ​In an optional embodiment of the embodiment, a second statistical value of the target user having a subscription cancellation operation for the third push strategy in a preset time period is acquired, and a probability value of the target user having a state change under the third push strategy is determined according to the second statistical value.

[0119] Specifically, the server sends push information of different contents to the terminal device to which the user belongs based on the third push strategy, and the total number of categories of push information that the user cancels is counted in a preset time period (for example, a week), and the ratio of the total number of categories of push information that the user cancels to the total number of categories of push information is taken as the probability value of the user having a state change (state change caused by a cancellation behavior) under the third push strategy. For example, it is assumed that a strategy of sending different types of push information to the user at a fixed time every day is adopted, the total number of categories of push information in a week is 7, and if the user cancels only one type of push information in a week, the probability value of the user having a cancellation behavior under the push strategy is 1 / 7.

[0120] In an optional embodiment of the embodiment, a third statistical value of the target user having no operation for the third push strategy in a preset time period is acquired, and a probability value of the target user having no state change under the third push strategy is determined according to the third statistical value.

[0121] Specifically, the server sends push information to the terminal device to which the user belongs based on the third push strategy, and the total number of times of ignoring push information (that is, no response) of the user in a preset time period (for example, a week) is counted, and the ratio of the total number of times of ignoring push information of the user to the total number of times of sending push information is taken as the probability value of the user having no state change (that is, maintaining a survival state) under the third push strategy. For example, it is assumed that a strategy of sending push information to the user at two time points every day is adopted, the total number of times of sending push information in a week is 2*7=14, and if the user has 10 times of no response, the probability value of the user having no response under the push strategy is 5 / 7.

[0122] The above several embodiments are only examples, and other statistical values can be defined according to actual needs as analysis parameters of user behavior data to provide data support for subsequent selection of appropriate push strategies.

[0123] In step 602, a state transition matrix of the target user is determined according to historical feedback data, and the state transition matrix is used to indicate a probability value of state change of the target user under different push strategies.

[0124] Step 602 of the embodiment is similar to step 502 of the above embodiment, and the related description of the above embodiment can be referred to, and details are not described herein.

[0125] In step 603, an intermediate value is selected from a preset interval of long-term value as an initial value of the long-term value.

[0126] Step 604, based on the initial value of the long-term value, the bisection method is used to determine the benchmark long-term value from the preset interval of the long-term value.

[0127] It should be noted that the right end value of the preset interval of the long-term value can be determined according to experiments or empirical values, for example, the preset interval of the long-term value is set to [0, 100]. The benchmark long-term value can be regarded as an optimal solution of the long-term value. In order to quickly solve the benchmark long-term value, the bisection method is used to determine the benchmark long-term value from the preset interval of the long-term value. The processing process specifically includes the following steps:

[0128] Step 6041, based on the initial value of the long-term value, the state transition matrix of the target user and the preset short-term value, a first push strategy is determined from the push strategy set when the long-short-term value is maximum.

[0129] Step 6042, if the long-short-term value corresponding to the first push strategy is not equal to the initial value of the long-term value, the preset interval of the long-term value is shortened.

[0130] In one possible case, if the long-short-term value corresponding to the first push strategy is greater than the initial value of the long-term value, the left end value of the preset interval of the long-term value is replaced by the initial value.

[0131] In another possible case, if the long-short-term value corresponding to the first push strategy is less than the initial value of the long-term value, the right end value of the preset interval of the long-term value is replaced by the initial value.

[0132] Step 6043, an intermediate value is selected from the shortened preset interval of the long-term value as a new long-term value.

[0133] Step 6044, based on the new long-term value, the state transition matrix of the target user and the preset short-term value, a second push strategy is determined from the push strategy set when the long-short-term value is maximum.

[0134] Step 6045, if the long-short-term value corresponding to the second push strategy is equal to the new long-term value, the new long-term value is taken as the benchmark long-term value.

[0135] With a fixed long-term value, a value of 50 is selected, for example, the midpoint of the preset long-term value range [0, 100] in step 604. This 50 is used as the initial long-term value. Based on the initial long-term value, the target user's state transition matrix, and the preset short-term value, the long-term and short-term values ​​corresponding to each push strategy are determined by traversing the set of push strategies. The first push strategy that maximizes both the long-term and short-term values ​​is then identified. This first push strategy is the optimal push strategy when the long-term value is 50. If the long-term value is adjusted, the optimal push strategy for the new long-term value can be determined using the same method. In other words, different long-term values ​​may correspond to different optimal push strategies.

[0136] In this embodiment, the process of finding the optimal solution for the long-term value g involves the following judgment formula:

[0137] p×g+r=g

[0138] Wherein, the user's short-term and long-term value p×g+r can be denoted as output(g), and correspondingly, the above formula can be written as output(g)=g. p and r are parameters related to the push strategy. Assume the following two conclusions (the specific proof of the conclusions will be provided later):

[0139] 1) As g increases, output(g) increases continuously.

[0140] 2) If output(g) is p×g+r given g, then output(g) is continuous and there exists g′ such that g′=r / (1-p).

[0141] This embodiment uses the bisection method to find the optimal solution g′, which can improve the solution speed. The following is a related example. Figure 7 The process of solving for the benchmark long-term value is explained.

[0142] Figure 8 This is a schematic diagram illustrating the use of the bisection method to solve for the baseline long-term value, as provided in an embodiment of this application. Figure 8 As shown, the preset range for long-term value is [0, 100]. The midpoint of this preset range, 50, is taken as the initial value for long-term value. Based on this initial value, the push strategy that maximizes the user's long-term and short-term value output(g) is determined from the push strategy set. The relationship between output(g) and g is judged: in one case, if output(g) is less than g, it indicates that the value of g is too large; in another case, if output(g) is greater than g, it indicates that the value of g is too small.

[0143] In the first case, the preset interval of the long-term value can be adjusted to [0, 50], and the middle value 25 of the adjusted preset interval can be taken as a new long-term value. Based on the new long-term value, the push strategy with the maximum output (g) is determined again from the push strategy set until output (g) is equal to g.

[0144] In the second case, the preset interval of the long-term value can be adjusted to [50, 100], and the middle value 75 of the adjusted preset interval can be taken as a new long-term value. Based on the new long-term value, the push strategy with the maximum output (g) is determined again from the push strategy set until output (g) is equal to g.

[0145] As can be seen from the above description, the bisection method can quickly shorten the preset interval of the long-term value and find the optimal solution of the long-term value. The process of solving the optimal solution of the long-term value includes solving the optimal push strategy. The above solving process does not need to solve the optimal push strategy for each value in the preset interval of the long-term value, greatly reducing the calculation amount and shortening the time of solving the target push strategy, and achieving fast and accurate push effect.

[0146] In step 605, the push strategy used when the long-term value and the short-term value are maximum is determined based on the benchmark long-term value, the state transition matrix of the target user, and the preset short-term value, and is taken as the target push strategy.

[0147] Based on steps 6041 to 6045, when the new long-term value is the benchmark long-term value, the second push strategy used when the long-term value and the short-term value are maximum is determined based on the new long-term value, the state transition matrix of the target user, and the preset short-term value, and is taken as the target push strategy.

[0148] The push strategy acquisition method illustrated in this application obtains historical feedback data of the target user for different push strategies. First, it determines the target user's state transition matrix based on the historical feedback data. The state transition matrix indicates the probability value of the target user's state changes under different push strategies. Then, based on the target user's state transition matrix, the preset short-term value corresponding to the state changes under different push strategies, and the initial value of the long-term value, it determines the short- and long-term value corresponding to the current optimal push strategy based on the initial value, where the initial value of the long-term value is the median value of its preset interval. By comparing the magnitude of the initial value with the short- and long-term value corresponding to the current optimal push strategy, the length of the preset interval is continuously narrowed, and finally, the optimal solution for the long-term value is determined from the preset interval. Finally, the optimal push strategy determined based on the optimal solution for the long-term value is used as the final target push strategy. On the one hand, the above scheme considers the user's value in both the short-term and long-term directions, resulting in higher reliability of value assessment and better selected push strategy performance. On the other hand, the above scheme uses a bisection method to solve for the optimal solution of the long-term value, shortening the time to solve for the target push strategy and achieving fast and accurate push effects.

[0149] The theoretical knowledge and argumentation process involved in the technical solutions of the embodiments of this application will be explained below.

[0150] The above embodiments involve the long-short splitting of state values. Observe the state values ​​of the initial state. By capturing the path of its optimal value within the current period, it will inevitably transition to a certain probability p. The remaining probability is lost, resulting in a certain expected short-term gain r. The scheme is set with... Then there is =p Simplifying, we get Therefore, an equivalent description of the f-value can be obtained, which depends only on the short-term value (return within a single period) r and the long-term value (survival probability within a single period) p within a single period.

[0151] At this point, f can be redefined as an intermediate state. State value It can be divided into two parts: ( ), indicating that the state is reached at time t. Afterwards, a certain state transition path that can reach the optimal value (choose one if there are multiple solutions) will have the following at the end of the current cycle: The probability is carried over to the next cycle, and the expected gain in the current cycle is... The benefits.

[0152] At this point, the decomposition of the uniquely deterministic termination state must also be defined:

[0153] Lost state (denoted as) ): the state of permanent churn, where ;

[0154] Single-period survival state (denoted as ): the state of single-period survival, where ;

[0155] Based on the redefinition of f, a dynamic programming solution is proposed for the optimization of user lifetime value under the condition of no-memory repeated experiments. The core idea is to divide the initial state value, and then verify whether the answer is too large or too small through the recursion of single value, so as to finally divide the actual state value. The division of the initial state value described here corresponds to the long-term value of the above embodiment.

[0156] After dividing the answer, the main difference between the recursion process and the single-period sub-problem is the change of the recursion formula:

[0157]

[0158] Here max is a self-defined optimizer about , and the comparison rule is:

[0159] )

[0160] The above formula expresses the more optimal combination of long-term and short-term under the condition that the long-term value is determined as . Specifically, if a decision gets a short-term value r and has a probability of survival p, it will additionally obtain long-term value, and the total is . At this time, each combination of results (long-term, short-term) is mapped to a single value, so the optimal value of can be solved according to the conventional recursion method.

[0161] The following describes the lemmas, theorems and their proof processes involved in the solving process.

[0162] Lemma 1: In the above solving process, as the guess value increases, increases and is continuous.

[0163] Proof: Considering the process of monotonically increasing, if there is no change in the optimal strategy, then and will not change, and the monotonicity and continuity of are obviously satisfied. If a change is about to occur, i.e. the condition of occurs. Consider exchanging the two in advance, and the resulting change in results, such as Figure 9shown, Figure 9 p , r ) represents the survival probability p and the cumulative short-term income r generated from the initial state to the branching point, the income generated by the two decisions satisfies:

[0164]

[0165] Therefore, switching strategies on will not cause the value to change, and the local change after satisfies continuity and monotonicity, for the same reason as the first part of this proof, and the lemma is proved.

[0166] Lemma 2, let output( ) be the when given , then output( ) is continuous and there exists such that

[0167]

[0168] Proof: First, it is clear that is the actual life cycle value produced by the optimal decision, that is,

[0169] +…=

[0170] And in the process, as a guess life cycle value, only affects the optimal strategy when state transition. Therefore, the proposition of this lemma can be intuitively understood as there is a guess value that makes the real value and the guess value consistent, thus completing the layout of the dichotomy answer.

[0171] When , output( , therefore ; when , since any sending will make the value of p less than 1, therefore . By adding Lemma 1, we know about the monotonicity and continuity of , by the mean value theorem, there exists satisfying output( . At this time,

[0172] output( →

[0173] That is, the expression in the proposition.

[0174] Theorem 1: The bisection algorithm can optimize user lifetime value offline and find the optimal action combination.

[0175] Proof: First, use binary search to guess one... Next, calculate the output ( Finally, based on the output ( and The relationship determines whether the guessed value is too large or too small. By Lemma 1, we know that output( about Given the monotonicity of the expression, and combining it with Lemma 2, there is one and only one interval that satisfies the output( From this interval No optimal strategy is generated. The estimation bias, while the solution outside this interval (named policy) out Because estimation bias can lead to incorrect decisions during the transition, optimality cannot be guaranteed. Constructively, a policy can be used... out Actual long-term strategies replace At this point, a better solution will be obtained. This process is repeated until the solution converges to the target solution set output. Within the interval. The convergence property can be combined with monotonicity (the algorithm guarantees [a certain value] per iteration). It will definitely become closer to the output ( Given the finiteness of the interval (the range) and the discrete finiteness of the choices (the strategy is discrete and finite), the bisection algorithm can optimize the user lifetime value, thereby obtaining the optimal push strategy (action combination).

[0176] The algorithm's offline performance guarantees that user feedback will not affect the original optimal strategy if users do not churn; and in the case of churn, subsequent user behavior can be ignored, thus proving the theorem.

[0177] The foregoing described the method for obtaining push strategies provided in the embodiments of this application. The following describes the apparatus for obtaining push strategies provided in the embodiments of this application. The embodiments of this application can divide the apparatus for obtaining push strategies into functional modules according to the above method embodiments. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in the embodiments of this application is illustrative and only represents one logical functional division; other division methods may exist in actual implementation. The following description uses the division of functional modules according to each function as an example.

[0178] Figure 10 A structure diagram of a push strategy acquisition device provided by an embodiment of the present application is shown in FIG. 9. As shown in FIG. 9, the push strategy acquisition device 900 of the embodiment includes an acquisition module 901 and a processing module 902. Figure 10

[0179] The acquisition module 901 is configured to acquire historical feedback data of a target user for different push strategies, where the historical feedback data is used to indicate behavior statistical data of the target user for push information of different push strategies.

[0180] The processing module 902 is configured to determine a state transition matrix of the target user according to the historical feedback data, where the state transition matrix is used to indicate probability values of state changes of the target user under different push strategies.

[0181] The acquisition module 901 is configured to acquire a preset short-term value corresponding to a state change under different push strategies.

[0182] The processing module 902 is configured to determine a target push strategy at which a long-short-term value is maximum from a push strategy set according to the state transition matrix of the target user, the preset short-term value, and a benchmark long-term value.

[0183] In an optional embodiment of the embodiment, the long-short-term value includes a user long-term value of user state retention to a next period and a user short-term value of a current period, and the user long-term value is determined according to a product of a probability value of user state retention to the next period and the benchmark long-term value.

[0184] In an optional embodiment of the embodiment, the processing module 902 is configured to:

[0185] select a middle value as an initial value of the long-term value from a preset interval of the long-term value;

[0186] determine the benchmark long-term value from the preset interval of the long-term value based on the initial value of the long-term value by using a dichotomy method;

[0187] determine a push strategy at which a long-short-term value is maximum based on the benchmark long-term value, the state transition matrix of the target user, and the preset short-term value as the target push strategy.

[0188] In an optional embodiment of the embodiment, the processing module 902 is configured to:

[0189] determine a first push strategy at which a long-short-term value is maximum from the push strategy set based on the initial value of the long-term value, the state transition matrix of the target user, and the preset short-term value; ​

[0190] if the long-short term value corresponding to the first push strategy is not equal to the initial value of the long term value, shorten the preset interval of the long term value;

[0191] select a middle value from the shortened preset interval of the long term value as a new long term value;

[0192] determine a second push strategy with the maximum long-short term value from the push strategy set based on the new long term value, the state transition matrix of the target user, and the preset short term value;

[0193] if the long-short term value corresponding to the second push strategy is equal to the new long term value, take the new long term value as a reference long term value.

[0194] In an optional embodiment of the present embodiment, the processing module 902 is configured to:

[0195] take the second push strategy with the maximum long-short term value determined based on the new long term value, the state transition matrix of the target user, and the preset short term value as the target push strategy.

[0196] In an optional embodiment of the present embodiment, the processing module 902 is configured to:

[0197] if the long-short term value corresponding to the first push strategy is greater than the initial value of the long term value, replace the left end point value of the preset interval of the long term value with the initial value; or

[0198] if the long-short term value corresponding to the first push strategy is less than the initial value of the long term value, replace the right end point value of the preset interval of the long term value with the initial value.

[0199] In an optional embodiment of the present embodiment, the processing module 902 is configured to:

[0200] obtain a first statistical value of the target user having a click operation on a third push strategy within a preset time period, and determine a probability value of the target user having a state change under the third push strategy according to the first statistical value; or

[0201] obtain a second statistical value of the target user having an unsubscribe operation on the third push strategy within a preset time period, and determine a probability value of the target user having a state change under the third push strategy according to the second statistical value; or

[0202] obtain a third statistical value of the target user having no operation on the third push strategy within a preset time period, and determine a probability value of the target user having no state change under the third push strategy according to the third statistical value.

[0203] The third push strategy is any push strategy in the push strategy set.

[0204] In an optional embodiment of the present embodiment, the target push strategy comprises at least one of the following:

[0205] sending push information at a specified time within a preset period;

[0206] a number of times of sending push information within a preset period;

[0207] an information type of sending push information within a preset period;

[0208] sending push information containing specified content within a preset period.

[0209] The push strategy acquisition apparatus provided in the present embodiment can execute the technical solutions of any of the preceding method embodiments, and has similar implementation principles and technical effects, which will not be described herein again.

[0210] ​ A hardware structure diagram of an electronic device provided in the present embodiment is shown in FIG. 1. ​ As shown in FIG. 1, the electronic device 1000 provided in the present embodiment comprises:

[0211] a memory 1001, a processor 1002, and a computer program; wherein the computer program is stored in the memory 1001 and is configured to be executed by the processor 1002 to implement the technical solutions of any of the preceding method embodiments, and has similar implementation principles and technical effects, which will not be described herein again.

[0212] Optionally, the memory 1001 can be independent or integrated with the processor 1002. When the memory 1001 is independent of the processor 1002, the electronic device 1000 further comprises a bus 1003 for connecting the memory 1001 and the processor 1002.

[0213] The present embodiment provides a computer readable storage medium having a computer program stored thereon, and the computer program is executed by the processor 1002 to implement the technical solutions of any of the preceding method embodiments.

[0214] The present embodiment provides a computer program product comprising a computer program, and the computer program is executed by the processor to implement the technical solutions of any of the preceding method embodiments.

[0215] The present embodiment provides a chip comprising a processing module and a communication interface, and the processing module can execute the technical solutions of any of the preceding method embodiments.

[0216] Optionally, the chip further comprises a storage module (e.g., a memory), the storage module is configured to store instructions, the processing module is configured to execute the instructions stored in the storage module, and the execution of the instructions stored in the storage module causes the processing module to execute the technical solutions of any of the method embodiments.

[0217] It should be understood that the processor described above can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as execution completed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0218] The memory can include a high-speed RAM memory, and can also include a non-volatile storage NVM, for example, at least one disk memory, and can also be a U disk, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk, etc.

[0219] The bus can be an industry standard architecture (ISA) bus, a peripheral component (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus.

[0220] The storage medium described above can be realized by any type of volatile or non-volatile storage device or their combination, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0221] An example storage medium is coupled to the processor such that the processor can read information from, and can write information to, the storage medium. Of course, the storage medium can be a part of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the storage medium can be located in discrete components as well.

[0222] Finally, it should be noted that the above-described embodiments are merely intended to illustrate the technical solutions of the present application, but not to limit the same; even though the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can still be modified, or some or all of the technical features thereof can be replaced equivalently; and these modifications or replacements do not cause the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for acquiring a push policy, characterized by, The method comprises: acquiring historical feedback data of a target user for different push strategies, the historical feedback data being used to indicate behavior statistics of the target user for push information under different push strategies; determining a state transition matrix of the target user according to the historical feedback data, the state transition matrix being used to indicate probability values of state changes of the target user under different push strategies; acquiring preset short-term values corresponding to state changes under different push strategies; wherein the preset short-term values are obtained according to operations of the user for the push information; determining a target push strategy from a push strategy set when a long-short-term value is maximum according to the state transition matrix of the target user, the preset short-term values, and a benchmark long-term value; wherein the benchmark long-term value is determined from a preset interval of long-term values. The long-short-term value comprises a user long-term value obtained by retaining the user state to the next period, and a user short-term value of the current period; and the user long-term value is determined according to the product of the probability value of retaining the user state to the next period and the benchmark long-term value.

2. The method of claim 1, wherein: determining the target push strategy from the push strategy set when the long-short-term value is maximum according to the state transition matrix of the target user, the preset short-term values, and the benchmark long-term value comprises: selecting a middle value from the preset interval of long-term values as an initial value of the long-term value; determining the benchmark long-term value from the preset interval of long-term values by using a dichotomy method based on the initial value of the long-term value; determining the target push strategy as the push strategy used when the long-short-term value is maximum based on the benchmark long-term value, the state transition matrix of the target user, and the preset short-term values.

3. The method of claim 2, wherein, The method of claim 2, wherein: determining the benchmark long-term value from the preset interval of long-term values by using a dichotomy method based on the initial value of the long-term value comprises: determining a first push strategy from the push strategy set when the long-short-term value is maximum based on the initial value of the long-term value, the state transition matrix of the target user, and the preset short-term values; shortening the preset interval of long-term values if the long-short-term value corresponding to the first push strategy is not equal to the initial value of the long-term value; selecting a middle value from the shortened preset interval of long-term values as a new long-term value; determining a second push strategy from the push strategy set when the long-short-term value is maximum based on the new long-term value, the state transition matrix of the target user, and the preset short-term values; 4. The method of claim 3, wherein, determining the target push strategy as the second push strategy used when the long-short-term value is maximum based on the new long-term value, the state transition matrix of the target user, and the preset short-term values. The method of claim 2, wherein: determining the target push strategy as the second push strategy used when the long-short-term value is maximum based on the new long-term value, the state transition matrix of the target user, and the preset short-term values.

5. The method of claim 3, wherein, If the short-term and long-term values ​​corresponding to the first push strategy are not equal to the initial value of the long-term value, the preset range of the long-term value is shortened, including: If the short-term and long-term values ​​corresponding to the first push strategy are greater than the initial value of the long-term value, the left endpoint of the preset interval of the long-term value is replaced with the initial value; or If the short-term value corresponding to the first push strategy is less than the initial value of the long-term value, the right endpoint of the preset interval of the long-term value is replaced with the initial value.

6. The method according to any one of claims 1 to 5, characterized in that, Determining the state transition matrix of the target user based on the historical feedback data includes: Obtain a first statistical value of the target user's click operations in response to the third push strategy within a preset time period, and determine the probability value of the target user's state change under the third push strategy based on the first statistical value; or Obtain a second statistical value showing that the target user has unsubscribed from the third push strategy within a preset time period; and determine the probability value of the target user's status change under the third push strategy based on the second statistical value; or Obtain a third statistical value indicating that the target user has no operation in response to the third push strategy within a preset time period, and determine the probability value that the target user has no state change under the third push strategy based on the third statistical value. The third push strategy is any push strategy in the set of push strategies.

7. The method according to any one of claims 1 to 5, characterized in that, The target push strategy includes at least one of the following: Send push notifications at specified times within a preset time period; The number of times push notifications are sent within a preset time period; The types of push notifications to be sent within a preset time period; Send push notifications containing specified content within a preset time period.

8. A device for acquiring a push strategy, characterized in that, include: The acquisition module is used to acquire historical feedback data of the target user for different push strategies. The historical feedback data is used to indicate the behavioral statistics of the target user for push information under different push strategies. The processing module is used to determine the state transition matrix of the target user based on the historical feedback data. The state transition matrix is ​​used to indicate the probability value of the target user's state change under different push strategies. The acquisition module is used to acquire a preset short-term value corresponding to the state changes under different push strategies; wherein, the preset short-term value is obtained based on the user's operation on the push information; The processing module is used to determine the target push strategy that maximizes the short-term and long-term values ​​from the push strategy set based on the target user's state transition matrix, the preset short-term value, and the benchmark long-term value; wherein, the benchmark long-term value is determined from a preset range of long-term values; the short-term and long-term values ​​include the user's long-term value based on user state retention to the next cycle, and the user's short-term value in the current cycle; the user's long-term value is determined based on the product of the probability value of user state retention to the next cycle and the benchmark long-term value.

9. An electronic device, comprising: include: A memory, a processor, and a computer program; the computer program is stored in the memory and configured to be executed by the processor to implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, A computer program product comprising a computer program stored thereon for execution by a processor to implement the method of any of claims 1-7.

11. A computer program product, characterised in that, A computer program product comprising a computer program stored thereon for execution by a processor to implement the method of any of claims 1-7.

Citation Information

Patent Citations

  • Data pushing method and device and storage medium

    CN111401937A

  • User information determination method and device, electronic equipment and storage medium

    CN112734454A