Content pushing method and device, equipment, storage medium and product

By clustering and generating hierarchical tags from user account behavior data, and combining pre-trained recommendation models and Fogg models, the problem of low accuracy in content push in existing technologies has been solved, achieving precise push of points-related content, thereby improving user engagement and operational efficiency.

CN121836799APending Publication Date: 2026-04-10SHANGHAI PUDONG DEVELOPMENT BANK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI PUDONG DEVELOPMENT BANK
Filing Date
2025-11-26
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, the content push methods of business systems are not very accurate and cannot differentiate incentives based on user lifetime value. This results in high-value users lacking motivation to advance, low-value users having low activity levels, complex task design or untimely rewards, single content recommendation dimensions, inability to respond to changes in user behavior in real time, and operational costs increasing linearly with the scale of the activity.

Method used

By clustering user account behavior data to generate hierarchical labels, a pre-trained recommendation model is used to determine target content to push from a candidate content pool. The Fogg model and the Twin Towers model are combined for content instantiation and strategy scoring to ensure that the points-related content matches the user attributes. A reinforcement learning decision model is used to optimize the content push strategy.

Benefits of technology

It improved the accuracy and diversity of content delivery, enhanced users' enthusiasm and persistence in participating in targeted content delivery, improved operational efficiency, adapted to the automated segmentation needs of a large user group, and achieved differentiated content recommendation and precise operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121836799A_ABST
    Figure CN121836799A_ABST
Patent Text Reader

Abstract

The invention relates to a content pushing method and device, equipment, a storage medium and a product, and relates to the technical field of computer data processing, and the method comprises the steps: carrying out the clustering processing of each user account according to the behavior data of each user account in a business system, and obtaining the hierarchical label of each user account; generating candidate content pools corresponding to the hierarchical labels of the user accounts according to the hierarchical labels and the behavior data; the candidate content pool comprises a plurality of point-related contents; each integral related content is an integral related task and / or an integral replaceable resource; based on the current state of the user account, the behavior data, the hierarchical label and the candidate content pool, determining target push content from the candidate content pool through a pre-trained recommendation model; and pushing the target push content to the terminal of the user account. By adopting the method, the content pushing accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer data processing, and in particular, relates to a content pushing method and device, equipment, a storage medium and a product. BACKGROUND

[0002] With the development of online business, some platforms set up an integral system in the business system, and the user account accumulates points by completing tasks, and the points can be used to exchange resources. In related technologies, the traditional integral system in the business system usually uses static rules to obtain points, such as converting a fixed proportion of points according to the consumed resources or obtaining points by completing fixed tasks.

[0003] However, in the traditional technology, the content pushed to all user accounts is fixed, and there is a problem of low accuracy of content pushing. SUMMARY

[0004] Therefore, it is necessary to provide a content pushing method, device, computer equipment, computer readable storage medium and computer program product capable of improving the pushing accuracy in view of the above technical problems.

[0005] In a first aspect, the present application provides a content pushing method, comprising:

[0006] According to the behavior data of each user account in the business system, the user accounts are clustered to obtain the hierarchical labels of the user accounts;

[0007] According to the hierarchical labels and the behavior data, a candidate content pool corresponding to the hierarchical labels of each user account is generated; the candidate content pool includes a plurality of integral-related contents; each integral-related content is an integral-related task and / or an integral exchangeable resource;

[0008] Based on the current state of the user account, the behavior data, the hierarchical labels and the candidate content pool, a target pushing content is determined from the candidate content pool by a pre-trained recommendation model;

[0009] The target pushing content is pushed to the terminal of the user account.

[0010] In one embodiment, the clustering of each user account according to the behavior data of each user account in the business system to obtain the hierarchical labels of each user account comprises:

[0011] According to the preset time window and the behavior data, the latest operation time, operation frequency and operation resource quota of each user account are determined;

[0012] discretize the recent operation time, the operation frequency and the operation resource quota of all user accounts respectively to obtain discrete user accounts;

[0013] perform clustering processing on the discrete user accounts by using a clustering algorithm to obtain a hierarchical label of each user account.

[0014] In one of the embodiments, the generating, according to the hierarchical label and the behavior data, of a candidate content pool corresponding to each user account hierarchical label comprises:

[0015] configuring a trigger parameter, a motivation parameter and a capability parameter of a pre-constructed Fogg Model according to the hierarchical label and the behavior data to obtain a configured Fogg Model;

[0016] performing content instantiation processing according to the configured Fogg Model to obtain a candidate content pool corresponding to each user account hierarchical label.

[0017] In one of the embodiments, the pre-trained recommendation model comprises a pre-trained double-tower model and a pre-trained reinforcement learning decision model, and the determining, by the pre-trained recommendation model, of target push content from the candidate content pool based on the current state of the user account, the behavior data, the hierarchical label and the candidate content pool comprises:

[0018] processing the current state of the user account and the candidate content pool by using the pre-trained reinforcement learning decision model to obtain a strategy score of each integral-related content;

[0019] processing the behavior data, the hierarchical label and the candidate content pool thereof by using the pre-trained double-tower model to obtain a matching degree score of each integral-related content;

[0020] determining, from the candidate content pool, target push content of each user account according to the strategy score and the matching degree score of the integral-related content.

[0021] In one of the embodiments, the processing, by the pre-trained reinforcement learning decision model, of the current state of the user account and the candidate content pool to obtain a strategy score of each integral-related content comprises:

[0022] processing, by the pre-trained reinforcement learning decision model, the current state of the user account and the candidate content pool to obtain a strategy action of each integral-related content;

[0023] obtaining a strategy score of each integral-related content according to the strategy action of each integral-related content.

[0024] In one embodiment, the pre-trained dual-tower model is used to process the behavioral data, the hierarchical labels, and the candidate content pool to obtain a matching score for each of the score-related contents, including:

[0025] The pre-trained dual-tower model is used to process the behavioral data, the hierarchical labels, and the candidate content pool to obtain user account features and points-related content features.

[0026] Determine the similarity between the user account features and the points-related content features, and use the similarity as a matching score for each of the points-related content features.

[0027] In one embodiment, determining the target push content for each user account from the candidate content pool based on the strategy score and the matching score related to the points includes:

[0028] Based on the preset recall quantity and the matching degree score of each of the points-related contents, the candidate content pool is searched by nearest neighbor query to obtain multiple candidate points-related contents;

[0029] Based on the strategy score and the matching score of each candidate points-related content, the target push content for each user account is determined from the candidate points-related content.

[0030] Secondly, this application also provides a content push device, including:

[0031] The user clustering module is used to cluster the user accounts based on the behavioral data of each user account in the business system to obtain the hierarchical labels of each user account.

[0032] The content pool generation module is used to generate a candidate content pool corresponding to each user account level tag based on the level tags and the behavior data; the candidate content pool includes multiple points-related content; each point-related content is a points-related task and / or a point-exchangeable resource;

[0033] The target determination module is used to determine the target push content from the candidate content pool based on the current state of the user account, the behavioral data, the hierarchical tags, and the candidate content pool, using a pre-trained recommendation model.

[0034] The content push module is used to push the target content to the user's account terminal.

[0035] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0036] According to the behavior data of each user account in a business system, each user account is subjected to clustering processing to obtain a hierarchical label of each user account;

[0037] According to the hierarchical label and the behavior data, a candidate content pool corresponding to the hierarchical label of each user account is generated; the candidate content pool comprises a plurality of integral-related contents; each integral-related content is an integral-related task and / or an integral exchangeable resource;

[0038] Based on the current state of the user account, the behavior data, the hierarchical label, and the candidate content pool, a target push content is determined from the candidate content pool by a pre-trained recommendation model;

[0039] The target push content is pushed to a terminal of the user account.

[0040] In a fourth aspect, the present application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the following steps:

[0041] According to the behavior data of each user account in a business system, each user account is subjected to clustering processing to obtain a hierarchical label of each user account;

[0042] According to the hierarchical label and the behavior data, a candidate content pool corresponding to the hierarchical label of each user account is generated; the candidate content pool comprises a plurality of integral-related contents; each integral-related content is an integral-related task and / or an integral exchangeable resource;

[0043] Based on the current state of the user account, the behavior data, the hierarchical label, and the candidate content pool, a target push content is determined from the candidate content pool by a pre-trained recommendation model;

[0044] The target push content is pushed to a terminal of the user account.

[0045] In a fifth aspect, the present application further provides a computer program product comprising a computer program, wherein the computer program is executed by a processor to implement the following steps:

[0046] According to the behavior data of each user account in a business system, each user account is subjected to clustering processing to obtain a hierarchical label of each user account;

[0047] According to the level label and the behavior data, a candidate content pool corresponding to each user account level label is generated; the candidate content pool includes a plurality of integral related contents; each integral related content is an integral related task and / or an integral exchangeable resource;

[0048] Based on the current state of the user account, the behavior data, the level label, and the candidate content pool, a target push content is determined from the candidate content pool by a pre-trained recommendation model;

[0049] The target push content is pushed to the terminal of the user account.

[0050] The content push method, device, computer equipment, computer readable storage medium, and computer program product accurately identify the characteristics and demand differences of different value level user accounts by clustering each user account in a business system according to behavior data of each user account to obtain a level label of each user account. According to the level label and the behavior data, a candidate content pool corresponding to each user account level label is generated; the candidate content pool includes a plurality of integral related contents, wherein each integral related content is an integral related task and / or an integral exchangeable resource, ensuring that the integral related content matches the attributes of the user account of different level labels. Further, based on the current state of the user account, the behavior data, the level label, and the candidate content pool, a target push content is determined from the candidate content pool by a pre-trained recommendation model, and the target push content is pushed to the terminal of the user account. This not only improves the accuracy of integral related content push, but also enhances the enthusiasm and persistence of user accounts in participating in target push content activities, avoiding the defects of single and low accuracy of related technical content push methods. BRIEF DESCRIPTION OF DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other related drawings can also be obtained without creative labor.

[0052] Figure 1 It is a flowchart of the content push method in one embodiment;

[0053] Figure 2 It is a step flowchart of user level label construction in one embodiment;

[0054] Figure 3 It is a Fogg model principle diagram of a daily sign-in task in one embodiment;

[0055] Figure 4 a training step flowchart of a reinforcement learning decision model in one embodiment;

[0056] Figure 5 a structural diagram of a content pushing system in one embodiment;

[0057] Figure 6 a structural block diagram of a content pushing device in one embodiment;

[0058] Figure 7 an internal structural diagram of a computer device in one embodiment. DETAILED DESCRIPTION

[0059] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0060] As described in the background, the related art content pushing method has the problem of low pushing accuracy. The inventors have found that the reason for this problem is that, with the development of online business, some platforms set up a point system in their business system. The user account accumulates points by completing tasks, and the points can be used to exchange resources. The traditional point system usually adopts static rules, and the consumed resources are converted into fixed proportion points, and the points can be used to exchange fixed resources. However, the point management system of the related art cannot provide differential incentives according to the user life cycle value, resulting in a lack of further transition motivation for high-value users and low activity for low-value users; complex task design or delayed rewards, resulting in low user participation; single content recommendation dimension, mismatch between pushed content and user needs, low exchange rate and participation rate; unable to dynamically adjust the content pushing strategy according to the real-time behavior feedback of the user, and the operation cost increases linearly with the size of the activity. In the prior art, some business platforms attempt to push content based on simple tags or single machine learning models, such as collaborative filtering, but there are the following deficiencies: unable to balance the long-term value and short-term conversion of users; the update cycle of the pushed content and the point strategy is long, and it cannot respond to changes in user behavior in real time; the pushed content and the point reward design lack the support of behavioral psychology, resulting in limited user participation.

[0061] Based on the above reasons, the present application provides a content pushing method, which aims to improve the accuracy of content pushing.

[0062] In one embodiment, as Figure 1As shown, a content pushing method is provided, and the embodiment takes the method applied to a server as an example. It can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is realized through the interaction of the terminal and the server. The server can be a stand-alone physical server, a server cluster or a distributed system formed by multiple physical servers, or a cloud server providing cloud computing services. The terminal can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. In the embodiment, the method includes the following steps:

[0063] In step S102, the user accounts in the business system are clustered according to the behavior data of the user accounts, and the hierarchical labels of the user accounts are obtained.

[0064] The business system is a comprehensive information system for managing user accounts, transaction behaviors, credit systems, and related operation activities in a financial service platform of the business system. The user account can be a personal or enterprise account registered in the business system and having a unique identifier, which can perform operation behaviors such as login, transaction, browsing, etc., and participate in the credit system interaction. The behavior data can be various types of recordable operation behavior information generated by the user account in the business system, including but not limited to transferring resource quota, login frequency, page browsing, task completion, sharing behavior, etc.

[0065] The clustering process can be an unsupervised machine learning method, which divides users with similar behavior patterns into the same category by similarity measurement of the behavior feature vectors of the users.

[0066] The hierarchical label can be a user grouping result label obtained by clustering analysis, which represents the value level or behavior stage of the user, such as high value (HV), medium-high value (MH), medium value (M), medium-low value (ML), and low value (LV).

[0067] Optionally, the server uses a preset clustering algorithm to cluster the user accounts in the business system according to the similarity between the behavior data of the user accounts, clusters similar user accounts into clusters, sets a corresponding label for each cluster, and obtains the hierarchical labels of the user accounts.

[0068] In step S104, a candidate content pool corresponding to the hierarchical label of each user account is generated according to the hierarchical label and the behavior data.

[0069] The candidate content pool includes a plurality of points-related contents; each points-related content is a points-related task and / or a points-exchangeable resource; the points-related task can be a task activity that requires active participation of the user to obtain points, such as check-in, clicking on a predetermined page, etc.

[0070] The points-exchangeable resource can be a resource that can be exchanged by the user's points, such as goods, discount coupons, service benefits, or other virtual / physical resources.

[0071] Optionally, the server first analyzes the preferences, active features, and task completion of different level user accounts according to the level tags of the user accounts and the behavior data thereof, identifies high-response behavior patterns such as high-frequency participation in check-in, preference for high-point tasks, etc.; then configures differentiated content strategy rules for each level tag of the user account in combination with the business target, for example, the server matches high-order tasks and higher-value resources for the user account with a high-value user level tag, and configures low-threshold incentive tasks for a low-value user, and filters and combines a candidate content subset suitable for each level from the full points-related content library based on these rules to form a candidate content pool corresponding to the level tag.

[0072] Step S106, based on the current state of the user account, the behavior data, the level tag, and the candidate content pool, the target push content is determined from the candidate content pool by the pre-trained recommendation model.

[0073] The current state can be a dynamic environment state of the user account at the current time, including a real-time behavior, a recent operation, a balance of points, a level tag, device information, a geographic location, a time context, and other information sets that can be used for decision-making.

[0074] The pre-trained recommendation model can be a recommendation system model that has been preliminarily trained based on the historical behavior data of each user, and can quickly generate personalized recommendation results under new input.

[0075] Optionally, the server determines the target push content that matches the current state of the user account, the behavior data, and the level tag from the candidate content pool by the pre-trained recommendation model based on the current state of the user account, the behavior data, the level tag, and the candidate content pool.

[0076] Step S108, the target push content is pushed to the terminal of the user account.

[0077] The target push content can be the specific content finally decided to be pushed to a certain user account, which is usually one or more high-priority points-related tasks or exchangeable resources, and aims to improve user activity or conversion rate.

[0078] Optionally, the server pushes the target push content to the terminal of the user account in a preset form, for example, in a list form to the terminal of the user account, and displays the target push content on the page related to the point system of the application software.

[0079] In the content pushing method, the behavior data of each user account in the business system is used to cluster each user account to obtain a hierarchical label of each user account, so that the characteristics and demand differences of user accounts in different value levels can be accurately identified. A candidate content pool corresponding to the hierarchical label of each user account is generated according to the hierarchical label and the behavior data. The candidate content pool includes a plurality of point-related contents, wherein each point-related content is a point-related task and / or a point-exchangeable resource, so as to ensure that the point-related content matches the attributes of the user account with different hierarchical labels. Further, based on the current state of the user account, the behavior data, the hierarchical label, and the candidate content pool, a pre-trained recommendation model is used to determine the target push content from the candidate content pool, and the target push content is pushed to the terminal of the user account. The accuracy of the point-related content pushing is improved, the enthusiasm and continuity of the user account in participating in the target push content activity are enhanced, and the defects of the related art, such as single content pushing and low accuracy, are avoided.

[0080] In an exemplary embodiment, step S102 clusters each user account according to the behavior data of each user account in the business system to obtain a hierarchical label of each user account, including:

[0081] According to the preset time window and the behavior data, the recent operation time, operation frequency, and operation resource amount of each user account are determined. The recent operation time, operation frequency, and operation resource amount of all user accounts are discretized by binning, and the discretized user accounts are obtained. A clustering algorithm is used to cluster the discretized user accounts to obtain a hierarchical label of each user account.

[0082] The preset time window can be a time range for statistical user behavior, which is usually a fixed period.

[0083] The recent operation time (Recency) can be the time interval between the last time the user account performs a key operation, such as login, transaction, etc., and the current time, reflecting the recent activity level. The operation frequency (Frequency) can be the number of times the user account performs a specific operation within the preset time window, reflecting the frequency of its behavior habits. The operation resource amount (Monetary) can be the cumulative resource consumption or contribution value of the user account within the preset time window, such as the total transaction amount, which can also be extended to asset size, financial purchase amount, etc.

[0084] Among them, binning discretization can be used to divide continuous numerical variables into several intervals (called bins) and assign them a discrete level, such as 1 to 5, to facilitate subsequent modeling and clustering.

[0085] Optionally, the server performs statistical processing on the behavioral data according to a preset time window to obtain the recent operation time, operation frequency, and operation resource limit of each user account. The recent operation time, operation frequency, and operation resource limit of all user accounts are binned and discretized according to preset quantiles to obtain discretized user accounts. A clustering algorithm and a preset number of user levels (e.g., if five levels of users are required, then the number of user levels is five) are used to cluster the discretized user accounts to obtain multiple clusters corresponding to the preset number of user levels. After assigning labels to each cluster, the level labels of each user account are obtained.

[0086] For example, such as Figure 2 The diagram illustrates the steps involved in constructing user-level tags. The server, operating within a preset time window ΔT = 90 days, calculates the recent operation time, frequency, and resource allocation for each user account. Further, it discretizes these three dimensions (recent operation time, frequency, and resource allocation) using quintiles, resulting in 3 × 5 = 125 initial cells. Specifically, the K-means clustering algorithm is used to cluster these 125 cells into N = 5 levels based on the recent operation time. The user-level tags are set as follows: High Value (HV), Medium-High Value (MH), Medium Value (M), Medium-Low Value (ML), and Low Value (LV). It's worth noting that after obtaining the user-level tags, a points reward coefficient α and a cost control coefficient β can be configured for each tag. The points reward coefficient is a multiplier used to amplify points rewards; a higher α value results in more points earned by the user for the same action. The cost control coefficient can be a discount rate used to control or reduce points costs, affecting the value of points redemption or the consumption of points themselves; a higher β value indicates lower points usage value or consumption rate.

[0087] In this embodiment, by introducing a preset time window to dynamically count user behavior data, combining the three dimensions of recent operation time, operation frequency and operation resource quota, and using quantile binning discretization processing, the continuous behavior indicators are effectively converted into comparable level features, enhancing the robustness and interpretability of the data. On this basis, the clustering algorithm is used to automatically cluster the discretized users according to the preset number of levels, realizing the scientific division and label management of user value. Not only does it overcome the subjectivity and poor flexibility of traditional manual stratification, but it also adapts to the automated stratification needs of large-scale user groups, improving the accuracy and stability of user level division, providing a reliable data foundation for subsequent differentiated content recommendation and precise operation, and significantly enhancing the pertinence and operational efficiency of the point incentive system.

[0088] In an exemplary embodiment, step S104 generates a candidate content pool corresponding to each user account level label according to the level label and behavior data, including:

[0089] According to the level label and behavior data, the trigger parameters, motivation parameters and ability parameters of the pre-constructed Fogg model are configured to obtain the configured Fogg model; and according to the configured Fogg model, content instantiation processing is performed to obtain the candidate content pool corresponding to each user account level label.

[0090] The Fogg Behavior Model can be a behavior psychology model that believes that when motivation, ability and trigger are met at the same time, users will produce target behavior; the trigger parameter can be an external stimulus setting for starting user behavior, including push channels, trigger time periods (8 o'clock / 12 o'clock / 18 o'clock, etc.); the motivation parameter can be a factor affecting the user's willingness to perform tasks, mainly including the strength of integral reward incentives, such as the number of points, the probability of randomly falling points for exchangeable resources, etc.; the ability parameter can be an index for measuring the difficulty of completing a task, including the number of task steps, required skills, time cost, etc., the simpler the easier to complete.

[0091] Optionally, the server dynamically configures three core parameters in the pre-built Fogg behavior model based on the hierarchical labels and historical behavior data of each user account. Specifically, in the trigger parameter configuration, according to the active period and preferences of users at different levels, the optimal push channel and trigger time are set, such as application software push, short message, WeChat message, etc. For example, high-value users prefer to push at 8 pm, and low-active users increase the noon reminder to ensure that the task reminder appears at the time when the user is most likely to respond. In the motivation parameter configuration, combined with the user's sensitivity to points and reward feedback behavior, different levels of users are set with different incentive intensity, for example, high points are provided for low-value users to improve participation willingness, and scarce rights or level-specific benefits are introduced for high-value users to enhance the sense of belonging. In the ability parameter configuration, according to the user's operation habit and task completion difficulty tolerance, the task complexity is adjusted, such as designing a simple task that can be completed in a single step for new users or low-active users, and opening a multi-step but high-return task path for high-value users. After the server completes the parameter configuration, the abstract Fogg model output is converted into specific executable points-related content through the content instantiation engine, for example, generating a complete task instance of "pushing the application pop-up window at 18:00 today, and completing the check-in to get 50 points", and collecting it according to the hierarchical label to form a structured candidate content pool, so as to realize the accurate landing from psychological motivation modeling to actual operation content, so that each type of user receives content that meets their behavior ability and can be effectively motivated to participate, and obtains appropriate triggers at the right time, thereby comprehensively improving task accessibility and conversion rate.

[0092] For example, as shown in Figure 3 , a Fogg model schematic diagram of a daily check-in task is provided. Taking the construction of the daily check-in points task as an example, the parameter configuration of the trigger T is that the push channel c∈{app_push,sms,wechat,pop_up}, the trigger time period h∈{8,12,18,21}, the parameter configuration of the motivation M is that the reward points r∈{10,20,50,100}, the random drop coupon probability p∈{0.1,0.2,0.4}, the parameter configuration of the ability A is that the task step n_steps∈{1,2,3}, the estimated completion time t_cost∈{5s,30s,120s}, and the parameter configuration of the reward R is that the continuous 7-day extra reward is 100 points. The parameter configuration of the reminder P is that if the user does not check in for 2 consecutive days, the application pop-up window will be pushed again at noon on the third day.

[0093] In an exemplary embodiment, the pre-trained recommendation model of step S106 includes a pre-trained dual-tower model and a pre-trained reinforcement learning decision model, and the target push content is determined from the candidate content pool based on the current state of the user account, the behavior data, the hierarchical label, and the candidate content pool through the pre-trained recommendation model, including:

[0094] The pre-trained reinforcement learning decision model is used to process the current state of the user account and the candidate content pool to obtain a strategy score of each integral related content; the pre-trained double tower model is used to process the behavior data, the hierarchical label and the candidate content pool to obtain a matching degree score of each integral related content; and the target push content of each user account is determined from the candidate content pool according to the strategy score and the matching degree score of the integral related content.

[0095] The content instantiation processing can be based on the parameter configuration of the Forrester model to convert an abstract task template into a specific, executable and pushable actual task content.

[0096] The pre-trained reinforcement learning decision model can be a pre-trained policy model based on a reinforcement learning framework (such as DQN or Policy Gradient), which can select the best action according to the current state of the user, such as pushing which task, to maximize the long-term benefit (such as user retention and integral activity).

[0097] The strategy score can be a quantitative evaluation score of the expected long-term benefit (such as user transition probability, integral growth and retention rate) of the candidate content in the current state by the pre-trained reinforcement learning decision model.

[0098] The pre-trained double tower model can be a classic recommendation system structure, which includes two independent neural network towers: a user tower and an item tower, which respectively encode user features and content features, and finally calculate the matching score through inner product or similarity function.

[0099] The matching degree score can be a numerical value that measures the degree of fit between a certain integral related content and a specific user, which is usually determined by the similarity between the user features and the content features output by the double tower model.

[0100] Optionally, the server uses the pre-trained reinforcement learning decision model to process the current state of the user account and the candidate content pool to obtain a strategy score of each integral related content, and uses the pre-trained double tower model to process the behavior data, the hierarchical label and the candidate content pool to obtain a matching degree score of each integral related content. The integral related content in the candidate content pool is scored and ranked according to the score and value of the strategy score and the matching degree score of the integral related content, so as to determine the target push content of each user account from the candidate content pool according to the score and value order and the number of push content required.

[0101] In this embodiment, the reinforcement learning model evaluates the long-term benefits of each integral-related content in the future, such as user behavior transition and retention improvement, based on the current state of the user and the candidate content pool, and outputs a strategy score reflecting the strategy value; the dual tower model calculates the matching degree score reflecting the immediate preference based on the deep matching between the user historical behavior data, the hierarchical label and the candidate content pool, and the two are weighted and fused for comprehensive sorting, overcoming the slow convergence and poor real-time performance of single reinforcement learning, so that the final selected target push content not only fits the user's interest, but also effectively guides the user to complete high-value behavior. Further improve the accuracy, diversity and operation efficiency of content push.

[0102] In an exemplary embodiment, the steps of the above embodiments employ a pre-trained reinforcement learning decision model to process the current state of the user account and the candidate content pool to obtain the strategy score of each integral-related content, including:

[0103] A pre-trained reinforcement learning decision model is used to process the current state of the user account and the candidate content pool to obtain the strategy action of each integral-related content; and the strategy score of each integral-related content is obtained according to the strategy action of each integral-related content.

[0104] Among them, the strategy action can be a specific decision behavior output by the pre-trained reinforcement learning decision model, such as pushing the task of "checking in to get 50 points" to the user, delaying the push, not pushing, etc.

[0105] It should be noted that the pre-trained reinforcement learning decision model can be a MDP (Markov Decision Process, Markov Decision Model), and the definition includes: state s_t = [user level, near 7-day activity, remaining points, last time to exchange category]; action a_t ∈ {improve a, reduce b, push high-value goods, push low-threshold task, null operation}; reward r_t = ΔLTV_t - λ·cost_t; where a is the integral reward coefficient, b is the cost control coefficient, ΔLTV_t is the measure of the increase in the long-term value of the user account due to the current action, i.e. the user account value increment, and λ is the cost weight coefficient, which is a hyperparameter. For example Figure 4As shown, a training step flowchart of the reinforcement learning decision model is provided, and the reinforcement learning decision model training flow adopts offline pre-training and online fine-tuning synchronous training. In the offline pre-training process, the server uses the behavior data and corresponding experience data of the user account in the historical time period, such as data in 180 days, and replays 50 million experiences, and trains Double DQN (Double Deep Q-Network, an improved deep reinforcement learning algorithm) based on the behavior data and corresponding experience data in the historical time period. In the online training process, every preset time period, new experiences are obtained from the experience replay buffer (Replay Buffer), such as every 5 minutes, the server collects 512 new experiences from the Replay Buffer, and the Replay Buffer is a FIFO (First In, First Out, First In First Out Buffer) with a capacity of 1M. Further, the server triggers the update training of the Double DQN policy network θ according to the preset time period. In addition, during the training process, the server also needs to comply with the safety policy, and the safety policy includes action filtering, if α or β will cause the budget to exceed, the action is shielded; manual bottom, generate budget report regularly every day, and the operation can one-key rollback to yesterday's strategy in the background.

[0106] Optionally, the server pre-trains the reinforcement learning decision model to take the current state of the user account, such as the balance of points, the recent active time, the current level label, the history task completion condition, etc. as the input state, and the strategy action of pushing or not pushing each point-related content in the candidate content pool is regarded as the optional action set. The pre-trained reinforcement learning decision model evaluates each point-related content, outputs the probability of it as the optimal action, that is, judges whether executing the action is more beneficial to improving the long-term goal, such as user retention, point activity, level transition, etc. in the current state, thereby generating the strategy action for the content, for example, immediately pushing, delaying to push in the evening or not pushing for the time being. Subsequently, the server calculates the expected return (or action value function value, quantifies the strategy score of the point-related content according to the strategy action corresponding to the expected return (or action value function value, quantifies the strategy score of the point-related content.

[0107] In this embodiment, the pre-trained reinforcement learning decision model takes the current state of the user account as input, considers the push or non-push of candidate content as a set of optional behaviors, evaluates the contribution of each action to long-term business goals such as user retention, point activity, and level transition based on the pre-trained policy network, outputs the probability of it as the optimal action, and generates specific policy actions such as immediate push, delayed push, or no push. Further, according to the action value function (such as Q value) or expected return corresponding to the action, the policy score is quantified, which not only improves the timing rationality and decision intelligence level of content push, but also effectively avoids frequent disturbance or missed opportunities, enhances user experience and operational efficiency, and maximizes the value of the user's life cycle while achieving dynamic and personalized closed-loop optimization of the point incentive system.

[0108] In one exemplary embodiment, the steps of the above embodiment employ a pre-trained dual tower model to process behavior data, level labels, and their candidate content pool to obtain matching score of each point-related content, including:

[0109] The pre-trained dual tower model is used to process behavior data, level labels, and their candidate content pool to obtain user account features and point-related content features; determine the similarity between user account features and point-related content features, and use the similarity as the matching score of each point-related content.

[0110] Wherein, the user account features can be a low-dimensional dense vector formed by embedding or encoding the user's historical behavior data, level labels, demographic attributes and other information, which is used to represent the user's preferences and behavior patterns.

[0111] Wherein, the point-related content features can be a feature vector formed by structuring and encoding point-related tasks or redeemable resources, containing task type, reward size, difficulty level, target audience and other information.

[0112] Wherein, the similarity can be the distance measure between the user account feature vector and the point-related content feature vector, such as cosine similarity, dot product, reflecting the association strength between the two.

[0113] It should be noted that the pre-trained dual tower model is obtained through offline training and online inference. In the offline training process, training samples are obtained, the training samples are historical behavior data of user accounts in a historical time period, such as recent 90-day interaction logs (browsing, adding to cart, redeeming, sharing), feature extraction is performed on the training samples to obtain user-side features, commodity-side features (i.e., task-related content features), and context features. The user-side features can include user RFM vectors, 30-day activity, interest label embedding (128 dimensions), and the like. The commodity-side features include category one-hot, price bucketing, inventory, and point value. The context features include time, location, and weather. According to preset training parameters, such as batch_size (sample batch size) = 8192, learning_rate (learning rate) = 1e-3, and epoch (training rounds) = 10, the initial dual tower model is trained through the features of the training samples. In the training process, the conversion rate regression function is used as the loss function, and the process is repeated until the preset training rounds are reached, to obtain the pre-trained dual tower model.

[0114] Optionally, the server uses the pre-trained dual tower model to process the behavior data, the hierarchical label, and the candidate content pool to obtain user account features and point-related content features, calculates the similarity between the user account features and the point-related content features, and uses the similarity as a matching degree score of each point-related content.

[0115] In this embodiment, by using the pre-trained dual tower model, the user-side information and the point-related content-side information are respectively deep feature encoded to generate user account feature vectors and point-related content feature vectors in a high-dimensional semantic space. The similarity, such as the cosine similarity, between the two is calculated to quantify the preference degree of the user for each item of content, which is then used as a matching degree score. The potential association between the user behavior pattern and the point-related content attribute is fully mined, cross-modal and fine-grained personalized matching is achieved, the relevance and accuracy of the recommendation result are improved, and the generalization recommendation capability for new content and long-tail content is enhanced. Compared with traditional rule matching or collaborative filtering, the dual tower model supports large-scale real-time retrieval and efficient online inference, improves the system response speed and recommendation diversity, provides high-quality matching basis for subsequent multi-model fusion sorting, and effectively promotes the improvement of user participation willingness and point system activity.

[0116] In one exemplary embodiment, step S106 determines the target push content for each user account from the candidate content pool according to the strategy score and the matching degree score of the point-related content, including:

[0117] According to the preset recall quantity and the matching degree score of each integral related content, nearest neighbor query retrieval is performed on the candidate content pool to obtain a plurality of candidate integral related contents; and according to the strategy score of each candidate integral related content and the matching degree score of each candidate integral related content, target push contents of each user account are determined from the candidate integral related contents.

[0118] The preset recall quantity can be the number of integral related contents most likely to be accepted by the user selected from the large-scale candidate content pool in the recall stage of the recommendation process, which is used to reduce the calculation burden in the sorting stage.

[0119] The nearest neighbor query retrieval can be a kind of efficient retrieval technology, which is used to find a plurality of integral related content items most similar to the target user features in a high-dimensional feature space, i.e. neighbors.

[0120] Optionally, the server performs nearest neighbor query retrieval on the candidate content pool according to the preset recall quantity and the matching degree score of each integral related content, recalls a plurality of candidate integral related contents corresponding to the preset recall quantity from the candidate content pool, further sorts according to the score sum of the strategy score of each candidate integral related content and the matching degree score of each candidate integral related content, and determines target push contents of a preset number of user accounts with higher scores from the candidate integral related contents, such as ten candidate integral related contents for each user account as target push contents.

[0121] In the embodiment, the server first recalls a preset number of candidate integral related contents most relevant to the user from the large-scale candidate content pool by using nearest neighbor query retrieval according to the matching degree score, effectively reduces the calculation complexity, and improves the system response speed; then, on the basis of the recalled high-quality candidate contents, the matching degree score and the strategy score of the contents are comprehensively considered, a comprehensive score is generated by weighted fusion and sorting, and a plurality of contents with higher scores are finally determined as target push contents. The advantages of the double tower model in efficiently recalling relevant items in high-dimensional space are exerted, and the guiding ability of reinforcement learning to long-term user value is also integrated, realizing the collaborative optimization of precise matching and intelligent guidance, improving the relevance of push contents and user click conversion rate, and enhancing the positive incentive effect of the integral system on user behavior, significantly improving the operation efficiency and user experience.

[0122] In an exemplary embodiment, as shown in Figure 5 a content push system is provided, comprising:

[0123] a data collection layer for collecting transaction, browsing, sharing, login and other behavior logs;

[0124] A user stratification module performs RFM clustering and outputs user level labels L1…LN;

[0125] A behavior incentive module is configured to invoke a Forger model engine according to the level labels to generate a task pool TaskPool (a candidate content pool).

[0126] A personalized recommendation module is configured to fuse ALS (Alternating Least Squares) collaborative filtering and a Wide&Deep deep model (a combination of a wide and deep learning model) to generate a recommendation task TaskRec (an integral task) and a commodity ItemRec (an integral exchangeable resource).

[0127] A reinforcement learning decision module is configured to output an action a_t according to a state s_t by an online policy network πθ (a pre-trained reinforcement learning decision model) and deliver the action to a client (a terminal of a user account) through an A / B test gateway.

[0128] A feedback loop is configured to return a reward r_t from the client to a Replay Buffer training network.

[0129] In this embodiment, the key points of the embodiment are a multi-model fusion core architecture, a user stratification and differential rule linkage design, a behavior science-based task incentive mechanism, a data-driven personalized recommendation logic, and a real-time dynamic optimization capability. The multi-model collaboration breaks through the limitations of a single model. The fine-grained operation improves the user value mining efficiency. The scientifically driven user engagement improvement mechanism. The data-driven personalized matching improves the exchange rate and satisfaction. The dynamic self-optimization ensures long-term effectiveness, and finally realizes the dual improvement of the integral system operation effect and user satisfaction.

[0130] It should be understood that although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or steps or stages in other steps.

[0131] Based on the same inventive concept, the embodiments of the present application also provide a content pushing device for implementing the content pushing method described above. The implementation scheme of the device for solving the problem is similar to the implementation scheme described in the above method, so the specific limitations in one or more content pushing device embodiments provided below can refer to the limitations of the content pushing method in the above, which will not be described here.

[0132] In one exemplary embodiment, as shown in Figure 6 a content pushing device 600 is provided, comprising a user clustering module 601, a content pool generation module 602, a target determination module 603 and a content pushing module 604, wherein:

[0133] The user clustering module 601 is configured to perform clustering processing on each user account according to the behavior data of each user account in the business system, to obtain a hierarchical label of each user account.

[0134] The content pool generation module 602 is configured to generate a candidate content pool corresponding to the hierarchical label of each user account according to the hierarchical label and the behavior data; the candidate content pool comprises a plurality of points-related contents; each points-related content is a points-related task and / or a points-exchangeable resource.

[0135] The target determination module 603 is configured to determine a target pushing content from the candidate content pool based on the current state of the user account, the behavior data, the hierarchical label and the candidate content pool through a pre-trained recommendation model.

[0136] The content pushing module 604 is configured to push the target pushing content to the terminal of the user account.

[0137] Further, in one embodiment, the user clustering module 601 is further configured to determine the latest operation time, operation frequency and operation resource quota of each user account according to a preset time window and the behavior data; perform binning discretization processing on the latest operation time, operation frequency and operation resource quota of all user accounts respectively, to obtain discrete user accounts; perform clustering processing on the discrete user accounts by using a clustering algorithm, to obtain the hierarchical label of each user account.

[0138] Further, in one embodiment, the content pool generation module 602 is further configured to configure the trigger parameter, motivation parameter and ability parameter of a pre-constructed Fogg model according to the hierarchical label and the behavior data, to obtain a configured Fogg model; perform content instantiation processing according to the configured Fogg model, to obtain the candidate content pool corresponding to the hierarchical label of each user account.

[0139] Further, in an embodiment, the content pool generation module 602 is further configured to process the current state of the user account and the candidate content pool by using a pre-trained reinforcement learning decision model to obtain a strategy score of each credit-related content; process the behavior data, the hierarchical label and the candidate content pool by using a pre-trained dual tower model to obtain a matching degree score of each credit-related content; and determine the target push content of each user account from the candidate content pool according to the strategy score and the matching degree score of each credit-related content.

[0140] Further, in an embodiment, the content pool generation module 602 is further configured to process the current state of the user account and the candidate content pool by using a pre-trained reinforcement learning decision model to obtain a strategy action of each credit-related content; and obtain a strategy score of each credit-related content according to the strategy action of each credit-related content.

[0141] Further, in an embodiment, the content pool generation module 602 is further configured to process the behavior data, the hierarchical label and the candidate content pool by using a pre-trained dual tower model to obtain a user account feature and a credit-related content feature; determine a similarity between the user account feature and the credit-related content feature, and use the similarity as a matching degree score of each credit-related content.

[0142] Further, in an embodiment, the content push module 604 is further configured to perform a nearest neighbor query search on the candidate content pool according to a preset recall quantity and the matching degree score of each credit-related content to obtain a plurality of candidate credit-related contents; and determine the target push content of each user account from each candidate credit-related content according to the strategy score of each candidate credit-related content and the matching degree score of each candidate credit-related content.

[0143] Each module in the content push device 600 described above can be realized by software, hardware and a combination thereof in whole or in part. Each module described above can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform the operations corresponding to each module.

[0144] In an exemplary embodiment, a computer device is provided, which can be a server, and an internal structure diagram of the computer device can be as shown in FIG. 6. Figure 7As shown in the figure. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store behavior data, hierarchical labels, candidate content pool, target push content and the like. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement a content push method.

[0145] Those skilled in the art can understand that, Figure 7 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0146] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in each of the method embodiments described above.

[0147] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps in each of the method embodiments described above.

[0148] In one embodiment, a computer program product is provided, which includes a computer program, and the computer program is executed by a processor to implement the steps in each of the method embodiments described above.

[0149] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0150] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0151] The technical features of the above embodiments can be combined in any manner. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0152] The above-described embodiments are merely illustrative of several embodiments of the present application, which are described in more detail and in a specific manner, but should not be construed as limiting the scope of the patent of the present application. It should be noted that, for those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A content push method, characterized in that, The method includes: Based on the behavioral data of each user account in the business system, the user accounts are clustered to obtain hierarchical labels for each user account. Based on the hierarchical tags and the behavioral data, a candidate content pool is generated corresponding to each user account hierarchical tag; the candidate content pool includes multiple points-related content; each points-related content is a points-related task and / or a resource that can be exchanged for points. Based on the current state of the user account, the behavioral data, the hierarchical tags, and the candidate content pool, the target push content is determined from the candidate content pool using a pre-trained recommendation model; The target content is pushed to the user's account terminal.

2. The method according to claim 1, characterized in that, The step of clustering user accounts based on their behavioral data in the business system to obtain hierarchical labels for each user account includes: Based on the preset time window and the behavioral data, determine the most recent operation time, operation frequency, and operation resource limit for each user account; The most recent operation time, operation frequency, and operation resource quota of all user accounts are binned and discretized to obtain the discretized user accounts; Clustering algorithms are used to cluster the discrete user accounts to obtain hierarchical labels for each user account.

3. The method according to claim 1, characterized in that, The step of generating a candidate content pool corresponding to each user account hierarchical tag based on the hierarchical tags and the behavioral data includes: The trigger parameters, motivation parameters, and ability parameters of the pre-built Fogg model are configured based on the hierarchical labels and the behavioral data to obtain the configured Fogg model. Based on the configured Fogg model, content instantiation is performed to obtain the candidate content pool corresponding to each user account level tag.

4. The method according to claim 1, characterized in that, The pre-trained recommendation model includes a pre-trained dual-tower model and a pre-trained reinforcement learning decision model. Based on the user account's current state, the behavioral data, the hierarchical tags, and the candidate content pool, the pre-trained recommendation model determines the target content to be pushed from the candidate content pool, including: The pre-trained reinforcement learning decision model is used to process the current state of the user account and the candidate content pool to obtain the strategy score for each of the points-related contents. The pre-trained dual-tower model is used to process the behavioral data, the hierarchical labels, and the candidate content pool to obtain the matching score of each of the points-related contents; Based on the strategy score and the matching score of the points-related content, the target push content for each user account is determined from the candidate content pool.

5. The method according to claim 4, characterized in that, The pre-trained reinforcement learning decision model is used to process the current state of the user account and the candidate content pool to obtain policy scores for each of the points-related contents, including: The pre-trained reinforcement learning decision model is used to process the current state of the user account and the candidate content pool to obtain the policy actions for each of the points-related contents. Based on the strategy actions related to each of the points, a strategy score for each of the points-related contents is obtained.

6. The method according to claim 4, characterized in that, The pre-trained dual-tower model is used to process the behavioral data, the hierarchical labels, and the candidate content pool to obtain a matching score for each of the score-related contents, including: The pre-trained dual-tower model is used to process the behavioral data, the hierarchical labels, and the candidate content pool to obtain user account features and points-related content features. Determine the similarity between the user account features and the points-related content features, and use the similarity as a matching score for each of the points-related content features.

7. The method according to claim 4, characterized in that, The step of determining the target push content for each user account from the candidate content pool based on the strategy score and the matching score related to the points includes: Based on the preset recall quantity and the matching degree score of each of the points-related contents, the candidate content pool is searched by nearest neighbor query to obtain multiple candidate points-related contents; Based on the strategy score and the matching score of each candidate points-related content, the target push content for each user account is determined from the candidate points-related content.

8. A content push device, characterized in that, The device includes: The user clustering module is used to cluster the user accounts based on the behavioral data of each user account in the business system to obtain the hierarchical labels of each user account. The content pool generation module is used to generate a candidate content pool corresponding to each user account level tag based on the level tags and the behavior data; the candidate content pool includes multiple points-related content; each point-related content is a points-related task and / or a point-exchangeable resource; The target determination module is used to determine the target push content from the candidate content pool based on the current state of the user account, the behavioral data, the hierarchical tags, and the candidate content pool, using a pre-trained recommendation model. The content push module is used to push the target content to the user's account terminal.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.