A user retention prediction system and method based on behavior feature modeling

CN122777918APending Publication Date: 2026-09-18SHENZHEN DREAM WORKSHOP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610934738.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0003]现有留存预测方法在建模用户行为时难以捕捉用户行为随时间衰减的规律,传统方法通常基于固定时间窗口内的行为频次统计,忽略了用户操作间隔周期对留存意愿的衰减影响,导致对行为节奏不规律的用户预测偏差较大,同时忽视了用户注意力在页面间转移路径对留存的影响,传统方法多采用日活或访问深度等聚合特征,缺乏对注意力在焦点页面之间流动方向与强度的结构化建模,无法识别关键行为路径的断裂风险

Benefits of technology

[0041] This application provides a user retention prediction system and method based on behavioral feature modeling. The system collects login behavior sequences, function click sequences, and page dwell time of target users within a specified time window, thereby extracting user behavioral habit features and operation interval cycles. Based on the behavioral habit features, a user behavior rhythm vector is constructed. The user's behavior decay curve is fitted using the behavioral rhythm vector and the operation interval cycle to obtain the user's habitual retention potential at each preset time node. The user's attention shift features are extracted from the page dwell time. A user behavior retention feature map is established based on the attention shift features and the habitual retention potential. The user's layer-by-layer retention probability at different lifecycle stages is calculated using the behavior retention feature map. The layer-by-layer retention probability is compared with the actual retention label. The behavioral weights in the behavioral rhythm vector are adjusted based on the comparison results, and the predicted retention rate of the user in the next time window is output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122777918A_ABST
    Figure CN122777918A_ABST
Patent Text Reader

Abstract

The application provides a user retention prediction system and method based on behavior feature modeling, constructs a user behavior rhythm vector based on user behavior habit features, fits a user behavior decay curve through the behavior rhythm vector and a user operation interval period, obtains habit retention potential of the user at each preset time node, extracts a user attention transfer feature from a user page stay duration, establishes a user behavior retention feature map according to the attention transfer feature and the habit retention potential, calculates layer-by-layer retention probability of the user at different life cycle stages through the behavior retention feature map, compares the layer-by-layer retention probability with a real retention label, corrects a behavior weight in the behavior rhythm vector according to a comparison result, and outputs a predicted retention rate of the user in a next time window. Based on the above scheme, joint modeling of behavior rhythm and attention transfer can be realized, thereby improving long-term accuracy of retention prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of behavioral feature modeling technology, and more specifically, to a user retention prediction system and method based on behavioral feature modeling. Background Technology

[0002] User retention prediction refers to the process of estimating the probability of users continuing to use a product or service within a specified future time window by modeling user behavior patterns such as login, function clicks, and page dwell times based on historical user behavior data. This prediction serves product operation and user growth scenarios, helps identify high-value users, warns of churn risks, and provides a quantitative basis for personalized intervention strategies.

[0003] Existing retention prediction methods struggle to capture the decay patterns of user behavior over time when modeling user behavior. Traditional methods typically rely on frequency statistics within fixed time windows, neglecting the impact of user action intervals on retention intention. This leads to significant prediction biases for users with irregular behavior rhythms. Furthermore, they ignore the influence of user attention shifts between pages on retention. Traditional methods often employ aggregated features such as daily active users (DAU) or visit depth, lacking structured modeling of the direction and intensity of attention flow between focused pages, and failing to identify the risk of breakage in key behavioral paths. These two limitations combine to make it difficult for prediction models to distinguish between users with stable rhythms and those with fluctuating intervals, and they also lack utilization of the topological structure of behavioral sequences, ultimately reducing the accuracy of retention predictions over long periods and under complex behavioral patterns. Therefore, how to achieve joint modeling of behavioral rhythms and attention shifts to improve the long-term accuracy of retention predictions has become a challenge for the industry. Summary of the Invention

[0004] This application provides a user retention prediction system and method based on behavioral feature modeling, which can achieve joint modeling of behavioral rhythm and attention shift, thereby improving the long-term accuracy of retention prediction.

[0005] Firstly, this application provides a user retention prediction method based on behavioral feature modeling, including:

[0006] Collect the login behavior sequence, function click sequence and page dwell time of target users within a specified time window, and then extract the user's behavioral habit characteristics and operation interval cycle;

[0007] Based on the behavioral habit characteristics, a user's behavior rhythm vector is constructed. The user's behavior decay curve is fitted by the behavior rhythm vector and the operation interval period to obtain the user's habit retention potential at each preset time node.

[0008] Extract user attention shift characteristics from the page dwell time, establish user behavior retention feature map based on the attention shift characteristics and the habit retention potential, and calculate the layer-by-layer retention probability of users at different life cycle stages through the behavior retention feature map;

[0009] The layer-by-layer retention probability is compared with the actual retention label. Based on the comparison result, the behavior weights in the behavior rhythm vector are adjusted, and the predicted retention rate of the user in the next time window is output.

[0010] In some embodiments, extracting user behavioral habit characteristics and operation intervals specifically includes:

[0011] The login frequency of users in different time periods is counted from the login behavior sequence to obtain the user's time period preference characteristics;

[0012] The time difference between adjacent clicks is determined by the function click sequence, thereby obtaining the user's operation rhythm characteristics;

[0013] The time period preference feature and the operation rhythm feature are combined into the user's behavioral habit feature, and the average click interval is used as the operation interval period.

[0014] In some embodiments, constructing a user's behavioral rhythm vector based on the behavioral habit features specifically includes:

[0015] Encode the time-period preference feature in the behavioral habit features into a time-period distribution vector;

[0016] The operation rhythm feature in the behavioral habit features is normalized and then appended to the end of the time period distribution vector to obtain the original behavior vector;

[0017] The original behavior vector is normalized by its magnitude to obtain the user's behavior rhythm vector.

[0018] In some embodiments, fitting the user's behavior decay curve with the behavior rhythm vector and the operation interval period to obtain the user's habit retention potential at each preset time node specifically includes:

[0019] Set multiple preset time nodes along the positive time axis;

[0020] Using the behavior rhythm vector as the initial intensity, the user's attenuation coefficient at each preset time node is calculated using the operation interval period;

[0021] The user's habitual retention potential at each preset time point is determined by the initial intensity and various attenuation coefficients.

[0022] In some embodiments, extracting user attention shift features from the page dwell time specifically includes:

[0023] Based on the page dwell time, multiple focus pages are identified from all the pages the user stays on, and then the jump frequency between each focus page is counted.

[0024] All jump frequencies are constructed into a directed transition matrix to obtain the user's attention shift characteristics.

[0025] In some embodiments, establishing a user behavior retention feature map based on the attention shift characteristics and the habit retention potential specifically includes:

[0026] A directed graph is constructed using the user's various focused pages as nodes and the jump probability as the edge weight;

[0027] The initial weights of each node are extracted from the habitual retention potential energy to obtain the weights of multiple nodes in the directed graph.

[0028] By constructing a weighted directed graph from all node weights and edge weights, we can obtain the user behavior retention feature map.

[0029] In some embodiments, calculating the tiered retention probability of a user at different lifecycle stages using the behavioral retention feature map specifically includes:

[0030] The behavior retention feature map is sliced ​​using a preset time window, with each time window corresponding to a life cycle stage, to obtain the retention potential transfer value for each life cycle stage.

[0031] The retention potential value is accumulated along the directed edge starting from the starting node in the behavior retention feature graph;

[0032] The retention probability of users at different lifecycle stages is determined by accumulating potential energy across all stages.

[0033] Secondly, this application provides a user retention prediction system based on behavioral feature modeling, including a behavior prediction unit, the behavior prediction unit comprising:

[0034] The data collection module is used to collect the login behavior sequence, function click sequence and page dwell time of the target user within a specified time window, and then extract the user's behavioral habit characteristics and operation interval cycle.

[0035] The processing module is used to construct a user's behavior rhythm vector based on the behavioral habit characteristics, and to fit the user's behavior decay curve by the behavior rhythm vector and the operation interval period to obtain the user's habit retention potential at each preset time node.

[0036] The processing module is also used to extract the user's attention shift characteristics from the page dwell time, establish the user's behavior retention feature map based on the attention shift characteristics and the habit retention potential, and calculate the user's layer-by-layer retention probability at different life cycle stages through the behavior retention feature map.

[0037] The execution module is used to compare the layer-by-layer retention probability with the actual retention label, adjust the behavior weight in the behavior rhythm vector according to the comparison result, and output the predicted retention rate of the user in the next time window.

[0038] Thirdly, this application provides a computer device, the computer device including a memory and a processor, the memory for storing a computer program, and the processor for calling and running the computer program from the memory, so that the computer device performs the above-described user retention prediction method based on behavioral feature modeling.

[0039] Fourthly, this application provides a computer-readable storage medium storing instructions or code that, when executed on a computer, cause the computer to implement the aforementioned user retention prediction method based on behavioral feature modeling.

[0040] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects:

[0041] This application provides a user retention prediction system and method based on behavioral feature modeling. The system collects login behavior sequences, function click sequences, and page dwell time of target users within a specified time window, thereby extracting user behavioral habit features and operation interval cycles. Based on the behavioral habit features, a user behavior rhythm vector is constructed. The user's behavior decay curve is fitted using the behavioral rhythm vector and the operation interval cycle to obtain the user's habitual retention potential at each preset time node. The user's attention shift features are extracted from the page dwell time. A user behavior retention feature map is established based on the attention shift features and the habitual retention potential. The user's layer-by-layer retention probability at different lifecycle stages is calculated using the behavior retention feature map. The layer-by-layer retention probability is compared with the actual retention label. The behavioral weights in the behavioral rhythm vector are adjusted based on the comparison results, and the predicted retention rate of the user in the next time window is output.

[0042] Therefore, in this application, the layer-by-layer retention probability is compared with the actual retention label, and the behavioral weights in the behavioral rhythm vector are corrected based on the comparison results to output the predicted retention rate of the user in the next time window. First, by determining the habitual retention potential, a quantitative representation of the user's behavior intensity decaying with the operation interval and time can be obtained. The multi-dimensional behavior patterns and operation interval periods in the behavioral rhythm vector are fitted into a decay curve, so that the potential value at each preset time node can reflect the behavior decay law of the user under uneven access intervals. The habitual retention potential based on the decay mechanism can effectively distinguish between high-frequency but irregularly spaced users and low-frequency but stable rhythm users. In long-term retention prediction, the model can avoid misjudgment caused by short-term behavioral fluctuations, enhance the ability to capture the stability of user behavior habits, and thus improve the accuracy of cross-time window prediction. Then, by determining the layer-by-layer retention probability, the user's retention rate in different time windows can be obtained. This approach employs a hierarchical measurement of behavioral retention capabilities across different lifecycle stages, enabling the fusion and correction of behavioral rhythm and attention shifts in joint modeling. By using node and edge weights in the behavioral retention feature graph, the habitual retention potential is accumulated layer by layer along the attention shift path. This ensures that the retention probability at each lifecycle stage depends not only on the user's own rhythm stability but also on the direction and intensity of attention flow between focused pages. Graph-based probability calculations can characterize the inheritance relationship between different stages of user behavior, avoiding information breaks caused by modeling each stage in isolation. In long-term predictions, the model can adaptively adjust its attention to different behavioral paths, reducing sensitivity to noise jumps and thus improving the continuous accuracy of retention prediction across multiple lifecycle stages. In summary, this approach enables joint modeling of behavioral rhythm and attention shifts, thereby improving the long-term accuracy of retention prediction. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is an exemplary flowchart of a user retention prediction method based on behavioral feature modeling, according to some embodiments of this application;

[0045] Figure 2 It is a parallel logic diagram modeled according to some embodiments of this application, showing behavioral characteristics;

[0046] Figure 3 This is a flowchart illustrating the process of determining a behavioral retention feature map according to some embodiments of this application;

[0047] Figure 4 This is a schematic diagram of the structure of a behavior prediction unit according to some embodiments of this application;

[0048] Figure 5 This is a schematic diagram of the structure of a computer device that implements a user retention prediction method based on behavioral feature modeling, according to some embodiments of this application. Detailed Implementation

[0049] To better understand the technical solution of this application, the technical solution of this application will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0050] refer to Figure 1 The figure is an exemplary flowchart of a user retention prediction method based on behavioral feature modeling according to some embodiments of this application. The user retention prediction method based on behavioral feature modeling mainly includes the following steps:

[0051] In step 101, the login behavior sequence, function click sequence and page dwell time of the target user within a specified time window are collected, and then the user's behavioral habit characteristics and operation interval cycle are extracted.

[0052] It should be noted that, in this application, the login behavior sequence is an ordered set used to record each login action performed by the user in chronological order; the function click sequence is an ordered set used to record each function entry operation triggered by the user in chronological order; and the page dwell time is a set of values ​​used to measure the length of time a user spends from entering a page to leaving the page.

[0053] In practice, firstly, all access records of the target user within a specified time window are filtered from the system's backend logs and frontend event tracking data. This time window is set to the last thirty days by default to ensure that the amount of data is sufficient to reflect user habits. Secondly, the login behavior sequence is extracted from all access records. That is, in chronological order, the moment when the user clicks the login button or completes login through WeChat authorization is extracted. Arranging the login moments in sequence yields a list of login actions sorted by time, and this login action list is used as the login behavior sequence. Then, the function click sequence is extracted from the same access records. In short: Filter out non-functional operations such as page switching and scrolling, and only retain operation events with clear business functions, such as user-initiated clicks on buttons, links, and menus. Arrange the trigger times of each operation event and its corresponding function name in chronological order to obtain a function click list, which is used as the function click sequence. Then, extract the dwell time for each page: For each record of a user entering a page, record the entry timestamp; when the user leaves the page and navigates to the next page, record the exit timestamp. Subtract the entry timestamp from the exit timestamp to obtain a dwell time value. For the last page visited by the user, if there is no explicit exit record, use the time when the user exits the system or closes the page as the cutoff point to calculate the dwell time for all pages, thus obtaining a set of dwell time values. The collection of all dwell time values ​​is taken as the page dwell time.

[0054] It should be noted that in this application, Figure 2 The parallel logic diagram for behavioral feature modeling adopts a multi-link parallel advancement and node convergence linkage logical structure, which sequentially covers five parallel pre-processing stages: user behavior data collection, behavioral feature extraction, feature preprocessing, sample dataset partitioning, and feature correlation analysis. Each branch runs synchronously and independently before being uniformly merged into the behavioral feature modeling center module. The modeling stage relies on a multi-model parallel training mechanism to generate a standardized user retention prediction model. The model output is then split into three parallel application branches: user retention probability prediction, user group segmentation, and retention key influencing factor attribution analysis. The entire diagram fully presents the parallel business logic of the entire chain from the underlying behavioral data source, multi-dimensional feature engineering, parallel modeling training to retention prediction, user segmentation, and factor attribution. Each process node has a clear division of labor and works in parallel collaboration, realizing the standardization and structured implementation of the entire process of behavioral feature modeling and user retention prediction.

[0055] In some embodiments, extracting user behavioral habit characteristics and operation intervals can be achieved through the following steps:

[0056] The login frequency of users in different time periods is counted from the login behavior sequence to obtain the user's time period preference characteristics;

[0057] The time difference between adjacent clicks is determined by the function click sequence, thereby obtaining the user's operation rhythm characteristics;

[0058] The time period preference feature and the operation rhythm feature are combined into the user's behavioral habit feature, and the average click interval is used as the operation interval period.

[0059] It should be noted that, in this application, the time period preference feature is a set of values ​​used to represent the degree of a user's tendency to log in at different time periods within a day; the operation rhythm feature is two statistical measures used to reflect the concentration and fluctuation range of the time interval when a user performs adjacent function clicks; the behavior habit feature is a combination vector used to comprehensively describe when a user logs in and at what speed they click on functions; and the operation interval period is a value used to measure the average waiting time between two adjacent function clicks by a user.

[0060] In practical implementation, firstly, the 24 hours of a day are divided into 24 time periods, each hour representing a segment. Each login record in the login behavior sequence is retrieved one by one, and the hourly interval in which the login time falls is determined. A count is incremented for the corresponding hourly interval. After traversing all login records, each hourly interval has a cumulative login frequency. The 24 frequency values ​​are arranged in hourly order into a list containing 24 values, which serves as the time period preference feature. Then, the occurrence time of each function click event is extracted from the function click sequence and, according to the chronological order in the sequence, is used... The time difference is calculated by subtracting the time of the previous click from the time of the subsequent click. This subtraction operation is repeated until all adjacent records in the sequence have been calculated, resulting in a set of time differences. The set of average and standard deviation of this set of differences is taken as the user's operation rhythm feature. Finally, the list of time period preference features containing twenty-four values ​​is concatenated with the list of operation rhythm features containing only the average and standard deviation values ​​in chronological order to form a new list containing twenty-six values. This list is taken as the user's behavioral habit feature. The average of the set of time differences calculated above is taken as the operation interval period.

[0061] In step 102, a user's behavior rhythm vector is constructed based on the behavioral habit characteristics. The user's behavior decay curve is fitted by the behavior rhythm vector and the operation interval period to obtain the user's habit retention potential at each preset time node.

[0062] In some embodiments, constructing a user's behavioral rhythm vector based on the behavioral habit features can be achieved through the following steps:

[0063] Encode the time-period preference feature in the behavioral habit features into a time-period distribution vector;

[0064] The operation rhythm feature in the behavioral habit features is normalized and then appended to the end of the time period distribution vector to obtain the original behavior vector;

[0065] The original behavior vector is normalized by its magnitude to obtain the user's behavior rhythm vector.

[0066] It should be noted that in this application, the time period distribution vector is an encoding form used to convert the user's time period preference features into a fixed-dimensional numerical list; the original behavior vector is an intermediate numerical list that contains both user time period preference and operation rhythm information before length adjustment; and the behavior rhythm vector is a standard numerical list used to retain only the structural features of the behavior pattern.

[0067] In practice, firstly, time-period preference features are extracted from behavioral habit characteristics. This feature is a list of twenty-four values, each corresponding to the login frequency in one hour of the day. Maintaining the original order of these twenty-four values, they are sequentially extracted and placed into a new list, which serves as the time-period distribution vector. Next, operation rhythm features are extracted from the behavioral habit characteristics. This feature contains two values: mean and standard deviation. The mean is normalized by subtracting the minimum of all user averages from the mean, and then dividing the difference by the difference between the maximum and minimum of all user averages. The standard deviation is calculated using the same method to obtain the normalized mean. The standard deviation is used as the basis for the time period distribution vector. The normalized mean and standard deviation are appended to the end of the vector to form a new list containing twenty-six values. This new list is used as the original behavior vector. Finally, the twenty-six values ​​in the original behavior vector are recorded as the first to the twenty-sixth. The modulus of the original behavior vector is calculated by multiplying each value by itself to get the square value. All the square values ​​are added together to get a sum. The square root of the sum is then taken to get the modulus. Each value in the original behavior vector is divided by this modulus to get twenty-six new values, which are then arranged in the original order. This new list, after being normalized by the modulus, is used as the user's behavior rhythm vector.

[0068] In some embodiments, the user's habit retention potential at each preset time node can be obtained by fitting the user's behavior decay curve with the behavior rhythm vector and the operation interval period using the following steps:

[0069] Set multiple preset time nodes along the positive time axis;

[0070] Using the behavior rhythm vector as the initial intensity, the user's attenuation coefficient at each preset time node is calculated using the operation interval period;

[0071] The user's habitual retention potential at each preset time point is determined by the initial intensity and various attenuation coefficients.

[0072] It should be noted that, in this application, the preset time node is used to define the specific moment when the user's behavior retention probability is evaluated on the future timeline; the initial strength is a benchmark value used to represent the overall strength of the user's behavioral habits at the current moment; the decay coefficient is a proportional factor used to characterize the decrease in the user's behavior retention probability as time and operation intervals extend; and the habit retention potential is a value used to quantify the degree to which the user continues to maintain the original behavioral habits at a specified future time node.

[0073] In practice, firstly, starting from the current moment, multiple specific moments are set as preset time nodes along the time axis in the future direction. The interval between any two adjacent nodes is equal, and this interval is set to one day by default. A total of seven nodes are set, corresponding to the first day, the second day, and so on up to the seventh day in the future. Then, the value obtained after normalizing the behavior rhythm vector is used as the initial intensity. This value represents the overall intensity of the user's behavioral habits. For each preset time node, the number of days corresponding to that node is first calculated. Using the natural constant as the base, the ratio of the number of days to the operation interval period is negatively taken as the exponent. The value of the exponential function is then calculated, and this value is used as the... The longer the operation interval, the slower the decay coefficient decreases with the number of days. Finally, for each preset time node, the initial intensity is multiplied by the decay coefficient corresponding to that node to obtain a product value. All product values ​​at that node are summed, and the sum is taken as the habit retention potential at that node. The above multiplication and summation operations are performed on the seven preset time nodes in turn to obtain seven habit retention potential values, which correspond to the probability that the user will continue to maintain the behavior habit from the first day to the seventh day. The habit retention potential of the user at each preset time node can be obtained in the above way.

[0074] In step 103, user attention shift features are extracted from the page dwell time, and user behavior retention feature map is established based on the attention shift features and the habit retention potential. The user's layer-by-layer retention probability at different life cycle stages is calculated through the behavior retention feature map.

[0075] In some embodiments, extracting user attention shift features from the page dwell time can be achieved through the following steps:

[0076] Based on the page dwell time, multiple focus pages are identified from all the pages the user stays on, and then the jump frequency between each focus page is counted.

[0077] All jump frequencies are constructed into a directed transition matrix to obtain the user's attention shift characteristics.

[0078] It should be noted that, in this application, the focus page is used to identify the page where the user stays for more than a set threshold during browsing; the jump frequency is a value used to record the number of times the user jumps directly from one focus page to another; the directed transition matrix is ​​a two-dimensional table structure used to organize the number of jumps or probabilities between all focus pages in terms of direction; and the attention transfer feature is a feature representation used to describe the direction and intensity of the user's attention flow between various focus pages.

[0079] In practice, firstly, the names of all pages visited by the user and the corresponding duration of each page are extracted from the page dwell time. A duration threshold is set, which defaults to three seconds. Pages with a dwell time longer than three seconds are marked as focus pages. All pages marked as focus pages are then filtered out from all pages to obtain a focus page list. Next, the user's function click sequence is traversed to find all jump records occurring between two focus pages. For each jump from page A to page B, the corresponding jump count is incremented by one. After the traversal is complete, each pair of jumps from the source page to the target page is... A cumulative number of jumps is calculated, and the set of all jump counts is taken as the jump frequency. Then, the total number of pages in the focus page list is counted and denoted as N. An N-row N-column blank table is constructed, where each row represents the starting page of the jump and each column represents the target page of the jump. Each jump frequency obtained is filled into the corresponding row and column intersection position in the table. If no jump occurs between some pages, the value zero is filled into the corresponding position. After filling all positions, an N-row N-column two-dimensional numerical table is obtained. This table is used as a directed transition matrix, and this directed transition matrix is ​​directly used as the user's attention transfer feature.

[0080] In some embodiments, a user behavior retention feature map is established based on the attention shift characteristics and the habit retention potential, with reference to... Figure 3 The diagram is a flowchart illustrating the process of determining the behavioral retention feature map in some embodiments of this application. In this embodiment, determining the behavioral retention feature map can be achieved through the following steps:

[0081] In step 1031, a directed graph is constructed using each of the user's focused pages as nodes and the jump probability as the edge weight;

[0082] In step 1032, the initial weights of each node are extracted from the habitual retention potential energy to obtain the weights of multiple nodes in the directed graph.

[0083] In step 1033, all node weights and edge weights are constructed into a weighted directed graph, thereby obtaining the user behavior retention feature graph.

[0084] It should be noted that in this application, a directed graph is a set structure of nodes and edges used to represent the directional jump relationship between focus pages; node weight is a value used to represent the likelihood of a user's habitual retention at a specified time node for each focus page; a weighted directed graph is a graph structure used to simultaneously include the weight of the node itself and the weight of the directional jump between nodes; and a behavior retention feature graph is a feature graph representation that reflects the intensity of user behavior habits and the pattern of attention shift.

[0085] In practice, firstly, all distinct page names are retrieved from the focus page list, and each page name is treated as a node. The jump probability is extracted from the directed transition matrix. The jump probability is calculated as follows: for each row, the jump frequency of each page name in that row is divided by the sum of all jump frequencies in that row to obtain the probability value of jumping from the corresponding page in that row to the corresponding page in each column. Using each node as the starting and ending point, and the jump probability as the weight value of the connecting edge, directed edges are drawn according to the direction from the source node to the target node, forming a directed graph. Then, the user's potential energy value at the first preset time node in the future is retrieved from the habit retention potential energy. This value corresponds to the first page in the focus page list. Following the same... Based on the correspondence, the potential energy values ​​at subsequent time nodes in the habit retention potential energy are extracted sequentially and assigned to the second page, the third page, and so on in the focus page list, until each focus page is assigned a habit retention potential energy value. The assigned values ​​are used as the initial weights of each node, thus obtaining multiple node weights in the directed graph. Finally, all nodes and directed edges in the constructed directed graph, as well as the jump probability weights on each directed edge, are retained. Each node weight is attached to the corresponding node, so that each node carries a value representing habit retention potential energy in addition to its own identifier. The directed graph that simultaneously contains node weights and edge weights is used as the user's behavior retention feature graph.

[0086] In some embodiments, calculating the tiered retention probability of a user at different lifecycle stages using the behavioral retention feature map can be achieved through the following steps:

[0087] The behavior retention feature map is sliced ​​using a preset time window, with each time window corresponding to a life cycle stage, to obtain the retention potential transfer value for each life cycle stage.

[0088] The retention potential value is accumulated along the directed edge starting from the starting node in the behavior retention feature graph;

[0089] The retention probability of users at different lifecycle stages is determined by accumulating potential energy across all stages.

[0090] It should be noted that, in this application, the lifecycle stage is a continuous time window used to identify different periods in a user's usage cycle; the retention potential transfer value is a value used to represent the retention ability of user behavior habits transferred from one node to another along a directed edge; the stage cumulative potential is a value used to represent the sum of the retention potential transfer values ​​on all nodes within a lifecycle stage; and the layer-by-layer retention probability is a value used to represent the probability that a user will successfully transition from the current lifecycle stage to the next lifecycle stage.

[0091] In practice, firstly, the length of each time window is set to seven days. The entire timeline is divided into segments of seven days each. The first seven days correspond to the first lifecycle phase, the second seven days correspond to the second lifecycle phase, and so on. For each lifecycle phase, nodes in the behavior retention feature graph that fall within that phase are retained according to their original connections, while nodes and edges outside that phase are removed, forming a subgraph for that phase. For each node in each subgraph, the habit retention potential carried by that node is divided by the number of days in that phase to obtain the node's position in that phase. The average potential energy for each day is used as the retention potential energy transfer value. Then, the starting node in the behavior retention feature map is identified. The starting node is the node corresponding to the first focused page visited by the user within the first time window. The retention potential energy transfer value of the starting node is used as the initial accumulation value. All directly pointed neighbor nodes are found along the directed edges emanating from the starting node. The retention potential energy transfer value of the starting node is multiplied by the calculated jump probability on the directed edge as the potential energy transferred to each neighbor node, and this potential energy is added to the original retention potential energy transfer of the neighbor node. In terms of value, the potential energy is passed down from each neighboring node to the next level node along the directed edges it emanates from. For each directed edge traversed, the accumulated potential energy of the current node is multiplied by the jump probability of that edge and then added to the target node. This process is repeated until all reachable nodes have been traversed. Finally, the accumulated potential energy values ​​of all nodes in each lifecycle stage are summed to obtain the stage cumulative potential energy. The stage cumulative potential energy is calculated for the first, second, up to the seventh lifecycle stage, resulting in seven stage cumulative potential energy values. For the first lifecycle stage, its stage cumulative potential energy is used as the base value. For the second lifecycle stage, the stage cumulative potential energy of the second stage is divided by the stage cumulative potential energy of the first stage to obtain the layer-by-layer retention probability from the first stage to the second stage. For the third lifecycle stage, the stage cumulative potential energy of the third stage is divided by the stage cumulative potential energy of the second stage to obtain the layer-by-layer retention probability from the second stage to the third stage, and so on, until the layer-by-layer retention probability from the sixth stage to the seventh stage is calculated, resulting in six layer-by-layer retention probability values.

[0092] In step 104, the layer-by-layer retention probability and the actual retention label are compared, the behavior weights in the behavior rhythm vector are corrected according to the comparison results, and the predicted retention rate of the user in the next time window is output.

[0093] It should be noted that in this application, the real retention label is a numerical value used to indicate whether the user actually completes the retention behavior in the next time window; the behavior weight is an adjustable parameter used to control the degree of influence of each dimension feature in the behavior rhythm vector on the retention prediction result; and the predicted retention rate is a numerical value used to indicate the probability that the user will continue to use the system in the next time window.

[0094] In practice, firstly, the actual user retention rate for the next time window is obtained from historical data. If the user completes at least one login or function click within that window, the actual retention tag is marked with a value of one; otherwise, it is marked with a value of zero. Secondly, the probability value of the last stage in the tiered retention probability is used as the current predicted retention rate. This predicted retention rate is compared with the actual retention tag by calculating the difference between the predicted retention rate and the actual retention tag. If the difference is greater than zero, it indicates that the predicted value is higher than the actual value; if the difference is less than zero, it indicates that the predicted value is lower than the actual value; if the difference is equal to zero, it indicates that the prediction is completely accurate. Then, the rows in the behavior rhythm vector are adjusted based on the comparison results. The weighting and correction method is as follows: when the predicted retention rate is higher than the actual retention label, the weight value of the corresponding time period in the behavior rhythm vector is reduced to decrease the influence of that time period on the prediction result; when the predicted retention rate is lower than the actual retention label, the weight value of the corresponding time period in the behavior rhythm vector is increased to increase the influence of that time period on the prediction result. The correction magnitude of the weight is the result of multiplying the absolute value of the difference by a preset small step size coefficient. Finally, the corrected behavior rhythm vector is substituted back into the previous calculation process, from constructing the behavior decay curve to calculating the layer-by-layer retention probability, to obtain a new predicted retention rate. This new predicted retention rate is used as the predicted retention rate of the user in the next time window.

[0095] Furthermore, in another aspect of this application, in some embodiments, this application provides a user retention prediction system based on behavioral feature modeling, which includes a behavioral prediction unit, referencing... Figure 4 The figure is a schematic diagram of the structure of a behavior prediction unit according to some embodiments of this application. The behavior prediction unit includes: a data acquisition module 201, a processing module 202, and an execution module 203, which are described below:

[0096] The data collection module 201 in this application is mainly used to collect the login behavior sequence, function click sequence and page dwell time of the target user within a specified time window, and then extract the user's behavioral habit characteristics and operation interval cycle.

[0097] Processing module 202, in this application, is used to construct a user's behavior rhythm vector based on the behavioral habit characteristics, and to fit the user's behavior decay curve by the behavior rhythm vector and the operation interval period to obtain the user's habit retention potential at each preset time node.

[0098] It should be noted that the processing module 202 is also used to extract the user's attention shift characteristics from the page dwell time, establish the user's behavior retention feature map based on the attention shift characteristics and the habit retention potential, and calculate the user's layer-by-layer retention probability at different life cycle stages through the behavior retention feature map.

[0099] The execution module 203 in this application is mainly used to compare the layer-by-layer retention probability with the actual retention label, correct the behavior weight in the behavior rhythm vector according to the comparison result, and output the predicted retention rate of the user in the next time window.

[0100] The foregoing has detailed examples of user retention prediction systems and methods based on behavioral feature modeling provided in this application. It is understood that the corresponding apparatus, in order to achieve the above functions, includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specified application, but such implementation should not be considered beyond the scope of this application.

[0101] In some embodiments, this application also provides a computer device, the computer device including a memory and a processor, the memory for storing a computer program, and the processor for calling and running the computer program from the memory, so that the computer device performs the above-described user retention prediction method based on behavioral feature modeling.

[0102] In some embodiments, reference Figure 5 The dashed lines in the figure indicate that the unit or module is optional. This figure is a schematic diagram of the structure of a computer device implementing a user retention prediction method based on behavioral feature modeling according to an embodiment of this application. The user retention prediction method based on behavioral feature modeling described in the above embodiments can... Figure 5The computer device shown is used to implement this, and the computer device includes at least one processor 301, a memory 302 and at least one communication unit 305. The computer device may be a terminal device, a server or a chip.

[0103] Processor 301 can be a general-purpose processor or a special-purpose processor. For example, processor 301 can be a central processing unit (CPU), which can be used to control computer devices, execute software programs, and process data from software programs. The computer device may also include a communication unit 305 for inputting (receiving) and outputting (transmitting) signals.

[0104] For example, the computer device may be a chip, and the communication unit 305 may be the input and / or output circuit of the chip, or the communication unit 305 may be the communication interface of the chip, which may be a component of a terminal device, network device or other device.

[0105] For example, the computer device may be a terminal device or a server, and the communication unit 305 may be a transceiver of the terminal device or the server, or the communication unit 305 may be a transceiver circuit of the terminal device or the server.

[0106] The computer device may include one or more memories 302 storing a program 304. The program 304 can be executed by a processor 301 to generate instructions 303, causing the processor 301 to execute the method described in the above method embodiments according to the instructions 303. Optionally, the memory 302 may also store data (such as a target audit model). Optionally, the processor 301 may also read data stored in the memory 302, which may be stored at the same storage address as the program 304, or it may be stored at a different storage address than the program 304.

[0107] The processor 301 and memory 302 can be configured separately or integrated together, for example, integrated on the system on chip (SOC) of the terminal device.

[0108] It should be understood that each step of the above method embodiment can be completed by hardware logic circuits or software instructions in the processor 301. The processor 301 can be a CPU, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, such as discrete gate, transistor logic devices, or discrete hardware components.

[0109] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0110] For example, in some embodiments, this application also provides a computer-readable storage medium storing instructions or code that, when executed on a computer, cause the computer to implement the above-described user retention prediction method based on behavioral feature modeling.

[0111] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0112] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A user retention prediction method based on behavioral feature modeling, characterized in that, Includes the following steps: Collect the login behavior sequence, function click sequence and page dwell time of target users within a specified time window, and then extract the user's behavioral habit characteristics and operation interval cycle; Based on the behavioral habit characteristics, a user's behavior rhythm vector is constructed. The user's behavior decay curve is fitted by the behavior rhythm vector and the operation interval period to obtain the user's habit retention potential at each preset time node. Extract user attention shift characteristics from the page dwell time, establish user behavior retention feature map based on the attention shift characteristics and the habit retention potential, and calculate the layer-by-layer retention probability of users at different life cycle stages through the behavior retention feature map; The layer-by-layer retention probability is compared with the actual retention label. Based on the comparison result, the behavior weights in the behavior rhythm vector are adjusted, and the predicted retention rate of the user in the next time window is output.

2. The method as described in claim 1, characterized in that, Extracting user behavior patterns and operation intervals specifically includes: The login frequency of users in different time periods is counted from the login behavior sequence to obtain the user's time period preference characteristics; The time difference between adjacent clicks is determined by the function click sequence, thereby obtaining the user's operation rhythm characteristics; The time period preference feature and the operation rhythm feature are combined into the user's behavioral habit feature, and the average click interval is used as the operation interval period.

3. The method as described in claim 1, characterized in that, Constructing a user's behavioral rhythm vector based on the aforementioned behavioral habit characteristics specifically includes: Encode the time-period preference feature in the behavioral habit features into a time-period distribution vector; The operation rhythm feature in the behavioral habit features is normalized and then appended to the end of the time period distribution vector to obtain the original behavior vector; The original behavior vector is normalized by its magnitude to obtain the user's behavior rhythm vector.

4. The method as described in claim 1, characterized in that, By fitting the user's behavior decay curve using the behavior rhythm vector and the operation interval period, the user's habit retention potential at each preset time node is obtained, specifically including: Set multiple preset time nodes along the positive time axis; Using the behavior rhythm vector as the initial intensity, the user's attenuation coefficient at each preset time node is calculated using the operation interval period; The user's habitual retention potential at each preset time point is determined by the initial intensity and various attenuation coefficients.

5. The method as described in claim 1, characterized in that, Extracting user attention shift features from the page dwell time specifically includes: Based on the page dwell time, multiple focus pages are identified from all the pages the user stays on, and then the jump frequency between each focus page is counted. All jump frequencies are constructed into a directed transition matrix to obtain the user's attention shift characteristics.

6. The method as described in claim 1, characterized in that, The establishment of a user behavior retention feature map based on the attention shift characteristics and the habit retention potential specifically includes: A directed graph is constructed using the user's various focused pages as nodes and the jump probability as the edge weight; The initial weights of each node are extracted from the habitual retention potential energy to obtain the weights of multiple nodes in the directed graph. By constructing a weighted directed graph from all node weights and edge weights, we can obtain the user behavior retention feature map.

7. The method as described in claim 1, characterized in that, Calculating the tiered retention probability of users at different lifecycle stages using the behavioral retention feature map specifically includes: The behavior retention feature map is sliced ​​using a preset time window, with each time window corresponding to a life cycle stage, to obtain the retention potential transfer value for each life cycle stage. The retention potential value is accumulated along the directed edge starting from the starting node in the behavior retention feature graph; The retention probability of users at different lifecycle stages is determined by accumulating potential energy across all stages.

8. A user retention prediction system based on behavioral feature modeling, the user retention prediction system based on behavioral feature modeling includes a behavioral prediction unit, characterized in that, The behavior prediction unit includes: The data collection module is used to collect the login behavior sequence, function click sequence and page dwell time of the target user within a specified time window, and then extract the user's behavioral habit characteristics and operation interval cycle. The processing module is used to construct a user's behavior rhythm vector based on the behavioral habit characteristics, and to fit the user's behavior decay curve by the behavior rhythm vector and the operation interval period to obtain the user's habit retention potential at each preset time node. The processing module is also used to extract the user's attention shift characteristics from the page dwell time, establish the user's behavior retention feature map based on the attention shift characteristics and the habit retention potential, and calculate the user's layer-by-layer retention probability at different life cycle stages through the behavior retention feature map. The execution module is used to compare the layer-by-layer retention probability with the actual retention label, adjust the behavior weight in the behavior rhythm vector according to the comparison result, and output the predicted retention rate of the user in the next time window.

9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory for storing computer programs, and the processor for calling and running the computer programs from the memory, causing the computer device to perform the user retention prediction method based on behavioral feature modeling as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions or code that, when executed on a computer, cause the computer to implement the user retention prediction method based on behavioral feature modeling as described in any one of claims 1 to 7.