A method and apparatus for predicting user retention rate
By combining multiple regression prediction models and deep classification models, user data features are extracted to generate detailed user retention rate predictions, which solves the problem of coarse prediction results in existing technologies and improves the reference value and accuracy of the predictions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA UNITED NETWORK COMM GRP CO LTD
- Filing Date
- 2022-12-30
- Publication Date
- 2026-07-17
Smart Images

Figure CN116205668B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically to a method and apparatus for predicting user retention rates. Background Technology
[0002] With the increasing variety of business platforms, products, and applications (APPs), improving user stickiness is just as important as increasing the number of users. Predicting user retention intention (also known as user retention rate) is a key indicator for improving user stickiness.
[0003] Conventional user retention prediction methods can predict whether a user will continue to visit a product after a certain period of time. That is, they can only predict a rough user retention rate over a relatively large time range. This makes the predicted user retention rate less reliable and cannot provide any reference value for the formulation of user recall strategies. Summary of the Invention
[0004] This invention proposes a user retention rate prediction method and apparatus to solve the problem that the user retention rate predicted by existing user retention rate prediction methods is too rough and has poor reference value.
[0005] Firstly, this disclosure provides a method for predicting user retention rate, including:
[0006] Acquire target user data and the target user's access data to the target application;
[0007] Features are extracted from the target user data and the access data to obtain basic features and access features. The access features are features within a preset time period, which includes historical time periods and the time period to be predicted.
[0008] Based on the basic features, user retention rate is predicted in at least two ways to obtain at least two initial prediction durations, and the credibility probability of each duration is calculated based on the basic features and the access features.
[0009] A retention rate is generated based on the reliability probability of each duration obtained from the at least two initial prediction durations and the time period to be predicted. The retention rate is used to characterize the duration of access to the target application by the target user during the time period to be predicted.
[0010] Secondly, this disclosure provides a user retention rate prediction device, comprising:
[0011] The acquisition module is used to acquire target user data and the target user's access data to the target application;
[0012] The feature extraction module is used to extract features from the target user data and the access data to obtain basic features and access features. The access features are features within a preset time period, and the preset time period includes historical time periods and time periods to be predicted.
[0013] The prediction module is used to predict user retention rate in at least two ways based on the basic features, obtain at least two initial prediction durations respectively, and calculate the confidence probability of each duration obtained by dividing the time period to be predicted based on the basic features and the access features.
[0014] The generation module is used to generate a retention rate based on the reliability probability of each duration divided into at least two initial prediction durations and the time period to be predicted. The retention rate is used to characterize the access duration of the target user to the target application during the time period to be predicted.
[0015] Thirdly, this disclosure provides an electronic device, including:
[0016] At least one processor; and a memory communicatively connected to said at least one processor;
[0017] The memory stores one or more computer programs that can be executed by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the user retention rate prediction method described above.
[0018] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the user retention rate prediction method described above.
[0019] According to the user retention rate prediction method and apparatus proposed in this invention, the following steps are taken: First, target user data and access data of the target user to a target application are acquired. Then, features are extracted from the target user data and the access data to obtain basic features and access features. The access features are features within a preset time period, which includes historical time periods and a period to be predicted. Based on the basic features, user retention rate prediction is performed in at least two ways to obtain at least two initial prediction durations. The reliability probability of each duration divided into the period to be predicted is calculated based on the basic features and the access features. Finally, a retention rate is generated based on the at least two initial prediction durations and the reliability probabilities of each duration divided into the period to be predicted. The retention rate is used to characterize the access duration of the target user to the target application within the period to be predicted. Therefore, this implementation can predict the access duration of the target user to the target application within the period to be predicted. The granularity of this prediction is detailed within the range of the period to be predicted, making the predicted user retention rate highly reliable and providing significant reference value for the formulation of user recall strategies.
[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0021] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0022] Figure 1 A flowchart of the user retention rate prediction method provided by the present invention is shown;
[0023] Figure 2 This diagram illustrates the data flow of the user retention rate prediction method provided by the present invention.
[0024] Figure 3 A schematic diagram of the user retention rate prediction device provided by the present invention is shown;
[0025] Figure 4 A schematic diagram of the electronic device provided by the present invention is shown. Detailed Implementation
[0026] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0027] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0028] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0029] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0030] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0031] This invention provides a user retention rate prediction method, wherein the user retention rate of a product presented by an application can characterize the user's intention to retain that product. Compared to conventional user retention rate prediction methods, which can only roughly predict whether a user will remain, this invention provides a user retention rate prediction method that can predict the number of days, n, a user will continue to access the product within N days after a certain date. Here, n is less than or equal to N, and N is an integer greater than or equal to 1. This invention refers to the number of days, n, a user will continue to access the product within N days after a certain date as the user's "N-day retention intention" or "N-day retention rate".
[0032] For example, a user's "7-day retention intention" is 3 on November 1st, indicating that it is predicted that the user will visit a product for 3 out of the 7 days following November 1st (i.e., from November 2nd to November 8th). In this example, N is 7 and n is 3.
[0033] The user retention rate prediction method of the present invention will be described below with reference to embodiments.
[0034] like Figure 1 As shown, Figure 1 The present invention illustrates a flowchart of a user retention rate prediction method, which includes:
[0035] Step S101: Obtain target user data and target user access data to the target application.
[0036] Here, target users refer to the users whose retention rate is to be calculated, and target applications refer to the objects corresponding to the target users' retention intentions.
[0037] In some implementations, when the target user's access data to the target application is relatively limited, the access data obtained by this invention can be all historical access data. In other implementations, when the target user's access data to the target application is relatively abundant, since this invention does not require all historical data to predict user retention rate, the access data obtained by this invention can be access data within a preset time period. The implementation within the preset time period is described below.
[0038] In some implementations, the target user data includes at least one of the following: the target user's personal data, the target user's corresponding terminal data, etc. The target user's access data to the target application includes at least one of the following: the time the target user logs into the target application, the total number of logins, the login type, the duration of each access to the target application, and whether the target user logged into the target application that day, etc.
[0039] The target user's personal data may include, for example, their gender and location. The corresponding terminal data may include, for example, the number of devices the target user has logged into the target application on, and the type of device used to log into the target application. The total number of logins may refer to the total number of logins within a preset time period. The login type may refer to account / password login or QR code login.
[0040] In some implementations, embodiments of the present invention can read the aforementioned types of data from at least one of the random access memory (RAM) and read-only memory (ROM) of at least one terminal device used by the target user.
[0041] Step S102: Extract features from target user data and access data to obtain basic features and access features.
[0042] The access characteristics can be characteristics within a preset time period, which includes historical time periods and time periods to be predicted.
[0043] For example, the basic features include at least one of the following: the target user's login device feature, login information feature, access duration feature, and weekly login frequency feature. The target user's login device feature characterizes the type of login device; when the target user has multiple login devices, this feature can be the average of the multiple device type features. The login information feature includes the time difference between the most recent login and the current time, the total number of logins within a preset time period, the login type, the duration of each access to the target application, and whether the target application was accessed on the current day, etc.
[0044] For example, the access features include at least one of the following: login sequence features within the preset time period, usage duration sequence features within the preset time period, and tag sequence features of accessed content within the preset time period.
[0045] For example, the aforementioned period to be predicted can be M days before a specified date to N days after that specified date. The specified date can be, for example, the current date. N is as described above, and M can be an integer greater than or equal to 2.
[0046] Step S103: Based on the basic features, perform user retention rate prediction in at least two ways to obtain at least two initial prediction durations, and calculate the credibility probability of each duration obtained by dividing the time period to be predicted based on the basic features and access features.
[0047] The durations obtained by dividing the period to be predicted include the number of days corresponding to each value in the prediction period from 1. For example, if the length of the period to be predicted is 7 days, the durations that can be divided into the period to be predicted include 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, and 7 days.
[0048] Based on the basic characteristics, at least two methods are used to predict user retention rates, resulting in at least two initial prediction durations. Specifically, at least two regression prediction models are invoked to perform regression predictions based on the basic characteristics, and the prediction results of the at least two regression prediction models are used as the at least two initial prediction durations.
[0049] For example, taking a prediction period of 7 days as an example, each initial prediction duration is, for example, 4.3 days, meaning that the target user will log in to the target application again for 4.3 days. In some implementations, the obtained initial prediction duration can be rounded to obtain an integer prediction value in days, for example, 4.3 days is equivalent to 4 days.
[0050] For example, the at least two regression prediction models include at least two of the following: a Lightweight Gradient Boosting Model (Lightgbm) single regression model, an Extreme Gradient Boosting (XGBoost) model, and a Catalogical Boosting (Catboost) single regression model.
[0051] Using this implementation method, each regression prediction model can obtain the prediction time of the model's characteristics. Therefore, by using at least two regression prediction models to process the basic features, at least two predicted values of the model characteristics can be obtained.
[0052] The present invention relates to calculating the confidence probability of each duration obtained from the predicted time period based on the basic features and the access features. This can be achieved by calling a deep classification model to calculate the confidence probability of each duration obtained from the division of the predicted time period based on the basic features and the access features. Referring to the examples of the various durations mentioned above, each confidence probability represents the probability that the predicted duration is the corresponding duration. For example, the probability that the predicted duration is 4 days is 80%, and the probability that the predicted duration is 3 days is 90%.
[0053] For example, the deep classification model can be a Convolutional Neural Network (CNN) model or a Deep Neural Network (DNN) model. Based on this, calling the deep classification model to calculate the confidence probability of each duration obtained from the predicted time period segmentation according to the basic features and the access features may include: calling the CNN model to process the access features to extract partial features from the access features; concatenating the basic features and the partial features to obtain concatenated features; and calling the DNN model to predict the concatenated features to obtain the confidence probability of each duration.
[0054] Step S104: Generate retention rate based on the credibility probability of each duration obtained from at least two initial prediction durations and the time period to be predicted.
[0055] The retention rate is used to characterize the duration of access to the target application by the target user during the predicted period, i.e., the "N-day retention rate" when the predicted period is N, and the value of the retention rate is n.
[0056] In some implementations, the present invention can perform a harmonic average on the at least two initial prediction durations to obtain a harmonic average value, thereby integrating the prediction characteristics of multiple models, and then generating the retention rate based on the harmonic average value and the confidence probability.
[0057] For example, generating the retention rate based on the harmonic mean and the confidence probability can specifically include compatibility analysis of the confidence probability output by the DNN and the harmonic mean, using the compatible result as the retention rate. Taking N as 7 as an example, if the probability of 4 days appears most frequently in the confidence probability output by the DNN, and the harmonic mean is also 4 days, then the retention rate is 4 days. Similarly, if the probability of 3 days appears most frequently in the confidence probability output by the DNN, and the harmonic mean is 5 days, then the average of 3 days and 5 days is taken as the retention rate. This ensures both the generalization of the predicted retention rate and its accuracy.
[0058] According to the description of step S103, at least two initial prediction durations are predicted by at least two regression prediction models, and the confidence probability is processed by a deep classification model. Compared with the conventional use of regression or classification algorithms to predict user retention rate, the present invention uses a compatible and mutually reinforcing approach of multiple regression prediction models and deep classification models to predict user retention rate. This not only obtains the predicted values output by multiple regression prediction models, but also corrects the predicted values by using the probability obtained by the deep classification algorithm, thereby improving the prediction effect of user retention rate.
[0059] Therefore, by adopting the embodiments of the present invention, firstly, target user data and access data of the target user to the target application are acquired; features are extracted from the target user data and the access data to obtain basic features and access features. The access features are features within a preset time period, which includes historical time periods and a time period to be predicted; user retention rate is predicted in at least two ways based on the basic features, resulting in at least two initial prediction durations, and the credibility probability of each duration divided into the time period to be predicted is calculated based on the basic features and the access features; a retention rate is generated based on the at least two initial prediction durations and the credibility probability of each duration divided into the time period to be predicted. The retention rate is used to characterize the access duration of the target user to the target application within the time period to be predicted. It is evident that by adopting this implementation method, the access duration of the target user to the target application within the time period to be predicted can be predicted. The granularity of this prediction is detailed to the range of the time period to be predicted, making the predicted user retention rate highly referential and providing significant reference value for the formulation of user recall strategies. Furthermore, this invention employs a compatible and mutually reinforcing approach that combines multiple regression prediction models with deep classification models to predict user retention rates. This not only ensures the generalizability of the predicted retention rates but also guarantees their accuracy, thereby improving the prediction performance of user retention rates.
[0060] It should be noted that, for the target application, if the target user's login information is greater than or equal to a preset threshold of not logging in, then the user retention rate can be directly set to 0. The preset threshold can be, for example, 30-t days, where t = 1, 2, 3...25.
[0061] For ease of understanding, the following is combined with Figure 2 The data flow diagram shown illustrates the user retention rate prediction method provided by this invention.
[0062] Figure 2 Let's take the prediction of user retention rate n over N days as an example.
[0063] Step S201, data acquisition.
[0064] This can involve acquiring data from the target user's terminal, including, for example, user terminal data, user information data, user login data, and user usage data. The user terminal data and user information data can be... Figure 1 The target user data, user login data, and user usage data described in the embodiment may be... Figure 1 Access data as described in the embodiments.
[0065] For example, user terminal data includes: Ram data, ROM data, and system type data, etc.; user information data includes: user gender, login region, etc.; user login data includes: whether logged in on the current day, login type, and the total number of logins in the previous M days, etc.; user usage data includes: the duration of daily use of the target application in the previous M days, etc.
[0066] Step S202, Feature extraction.
[0067] Basic features and sequence features are extracted from the aforementioned user terminal data, user information data, user login data, and user usage data. These sequence features are... Figure 1 Access features as described in the embodiments.
[0068] The basic features extracted from the aforementioned user terminal data, user information data, user login data, and user usage data include: user personal information feature extraction, user login status feature extraction, application usage feature extraction, and periodic features.
[0069] User personal information feature extraction includes: for users logging in using multiple devices, averaging the parameters of multiple device types.
[0070] User login feature extraction includes: whether the user logged in on the current day, the login type, the total number of logins in the previous M days, the total number of logins in the previous M days, and the time difference between the most recent login time and the current time.
[0071] Application usage feature extraction includes: extracting the user's daily usage time and number of uses for the current day and the M days prior to that day.
[0072] Periodic feature extraction includes: obtaining the periodic pattern of the total number of user logins on a weekly basis.
[0073] The sequence features extracted from the aforementioned user terminal data, user information data, user login data, and user usage data include: login sequence feature extraction, usage duration sequence feature extraction, and access content tag sequence feature extraction.
[0074] Login sequence feature extraction: Login sequences are extracted in chronological order from early to late, within the time period [last day - NM, last day - N]. "Last day" refers to the last day corresponding to the retention rate prediction period.
[0075] Application usage time sequence feature extraction: Extract application usage time sequences within the time period [last day-NM, last day-N] in chronological order from morning to night.
[0076] Tag sequence feature extraction of accessed content: Extract the tag sequence of accessed content in the time period [last day-NM, last day-N] according to the time from early to late.
[0077] Step S203, model prediction.
[0078] The basic features are input into the pre-configured LightGBM single regression model, XGBoost single regression model, CatBoost single regression model, and DNN model, respectively.
[0079] The pre-configured LightGBM single regression model, XGBoost single regression model, and CatBoost single regression model are called respectively to perform regression prediction on the above basic features. Each of the three single regression models outputs an initial prediction duration, resulting in three initial prediction durations.
[0080] The login sequence, usage duration sequence, and access content tag sequence from the extracted sequence features are input into a CNN model with preset values, so that the CNN model can extract some features from these three sequence features and transmit these features to the DNN model.
[0081] The DNN is called to concatenate the basic features and the partial features obtained by the CNN, and the concatenated features are processed to obtain the predicted probabilities of n=1,2...N in "N-day retention intention", that is, the confidence probability of each n.
[0082] Step S204, model results are compatible.
[0083] The three initial prediction durations obtained from the LightGBM single regression model, XGBoost single regression model, and CatBoost single regression model in step S203 are harmonicly averaged to obtain the harmonic mean. Then, the harmonic mean is compatible with the prediction probability output by the DNN to generate the prediction result, which is the retention intention n of the target user within N days.
[0084] Step S205: Output the prediction results.
[0085] Output the value of n.
[0086] As can be seen, this implementation can predict the duration of continued access to the target application by target users within the predicted time period. The granularity of this prediction, down to the predicted time period, makes the predicted user retention rate highly reliable and valuable for developing user reactivation strategies. Furthermore, this invention employs a compatible and mutually reinforcing approach combining multiple regression prediction models and deep classification models for user retention rate prediction. This not only ensures the generalization and accuracy of the predicted retention rate but also enhances the prediction effectiveness.
[0087] Corresponding to the above implementation steps, the present invention also provides software, hardware, or a combination of software and hardware modules / units for performing the above implementation steps.
[0088] Figure 3 A schematic diagram of a user retention rate prediction device provided by the present invention is shown. (Refer to...) Figure 3 It includes: an acquisition module 31, a feature extraction module 32, a prediction module 33, and a generation module 34. Each module is used to perform... Figure 1 and Figure 2 Some or all of the steps of the method shown.
[0089] For example, the acquisition module 31 is used to acquire target user data and the target user's access data to the target application; the feature extraction module 32 is used to extract features from the target user data and the access data to obtain basic features and access features, wherein the access features are features within a preset time period, the preset time period includes historical time periods and the time period to be predicted; the prediction module 33 is used to predict user retention rate in at least two ways based on the basic features, to obtain at least two initial prediction durations, and to calculate the confidence probability of each duration divided by the time period to be predicted based on the basic features and the access features; the generation module 34 is used to generate a retention rate based on the at least two initial prediction durations and the confidence probability of each duration divided by the time period to be predicted, wherein the retention rate is used to characterize the access duration of the target user to the target application within the time period to be predicted.
[0090] Optionally, the prediction module 33 is also used to call at least two regression prediction models to perform regression prediction based on the basic features, and use the prediction results of the at least two regression prediction models as the at least two initial prediction durations.
[0091] Optionally, the at least two regression prediction models include at least two of the following: a lightweight gradient boosting tree (Light ghtgbm) single regression model, an extreme gradient boosting tree (Xgboost) single regression model, and a symmetric decision tree (Catboost) single regression model.
[0092] Optionally, the prediction module 33 is further configured to call a deep classification model to calculate the confidence probability of each duration obtained by dividing the time period to be predicted based on the basic features and the access features, wherein each duration obtained by dividing the time period to be predicted includes the number of days corresponding to each value in the prediction time period from 1.
[0093] Optionally, the deep classification model includes a CNN model and a DNN model. The prediction module 33 is further configured to call the CNN model to process the access features to extract some features from the access features; concatenate the basic features and the partial features to obtain concatenated features; and call the DNN model to predict the concatenated features to obtain the confidence probability of each duration.
[0094] Optionally, the generation module 34 is further configured to perform a harmonic average on the at least two initial prediction durations to obtain a harmonic average; and generate the retention rate based on the harmonic average and the confidence probability.
[0095] Optionally, the target user data includes at least one of the following: the target user's personal data, the target user's corresponding terminal data; the target user's access data to the target application includes at least one of the following: the time the target user logs into the target application, the total number of logins, the login type, and the duration of each access to the target application; the basic features include at least one of the following: the target user's login device features, login information features, access duration features, and weekly login frequency features; the access features include at least one of the following: login sequence features within the preset time period, usage duration sequence features within the preset time period, and accessed content tag sequence features within the preset time period.
[0096] Figure 4 A schematic diagram of an electronic device provided by the present invention is shown. Specific embodiments of the present invention do not limit the specific implementation of the electronic device. (Refer to...) Figure 4 The electronic device includes:
[0097] At least one processor 401; a memory 402 communicatively connected to at least one processor; a communication interface 403; and a communication bus 403.
[0098] in:
[0099] The processor 401, memory 402, and communication interface 403 communicate with each other through the communication bus 403.
[0100] Communication interface 403 is used to communicate with other network elements such as clients or other servers.
[0101] The memory 402 stores one or more computer programs 405 that can be executed by at least one processor 401, which enables the at least one processor 401 to perform the corresponding operations as described in the above-described user retention rate prediction method embodiment.
[0102] This invention provides a non-volatile computer storage medium storing at least one executable instruction that can execute the object loading method in the virtual scene of any of the above method embodiments. Specifically, the executable instruction can be used to cause the processor to perform the corresponding operations in the above method embodiments.
[0103] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital user retention predictor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0104] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0105] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0106] The computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), are personalized by utilizing state information of computer-readable program instructions. These electronic circuits can execute computer-readable program instructions to implement various aspects of this disclosure.
[0107] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0108] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0109] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0110] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0112] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A method for predicting user retention rate, characterized in that, include: Acquire target user data and target user access data to the target application; Features are extracted from the target user data and the access data to obtain basic features and access features. The access features are features within a preset time period, which includes historical time periods and the time period to be predicted. At least two regression prediction models are invoked to perform regression prediction based on the basic features. The prediction results of the at least two regression prediction models are used as at least two initial prediction durations. A deep classification model is invoked to perform multi-class prediction based on the basic features and the access features. The confidence probability of each duration obtained by dividing the period to be predicted is calculated. The initial prediction duration represents the number of access days predicted for the target user within the period to be predicted. Each duration obtained by dividing the period to be predicted includes the number of days corresponding to each value in the period to be predicted from 1. A retention rate is generated based on the reliability probability of each duration obtained from the at least two initial prediction durations and the time period to be predicted. The retention rate is used to characterize the duration of access to the target application by the target user during the time period to be predicted.
2. The method according to claim 1, characterized in that, The at least two regression prediction models include at least two of the following: LightGBM single regression model, XGBoost single regression model, and CatBoost single regression model based on symmetric decision tree.
3. The method according to claim 1, characterized in that, The deep classification model includes a convolutional neural network (CNN) model and a deep neural network (DNN) model. The process of calling the deep classification model to calculate the confidence probability of each duration obtained from the predicted time period segmentation based on the basic features and the access features includes: The CNN model is invoked to process the access features in order to extract some features from the access features. By combining the basic features and the partial features, a combined feature is obtained; The DNN model is invoked to predict the spliced features, thereby obtaining the credibility probability of each duration.
4. The method according to claim 1, characterized in that, The step of generating a retention rate based on the reliability probability of each duration obtained from the at least two initial prediction durations and the prediction period division includes: The harmonic mean is obtained by harmonic averaging the at least two initial prediction durations. The retention rate is generated based on the harmonic mean and the confidence probability.
5. The method according to claim 1, characterized in that, The target user data includes at least one of the following: the target user's personal data, and the target user's corresponding terminal data; The target user's access data to the target application includes at least one of the following: the time the target user logs into the target application, the total number of logins, the login type, and the duration of each access to the target application; The basic features include at least one of the following: the target user's login device features, login information features, access duration features, and login frequency features on a weekly basis; The access features include at least one of the following: login sequence features within the preset time period, usage duration sequence features within the preset time period, and access content tag sequence features within the preset time period.
6. A user retention rate prediction device, characterized in that, include: The acquisition module is used to acquire target user data and target user access data to the target application; The feature extraction module is used to extract features from the target user data and the access data to obtain basic features and access features. The access features are features within a preset time period, which includes historical time periods and the time period to be predicted. The prediction module is used to call at least two regression prediction models to perform regression prediction based on the basic features, use the prediction results of the at least two regression prediction models as at least two initial prediction durations, and call a deep classification model to perform multi-class prediction based on the basic features and the access features, and calculate the confidence probability of each duration obtained by dividing the period to be predicted, wherein the initial prediction duration represents the number of access days predicted for the target user within the period to be predicted, and each duration obtained by dividing the period to be predicted includes the number of days corresponding to each value in the period to be predicted; The generation module is used to generate a retention rate based on the reliability probability of each duration divided into at least two initial prediction durations and the time period to be predicted. The retention rate is used to characterize the access duration of the target user to the target application during the time period to be predicted.
7. An electronic device, characterized in that, include: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores one or more computer programs that can be executed by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1 to 5.