Method and device for detecting cost on game cloud, storage medium and electronic equipment

By employing masking enhancement techniques and contrastive learning methods, a balanced feature distribution is generated, which solves the problem of uneven features in the user lifetime value prediction model and achieves higher prediction accuracy.

CN122298024APending Publication Date: 2026-06-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2024-12-31
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

In existing technologies, user lifetime value prediction models suffer from uneven feature distribution during training, resulting in low accuracy of prediction results.

Method used

By employing masking enhancement techniques and contrastive learning methods, a neural network model is trained by generating masking enhancement features and positive and negative sample pairs to ensure balanced feature distribution and improve prediction accuracy.

Benefits of technology

By employing contrastive learning and masking-enhanced training, more robust supervisory signals are provided, ensuring uniform feature distribution and improving the accuracy of user lifetime value prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122298024A_ABST
    Figure CN122298024A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, storage medium, and electronic device for predicting user lifetime value. The method includes: acquiring target features corresponding to the user to be predicted; inputting the target features into a trained lifetime value model to obtain a predicted result for the lifetime value of the user to be predicted. This application solves the technical problem of low accuracy in predicting user lifetime value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically, to a method, apparatus, storage medium, and electronic device for predicting user lifetime value. Background Technology

[0002] In related technologies, when predicting user lifetime value, a trained lifetime value prediction model is often used. However, due to the specific nature of lifetime value prediction scenarios, it is impossible to guarantee the balance of samples with different lifetime values ​​during the training process of the lifetime value prediction model. This leads to uneven feature distribution during the training process, affecting the prediction results and resulting in low accuracy in user lifetime value prediction. Therefore, there is a problem of low accuracy in user lifetime value prediction.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a method, apparatus, storage medium, and electronic device for predicting user lifetime value, to at least address the technical problem of low accuracy in predicting user lifetime value.

[0005] According to one aspect of the embodiments of this application, a method for predicting user lifetime value is provided, comprising: acquiring target features corresponding to a user to be predicted; inputting the target features into a trained lifetime value model to obtain a prediction result corresponding to the lifetime value of the user to be predicted; wherein the lifetime value model is a neural network model for predicting user lifetime value, the trained lifetime value model is an initial lifetime value model, obtained after training based on a first positive sample pair, a first negative sample pair, a second positive sample pair, and a second negative sample pair, the first positive sample pair including a first sample target feature and a first masking enhancement feature, the first negative sample pair including a second sample target feature and a second masking enhancement feature, and the second positive sample pair including the second sample target feature and a first masking enhancement feature. The aforementioned second masking enhancement feature, the aforementioned second negative sample pair includes the aforementioned first sample target feature and the aforementioned first masking enhancement feature. The aforementioned first masking enhancement feature is obtained by masking and enhancing the first preliminary enhancement feature corresponding to the aforementioned first sample target feature using the first masking probability vector. The aforementioned second masking enhancement feature is obtained by masking and enhancing the second preliminary enhancement feature corresponding to the aforementioned second sample target feature using the second masking probability vector. The aforementioned first masking probability vector corresponds to the aforementioned first preliminary enhancement feature, and the aforementioned second masking probability vector corresponds to the aforementioned second preliminary enhancement feature. The aforementioned first masking probability vector is used to represent the probability that the aforementioned first preliminary enhancement feature is masked or enhanced, and the aforementioned second masking probability vector is used to represent the probability that the aforementioned second preliminary enhancement feature is masked or enhanced.

[0006] According to another aspect of the embodiments of this application, a user lifetime value prediction apparatus is also provided, comprising: a first acquisition unit, configured to acquire target features corresponding to a user to be predicted; and a prediction unit, configured to input the target features into a trained lifetime value model to obtain a prediction result corresponding to the lifetime value of the user to be predicted; wherein the lifetime value model is a neural network model for predicting user lifetime value, the trained lifetime value model is an initial lifetime value model, obtained after training based on a first positive sample pair, a first negative sample pair, a second positive sample pair, and a second negative sample pair, the first positive sample pair including a first sample target feature and a first masking enhancement feature, the first negative sample pair including a second sample target feature and a second masking enhancement feature, and the second positive sample pair including the first positive sample feature and a second masking enhancement feature. The second sample target feature and the aforementioned second masking enhancement feature, the aforementioned second negative sample pair includes the aforementioned first sample target feature and the aforementioned first masking enhancement feature, the aforementioned first masking enhancement feature is obtained by masking and enhancing the first preliminary enhancement feature corresponding to the aforementioned first sample target feature using the first masking probability vector, the aforementioned second masking enhancement feature is obtained by masking and enhancing the second preliminary enhancement feature corresponding to the aforementioned second sample target feature using the second masking probability vector, the aforementioned first masking probability vector corresponds to the aforementioned first preliminary enhancement feature, the aforementioned second masking probability vector corresponds to the aforementioned second preliminary enhancement feature, the aforementioned first masking probability vector is used to represent the probability that the aforementioned first preliminary enhancement feature is masked or enhanced, the aforementioned second masking probability vector is used to represent the probability that the aforementioned second preliminary enhancement feature is masked or enhanced.

[0007] As an optional solution, the first acquisition unit includes: a first acquisition module, used to acquire the user basic features corresponding to the user to be predicted and to acquire the target behavior features corresponding to the user to be predicted, wherein the user basic features are used to represent the basic user information of the user to be predicted and the target behavior features are used to represent the historical behavior information of the user to be predicted; and an integration module, used to integrate the user basic features and the target behavior features to obtain the target features.

[0008] As an optional solution, the first acquisition module includes at least one of the following: a first acquisition submodule, configured to acquire first behavioral features of the user to be predicted in various game applications when the user to be predicted is a game user and the lifetime value model is a neural network model for predicting the lifetime value of the game user, wherein the first behavioral features represent statistical information of the user's historical behavior in the various game applications, and the target behavioral features include the first behavioral features; and a second acquisition submodule, configured to acquire first behavioral features when the user to be predicted is a game user and the lifetime value model is a neural network model for predicting the lifetime value of the game user. The following steps are taken: First, a second behavioral feature of the user to be predicted in a specific game application is obtained, wherein the second behavioral feature represents the behavioral information generated by the user to be predicted in the specific game application, and the target behavioral feature includes the second behavioral feature. Second, a third acquisition submodule is used to obtain a third behavioral feature of the user to be predicted within a recent preset time period, provided that the user to be predicted is a game user and the lifetime value model is a neural network model used to predict the lifetime value of the game user. The third behavioral feature represents the time series information of the game behavior generated by the user to be predicted within the recent preset time period, and the target behavioral feature includes the third behavioral feature.

[0009] As an optional solution, the aforementioned third acquisition submodule includes at least one of the following: a first acquisition subunit, configured to acquire a first time-series feature of the user to be predicted within the aforementioned recent preset time period, wherein the first time-series feature represents the game registration sequence of the user to be predicted within the aforementioned recent preset time period, and the third behavioral feature includes the first time-series feature; a second acquisition subunit, configured to acquire a second time-series feature of the user to be predicted within the aforementioned recent preset time period, wherein the second time-series feature represents the game activity sequence of the user to be predicted within the aforementioned recent preset time period, and the third behavioral feature includes the second time-series feature; and a third acquisition subunit, configured to acquire a third time-series feature of the user to be predicted within the aforementioned recent preset time period. The third time series feature is used to represent the game activity duration sequence of the user to be predicted within the recent preset time period, and the third behavioral feature includes the third time series feature; the fourth acquisition subunit is used to acquire the fourth time series feature of the user to be predicted within the recent preset time period, wherein the fourth time series feature is used to represent the game payment sequence of the user to be predicted within the recent preset time period, and the third behavioral feature includes the fourth time series feature; the fifth acquisition subunit is used to acquire the fifth time series feature of the user to be predicted within the recent preset time period, wherein the fifth time series feature is used to represent the game payment amount sequence of the user to be predicted within the recent preset time period, and the third behavioral feature includes the fifth time series feature.

[0010] According to another aspect of the embodiments of this application, a training method for a lifecycle value model is provided, comprising: obtaining a first preliminary enhancement feature corresponding to a first sample target feature and a second preliminary enhancement feature corresponding to a second sample target feature; obtaining a first masking probability vector corresponding to the first preliminary enhancement feature and a second masking probability vector corresponding to the second preliminary enhancement feature, wherein the first masking probability vector is used to represent the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector is used to represent the probability that the second preliminary enhancement feature is masked or enhanced; masking and enhancing the first preliminary enhancement feature using the first masking probability vector to obtain a first masking enhancement feature, and masking and enhancing the second preliminary enhancement feature using the second masking probability vector to obtain a second masking enhancement feature; obtaining the first preliminary enhancement feature corresponding to a first sample target feature and a second preliminary enhancement feature corresponding to a second sample target feature. A first positive sample pair and a first negative sample pair corresponding to a sample target feature, and a second positive sample pair and a second negative sample pair corresponding to the second sample target feature, wherein the first positive sample pair includes the first sample target feature and the first masking enhancement feature, the first negative sample pair includes the second sample target feature and the second masking enhancement feature, the second positive sample pair includes the second sample target feature and the second masking enhancement feature, and the second negative sample pair includes the first sample target feature and the first masking enhancement feature; based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair, the initial lifetime value model is trained to obtain a trained lifetime value model, wherein the lifetime value model is a neural network model used to predict user lifetime value.

[0011] According to another aspect of the embodiments of this application, a training apparatus for a lifecycle value model is also provided, comprising: a second acquisition unit, configured to acquire a first preliminary enhancement feature corresponding to a first sample target feature, and a second preliminary enhancement feature corresponding to a second sample target feature; a third acquisition unit, configured to acquire a first masking probability vector corresponding to the first preliminary enhancement feature, and a second masking probability vector corresponding to the second preliminary enhancement feature, wherein the first masking probability vector represents the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector represents the probability that the second preliminary enhancement feature is masked or enhanced; and a masking enhancement unit, configured to mask and enhance the first preliminary enhancement feature using the first masking probability vector to obtain a first masking enhancement feature, and mask and enhance the second preliminary enhancement feature using the second masking probability vector to obtain a second masking enhancement feature. The fourth acquisition unit is used to acquire the first positive sample pair and the first negative sample pair corresponding to the first sample target feature, and the second positive sample pair and the second negative sample pair corresponding to the second sample target feature, wherein the first positive sample pair includes the first sample target feature and the first masking enhancement feature, the first negative sample pair includes the second sample target feature and the second masking enhancement feature, the second positive sample pair includes the second sample target feature and the second masking enhancement feature, and the second negative sample pair includes the first sample target feature and the first masking enhancement feature; the training unit is used to train the initial lifetime value model based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair to obtain a trained lifetime value model, wherein the lifetime value model is a neural network model used to predict user lifetime value.

[0012] As an optional solution, the second acquisition unit includes: a first mapping module, used to map the first sample target features to a low-dimensional hidden space to obtain a first low-dimensional embedding, and to map the second sample target features to the low-dimensional hidden space to obtain a second low-dimensional embedding; and a second acquisition module, used to acquire a first enhanced feature corresponding to the first low-dimensional embedding and a second enhanced feature corresponding to the second low-dimensional embedding, wherein the first preliminary enhanced feature includes the first enhanced feature, and the second preliminary enhanced feature includes the second enhanced feature.

[0013] As an optional solution, the second acquisition module includes: a first nonlinear submodule, used to perform a nonlinear transformation on the first low-dimensional embedding through a multilayer perceptron to obtain the first enhanced feature; and a second nonlinear submodule, used to perform a nonlinear transformation on the second low-dimensional embedding through the multilayer perceptron to obtain the second enhanced feature.

[0014] As an optional solution, the third acquisition unit includes: an adding module, used to add corresponding noise to each feature element in the first preliminary enhancement feature, wherein the noise is sampled from an extreme value distribution; and an input module, used to input the first preliminary enhancement feature after adding the noise into a normalized exponential function to obtain a probability distribution vector, wherein the probability distribution vector is used to represent the probability that each feature element in the first preliminary enhancement feature after adding the noise is masked or enhanced, and the first masking probability vector includes the probability distribution vector.

[0015] As an optional solution, the masking enhancement unit includes: an embedding module for embedding the probability distribution vector and the first low dimension into the input masking enhancement function to obtain masking enhancement hidden space features; and a transformation module for converting the hidden space features into masking enhancement features of the original dimension, wherein the original dimension is the dimension of the first sample target feature, and the first masking enhancement feature includes the masking enhancement features of the original dimension.

[0016] As an optional approach, the training unit includes: a pre-training module for pre-training the initial lifecycle value model based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair to obtain a first lifecycle value model; and a fine-tuning module for fine-tuning the first lifecycle value model using the sample data from the target domain to obtain a second lifecycle value model, wherein the trained lifecycle value model includes the second lifecycle value model.

[0017] As an optional approach, the aforementioned pre-training module includes: a fourth acquisition submodule, used to acquire current positive and negative sample pairs, wherein the first positive sample pair and the first negative sample pair belong to the positive and negative sample pairs, and the second positive sample pair and the second negative sample pair belong to the positive and negative sample pairs; a fifth acquisition submodule, used to acquire the first similarity between positive sample pairs in the current positive and negative sample pairs, and the second similarity between negative sample pairs in the current positive and negative sample pairs; and a sixth acquisition submodule, used to acquire the current contrastive loss of the current initial lifecycle value model based on the first similarity and the second similarity, using the current contrastive loss function, wherein the contrastive loss is used to measure the lifecycle value. The value model is sensitive to distinguishing between positive and negative samples; a determination submodule is used to determine the current initial lifecycle value model as the first lifecycle value model when the current contrastive loss meets the convergence condition of the pre-training stage; an adjustment submodule is used to adjust the temperature parameter in the current contrastive loss function when the current contrastive loss does not meet the convergence condition of the pre-training stage, wherein the temperature parameter is used to control the sensitivity of the lifecycle value model to distinguish between positive and negative samples; a seventh acquisition submodule is used to acquire the next positive and negative sample pair and use the next positive and negative sample pair as the current positive and negative sample pair, until the first lifecycle value model is obtained.

[0018] As an optional approach, the training unit includes a training module for training the initial lifetime value model by combining cross-entropy loss and log-normal loss, wherein the cross-entropy loss is used to determine whether a user is a high-value user or a low-value user, and the log-normal loss is used to make the predicted value and the actual value of the lifetime value model conform to the log-normal distribution parameters.

[0019] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform a user lifetime value prediction method or a lifetime value model training method as described above.

[0020] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described user lifetime value prediction method or lifetime value model training method through the computer program.

[0021] According to another aspect of the embodiments of this application, a computer program / instruction is also provided, wherein when the computer program / instruction is executed by a processor, it implements the above-mentioned method for predicting user lifetime value or the training method for the above-mentioned lifetime value model.

[0022] In this embodiment, target features corresponding to the user to be predicted are obtained; the target features are input into a trained lifetime value model to obtain the prediction result corresponding to the lifetime value of the user to be predicted; wherein, the lifetime value model is a neural network model used to predict the lifetime value of a user, and the trained lifetime value model is an initial lifetime value model, obtained after training based on a first positive sample pair, a first negative sample pair, a second positive sample pair, and a second negative sample pair. The first positive sample pair includes a first sample target feature and a first masking enhancement feature, the first negative sample pair includes a second sample target feature and a second masking enhancement feature, and the second positive sample pair includes a second sample target feature and a second masking enhancement feature. The strong feature, the second negative sample pair includes the first sample target feature and the first masking enhancement feature. The first masking enhancement feature is obtained by masking and enhancing the first preliminary enhancement feature corresponding to the first sample target feature using the first masking probability vector. The second masking enhancement feature is obtained by masking and enhancing the second preliminary enhancement feature corresponding to the second sample target feature using the second masking probability vector. The first masking probability vector corresponds to the first preliminary enhancement feature, and the second masking probability vector corresponds to the second preliminary enhancement feature. The first masking probability vector is used to represent the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector is used to represent the probability that the second preliminary enhancement feature is masked or enhanced.

[0023] Specifically, after acquiring the target features, these features are input into a model trained through contrastive learning and masking enhancement. Contrastive learning provides a more robust supervision signal for the model's training process. During the contrastive learning phase, masking probability vectors are used for the features to enhance those with low or high masking probabilities. This masks features with low or high masking probabilities, thus preserving or enhancing relatively important features while masking relatively unimportant features. This achieves the goal of uniform feature distribution during the training process of the lifetime value prediction model, thereby improving the accuracy of user lifetime value prediction and solving the technical problem of low prediction accuracy. Attached Figure Description

[0024] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0025] Figure 1This is a schematic diagram of an application environment for an optional method for predicting user lifetime value according to an embodiment of this application;

[0026] Figure 2 This is a schematic diagram of the flow of an optional method for predicting user lifetime value according to an embodiment of this application;

[0027] Figure 3 This is a schematic diagram of an optional method for predicting user lifetime value according to an embodiment of this application;

[0028] Figure 4 This is a schematic diagram of another optional method for predicting user lifetime value according to an embodiment of this application;

[0029] Figure 5 This is a schematic diagram of another optional method for predicting user lifetime value according to an embodiment of this application;

[0030] Figure 6 This is a schematic diagram of another optional method for predicting user lifetime value according to an embodiment of this application;

[0031] Figure 7 This is a schematic diagram of another optional method for predicting user lifetime value according to an embodiment of this application;

[0032] Figure 8 This is a schematic diagram of another optional method for predicting user lifetime value according to an embodiment of this application;

[0033] Figure 9 This is a schematic diagram of another optional method for predicting user lifetime value according to an embodiment of this application;

[0034] Figure 10 This is a schematic diagram of an optional user lifetime value prediction device according to an embodiment of this application;

[0035] Figure 11 This is a schematic diagram of an optional lifecycle value model training device according to an embodiment of this application;

[0036] Figure 12 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application. Detailed Implementation

[0037] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0038] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0039] According to one aspect of the embodiments of this application, a method for predicting user lifetime value is provided. Optionally, as an optional implementation, the above-described method for predicting user lifetime value can be applied to, but is not limited to, [examples of other methods]. Figure 1 The environment shown may include, but is not limited to, user equipment 102 and server 112. User equipment 102 may include, but is not limited to, a display 104, a processor 106 and a memory 108. Server 112 includes a database 114 and a processing engine 116.

[0040] The specific process can be summarized in the following steps:

[0041] Step S102, user equipment 102 obtains a prediction request for user lifetime value;

[0042] Step S104: Send the user lifetime value prediction request to server 112 via network 110;

[0043] In steps S106-S108, server 112 obtains the target features corresponding to the user to be predicted through processing engine 116, and then inputs the target features into the trained lifetime value model to obtain the prediction result corresponding to the lifetime value of the user to be predicted.

[0044] In step S110, the prediction result corresponding to the lifetime value of the user to be predicted is sent to the user equipment 102 via the network 110. The user equipment 102 receives the prediction result corresponding to the lifetime value of the user to be predicted via the processor 106, displays the prediction result on the display 104, and stores the prediction result corresponding to the lifetime value of the user to be predicted in the memory 108.

[0045] remove Figure 1 Beyond the examples shown, the terminal devices described above can be terminal devices configured with a target client, including but not limited to at least one of the following: mobile phones (such as Android phones, iOS phones, etc.), laptops, tablets, PDAs, MIDs (Mobile Internet Devices), PADs, desktop computers, smart TVs, etc. The target client can be a video client, instant messaging client, browser client, educational client, etc. The networks described above can include, but are not limited to, wired networks and wireless networks. The wired networks include local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). The wireless networks include Bluetooth, Wi-Fi, and other networks that enable wireless communication. The server described above can be a single server, a server cluster consisting of multiple servers, or a cloud server. The above is merely an example, and no limitations are imposed in this embodiment.

[0046] Alternatively, as an alternative implementation method, such as Figure 2 As shown, the method for predicting user lifetime value can be performed by an electronic device, such as... Figure 1 The user equipment or server shown includes the following specific steps:

[0047] S202, Obtain the target features corresponding to the user to be predicted;

[0048] S204, Input the target features into the trained lifetime value model to obtain the prediction result corresponding to the lifetime value of the user to be predicted;

[0049] The Lifetime Value Model is a neural network model used to predict user lifetime value. The trained Lifetime Value Model is the initial Lifetime Value Model, obtained after training based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair. The first positive sample pair includes the first sample target feature and the first masking enhancement feature. The first negative sample pair includes the second sample target feature and the second masking enhancement feature. The second positive sample pair includes the second sample target feature and the second masking enhancement feature. The second negative sample pair includes the first sample target feature and the first masking enhancement feature. The first masking enhancement feature is obtained by masking and enhancing the first preliminary enhancement feature corresponding to the first sample target feature using the first masking probability vector. The second masking enhancement feature is obtained by masking and enhancing the second preliminary enhancement feature corresponding to the second sample target feature using the second masking probability vector. The first masking probability vector corresponds to the first preliminary enhancement feature, and the second masking probability vector corresponds to the second preliminary enhancement feature. The first masking probability vector is used to represent the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector is used to represent the probability that the second preliminary enhancement feature is masked or enhanced.

[0050] In optional embodiments, the above-mentioned user lifetime value prediction method can be applied, but is not limited to, in the scenario of predicting the lifetime value of game users. By using a trained lifetime value prediction model, the acquired user data is input into the prediction model to predict the total value of the user within a specific lifetime. This helps game operators to place corresponding advertisements for users in the game based on the predicted total value of the user within a specific lifetime, thereby achieving the technical effect of improving the accuracy of advertising placement.

[0051] In optional embodiments, the above-mentioned user lifetime value prediction method can also be applied, but is not limited to, to the user lifetime value prediction scenario of short video playback platforms. By using a trained lifetime value prediction model, the acquired user data is input into the prediction model to predict the total value of the user in a specific lifetime. Therefore, based on the total value of the user in a specific lifetime, the amount that the short video platform needs to charge the advertiser for placing ads on that user can be adjusted. For example, for high-value users, a larger coefficient is used to increase the bid, while for low-value users, a smaller coefficient is used to decrease the bid, thereby achieving the technical effect of improving the accuracy of ad placement management.

[0052] Optionally, the target features may be, but are not limited to, a series of data points used to describe user behavior and attributes, and may be, but are not limited to, used to predict the value of a user within a specific time period. For example, in the scenario of predicting the lifetime value of a game user, the target features may include, but are not limited to, basic user characteristics, target behavioral characteristics, and game behavior time series.

[0053] Among them, the time series features of game behavior may include, but are not limited to: game registration sequence, game activity sequence, game activity duration sequence, game payment sequence, game payment amount sequence, etc.

[0054] Optionally, User Lifetime Value (VWV) can be understood, but is not limited to, as the value of all paid activities generated by a user during a specific lifecycle of using an application or service.

[0055] To illustrate further, suppose user A spends a total of X yuan to acquire items or services within game B over a specific lifecycle. Understandably, user A's lifetime value within this lifecycle is X yuan.

[0056] In optional embodiments, the first positive sample pair may include, but is not limited to, a first sample target feature and a first masking enhancement feature; the first negative sample pair may include, but is not limited to, a second sample target feature and a second masking enhancement feature; the second positive sample pair may include, but is not limited to, a second sample target feature and a second masking enhancement feature; and the second negative sample pair may include, but is not limited to, a first sample target feature and a first masking enhancement feature.

[0057] To further illustrate, if we assume that there are users A and B, then the first positive sample pair is the target feature and masking enhancement feature of user A, the first negative sample pair is the target feature and masking enhancement feature of user B, the second positive sample pair is the target feature and masking enhancement feature of user B, and the second negative sample pair is the target feature and masking enhancement feature of user A.

[0058] It should be noted that in this embodiment, the training of the life cycle value model is actually performed using multiple sample data. Therefore, the first and second examples mentioned above are only illustrative and should not be construed as limiting the scope of protection.

[0059] For example, this embodiment can also use more than two sample data for training. The positive sample pairs of the current sample data represent the target features and masking enhancement features of the current sample data, while the negative samples of the current sample data correspond to the target features and masking enhancement features of other sample data besides the current sample data.

[0060] In an optional embodiment, the first sample target feature and the second sample target feature may be, but are not limited to, basic feature information representing two different users, and may include, but are not limited to, basic profiles, historical game behavior statistics, etc.

[0061] It should be noted that in the scenario of user lifecycle prediction, high-value sample users often have richer feature representations, while low-value sample users often provide simpler feature representations. Therefore, when using sample users with large differences to train the model, it will lead to the problem of feature imbalance. When using a model trained with imbalanced features to predict user lifecycle value, the accuracy of user lifecycle value prediction will be low.

[0062] In this embodiment, through comparative learning, the first sample target feature of the first sample user and the enhanced first masking feature are taken as the first positive sample pair, while the corresponding first negative sample pair are the second sample target feature of the second sample user other than the first sample user and the enhanced second masking feature.

[0063] When the second positive sample pair is the second sample target feature and the second masking enhancement feature corresponding to the second sample user, the corresponding second negative sample pair is the first sample target feature and the enhanced first masking enhancement feature corresponding to the first sample user.

[0064] By using positive and negative sample pairs, i.e., adopting a contrastive learning paradigm, more robust supervision signals are provided during the model training process. This effectively solves the problems of unbalanced and sparse distribution of sample target features, and ensures that the model can effectively distinguish between similar and different user representations after learning, thereby improving the accuracy of user lifetime value prediction.

[0065] In optional embodiments, the first masking enhancement feature may be, but is not limited to, obtained by masking and enhancing the first preliminary enhancement feature corresponding to the first sample target feature using the first masking probability vector. The second masking enhancement feature may be, but is not limited to, obtained by masking and enhancing the second preliminary enhancement feature corresponding to the second sample target feature using the second masking probability vector.

[0066] In optional embodiments, the first preliminary enhancement feature may be, but is not limited to, the preliminary enhancement feature obtained by transforming the sample target feature through a feature encoder and mapping the sample target feature to a low-dimensional hidden feature. The second preliminary enhancement feature may be, but is not limited to, the preliminary enhancement feature obtained by transforming the sample target feature through a feature encoder and mapping the sample target feature to a low-dimensional hidden feature.

[0067] In an optional embodiment, the first masking probability vector corresponds to the first preliminary enhancement feature and may be obtained through the first preliminary enhancement feature, and may be used to represent the probability that the first preliminary enhancement feature is masked or enhanced; the second masking probability vector corresponds to the second preliminary enhancement feature and may be obtained through the second preliminary enhancement feature, and may be used to represent the probability that the second preliminary enhancement feature is masked or enhanced.

[0068] It should be noted that before conducting contrastive learning, when masking or enhancing the initial enhancement features, random masking or random enhancement can be selected. However, considering that random masking or enhancement may cause the loss of some important original features or the enhancement of some unimportant features, thus affecting the sample distribution in the contrastive learning stage, making the sample distribution relatively sparse and unbalanced, thereby reducing the accuracy of the trained model in predicting user lifetime value.

[0069] In this embodiment, a masking probability vector is assigned to each feature using a masking matrix. This vector is then used to mask or enhance the initial enhanced features, generating masked enhanced features. This process enhances features that should be retained and masks features that should not be retained. Based on this, masked enhanced features are generated. During the contrastive learning phase, the masked enhanced features of the first sample user form positive sample pairs with the original sample features, while the masked enhanced features of other users form negative sample pairs with the original sample features. In other words, by assigning masking probability vectors to features, features that should be retained are enhanced, and features that should not be retained are masked, thereby further ensuring a balanced feature distribution and improving the accuracy of user lifetime value prediction.

[0070] In optional embodiments, the first preliminary enhancement feature may be, but is not limited to, obtaining a first low-dimensional embedding by mapping the first sample target feature to a low-dimensional hidden space through a feature encoder, and then obtaining the first enhancement feature based on the first low-dimensional embedding; the second preliminary enhancement feature may be, but is not limited to, obtaining a second low-dimensional embedding by mapping the second sample target feature to a low-dimensional hidden space through a feature encoder, and then obtaining the second enhancement feature based on the second low-dimensional embedding.

[0071] By mapping the target features of samples to a low-dimensional hidden space, the complexity and storage requirements of user sample feature data can be reduced by lowering the dimensionality, thereby accelerating the acquisition of initial enhanced features and improving the efficiency of user lifecycle prediction.

[0072] It should be noted that this embodiment uses positive and negative sample pairs, i.e., a contrastive learning paradigm, to provide more robust supervision signals during the model training process. This effectively solves the problem of unbalanced and sparse distribution of sample target features, and ensures that the model can effectively distinguish between similar and different user representations after learning, thereby improving the accuracy of user lifetime value prediction.

[0073] Furthermore, in this embodiment, before performing comparative learning, random masking or random enhancement can be selected when masking or enhancing the initial enhancement features. However, considering that random masking or enhancement may cause the loss of some important original features or the enhancement of some unimportant features, thereby affecting the sample distribution in the comparative learning stage, making the sample distribution relatively sparse and unbalanced, and thus reducing the accuracy of the trained model in predicting user lifetime value.

[0074] As further examples, optional ones include Figure 3 As shown, after receiving the prediction request 304 sent by the user equipment 302, the prediction target feature 306 that needs to be predicted is obtained, and the prediction target feature 306 is input into the life cycle value model 308. The life cycle value model 308 obtains the prediction result 310, wherein the prediction result 310 is related to the prediction target feature 306, and the prediction target feature 306 contains multiple features such as feature 1 and feature 2.

[0075] Through the embodiments of this application, after obtaining the target features, the target features are input into the model trained through contrastive learning and masking enhancement. Since contrastive learning can provide more robust supervision signals for the model training process, and the features are assigned masking probability vectors during the contrastive learning stage, the features that need to be retained are enhanced and the features that do not need to be retained are masked, which can ensure the balance of feature distribution and solve the problem of imbalanced feature distribution. Thus, the technical objective of predicting user lifetime value by solving the problem of imbalanced feature distribution is achieved, thereby realizing the technical effect of improving the prediction accuracy of user lifetime value.

[0076] As an optional approach, the target features corresponding to the user to be predicted are obtained, including:

[0077] S1-1, Obtain the basic user features corresponding to the user to be predicted, and obtain the target behavior features corresponding to the user to be predicted. The basic user features are used to represent the basic user information of the user to be predicted, and the target behavior features are used to represent the historical behavior information of the user to be predicted.

[0078] S1-2 integrates and processes user basic characteristics and target behavioral characteristics to obtain target characteristics.

[0079] In optional embodiments, the user basic features may be, but are not limited to, basic user information representing the user to be predicted, and may include, but are not limited to, the user's static information, such as gender, age, geographical location, etc.

[0080] In optional embodiments, the target behavioral features may, but are not limited to, be used to represent the historical behavioral information of the user to be predicted, and may, but are not limited to, cover the user's dynamic behavioral data and historical behavioral data.

[0081] It should be noted that, in the context of games, target behavioral characteristics may include, but are not limited to, game activity frequency, amount of money spent, game level, amount of gold coins, number of times in-game virtual tasks are completed, number of friends, rank level, and purchased gift packs or services.

[0082] In optional embodiments, the integration process may be, but is not limited to, the process of integrating user basic features and target behavioral features to form a comprehensive feature vector, and may be, but is not limited to, obtaining user behavior patterns and consumption tendencies based on the comprehensive feature vector.

[0083] In this embodiment, by integrating user basic features and target behavioral features to form a comprehensive feature vector, the target features input into the prediction model are the comprehensive features of the user, thereby improving the prediction accuracy of user lifetime value.

[0084] This application's embodiments obtain basic user features and target behavioral features corresponding to the user to be predicted. The basic user features represent the user's basic information, while the target behavioral features represent the user's historical behavioral information. The basic user features and target behavioral features are integrated to obtain the target features. By integrating the basic user features and target behavioral features, the technical objective of determining the target features input into the prediction model as the user's comprehensive features is achieved, thereby improving the accuracy of user lifetime value prediction.

[0085] As an optional approach, when the user to be predicted is a game user and the lifetime value model is a neural network model used to predict the lifetime value of game users, the target behavioral features corresponding to the user to be predicted are obtained, including at least one of the following:

[0086] S2-1, Obtain the first behavioral features of the user to be predicted in each game application, wherein the first behavioral features are used to represent the statistical information of the historical behavior of the user to be predicted in each game application, and the target behavioral features include the first behavioral features.

[0087] S2-2, Obtain the second behavioral feature of the user to be predicted in a specific game application, wherein the second behavioral feature is used to represent the behavioral information generated by the user to be predicted in a specific game application, and the target behavioral feature includes the second behavioral feature;

[0088] S2-3, Obtain the third behavioral feature of the user to be predicted within a recent preset time period. The third behavioral feature is used to represent the time series information of the user's game behavior within a recent preset time period. The target behavioral feature includes the third behavioral feature.

[0089] In an optional embodiment, the first behavioral feature may be, but is not limited to, the historical behavioral statistics of the user to be predicted in various game applications, and may include, but is not limited to, information such as total activity frequency, total payment amount, number of active users in game A, payment amount in game A, number of active users in game B, and payment amount in game B.

[0090] In an optional embodiment, the second behavioral feature may refer to, but is not limited to, the behavioral information of the user to be predicted in a specific game application, and may include, but is not limited to, information such as game level, amount of gold coins, number of virtual tasks completed, number of friends, and rank level.

[0091] In an optional embodiment, the third behavioral feature may, but is not limited to, represent time-series information of the user's gaming behavior in a recent preset time period, and may, but is not limited to, include information such as the user's game registration time, active time, and payment time.

[0092] It's important to note that the first behavioral feature allows us to obtain users' overall behavioral patterns and consumption history across different game applications, thus revealing their overall activity and spending tendencies in the gaming field. The second behavioral feature reflects the depth of user engagement in a specific game, enabling the model to predict the user's future willingness to continue playing that game. The third behavioral feature provides time-series information on how the target behavioral feature changes over time, allowing us to predict the user's spending intentions at different points in time. In short, by segmenting users' different behavioral features, we can accurately obtain the information implied by these features, thereby improving the accuracy of user lifetime value prediction.

[0093] This application's embodiments obtain first behavioral features of a user to be predicted in various game applications, wherein the first behavioral features represent statistical information of the user's historical behavior in each game application, and the target behavioral features include the first behavioral features; second behavioral features of a user to be predicted in a specific game application, wherein the second behavioral features represent behavioral information generated by the user in the specific game application, and the target behavioral features include the second behavioral features; and third behavioral features of a user to be predicted within a recent preset time period, wherein the third behavioral features represent time-series information of the user's game behavior within the recent preset time period, and the target behavioral features include the third behavioral features. By subdividing different user behavioral features, the technical objective of accurately obtaining the information implied by different user behavioral features is achieved, thereby improving the accuracy of user lifetime value prediction.

[0094] As an optional approach, obtain the third behavioral characteristics of the user to be predicted within a recent preset time period, including at least one of the following:

[0095] S3-1, Obtain the first time series features of the user to be predicted within a recent preset time period, wherein the first time series features are used to represent the game registration sequence of the user to be predicted within a recent preset time period, and the third behavioral features include the first time series features.

[0096] S3-2, Obtain the second time series features of the user to be predicted within a recent preset time period, wherein the second time series features are used to represent the game activity sequence of the user to be predicted within a recent preset time period, and the third behavioral features include the second time series features;

[0097] S3-3, Obtain the third time series features of the user to be predicted within a recent preset time period, wherein the third time series features are used to represent the game activity duration sequence of the user to be predicted within a recent preset time period, and the third behavioral features include the third time series features;

[0098] S3-4, Obtain the fourth time series feature of the user to be predicted within a recent preset time period, wherein the fourth time series feature is used to represent the game payment sequence of the user to be predicted within a recent preset time period, and the third behavioral feature includes the fourth time series feature.

[0099] S3-5, Obtain the fifth time series feature of the user to be predicted within a recent preset time period, wherein the fifth time series feature is used to represent the game payment amount sequence of the user to be predicted within a recent preset time period, and the third behavioral feature includes the fifth time series feature.

[0100] In an optional embodiment, the first time series feature may, but is not limited to, be used to represent the game registration sequence of the user to be predicted within a recent preset time period, and may, but is not limited to, be represented as an N-dimensional vector. This vector is used to record the N most recently registered games, and the vector elements are the identifiers of the games registered by the user. If there are fewer than N registered games, zeros are added to the end of the vector.

[0101] To illustrate further, suppose a user registers for two games with IDs 2 and 3, and the registration order is 2 first and then 3. The registration sequence would be 320000…0, where the number of 0s in the sequence is N-2.

[0102] In an optional embodiment, the second time series may, but is not limited to, represent the game activity sequence of the user to be predicted within a recent preset time period, and may, but is not limited to, be represented as an N-dimensional vector. This vector is used to record the most recent N game activity behaviors, and the vector elements are the identifiers of the user's active games. If the number of registered games is less than N, zeros are added to the end of the vector.

[0103] In an optional embodiment, the third time series may, but is not limited to, represent the sequence of game activity durations of the user to be predicted within a recent preset time period, and may, but is not limited to, be represented as an N-dimensional vector. This vector is used to record the duration of the most recent N game behaviors, and the vector elements are the user's activity duration. If the number of registered games is less than N, zeros are added to the end of the vector.

[0104] In an optional embodiment, the fourth time series may, but is not limited to, be used to represent the game payment sequence of the user to be predicted within a recent preset time period, and may, but is not limited to, be represented as an N-dimensional vector. This vector is used to record the most recent N game payment behaviors, and the vector elements are the game identifiers of the user who made the payment. If the number of registered games is less than N, zeros are added to the end of the vector.

[0105] In an optional embodiment, the fifth time series may, but is not limited to, be used to represent the game activity sequence of the user to be predicted in the recent preset time period, and may, but is not limited to, be represented as an N-dimensional vector. This vector is used to record the amount of money paid by the user in the most recent N games. The vector elements are the amount of money paid by the user. If the number of registered games is less than N, zeros are added to the end of the vector.

[0106] It's important to note that the first time-series feature, such as the game registration sequence, helps the model understand a user's recent gaming tendencies, indicating their willingness to participate in new games within a specific timeframe. The second time-series feature, the game activity sequence, shows the distribution of user activity across multiple games, helping the model understand the user's game genre preferences. The third time-series feature, the game activity duration sequence, provides the user's level of engagement in each game, helping the model determine the user's retention intentions across different games. The fourth and fifth time-series features, namely the game payment sequence and game payment amount sequence, directly reflect the user's consumption behavior, including preferences for paid games and spending levels, helping the model understand the user's consumption behavior preferences and levels. By distinguishing different time-series features of users, the model can obtain accurate user time-series characteristics, thereby improving the accuracy of user lifetime value prediction.

[0107] Through the embodiments of this application, a first time series feature of a user to be predicted within a recent preset time period is obtained, wherein the first time series feature represents the game registration sequence of the user to be predicted within the recent preset time period, and a third behavioral feature includes the first time series feature; a second time series feature of a user to be predicted within a recent preset time period is obtained, wherein the second time series feature represents the game activity sequence of the user to be predicted within the recent preset time period, and a third behavioral feature includes the second time series feature; a third time series feature of a user to be predicted within a recent preset time period is obtained, wherein the third time series feature represents the game activity duration sequence of the user to be predicted within the recent preset time period, and a third behavioral feature includes the third time series feature; a fourth time series feature of a user to be predicted within a recent preset time period is obtained, wherein the fourth time series feature represents the game payment sequence of the user to be predicted within the recent preset time period, and a third behavioral feature includes the fourth time series feature; a fifth time series feature of a user to be predicted within a recent preset time period is obtained, wherein the fifth time series feature represents the game payment amount sequence of the user to be predicted within the recent preset time period, and a third behavioral feature includes the fifth time series feature. By distinguishing different time series characteristics of users, the technical objective of helping the model obtain accurate user time series characteristics is achieved, thereby improving the technical effect of improving the accuracy of user lifetime value prediction.

[0108] According to another aspect of the embodiments of this application, a method for training a lifecycle value model is provided, characterized in that it includes:

[0109] S4-1, Obtain the first preliminary enhancement feature corresponding to the target feature of the first sample, and the second preliminary enhancement feature corresponding to the target feature of the second sample;

[0110] S4-2, obtain the first masking probability vector corresponding to the first preliminary enhancement feature and the second masking probability vector corresponding to the second preliminary enhancement feature, wherein the first masking probability vector is used to represent the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector is used to represent the probability that the second preliminary enhancement feature is masked or enhanced.

[0111] S4-3, using the first masking probability vector to mask and enhance the first preliminary enhancement feature to obtain the first masked enhancement feature, and using the second masking probability vector to mask and enhance the second preliminary enhancement feature to obtain the second masked enhancement feature;

[0112] S4-4, obtain the first positive sample pair and the first negative sample pair corresponding to the first sample target feature, and the second positive sample pair and the second negative sample pair corresponding to the second sample target feature, wherein the first positive sample pair includes the first sample target feature and the first masking enhancement feature, the first negative sample pair includes the second sample target feature and the second masking enhancement feature, the second positive sample pair includes the second sample target feature and the second masking enhancement feature, and the second negative sample pair includes the first sample target feature and the first masking enhancement feature;

[0113] S4-5, based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair, the initial lifetime value model is trained to obtain a trained lifetime value model, wherein the lifetime value model is a neural network model used to predict the lifetime value of users.

[0114] In an optional embodiment, the first sample target feature and the second sample target feature may be, but are not limited to, basic feature information representing two different users, and may include, but are not limited to, basic profiles, historical game behavior statistics, etc.

[0115] It should be noted that in this embodiment, the training of the life cycle value model is actually performed using multiple sample data. Therefore, the first and second examples mentioned above are only illustrative and should not be construed as limiting the scope of protection.

[0116] For example, this embodiment can also use more than two sample data for training. The positive sample pairs of the current sample data represent the target features and masking enhancement features of the current sample data, while the negative samples of the current sample data correspond to the target features and masking enhancement features of other sample data besides the current sample data.

[0117] In optional embodiments, the first preliminary enhancement feature may be, but is not limited to, the preliminary enhancement feature obtained by transforming the sample target feature through a feature encoder and mapping the sample target feature to a low-dimensional hidden feature. The second preliminary enhancement feature may be, but is not limited to, the preliminary enhancement feature obtained by transforming the sample target feature through a feature encoder and mapping the sample target feature to a low-dimensional hidden feature.

[0118] By mapping the target features of samples to a low-dimensional hidden space, the complexity and storage requirements of user sample feature data can be reduced by lowering the dimensionality, thereby accelerating the acquisition of initial enhanced features and improving the efficiency of model training.

[0119] In optional embodiments, the first masking enhancement feature may be, but is not limited to, obtained by masking and enhancing the first preliminary enhancement feature corresponding to the first sample target feature using the first masking probability vector. The second masking enhancement feature may be, but is not limited to, obtained by masking and enhancing the second preliminary enhancement feature corresponding to the second sample target feature using the second masking probability vector.

[0120] In an optional embodiment, the first masking probability vector corresponds to the first preliminary enhancement feature and may be obtained through the first preliminary enhancement feature, and may be used to represent the probability that the first preliminary enhancement feature is masked or enhanced; the second masking probability vector corresponds to the second preliminary enhancement feature and may be obtained through the second preliminary enhancement feature, and may be used to represent the probability that the second preliminary enhancement feature is masked or enhanced.

[0121] In optional embodiments, the first positive sample pair may include, but is not limited to, a first sample target feature and a first masking enhancement feature; the first negative sample pair may include, but is not limited to, a second sample target feature and a second masking enhancement feature; the second positive sample pair may include, but is not limited to, a second sample target feature and a second masking enhancement feature; and the second negative sample pair may include, but is not limited to, a first sample target feature and a first masking enhancement feature.

[0122] To further illustrate, if we assume that there are users A and B, then the first positive sample pair is the target feature and masking enhancement feature of user A, the first negative sample pair is the target feature and masking enhancement feature of user B, the second positive sample pair is the target feature and masking enhancement feature of user B, and the second negative sample pair is the target feature and masking enhancement feature of user A.

[0123] It should be noted that in the scenario of user lifecycle prediction, high-value sample users often have richer feature representations, while low-value sample users often provide simpler feature representations. Therefore, when using sample users with large differences to train the model, it will lead to the problem of feature imbalance. When using a model trained with imbalanced features to predict user lifecycle value, the accuracy of user lifecycle value prediction will be low.

[0124] In this embodiment, through contrastive learning, the first sample target feature of the first sample user and the enhanced first masking feature are used as the first positive sample pair, while the corresponding first negative sample pair consists of the second sample target feature of the second sample user (excluding the first sample user) and the enhanced second masking feature. When the second positive sample pair consists of the second sample target feature and the enhanced second masking feature corresponding to the second sample user, the corresponding second negative sample pair consists of the first sample target feature of the first sample user and the enhanced first masking feature. By using positive and negative sample pairs, i.e., employing a contrastive learning paradigm, a more robust supervision signal is provided during model training. This effectively addresses the problems of uneven and sparse distribution of sample target features, ensuring that the model, after learning, can effectively distinguish between similar and different user representations, thereby improving the accuracy of the model in predicting user lifetime value.

[0125] In addition, before conducting contrastive learning, when masking or enhancing the initial enhancement features, random masking or random enhancement can be selected. However, considering that random masking or enhancement may cause the loss of some important original features or the enhancement of some unimportant features, thus affecting the sample distribution in the contrastive learning stage, making the sample distribution relatively sparse and unbalanced, thereby reducing the accuracy of the trained model in predicting user lifetime value.

[0126] In this embodiment, a masking probability vector is assigned to each feature using a masking matrix. This vector is then used to mask or enhance the initial enhanced features, generating masked enhanced features. This process enhances features that should be retained and masks features that should not be retained. Based on this, masked enhanced features are generated. During the contrastive learning phase, the masked enhanced features of the first sample user form positive sample pairs with the original sample features, while the masked enhanced features of other users form negative sample pairs with the original sample features. In other words, by assigning masking probability vectors to features, features that should be retained are enhanced, and features that should not be retained are masked, thereby further ensuring a balanced feature distribution and improving the accuracy of the model in predicting user lifetime value.

[0127] It should be noted that this embodiment uses positive and negative sample pairs, i.e., a contrastive learning paradigm, to provide more robust supervision signals during the model training process. This effectively solves the problem of unbalanced and sparse distribution of sample target features, and ensures that the model can effectively distinguish between similar and different user representations after learning, thereby improving the accuracy of user lifetime value prediction.

[0128] Furthermore, in this embodiment, before performing comparative learning, random masking or random enhancement can be selected when masking or enhancing the initial enhancement features. However, considering that random masking or enhancement may cause the loss of some important original features or the enhancement of some unimportant features, thereby affecting the sample distribution in the comparative learning stage, making the sample distribution relatively sparse and unbalanced, and thus reducing the accuracy of the trained model in predicting user lifetime value.

[0129] Through the embodiments of this application, after obtaining the target features, the target features are input into the model trained through contrastive learning and masking enhancement. Since contrastive learning can provide more robust supervision signals for the model training process, and the features are assigned masking probability vectors during the contrastive learning stage, the features that need to be retained are enhanced and the features that do not need to be retained are masked, which can ensure the balance of feature distribution and solve the problem of imbalanced feature distribution. Thus, the technical objective of predicting user lifetime value by solving the problem of imbalanced feature distribution is achieved, thereby realizing the technical effect of improving the prediction accuracy of user lifetime value.

[0130] As an optional approach, the first preliminary enhancement feature corresponding to the target feature of the first sample and the second preliminary enhancement feature corresponding to the target feature of the second sample are obtained, including:

[0131] S5-1, map the target features of the first sample to the low-dimensional hidden space to obtain the first low-dimensional embedding, and map the target features of the second sample to the low-dimensional hidden space to obtain the second low-dimensional embedding;

[0132] S5-2, obtain the first enhancement feature corresponding to the first low-dimensional embedding and the second enhancement feature corresponding to the second low-dimensional embedding, wherein the first preliminary enhancement feature includes the first enhancement feature and the second preliminary enhancement feature includes the second enhancement feature.

[0133] In an optional embodiment, mapping the sample target features to a low-dimensional hidden space can be understood, but is not limited to, as converting the target features into a low-dimensional representation space through a feature encoder, thereby reducing the risk of overfitting and improving the accuracy of model predictions.

[0134] In optional embodiments, the low-dimensional embedding can be, but is not limited to, a target feature representation generated by a feature encoder in a low-dimensional hidden space.

[0135] In this embodiment, the target features of a first sample are mapped into a low-dimensional hidden space to obtain a first low-dimensional embedding, and the target features of a second sample are mapped into the low-dimensional hidden space to obtain a second low-dimensional embedding. A first enhanced feature corresponding to the first low-dimensional embedding and a second enhanced feature corresponding to the second low-dimensional embedding are obtained, wherein the first preliminary enhanced feature includes the first enhanced feature, and the second preliminary enhanced feature includes the second enhanced feature. By mapping the sample users into the low-dimensional hidden space, the technical objective of reducing overfitting risk is achieved, thereby improving the model's prediction accuracy.

[0136] As an optional approach, the first enhanced feature corresponding to the first low-dimensional embedding and the second enhanced feature corresponding to the second low-dimensional embedding are obtained, including:

[0137] S6-1, the first low-dimensional embedding is nonlinearly transformed by a multilayer perceptron to obtain the first enhanced feature;

[0138] S6-2, the second low-dimensional embedding is nonlinearly transformed using a multilayer perceptron to obtain the second enhanced feature.

[0139] In optional embodiments, the multilayer perceptron may be, but is not limited to, a feedforward neural network, may consist of an input layer, one or more hidden layers and an output layer, and may be used to perform nonlinear transformations on low-dimensional embeddings to obtain enhanced features.

[0140] It should be noted that by performing nonlinear transformation on the low-dimensional embedding through a multilayer perceptron, features in the low-dimensional embedding can be further extracted to generate enhanced features, providing more detailed information for the generation of the masking probability vector. This helps to solve the problems of feature imbalance and sparsity, thereby improving the prediction accuracy of the model.

[0141] In this embodiment, a first enhanced feature is obtained by performing a nonlinear transformation on the first low-dimensional embedding using a multilayer perceptron; a second enhanced feature is obtained by performing a nonlinear transformation on the second low-dimensional embedding using a multilayer perceptron. By performing a nonlinear transformation on the low-dimensional embedding using a multilayer perceptron, more detailed information is provided for the generation of the masking probability vector, thereby helping to solve the problems of feature imbalance and sparsity, and ultimately improving the model's prediction performance.

[0142] As an optional approach, the first masking probability vector corresponding to the first preliminary enhanced feature is obtained, including:

[0143] S7-1, add corresponding noise to each feature element in the first preliminary enhancement feature, where the noise is sampled from the extreme value distribution;

[0144] S7-2, the first preliminary enhancement feature after adding noise is input into the normalized exponential function to obtain the probability distribution vector, where the probability distribution vector is used to represent the probability that each feature element in the first preliminary enhancement feature after adding noise is masked or enhanced, and the first masking probability vector includes the probability distribution vector.

[0145] In optional embodiments, the extreme value distribution may be used, but is not limited to, to sample noise to help generate a masking probability vector for automatic feature enhancement or masking.

[0146] In an optional embodiment, the normalized exponential function may, but is not limited to, perform an exponential transformation on the elements of the input vector and then process the sum of the exponents of all elements to obtain the probability of each element. This may, but is not limited to, be used to obtain a probability distribution vector after obtaining the first preliminary enhanced features with added noise.

[0147] In optional embodiments, the probability distribution vector may be, but is not limited to, a vector processed by a normalized exponential function, and may, but is not limited to, represent the probability that each feature element in the feature space is masked or enhanced.

[0148] It should be noted that by adding noise with extreme value distribution to the initial enhanced features and then using a normalized exponential function to generate a probability distribution vector, not only is the semantic structure of the original features preserved, but also additional supervision signals are provided for model training, i.e., a self-supervised mechanism. This allows the trained model to automatically mask or enhance feature elements, thereby improving the accuracy of the predictions made by the trained model.

[0149] In this embodiment, noise is added to each feature element in the first preliminary enhancement feature. The noise is sampled from an extreme value distribution. The noise-added first preliminary enhancement feature is then input into a normalized exponential function to obtain a probability distribution vector. This probability distribution vector represents the probability that each feature element in the noise-added first preliminary enhancement feature is masked or enhanced. The first masking probability vector includes the probability distribution vector. By adding noise to the preliminary enhancement feature and then obtaining the probability distribution vector from the noise-added feature, the technical objective of enabling the trained model to automatically mask or enhance feature elements is achieved, thereby improving the prediction accuracy of the trained model.

[0150] As an optional approach, the first preliminary enhancement feature is masked and enhanced using the first masking probability vector to obtain the first masked enhancement feature, including:

[0151] S8-1, embed the probability distribution vector and the first low dimension into the input masking enhancement function to obtain the hidden space features of the masking enhancement;

[0152] S8-2 transforms the hidden space features into masking enhancement features of the original dimension, where the original dimension is the dimension of the target feature of the first sample, and the first masking enhancement feature includes the masking enhancement feature of the original dimension.

[0153] In optional embodiments, the masking enhancement function may be used, but is not limited to, to selectively mask or enhance low-dimensional embeddings based on probability distribution vectors to obtain masked and enhanced hidden space features.

[0154] It should be noted that after masking enhancement of the low-dimensional embedding, the hidden space features of the mask enhancement are obtained, and then the hidden space features are converted into mask enhancement features of the original dimension by the decoder. That is, they are converted into mask enhancement features that can be directly used for the training and prediction process of the prediction model, thereby improving the learning efficiency of the model and the accuracy of model prediction.

[0155] In this embodiment, the probability distribution vector and a first low-dimensional embedding are input into the masking enhancement function to obtain masked and enhanced hidden space features. These hidden space features are then converted into masking enhancement features of the original dimension, where the original dimension is the dimension of the first sample target features, and the first masking enhancement features include the masking enhancement features of the original dimension. By performing masking enhancement on the low-dimensional embedding to obtain hidden space features, and then converting these hidden space features into masking enhancement features that can be directly used in the training and prediction processes of the prediction model, the technical effect of improving the model's learning efficiency is achieved.

[0156] As an optional approach, when the training phase of the lifecycle value model includes a pre-training phase and a fine-tuning phase, the sample data in the pre-training phase comes from the source domain, and the sample data in the fine-tuning phase comes from the target domain. The sample data from the source domain includes first sample target features and second sample target features. Based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair, the initial lifecycle value model is trained to obtain a trained lifecycle value model, including:

[0157] S9-1, based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair, the initial life cycle value model is pre-trained to obtain the first life cycle value model;

[0158] S9-2, using sample data from the target domain, fine-tunes the first life cycle value model to obtain the second life cycle value model, wherein the trained life cycle value model includes the second life cycle value model.

[0159] In optional embodiments, pre-training can be understood, but is not limited to, as using a large amount of rich source domain sample data to initially train the model in the initial stage of model training, thereby establishing the model's initial understanding of the data and feature learning capabilities.

[0160] In optional embodiments, the source domain sample data may, but is not limited to, data from a dataset with rich data and relatively complete feature distribution, and may, but is not limited to, include numerous positive and negative sample pairs, and may, but is not limited to, be used to pre-train the model to obtain the first lifecycle value model.

[0161] In an optional embodiment, the target domain sample data may, but is not limited to, data from a dataset with a small amount of data, and may, but is not limited to, be used to fine-tune the first lifecycle value model so that the fine-tuned second lifecycle value model can adapt to the feature distribution under a specific scenario and improve the accuracy of model prediction.

[0162] It should be noted that, in order to mitigate the distributional gap between the source and target domains, this embodiment explicitly encodes the user representation by minimizing the JS (Jensen-Shannon) divergence between the source and target domains. The JS divergence can be, but is not limited to, a symmetric smoothed version of the KL (Kullback-Leibler) divergence, and can be, but is not limited to, used to stabilize training and improve the model's cross-domain generalization.

[0163] In addition, the use of JS divergence depends on the similarity between the source domain and the target domain. If the similarity between the source domain and the target domain is less than a preset value, that is, the data and feature distributions between the source domain and the target domain are significantly different, then JS divergence is used. If the similarity between the source domain and the target domain is greater than a preset value, that is, the data and feature distributions between the source domain and the target domain are relatively similar, then JS divergence is not used, thereby preventing overfitting and improving the prediction accuracy of the model.

[0164] In this embodiment, an initial lifecycle value model is pre-trained based on a first positive sample pair, a first negative sample pair, a second positive sample pair, and a second negative sample pair to obtain a first lifecycle value model. Then, using sample data from the target domain, the first lifecycle value model is fine-tuned to obtain a second lifecycle value model. The trained lifecycle value model includes the second lifecycle value model. By fine-tuning the first lifecycle value model using sample data from the target domain to obtain the second lifecycle value model, the fine-tuned second lifecycle value model can adapt to the feature distribution in a specific scenario, thereby improving the accuracy of model predictions.

[0165] As an optional approach, the initial lifecycle value model is pre-trained based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair to obtain the first lifecycle value model, which includes:

[0166] Perform the following steps until you obtain the first lifecycle value model:

[0167] S10-1, Obtain the current positive and negative sample pairs, where the first positive sample pair and the first negative sample pair are positive and negative sample pairs, and the second positive sample pair and the second negative sample pair are positive and negative sample pairs;

[0168] S10-2, obtain the first similarity between positive sample pairs in the current positive-negative sample pair, and the second similarity between negative sample pairs in the current positive-negative sample pair;

[0169] S10-3, Based on the first similarity and the second similarity, the current contrastive loss of the current initial life cycle value model is obtained using the current contrastive loss function. The contrastive loss is used to measure the sensitivity of the life cycle value model in distinguishing between positive and negative samples.

[0170] S10-4, If the current contrastive loss satisfies the convergence condition of the pre-training stage, the current initial lifecycle value model is determined as the first lifecycle value model.

[0171] S10-5, If the current contrastive loss does not meet the convergence condition of the pre-training stage, adjust the temperature parameter in the current contrastive loss function, where the temperature parameter is used to control the sensitivity of the life cycle value model in distinguishing between positive and negative samples.

[0172] S10-6, obtain the next positive and negative sample pair, and use the next positive and negative sample pair as the current positive and negative sample pair, until the first life cycle value model is obtained.

[0173] In an alternative embodiment, the contrastive loss function may be, but is not limited to, a loss function used until model training, and may, but is not limited to, measuring the model’s ability to distinguish user representations by calculating the similarity between pairs of positive and negative samples.

[0174] In optional embodiments, the temperature parameter may be, but is not limited to, a parameter in the contrastive loss function, and may be, but is not limited to, used to control the sensitivity of positive and negative samples in the contrastive learning process.

[0175] In an optional embodiment, satisfying the convergence condition of the pre-training stage can be understood, but is not limited to, as the model is considered to have converged when the value of the contrastive loss function reaches a predetermined threshold or the rate of change is lower than a certain threshold during the pre-training process, and the model pre-training is considered to be complete.

[0176] It's important to note that during the model's pre-training phase, contrastive learning is used to improve the model's ability to discriminate user representations. This process involves calculating the similarity between positive and negative sample pairs, then using a contrastive loss function to obtain the current contrastive loss. If the contrastive loss does not meet the convergence condition, it means the model has not yet learned sufficient feature discrimination ability and needs to continue training, thus performing iterative training. Once the contrastive loss obtained from the contrastive loss function meets the convergence condition, the model pre-training is considered complete. By performing iterative training during the model's pre-training phase, the contrastive loss of the pre-trained model is made to meet the convergence condition, thereby ensuring the accuracy of the model's predictions.

[0177] Through the embodiments of this application, the current positive and negative sample pairs are obtained, wherein the first positive sample pair and the first negative sample pair are positive and negative sample pairs, and the second positive sample pair and the second negative sample pair are positive and negative sample pairs; the first similarity between positive sample pairs in the current positive and negative sample pairs and the second similarity between negative sample pairs in the current positive and negative sample pairs are obtained; based on the first and second similarities, the current contrastive loss function is used to obtain the current contrastive loss of the current initial lifecycle value model, wherein the contrastive loss is used to measure the sensitivity of the lifecycle value model in distinguishing positive and negative samples; if the current contrastive loss meets the convergence condition of the pre-training stage, the current initial lifecycle value model is determined as the first lifecycle value model; if the current contrastive loss does not meet the convergence condition of the pre-training stage, the temperature parameter in the current contrastive loss function is adjusted, wherein the temperature parameter is used to control the sensitivity of the lifecycle value model in distinguishing positive and negative samples; the next positive and negative sample pair is obtained and used as the current positive and negative sample pair, until the first lifecycle value model is obtained. By pre-training the model through iterative training, the technical objective of ensuring that the contrastive loss of the pre-trained model meets the convergence condition is achieved, thereby ensuring the accuracy of the model's predictions.

[0178] As an optional approach, in the process of training the initial lifecycle value model based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair to obtain the trained lifecycle value model, the method also includes:

[0179] The initial lifetime value model is trained by combining cross-entropy loss and log-normal loss. Cross-entropy loss is used to determine whether a user is a high-value user or a low-value user, while log-normal loss is used to ensure that the predicted value of the lifetime value model conforms to the log-normal distribution parameters of the actual value.

[0180] Optionally, cross-entropy loss can be understood, but is not limited to, as a loss function used for classification problems to measure the difference between the model's predictions and the actual results. It evaluates the model's performance by calculating the negative log probability of the true class. The smaller the loss value, the closer the model's predictions are to the true results.

[0181] Optionally, the log-normal loss can be understood, but is not limited to, as a loss function for continuous numerical prediction. It is suitable for log-normally distributed data and maps the data to a normal distribution space through logarithmic transformation, thereby simplifying the modeling process. When the target variable exhibits a log-normal distribution, the log-normal loss can effectively capture the distribution characteristics of the data and improve the prediction accuracy of the model.

[0182] It's important to note that cross-entropy loss assesses a model's ability to classify users effectively, categorizing them as high-value or low-value users. Log-normal loss, on the other hand, minimizes the distributional difference between predicted and actual values, ensuring the model's predictions align with the distribution characteristics of the actual data, thus improving prediction accuracy.

[0183] In this embodiment, an initial lifetime value model is trained by combining cross-entropy loss and log-normal loss. Cross-entropy loss is used to determine whether a user is a high-value or low-value user, while log-normal loss ensures that the predicted values ​​of the lifetime value model conform to the log-normal distribution parameters of the actual values. By training the model using cross-entropy loss and log-normal loss, the technical objective of classifying users into high-value or low-value users and ensuring that the model's predictions are consistent with the distribution characteristics of the actual data is achieved is realized, thereby improving the accuracy of model predictions.

[0184] As an alternative, the aforementioned user lifetime value prediction method can be applied to game advertising scenarios.

[0185] This embodiment proposes a cross-domain user lifetime value (LTV) prediction model based on the self-reinforcing contrastive learning paradigm (SLTV). During the pre-training phase of the source domain data, this embodiment employs an automatic reinforcement paradigm to learn the masking matrix of user representations in the hidden space. Compared to random masking, this method is more effective in generating meaningful feature views. To overcome the challenge of uneven feature distribution in LTV data, this embodiment introduces a contrastive learning strategy that more closely aligns the original and reinforced hidden features, thereby providing additional self-supervised signals for model training. Compared to supervised learning that relies solely on labels, this embodiment produces more robust hidden features. The pre-trained LTV prediction model from the source domain is then used to guide the training of the target model.

[0186] In an optional embodiment, the above method can be applied to advertising scenarios in the gaming industry based on Realtime API (RTA) advertising. When user traffic reaches the advertising platform, the platform transmits the user's device ID to the advertiser. The advertiser obtains the user's value score based on the user value prediction model and sends this score back to the advertising platform. The advertising platform adjusts the advertising bid based on the user's returned value score, that is, multiplying the original bid by a coefficient (a larger coefficient is used for users with high value scores, increasing the bid; a smaller coefficient is used for users with low value scores, resulting in a lower bid), thus achieving personalized and precise advertising.

[0187] In optional embodiments, the above method can also be applied to custom event advertising scenarios in the gaming industry. Currently, mainstream advertising platforms support not only conventional advertising targeting activation and payment events, but also advertisers' custom events as advertising optimization targets. Through the LTV prediction model, high-value predictions can be used as custom events; that is, users whose predicted LTV exceeds a certain threshold are sent back to the advertising platform to improve the return on investment (ROI) of advertising.

[0188] It should be noted that the goal of cross-domain scenarios is to leverage the abundant data available on certain platforms for pre-training, thereby aiding learning tasks in data-scarce domains. This paper defines the game with abundant samples as the source domain S, and other games as the target domain T. Each user example is represented by (x, y), where x∈x represents the target feature, and y∈y represents the consumer's LTV label. The source domain dataset S = {X...} s Y s} and the target domain dataset T = {X t Y t The LTV prediction task is reduced to a regression problem, where the model takes input from the original features x as input to the LTV prediction model. s (t) is processed to produce an output representation e s (t), and then used to predict y s (t).

[0189] First, target features are constructed using basic profiles and historical game behavior:

[0190] Features are obtained through user profiles and historical gaming behavior, mainly including the following dimensions:

[0191] (1) Basic user profile: gender, age, province, city;

[0192] (2) Historical behavioral statistics of the game market: total active frequency, total payment amount, number of active games of game A, payment amount of game A, number of active games of game B, payment amount of game B (this includes the number of active games and payment amount of all games under the company);

[0193] (3) In-game behavioral characteristics: The user’s behavior in the game to be estimated, including game level, amount of gold coins, number of tasks completed, ranking score, number of friends, etc. Different games have different behaviors, so the characteristics may be different.

[0194] (4) Time-series characteristics of game behavior:

[0195] 1) Game registration sequence: A 10-dimensional vector that records the 20 most recently registered games. Each element of the vector is the ID of the game registered by the user. If there are fewer than 20 registered games, 0 is added to the end of the vector. For example, if a user registered two games with IDs 2 and 3 (registering 2 first and then 3), the registration sequence would be 3 2 0 0 0 0 0 0 0.

[0196] 2) Game activity sequence: a 20-dimensional vector that records the 20 most recent game activity behaviors. Each element of the vector is the ID of the user's active games. If there are fewer than 20 active games, zeros are added to the end of the vector.

[0197] 3) Game active duration sequence: a 20-dimensional vector that records the duration of the last 20 game active behaviors. Each element of the vector is the user's active duration. If there are fewer than 20 active games, zeros are added to the end of the vector.

[0198] 4) Game payment sequence: a 20-dimensional vector that records the 20 most recent game payment behaviors. Each element of the vector is the ID of the game paid by the user. If there are fewer than 20 paid games, zeros are added to the end of the vector.

[0199] 5) Game payment amount sequence: a 20-dimensional vector that records the amount of the most recent 20 game payment behaviors. Each element of the vector is the amount paid by the user. If there are fewer than 20 paid games, zeros are added to the end of the vector.

[0200] Non-numeric features are processed as follows: gender is represented by 1 / 0 to indicate male / female; each region is replaced by an integer between 1 and 34; each city is replaced by an integer between 1 and 1000.

[0201] Finally, by combining all the above features, a numerical vector is generated for each user, which is the user's final feature.

[0202] In an optional embodiment, the comprehensive architecture of the model is divided into two distinct phases: a pre-training phase and a fine-tuning phase. In the pre-training phase, this embodiment utilizes rich data from the source domain to train the LTV prediction model (expert model). This embodiment employs a learnable strategy to enhance features in the latent space and applies a contrastive learning paradigm to provide supervisory signals for the source domain features, thereby addressing the problem of uneven feature distribution. The pre-trained expert model provides a solid foundation for training on the target domain data, effectively transferring knowledge from the source domain to guide the information and encoding of target features in the target domain. In the fine-tuning phase, this embodiment uses a recursive alignment technique on the sample features to approximate the distribution characteristics of the encoded features across the two domains.

[0203] Further examples, such as Figure 4 As shown, in the expert model stage, the source domain data is pre-trained by inputting it into the source Transformer, and then the encoded representation of the user data e is obtained. s Then, an automatic enhancement strategy is used, that is, the feature encoder Enc is used to enhance e. s Mapping to a low-dimensional hidden space to obtain the low-dimensional embedding (hidden space features) h s Next, the vector p and the enhancement function Mask are obtained through an MLP layer using the Gumbel-Softmax technique. Finally, the decoder Dec is used to convert the enhanced or masked features into the original dimension e'. s The feature vectors are obtained, and the loss function is acquired. Then according to e s ,e' s as well as Obtain the expected distribution probability P(s), and finally output the prediction result, (p, μ, σ), where p represents the probability of payment, and μ and σ are parameters of the log-normal distribution. Finally, the result is calculated using the loss function. and Optimize the expert model;

[0204] In the target model phase, source domain data and target domain data are input into the target Transformer transmitted from the expert model phase, and then the encoded representation of user data e from the source domain is obtained respectively. s And the user data e represented by the encoding of the target domain t Then, during the embedding alignment stage, hs and ht are obtained as low-dimensional embeddings, followed by the balanced divergence loss. The expected probability distribution P(s) of the source domain and the expected probability distribution P(t) of the target domain are obtained, and the probability distribution of (p, μ, σ) is obtained, as well as the prediction loss of the target domain. Finally, based on as well as Optimize the target model.

[0205] The purpose of pre-training is to improve the LTV prediction performance of the target domain by leveraging the abundant samples in the source domain. Firstly, this embodiment utilizes a transformer e... s Transformation to obtain an encoded user representation, for example:

[0206] e s =Tran(x) s ),x s ∈X s

[0207] The goal of this embodiment is to provide a learnable mechanism to facilitate the automatic augmentation of target features in the latent space. This embodiment proposes a feature masking-based data augmentation method. First, this embodiment uses a feature encoder (Enc) to map the initial target features into a low-dimensional latent space. Within this low-dimensional space, this embodiment performs user embedding data augmentation through the following implementation details:

[0208] h s =σ(Enc(e) s )),e s ∈E s

[0209] z s =MLP 2 (h s )

[0210] p s =GumbelSoftmax(z s )

[0211] Among them, h s For e s The low-dimensional embedding is obtained through the feature encoder Enc, where σ is a non-linear activation layer. The enhancement pooling for each node is masked and preserved; in this embodiment, the dimension of the representation is set to the same number of possible enhancements performed by the MLP layer. Vector p s Based on the final representative z s The Gumbel-Softmax technique extracts a hot vector from this distribution and ensures its differentiability by using reparameterization techniques. This embodiment employs the following method to achieve data augmentation:

[0212] h′ s =M(h) s ,p s ),e s ′=Dec(h′ s )

[0213] By using differentiable operations (such as multiplication), the enhancement function M can be combined with h. s and p sThis achieves feature masking while preserving the weight gradient of the augmentation probability. This allows for computation using the backpropagation algorithm. In this embodiment, the decoder Dec is used to augment (mask) the hidden space features h′. s Convert to original dimension e s The feature vector of ''. This embodiment employs a learnable strategy to obtain the masking probability and adds features in a differentiable manner.

[0214] It should be noted that, to address the issues of imbalanced and sparsity in source domain feature distribution, this embodiment employs a contrastive learning paradigm, providing more robust supervision signals during model training. In this embodiment, for each user, the original embedding and its corresponding augmented embedding are considered as a positive pair (e pos ,e′ pos Embedsions from other users in the same batch are considered negative pairs (e.g., embeddings from other users in the same batch). neg ,e′ neg This setup ensures that the model learns to effectively distinguish between similar and different user representations. Formally, according to the InfoNCE loss function,

[0215] User-represented contrast loss Defined as:

[0216]

[0217] Here, s(,) represents the cosine similarity function, measuring the similarity between two embeddings, and τ is an adjustable temperature parameter, an adjustable loss, which controls the sharpness of the probability distribution. This temperature parameter is a common setting that helps balance the trade-off between hard and simple samples during training. This self-supervised mechanism allows the original embeddings and augmented embeddings to collaboratively enhance each other, thus providing richer and more powerful user representations for the target domain.

[0218] Optionally, this embodiment assumes that the underlying LTV data follows a log-normal distribution and uses a variant of the ZILN loss to optimize the LTV prediction model. This approach captures the heavy-tailed nature of LTV data, which typically features a large number of low-value customers and a small number of high-value customers. Formally, the loss function takes a sample (x, y) as input and outputs a prediction result (p, μ, σ), where p represents the probability of payment, and μ and σ are parameters of the log-normal distribution. The model is optimized by minimizing the following loss function:

[0219] L ZILN (y;p,v,σ)=L CrossEntropy (1 {y>0} ;p)+1 {y>0} L Lognormal (y;μ,σ)

[0220] Here, 1 represents an indicator function used to determine whether a payment has been made (y>0). L cross Entropy maximizes the likelihood of a customer paying, effectively encouraging the model to differentiate between paying and non-paying customers. log The normal distribution maximizes the observation of the payoff y by following a log-normal distribution of the prediction parameters (μ, σ), specifically expressed as:

[0221]

[0222] Here, μ and σ represent the mean and standard deviation of the log-normal distribution, respectively. The logarithmic transformation of y allows the model to handle the skewed nature of the LTV data, ensuring that the predicted values ​​are consistent with the observed distribution. Notably, this embodiment uses ZILN loss to train both the source and target experts. This approach is advantageous because, compared to ordinary self-supervised learning in the source domain, incorporating partial labels provides guidance, enabling the source experts to acquire valuable knowledge about LTV predictions. This will result in more accurate LTV predictions in the target domain.

[0223] Furthermore, to mitigate the distributional discrepancy between the source and target domains, this embodiment explicitly aligns the encoded user representations by minimizing the Jensen-Shannon (JS) divergence between the two distributions. JS divergence is a symmetric, smoothed version of Kullback-Leibler (KL) divergence, which helps stabilize training and improve the model's cross-domain generalization. Formally, the source domain distribution P... s Distribution P of the target domain t The JS divergence between them is defined as:

[0224]

[0225] Where M is P s and P t The average distribution of is given by the following formula:

[0226]

[0227] KL divergence, d KL (p||q) measures how much a probability distribution q deviates from another expected probability distribution p. It is defined as:

[0228]

[0229] By minimizing JS divergence, this implementation enables the model to generate similar representations for users in both domains, thereby reducing distributional discrepancies. This alignment leads to more consistent and reliable predictions when the model moves from a source domain with a large amount of labeled data to a target domain with potentially less labeled data. The advantage of using JS divergence compared to using KL divergence alone is its ability to handle situations where one distribution assigns the probability of an event to zero, while the other does not. This property makes JS divergence more robust and less sensitive to extreme differences between the two distributions, which is particularly effective in cross-domain scenarios where feature distributions may differ significantly.

[0230] By adjusting the distribution in this way, the model in this embodiment can effectively utilize knowledge learned from the source domain to improve the prediction accuracy of the target domain, and ultimately improve the overall performance of LTV prediction across different games.

[0231] It should be noted that the method in this embodiment consists of two key stages: a pre-training process and a fine-tuning process. The pre-training process focuses on learning robust user representations and initial LTV predictions by leveraging a large amount of data from the source domain. In contrast, the fine-tuning process adapts these pre-trained representations to the target domain, ensuring that the model can generalize well to new data with potentially different distributions. During the pre-training stage, the goal of this embodiment is to combine contrastive learning with ZILN loss to jointly optimize the model. Contrastive loss The model is encouraged to learn to discriminate user representations by comparing positive and negative pairs, while the ZILN loss... The focus is on accurate LTV prediction. These two loss components are integrated into a unified objective.

[0232]

[0233] This joint optimization helps the model capture semantic similarity between users and the distribution characteristics of LTV data, thus laying a solid foundation for subsequent fine-tuning. After pre-training, the model undergoes a fine-tuning phase, where it is further refined to fit the target domain. The fine-tuning process involves optimizing the combined loss function, which includes the Jensen-Shannon (JS) divergence loss, ensuring consistency between the distributions of the source and target domains and the prediction losses for both domains.

[0234]

[0235] Where, λ j , λ s and λ t It is a balancing JS divergence loss Source domain prediction loss and target domain prediction loss The contributing hyperparameters.

[0236] In addition, in this embodiment, the proposed framework is evaluated using two key metrics:

[0237] The area under the ROC curve (AUC) and the normalized Gini coefficient are both measured over the target domain. The AUC, calculated as the area under the receiver operating character feature (ROC) curve, assesses the accuracy of consumption probability predictions. The normalized Gini coefficient is determined by the ratio of the Gini coefficient for LTV predictions (Gini 1997) to the Gini coefficient for LTV labels. This metric evaluates the model's ability to accurately identify high-spending consumers among all consumers. High AUC and normalized Gini scores indicate good predictive ability for LTV.

[0238] In optional embodiments, to verify the proposed framework, such as Figure 5 As shown, this embodiment compares its model with hybrid models of single DNN, DeepFM, DCN, GateNet, and CFAD in two experimental settings. Datasets 1, 2, and 3 exist, each containing AUC and Gini data, and both single-domain and dual-domain methods are tested. The "single-domain" setting refers to a model specifically trained on target domain data and evaluated within the same domain. The "cross-domain" setting applies to the model proposed in this embodiment, which is trained using data from the source domain and then tested on the target domain. It can be seen that the cross-domain model has an advantage over the single-domain model, indicating that the source domain is effective in guiding target domain training. SLTV's absolute advantage over other methods across all datasets is attributed to the introduction of an automatically augmented comparative learning paradigm that brings better training to the source domain, which in turn provides more positive guidance for fine-tuning the target domain.

[0239] It should be noted that, in order to evaluate the impact of the four components in the method of this embodiment, an ablation study was conducted on datasets G1 and G2, and the results are as follows: Figure 6 , Figure 7 As shown,

[0240] in, Figure 6 There are datasets 1, 2, and 3, as well as the AUC and Gini coefficients of the fine-tuned model and the SLTV model. Figure 7Different settings exist, such as the source domain excluded model setting (excluding the source domain prediction loss), the unsupervised model setting (removing the self-supervised learning loss), the source domain-aware fine-tuning model (the target fine-tuning model without source domain model guidance), and the reward model. The data clearly shows that each component plays a crucial role in the method of this embodiment. For example, excluding the source domain prediction loss results in a 1.13% decrease in AUC and a 3.08% decrease in the Gini coefficient on G2. Similarly, removing the self-supervised learning loss results in a 0.6% decrease in AUC and a 2.11% decrease in the Gini coefficient. SLTV also compared the target fine-tuning model without source domain model guidance, demonstrating that the model in this embodiment achieves the best performance across all datasets.

[0241] To analyze the impact of the hyperparameters corresponding to the four components of the method in this embodiment, experiments were conducted on datasets G1 and G2 using a mixed model. The results are as follows: Figure 8 As shown. To more effectively illustrate the impact of different hyperparameters, this embodiment reports experimental results relative to AUC and relative to Gini, calculated as follows:

[0242] relative AUC=(AUC-refAUC)×1000,

[0243] relative Gini=(Gini-refGini)×1000,

[0244] Among them, refAUC and refGini are selected based on the actual AUC and Gini, respectively. Specifically, as follows... Figure 8 As shown in (a), the hyperparameter λ can be obtained. j The impact on AUC and Gini, such as Figure 8 As shown in (b), the hyperparameter λ can be obtained. s The impact on AUC and Gini, such as Figure 8 As shown in (c), the hyperparameter λ can be obtained. t Regarding the impact on AUC and Gini, the effects of hyperparameters on method performance exhibit similar patterns across both datasets. This observation simplifies the process of selecting hyperparameters, ensuring robust performance of the model in this embodiment regardless of the dataset used. Specifically, this embodiment allows selecting the hyperparameter λ within the range of 0.1 to 0.5. j and λ s Furthermore, this embodiment can select λ. t The hyperparameters range from 0.5 to 2.

[0245] In an optional embodiment, to examine the effectiveness of this embodiment in optimizing LTV features in both domain-invariant and domain-specific scenarios, this embodiment employs t-SNE technology to visualize the embeddings of the G2 dataset in order to visually demonstrate the learning status of the target features at different stages of module training.

[0246] result Figure 9 As shown, the specific Figure 9 Images (a) and (b) in the diagram illustrate the visualization of the source domain feature distribution under different pre-trained expert models. Without SLTV in this embodiment, the user representation encoded by the model appears cluttered, making it difficult for the LTV predictor to achieve a proper feature distribution. In contrast, the source domain trained with SLTV exhibits a more uniform distribution. Figure 9 (c) and (d) in the diagram describe the target domain feature distribution with a CDAF pre-trained expert model and the target domain feature distribution with an SLTV pre-trained expert model, respectively. This embodiment attributes this improvement to the use of an automatically enhanced contrastive learning paradigm, which provides sufficient supervision during the learning process of the source domain data, thereby mitigating the problems associated with sparse and heterogeneous feature distributions.

[0247] This application introduces a novel cross-domain LTV prediction model, SLTV, based on an automatic augmentation paradigm, aiming to address the issues of data scarcity and feature sparsity in LTV prediction tasks. During the pre-training phase on the source domain data, this embodiment uses an automatic augmentation paradigm to learn the masking matrix of the hidden space user representations. In the fine-tuning phase, this embodiment explicitly aligns the source and target user representations by minimizing the Jensen-Shannon divergence between the target and source domains. Extensive experiments on three real-world datasets demonstrate that the method of this embodiment significantly improves the LTV prediction performance in the target domain.

[0248] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0249] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0250] According to another aspect of the embodiments of this application, a user lifetime value prediction apparatus is also provided for implementing the above-described user lifetime value prediction method. For example... Figure 10 As shown, the device includes:

[0251] The first acquisition unit 1002 is used to acquire the target features corresponding to the user to be predicted;

[0252] Prediction unit 1004 is used to input target features into the trained lifetime value model to obtain the prediction result corresponding to the lifetime value of the user to be predicted.

[0253] The Lifetime Value Model is a neural network model used to predict user lifetime value. The trained Lifetime Value Model is the initial Lifetime Value Model, obtained after training based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair. The first positive sample pair includes the first sample target feature and the first masking enhancement feature. The first negative sample pair includes the second sample target feature and the second masking enhancement feature. The second positive sample pair includes the second sample target feature and the second masking enhancement feature. The second negative sample pair includes the first sample target feature and the first masking enhancement feature. The first masking enhancement feature is obtained by masking and enhancing the first preliminary enhancement feature corresponding to the first sample target feature using the first masking probability vector. The second masking enhancement feature is obtained by masking and enhancing the second preliminary enhancement feature corresponding to the second sample target feature using the second masking probability vector. The first masking probability vector corresponds to the first preliminary enhancement feature, and the second masking probability vector corresponds to the second preliminary enhancement feature. The first masking probability vector is used to represent the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector is used to represent the probability that the second preliminary enhancement feature is masked or enhanced.

[0254] For specific implementation examples, please refer to the examples shown in the above-described method for predicting user lifetime value, which will not be repeated here.

[0255] As an optional solution, the first acquisition unit 1002 includes: a first acquisition module, used to acquire the user basic features corresponding to the user to be predicted, and to acquire the target behavior features corresponding to the user to be predicted, wherein the user basic features are used to represent the basic user information of the user to be predicted, and the target behavior features are used to represent the historical behavior information of the user to be predicted; and an integration module, used to integrate the user basic features and the target behavior features to obtain the target features.

[0256] For specific implementation examples, please refer to the examples shown in the above-described method for predicting user lifetime value, which will not be repeated here.

[0257] As an optional solution, the first acquisition module includes at least one of the following: a first acquisition submodule, used to acquire first behavioral features of the user to be predicted in various game applications, provided that the user to be predicted is a game user and the lifetime value model is a neural network model for predicting the lifetime value of game users, wherein the first behavioral features represent statistical information of the user's historical behavior in various game applications, and the target behavioral features include the first behavioral features; a second acquisition submodule, used to acquire second behavioral features of the user to be predicted in a specific game application, provided that the user to be predicted is a game user and the lifetime value model is a neural network model for predicting the lifetime value of game users, wherein the second behavioral features represent behavioral information generated by the user to be predicted in the specific game application, and the target behavioral features include the second behavioral features; a third acquisition submodule, used to acquire third behavioral features of the user to be predicted within a recent preset time period, provided that the user to be predicted is a game user and the lifetime value model is a neural network model for predicting the lifetime value of game users, wherein the third behavioral features represent time-series information of the user's game behavior generated within the recent preset time period, and the target behavioral features include the third behavioral features.

[0258] For specific implementation examples, please refer to the examples shown in the above-described method for predicting user lifetime value, which will not be repeated here.

[0259] As an optional solution, the third acquisition submodule includes at least one of the following: a first acquisition subunit, used to acquire a first time series feature of the user to be predicted within a recent preset time period, wherein the first time series feature represents the game registration sequence of the user to be predicted within the recent preset time period, and the third behavioral feature includes the first time series feature; a second acquisition subunit, used to acquire a second time series feature of the user to be predicted within a recent preset time period, wherein the second time series feature represents the game activity sequence of the user to be predicted within the recent preset time period, and the third behavioral feature includes the second time series feature; and a third acquisition subunit, used to acquire a third time series feature of the user to be predicted within a recent preset time period. The system comprises the following features: a third time series feature representing the game activity duration sequence of the user to be predicted within a recent preset time period, and a third behavioral feature including the third time series feature; a fourth acquisition subunit for acquiring the fourth time series feature of the user to be predicted within a recent preset time period, wherein the fourth time series feature represents the game payment sequence of the user to be predicted within a recent preset time period, and the third behavioral feature including the fourth time series feature; and a fifth acquisition subunit for acquiring the fifth time series feature of the user to be predicted within a recent preset time period, wherein the fifth time series feature represents the game payment amount sequence of the user to be predicted within a recent preset time period, and the third behavioral feature including the fifth time series feature.

[0260] For specific implementation examples, please refer to the examples shown in the above-described method for predicting user lifetime value, which will not be repeated here.

[0261] According to another aspect of the embodiments of this application, a training apparatus for a lifecycle value model for implementing the training method of the above-described lifecycle value model is also provided. For example... Figure 11 As shown, the device includes:

[0262] The second acquisition unit 1102 is used to acquire the first preliminary enhancement feature corresponding to the target feature of the first sample and the second preliminary enhancement feature corresponding to the target feature of the second sample.

[0263] The third acquisition unit 1104 is used to acquire the first masking probability vector corresponding to the first preliminary enhancement feature and the second masking probability vector corresponding to the second preliminary enhancement feature, wherein the first masking probability vector is used to represent the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector is used to represent the probability that the second preliminary enhancement feature is masked or enhanced.

[0264] The masking enhancement unit 1106 is used to mask and enhance the first preliminary enhancement feature using the first masking probability vector to obtain the first masking enhancement feature, and to mask and enhance the second preliminary enhancement feature using the second masking probability vector to obtain the second masking enhancement feature;

[0265] The fourth acquisition unit 1108 is used to acquire the first positive sample pair and the first negative sample pair corresponding to the first sample target feature, and the second positive sample pair and the second negative sample pair corresponding to the second sample target feature, wherein the first positive sample pair includes the first sample target feature and the first masking enhancement feature, the first negative sample pair includes the second sample target feature and the second masking enhancement feature, the second positive sample pair includes the second sample target feature and the second masking enhancement feature, and the second negative sample pair includes the first sample target feature and the first masking enhancement feature;

[0266] Training unit 1110 is used to train the initial lifetime value model based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair to obtain a trained lifetime value model, wherein the lifetime value model is a neural network model used to predict the lifetime value of users.

[0267] For specific implementation examples, please refer to the examples shown in the training method of the above life cycle value model, which will not be repeated here.

[0268] As an optional solution, the second acquisition unit 1102 includes: a first mapping module, used to map the target features of the first sample to a low-dimensional hidden space to obtain a first low-dimensional embedding, and to map the target features of the second sample to a low-dimensional hidden space to obtain a second low-dimensional embedding; and a second acquisition module, used to acquire a first enhanced feature corresponding to the first low-dimensional embedding and a second enhanced feature corresponding to the second low-dimensional embedding, wherein the first preliminary enhanced feature includes the first enhanced feature, and the second preliminary enhanced feature includes the second enhanced feature.

[0269] For specific implementation examples, please refer to the examples shown in the training method of the above life cycle value model, which will not be repeated here.

[0270] As an optional approach, the second acquisition module includes: a first nonlinear submodule, used to perform a nonlinear transformation on the first low-dimensional embedding through a multilayer perceptron to obtain a first enhanced feature; and a second nonlinear submodule, used to perform a nonlinear transformation on the second low-dimensional embedding through a multilayer perceptron to obtain a second enhanced feature.

[0271] For specific implementation examples, please refer to the examples shown in the training method of the above life cycle value model, which will not be repeated here.

[0272] As an optional solution, the third acquisition unit 1104 includes: an adding module, used to add corresponding noise to each feature element in the first preliminary enhancement feature, wherein the noise is sampled from an extreme value distribution; and an input module, used to input the first preliminary enhancement feature after adding noise into a normalized exponential function to obtain a probability distribution vector, wherein the probability distribution vector is used to represent the probability that each feature element in the first preliminary enhancement feature after adding noise is masked or enhanced, and the first masking probability vector includes the probability distribution vector.

[0273] For specific implementation examples, please refer to the examples shown in the training method of the above life cycle value model, which will not be repeated here.

[0274] As an optional approach, the masking enhancement unit 1106 includes: an embedding module for embedding the probability distribution vector and the first low dimension into the input masking enhancement function to obtain masking enhancement hidden space features; and a transformation module for transforming the hidden space features into masking enhancement features of the original dimension, wherein the original dimension is the dimension of the first sample target feature, and the first masking enhancement feature includes the masking enhancement feature of the original dimension.

[0275] For specific implementation examples, please refer to the examples shown in the training method of the above life cycle value model, which will not be repeated here.

[0276] As an optional approach, the training unit 1110 includes: a pre-training module for pre-training an initial lifecycle value model based on a first positive sample pair, a first negative sample pair, a second positive sample pair, and a second negative sample pair to obtain a first lifecycle value model; and a fine-tuning module for fine-tuning the first lifecycle value model using sample data from the target domain to obtain a second lifecycle value model, wherein the trained lifecycle value model includes the second lifecycle value model.

[0277] For specific implementation examples, please refer to the examples shown in the training method of the above life cycle value model, which will not be repeated here.

[0278] As an optional approach, the pre-training module includes: a fourth acquisition submodule for acquiring current positive and negative sample pairs, wherein the first positive sample pair and the first negative sample pair are positive and negative sample pairs, and the second positive sample pair and the second negative sample pair are positive and negative sample pairs; a fifth acquisition submodule for acquiring the first similarity between positive sample pairs in the current positive and negative sample pairs, and the second similarity between negative sample pairs in the current positive and negative sample pairs; and a sixth acquisition submodule for acquiring the current contrastive loss of the current initial lifecycle value model based on the first and second similarities and using the current contrastive loss function, wherein the contrastive loss is used to measure lifecycle value. The model's sensitivity in distinguishing positive and negative samples; the determination submodule, used to determine the current initial lifecycle value model as the first lifecycle value model when the current contrastive loss meets the convergence condition of the pre-training stage; the adjustment submodule, used to adjust the temperature parameter in the current contrastive loss function when the current contrastive loss does not meet the convergence condition of the pre-training stage, where the temperature parameter is used to control the sensitivity of the lifecycle value model in distinguishing positive and negative samples; the seventh acquisition submodule, used to acquire the next positive and negative sample pair and use the next positive and negative sample pair as the current positive and negative sample pair, until the first lifecycle value model is obtained.

[0279] For specific implementation examples, please refer to the examples shown in the training method of the above life cycle value model, which will not be repeated here.

[0280] As an optional approach, training unit 1110 includes: a training module for training an initial lifetime value model by combining cross-entropy loss and log-normal loss, wherein the cross-entropy loss is used to determine whether a user is a high-value user or a low-value user, and the log-normal loss is used to make the predicted value of the lifetime value model conform to the log-normal distribution parameters of the actual value.

[0281] For specific implementation examples, please refer to the examples shown in the training method of the above life cycle value model, which will not be repeated here.

[0282] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described method for predicting user lifetime value is also provided. This electronic device may, but is not limited to, […]. Figure 1 The user equipment 102 or server 112 shown in the figure, in this embodiment, is taken as an example of an electronic device, namely user equipment 102. Further, as shown in the figure... Figure 12 As shown, the electronic device includes a memory 1202 and a processor 1204. The memory 1202 stores a computer program, and the processor 1204 is configured to execute the steps of any of the above method embodiments through the computer program.

[0283] In an optional embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.

[0284] In an optional embodiment, the processor described above may be configured to perform the following steps via a computer program:

[0285] S1, Obtain the target features corresponding to the user to be predicted;

[0286] S2, input the target features into the trained lifetime value model to obtain the prediction result corresponding to the lifetime value of the user to be predicted;

[0287] The Lifetime Value Model is a neural network model used to predict user lifetime value. The trained Lifetime Value Model is the initial model, obtained after training on a first positive sample pair, a first negative sample pair, a second positive sample pair, and a second negative sample pair. The first positive sample pair includes a first target feature and a first masking enhancement feature; the first negative sample pair includes a second target feature and a second masking enhancement feature; the second positive sample pair includes a second target feature and a second masking enhancement feature; and the second negative sample pair includes a first target feature and a first masking enhancement feature. The first masking enhancement feature is obtained by masking and enhancing the first preliminary enhancement feature corresponding to the first target feature using a first masking probability vector; the second masking enhancement feature is obtained by masking and enhancing the second preliminary enhancement feature corresponding to the second target feature using a second masking probability vector. The first masking probability vector corresponds to the first preliminary enhancement feature, and the second masking probability vector corresponds to the second preliminary enhancement feature. The first masking probability vector represents the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector represents the probability that the second preliminary enhancement feature is masked or enhanced.

[0288] S1, obtain the first preliminary enhancement feature corresponding to the target feature of the first sample, and the second preliminary enhancement feature corresponding to the target feature of the second sample;

[0289] S2, obtain the first masking probability vector corresponding to the first preliminary enhancement feature and the second masking probability vector corresponding to the second preliminary enhancement feature, wherein the first masking probability vector is used to represent the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector is used to represent the probability that the second preliminary enhancement feature is masked or enhanced.

[0290] S3, using the first masking probability vector to mask and enhance the first preliminary enhancement feature to obtain the first masked enhancement feature, and using the second masking probability vector to mask and enhance the second preliminary enhancement feature to obtain the second masked enhancement feature;

[0291] S4, obtain the first positive sample pair and the first negative sample pair corresponding to the first sample target feature, and the second positive sample pair and the second negative sample pair corresponding to the second sample target feature, wherein the first positive sample pair includes the first sample target feature and the first masking enhancement feature, the first negative sample pair includes the second sample target feature and the second masking enhancement feature, the second positive sample pair includes the second sample target feature and the second masking enhancement feature, and the second negative sample pair includes the first sample target feature and the first masking enhancement feature;

[0292] S5, based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair, train the initial lifetime value model to obtain a trained lifetime value model, wherein the lifetime value model is a neural network model used to predict the lifetime value of users.

[0293] Alternatively, as those skilled in the art will understand, Figure 12 The structure shown is for illustrative purposes only. Figure 12 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 12 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 12 The different configurations shown.

[0294] The memory 1202 can be used to store software programs and modules, such as the program instructions / modules corresponding to the user lifetime value prediction method and apparatus in this embodiment. The processor 1204 executes various functional applications and data processing by running the software programs and modules stored in the memory 1202, thereby realizing the aforementioned user lifetime value prediction method. The memory 1202 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1202 may further include memory remotely located relative to the processor 1204, and these remote memories can be connected to electronic devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1202 may be used, but is not limited to, to store information such as target features and prediction results. As an example, such as... Figure 12 As shown, the memory 1202 may include, but is not limited to, the first acquisition unit 1002 and prediction unit 1004, or the second acquisition unit 1102, third acquisition unit 1104, masking enhancement unit 1106, fourth acquisition unit 1108, and training unit 1110 (not shown in the figure) of the user lifetime value prediction device. Furthermore, it may include, but is not limited to, other module units in the user lifetime value prediction device, which will not be elaborated upon in this example.

[0295] Optionally, the transmission device 1206 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1206 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1206 is a radio frequency (RF) module, used for wireless communication with the Internet.

[0296] In addition, the aforementioned electronic device also includes: a display 1208 for displaying information such as the target features and prediction results; and a connection bus 1210 for connecting the various module components in the aforementioned electronic device.

[0297] In other embodiments, the aforementioned user equipment or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer network, and any form of computing device, such as a server, user equipment, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.

[0298] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.

[0299] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0300] It should be noted that the computer system of the electronic device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0301] A computer system includes a Central Processing Unit (CPU), which performs various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) or loaded from RAM. ROM also stores various programs and data required for system operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output interfaces (I / O interfaces) are also connected to the bus.

[0302] The following components are connected to the input / output interface: input sections including keyboards, mice, etc.; output sections including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage sections including hard drives; and communication sections including network interface cards such as LAN cards and modems. The communication section performs communication processing via a network such as the Internet. Drives are also connected to the input / output interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage section as required.

[0303] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions defined in the system of this application.

[0304] According to one aspect of this application, a computer-readable storage medium is provided, wherein a processor of a computer device reads computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.

[0305] In an optional embodiment, the computer-readable storage medium described above may be configured to store a computer program for performing the following steps:

[0306] S1, Obtain the target features corresponding to the user to be predicted;

[0307] S2, input the target features into the trained lifetime value model to obtain the prediction result corresponding to the lifetime value of the user to be predicted;

[0308] The Lifetime Value Model is a neural network model used to predict user lifetime value. The trained Lifetime Value Model is the initial model, obtained after training on a first positive sample pair, a first negative sample pair, a second positive sample pair, and a second negative sample pair. The first positive sample pair includes a first target feature and a first masking enhancement feature; the first negative sample pair includes a second target feature and a second masking enhancement feature; the second positive sample pair includes a second target feature and a second masking enhancement feature; and the second negative sample pair includes a first target feature and a first masking enhancement feature. The first masking enhancement feature is obtained by masking and enhancing the first preliminary enhancement feature corresponding to the first target feature using a first masking probability vector; the second masking enhancement feature is obtained by masking and enhancing the second preliminary enhancement feature corresponding to the second target feature using a second masking probability vector. The first masking probability vector corresponds to the first preliminary enhancement feature, and the second masking probability vector corresponds to the second preliminary enhancement feature. The first masking probability vector represents the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector represents the probability that the second preliminary enhancement feature is masked or enhanced.

[0309] S1, obtain the first preliminary enhancement feature corresponding to the target feature of the first sample, and the second preliminary enhancement feature corresponding to the target feature of the second sample;

[0310] S2, obtain the first masking probability vector corresponding to the first preliminary enhancement feature and the second masking probability vector corresponding to the second preliminary enhancement feature, wherein the first masking probability vector is used to represent the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector is used to represent the probability that the second preliminary enhancement feature is masked or enhanced.

[0311] S3, using the first masking probability vector to mask and enhance the first preliminary enhancement feature to obtain the first masked enhancement feature, and using the second masking probability vector to mask and enhance the second preliminary enhancement feature to obtain the second masked enhancement feature;

[0312] S4, obtain the first positive sample pair and the first negative sample pair corresponding to the first sample target feature, and the second positive sample pair and the second negative sample pair corresponding to the second sample target feature, wherein the first positive sample pair includes the first sample target feature and the first masking enhancement feature, the first negative sample pair includes the second sample target feature and the second masking enhancement feature, the second positive sample pair includes the second sample target feature and the second masking enhancement feature, and the second negative sample pair includes the first sample target feature and the first masking enhancement feature;

[0313] S5, based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair, train the initial lifetime value model to obtain a trained lifetime value model, wherein the lifetime value model is a neural network model used to predict the lifetime value of users.

[0314] Optionally, in embodiments of this application, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0315] In optional embodiments, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware of an electronic device. The program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0316] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0317] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0318] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0319] In the several embodiments provided in this application, it should be understood that the disclosed user equipment can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0320] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0321] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0322] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for predicting a user lifetime value, characterized by, include: Obtain the target features corresponding to the user to be predicted; The target features are input into the trained lifetime value model to obtain the prediction result corresponding to the lifetime value of the user to be predicted. The lifetime value model is a neural network model used to predict user lifetime value. The trained lifetime value model is an initial lifetime value model, obtained after training based on a first positive sample pair, a first negative sample pair, a second positive sample pair, and a second negative sample pair. The first positive sample pair includes a first sample target feature and a first masking enhancement feature. The first negative sample pair includes a second sample target feature and a second masking enhancement feature. The second positive sample pair includes a second sample target feature and a second masking enhancement feature. The second negative sample pair includes the first sample target feature and the first masking enhancement feature. The first masking enhancement feature is obtained by masking and enhancing a first preliminary enhancement feature corresponding to the first sample target feature using a first masking probability vector. The second masking enhancement feature is obtained by masking and enhancing a second preliminary enhancement feature corresponding to the second sample target feature using a second masking probability vector. The first masking probability vector corresponds to the first preliminary enhancement feature, and the second masking probability vector corresponds to the second preliminary enhancement feature. The first masking probability vector represents the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector represents the probability that the second preliminary enhancement feature is masked or enhanced.

2. The method of claim 1, wherein, The step of obtaining the target features corresponding to the user to be predicted includes: Obtain the basic user features corresponding to the user to be predicted, and obtain the target behavior features corresponding to the user to be predicted, wherein the basic user features are used to represent the basic user information of the user to be predicted, and the target behavior features are used to represent the historical behavior information of the user to be predicted; The user's basic characteristics and the target behavior characteristics are integrated to obtain the target characteristics.

3. The method of claim 2, wherein, When the user to be predicted is a game user, and the lifetime value model is a neural network model used to predict the lifetime value of the game user, obtaining the target behavioral features corresponding to the user to be predicted includes at least one of the following: Obtain the first behavioral features of the user to be predicted in each game application, wherein the first behavioral features are used to represent statistical information of the historical behavior of the user to be predicted in each game application, and the target behavioral features include the first behavioral features. Obtain a second behavioral feature of the user to be predicted in a specific game application, wherein the second behavioral feature is used to represent the behavioral information generated by the user to be predicted in the specific game application, and the target behavioral feature includes the second behavioral feature; Obtain the third behavioral feature of the user to be predicted within a recent preset time period, wherein the third behavioral feature is used to represent the time sequence information of the user's game behavior within the recent preset time period, and the target behavioral feature includes the third behavioral feature.

4. The method of claim 3, wherein, The acquisition of the third behavioral characteristics of the user to be predicted within a recent preset time period includes at least one of the following: Obtain the first time series features of the user to be predicted within the recent preset time period, wherein the first time series features are used to represent the game registration sequence of the user to be predicted within the recent preset time period, and the third behavioral features include the first time series features; Obtain the second time series features of the user to be predicted within the recent preset time period, wherein the second time series features are used to represent the game activity sequence of the user to be predicted within the recent preset time period, and the third behavioral features include the second time series features; Obtain the third time series features of the user to be predicted within the recent preset time period, wherein the third time series features are used to represent the game activity duration sequence of the user to be predicted within the recent preset time period, and the third behavioral features include the third time series features; Obtain the fourth time series feature of the user to be predicted within the recent preset time period, wherein the fourth time series feature is used to represent the game payment sequence of the user to be predicted within the recent preset time period, and the third behavioral feature includes the fourth time series feature; Obtain the fifth time series feature of the user to be predicted within the recent preset time period, wherein the fifth time series feature is used to represent the game payment amount sequence of the user to be predicted within the recent preset time period, and the third behavioral feature includes the fifth time series feature.

5. A method of training a life cycle value model, characterized by, include: Obtain the first preliminary enhancement feature corresponding to the target feature of the first sample, and the second preliminary enhancement feature corresponding to the target feature of the second sample; Obtain a first masking probability vector corresponding to the first preliminary enhancement feature and a second masking probability vector corresponding to the second preliminary enhancement feature, wherein the first masking probability vector is used to represent the probability that the first preliminary enhancement feature is masked or enhanced, and the second masking probability vector is used to represent the probability that the second preliminary enhancement feature is masked or enhanced. The first preliminary enhancement feature is masked and enhanced using the first masking probability vector to obtain the first masked enhancement feature, and the second preliminary enhancement feature is masked and enhanced using the second masking probability vector to obtain the second masked enhancement feature; Obtain the first positive sample pair and the first negative sample pair corresponding to the first sample target feature, and the second positive sample pair and the second negative sample pair corresponding to the second sample target feature, wherein the first positive sample pair includes the first sample target feature and the first masking enhancement feature, the first negative sample pair includes the second sample target feature and the second masking enhancement feature, the second positive sample pair includes the second sample target feature and the second masking enhancement feature, and the second negative sample pair includes the first sample target feature and the first masking enhancement feature; Based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair, the initial lifetime value model is trained to obtain a trained lifetime value model, wherein the lifetime value model is a neural network model used to predict the lifetime value of users.

6. The method of claim 5, wherein, The step of obtaining the first preliminary enhancement feature corresponding to the target feature of the first sample and the second preliminary enhancement feature corresponding to the target feature of the second sample includes: The first sample target features are mapped into a low-dimensional hidden space to obtain a first low-dimensional embedding, and the second sample target features are mapped into the low-dimensional hidden space to obtain a second low-dimensional embedding. Obtain a first enhancement feature corresponding to the first low-dimensional embedding and a second enhancement feature corresponding to the second low-dimensional embedding, wherein the first preliminary enhancement feature includes the first enhancement feature and the second preliminary enhancement feature includes the second enhancement feature.

7. The method of claim 6, wherein, The step of obtaining the first enhanced feature corresponding to the first low-dimensional embedding and the second enhanced feature corresponding to the second low-dimensional embedding includes: The first enhanced feature is obtained by performing a nonlinear transformation on the first low-dimensional embedding using a multilayer perceptron. The second enhanced feature is obtained by performing a nonlinear transformation on the second low-dimensional embedding using the multilayer perceptron.

8. The method of claim 6, wherein, The step of obtaining the first masking probability vector corresponding to the first preliminary enhanced feature includes: Add corresponding noise to each feature element in the first preliminary enhanced feature, wherein the noise is sampled from an extreme value distribution; The first preliminary enhanced feature after adding the noise is input into a normalized exponential function to obtain a probability distribution vector, wherein the probability distribution vector is used to represent the probability that each feature element in the first preliminary enhanced feature after adding the noise is masked or enhanced, and the first masking probability vector includes the probability distribution vector.

9. The method of claim 8, wherein, The step of masking and enhancing the first preliminary enhancement feature using the first masking probability vector to obtain the first masked enhancement feature includes: The probability distribution vector and the first low-dimensional embedding are used to input the masking enhancement function to obtain the masking enhancement hidden space features; The hidden space features are converted into masking enhancement features of the original dimension, wherein the original dimension is the dimension of the target feature of the first sample, and the first masking enhancement feature includes the masking enhancement feature of the original dimension.

10. The method of claim 5, wherein, In the case where the training phase of the lifecycle value model includes a pre-training phase and a fine-tuning training phase, the sample data in the pre-training phase is sample data from the source domain, and the sample data in the fine-tuning training phase is sample data from the target domain. The sample data from the source domain includes the first sample target feature and the second sample target feature. The step of training the initial lifecycle value model based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair to obtain the trained lifecycle value model includes: Based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair, the initial life cycle value model is pre-trained to obtain the first life cycle value model; Using the sample data from the target domain, the first life cycle value model is fine-tuned to obtain a second life cycle value model, wherein the trained life cycle value model includes the second life cycle value model.

11. The method of claim 10, wherein, The step of pre-training the initial lifecycle value model based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair to obtain the first lifecycle value model includes: Perform the following steps until the first lifecycle value model is obtained: Obtain the current positive and negative sample pairs, wherein the first positive sample pair and the first negative sample pair belong to the positive and negative sample pairs, and the second positive sample pair and the second negative sample pair belong to the positive and negative sample pairs; Obtain the first similarity between positive sample pairs in the current positive-negative sample pair, and the second similarity between negative sample pairs in the current positive-negative sample pair; Based on the first similarity and the second similarity, the current contrastive loss of the current initial life cycle value model is obtained using the current contrastive loss function, wherein the contrastive loss is used to measure the sensitivity of the life cycle value model in distinguishing between positive and negative samples; If the current contrastive loss satisfies the convergence condition of the pre-training phase, the current initial lifecycle value model is determined as the first lifecycle value model. If the current contrastive loss does not meet the convergence condition of the pre-training stage, the temperature parameter in the current contrastive loss function is adjusted, wherein the temperature parameter is used to control the sensitivity of the lifetime value model in distinguishing between positive and negative samples. Obtain the next positive and negative sample pair, and use the next positive and negative sample pair as the current positive and negative sample pair, until the first life cycle value model is obtained.

12. The method according to any one of claims 5 to 11, characterized in that, In the process of training the initial lifecycle value model based on the first positive sample pair, the first negative sample pair, the second positive sample pair, and the second negative sample pair to obtain the trained lifecycle value model, the method further includes: The initial lifetime value model is trained by combining cross-entropy loss and log-normal loss. The cross-entropy loss is used to determine whether a user is a high-value user or a low-value user, and the log-normal loss is used to make the predicted value of the lifetime value model conform to the log-normal distribution parameters of the actual value.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed by an electronic device, performs the method according to any one of claims 1 to 4, or 5 to 12.

14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, they implement the steps of the method described in any one of claims 1 to 4, or 5 to 12.

15. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 4, or 5 to 12, through the computer program.