Deep learning user value estimation method and device, medium and equipment
Through deep learning methods combined with the A-GNN model, policy network and time series analysis model, the problem that the mean statistical algorithm cannot accurately estimate user value, realize the accurate estimate of user value and the accurate decision-making of advertising delivery, and improve the advertising delivery effect and monetization income.
Patent Information
- Application Number
- CN202510598784.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-19
Smart Images

Figure CN120509938A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of advertising delivery technology, and in particular to a deep learning user value estimation method, device, medium and equipment. Background Art
[0002] In today's booming digital advertising landscape, advertising placement and monetization have become important revenue streams for many companies. The value of users, the audience for advertising, directly determines the monetization revenue generated by advertising. Therefore, accurately and real-timely estimating user value is crucial in the advertising and monetization process. Specifically, it's necessary to estimate the monetization revenue from a single user's ad placement into an internal product. This revenue encompasses the cumulative monetization revenue generated by multiple clicks across different ad scenarios and code locations within the product. Accurate user value estimation helps companies optimize advertising strategies, improve the targeting and effectiveness of advertising, and ultimately achieve efficient use of advertising resources and maximize monetization revenue.
[0003] The mean statistical algorithm was a common method for estimating user value in the early days. This method collects user monetization revenue data over a certain period of time, calculates its average, and uses this average to estimate future user value. For example, the total monetization revenue of all users who joined an internal product over the past month is calculated, then divided by the number of users to obtain the average monetization revenue. This average is then used as a reference for estimating the value of subsequent new users.
[0004] The mean-based statistical algorithm is overly simplistic and crude, ignoring the diversity and complexity of user behavior. Different users have vastly different interests, spending power, and behavioral habits. Using the mean to generate estimates fails to reflect these individual differences, leading to significant discrepancies between the estimated results and the actual situation. For example, the mean-based statistical algorithm will produce the same estimate for both high-value and low-value users, failing to accurately distinguish the value of different users and, therefore, failing to provide accurate decision-making for advertising. Summary of the Invention
[0005] In view of this, the present invention provides a deep learning user value estimation method, device, medium and equipment, which can accurately distinguish the value of different users and provide accurate decision-making basis for advertising delivery.
[0006] In a first aspect, an embodiment of the present invention provides a method for estimating user value through deep learning, the method comprising:
[0007] Collecting user behavior and environment data in advertising business scenarios and preprocessing the user behavior and environment data;
[0008] The pre-processed user behavior and environment data are used to train the pre-built A-GNN model, the reinforcement learning-based policy network, and the time series analysis model, respectively, to obtain the trained A-GNN model, the policy network, and the time series analysis model;
[0009] The user behavior and environmental data to be estimated are respectively input into the trained A-GNN model, policy network and time series analysis model to calculate the overall monetization benefits after the user enters the internal product, thereby realizing real-time estimation of user value.
[0010] Furthermore, the user behavior and environment data includes at least one of the following data: user basic information, historical behavior data, advertising feature data, and advertising scenario data.
[0011] Furthermore, the user behavior and environment data to be estimated are input into the trained A-GNN model, policy network, and time series analysis model to calculate the overall monetization benefits after the user enters the internal product, achieving real-time estimation of user value, including:
[0012] Preprocess the estimated user behavior and environment data;
[0013] The pre-processed user behavior and environment data are input into the trained A-GNN model, policy network and time series analysis model respectively;
[0014] The policy network is used to predict the possibility of a user clicking on a specific advertisement. If the prediction is a possible click, the time series analysis model is used to estimate the value of the click. Combined with the learning results of the A-GNN model on the user-advertising-scenario relationship, the overall monetization benefits after the user enters the internal product are comprehensively calculated to complete the real-time estimation of user value.
[0015] Furthermore, the method further comprises:
[0016] The possibility of a user clicking on a specific advertisement is predicted by the policy network. If the prediction is that the user is unlikely to click, other advertisements with a higher relevance to the current user's interest tags or browsing history are selected from the advertisement library for display.
[0017] Furthermore, the method further comprises:
[0018] The policy network is used to predict the likelihood of a user clicking on a specific ad. If the prediction is that the user is unlikely to click, the ad placement is changed to an area that is more likely to attract the user's attention based on the historical click-through rate data of different ad positions and the user's behavioral characteristics on the current page.
[0019] Furthermore, the method further comprises:
[0020] The policy network is used to predict the likelihood of a user clicking on a specific ad. If the prediction is that the user is unlikely to click, the real-time user behavior data within the product is further collected, and the user's interest tags are recalculated and calibrated in combination with the machine learning algorithm. Based on the new interest tags, ads that better match the user's interests are screened from the ad library.
[0021] Furthermore, the method further comprises:
[0022] Predicting the likelihood of a user clicking on a specific ad using the policy network; if the likelihood of a click is predicted to be unlikely, evaluating whether there are new, undeveloped advertising scenarios that may be suitable for the user based on the content of the page the user is currently on and the user's behavior patterns;
[0023] For potential new advertising scenarios that have been evaluated, advertising delivery tests are conducted among user groups with similar characteristics.
[0024] In a second aspect, an embodiment of the present invention provides a user value estimation device for deep learning, the device comprising:
[0025] A collection module, used to collect user behavior and environment data in advertising business scenarios and pre-process the user behavior and environment data;
[0026] The construction module is used to train the pre-built A-GNN model, the reinforcement learning-based policy network, and the time series analysis model using the pre-processed user behavior and environment data to obtain the trained A-GNN model, policy network, and time series analysis model;
[0027] The analysis module is used to input the user behavior and environmental data to be estimated into the trained A-GNN model, policy network and time series analysis model respectively to calculate the overall monetization benefits after the user enters the internal product, thereby realizing real-time estimation of user value.
[0028] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program is configured to execute any one of the methods in the first aspect when run.
[0029] In a fourth aspect, an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute any one of the methods described in the first aspect.
[0030] The technical solution provided by the present invention first collects user behavior and environmental data (such as user basic information, historical behavior data, etc.) in the advertising business scenario and pre-processes it, and then uses the pre-processed data to train the pre-built A-GNN model, the policy network based on reinforcement learning and the time series analysis model. The user data to be estimated is then input into the trained model, and the policy network is first used to predict the possibility of the user clicking on a specific advertisement. If the user is likely to click, the overall monetization revenue is calculated by combining the results of the time series analysis model and the A-GNN model; if the user is unlikely to click, there are a variety of response strategies, such as changing the advertising material, adjusting the advertising position, recalibrating the user interest tags, exploring new advertising scenarios and testing, etc. By collecting multiple types of user behavior and environmental data, such as basic user information, historical behavior data, ad feature data, and ad scenario data, the algorithm comprehensively considers factors such as individual user characteristics, behavioral habits, ad characteristics, and the context in which they are used. This allows for a more comprehensive and detailed characterization of the interactive relationship between users and ads, thereby achieving more accurate user value estimation. Furthermore, a reinforcement learning-based policy network can dynamically optimize ad display strategies based on real-time user feedback. When it predicts that a user is unlikely to click on a particular ad, it can adopt various strategies, such as replacing ad creatives and adjusting ad placement. It can also recalibrate user interest tags and explore new ad scenarios. This algorithm has strong dynamic adaptability and can promptly respond to changes in different users and scenarios, improving the effectiveness of ad placement and the accuracy of user value estimation. Furthermore, the embodiments of the present invention utilize an A-GNN model, a reinforcement learning-based policy network, and a time series analysis model. The A-GNN model can mine the user-advertising-scenario relationship, while the time series analysis model can capture the temporal changes in user click behavior. These models can deeply mine complex information and relationships in the data, assessing user value from multiple dimensions, with higher accuracy and reliability than mean statistical algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 This is a flowchart of a method for estimating user value through deep learning provided by an embodiment of the present invention.
[0032] Figure 2 This is a flowchart of a method for estimating user value through deep learning disclosed in another embodiment of the present invention.
[0033] Figure 3 This is a flowchart of a method for estimating user value through deep learning disclosed in another embodiment of the present invention.
[0034] Figure 4 This is a flowchart of a method for estimating user value through deep learning disclosed in another embodiment of the present invention.
[0035] Figure 5This is a flowchart of a method for estimating user value through deep learning disclosed in another embodiment of the present invention.
[0036] Figure 6 2 is a schematic diagram of the structure of a user value estimation device based on deep learning provided by an embodiment of the present invention.
[0037] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0038] To make the objectives, technical solutions, and advantages of the present invention more apparent, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the embodiments described herein are merely some, rather than all, of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0039] See also Figure 1 , Figure 1 : is a flowchart of a method for estimating user value through deep learning provided by an embodiment of the present invention, the method comprising the following steps:
[0040] Step 11: Collect user behavior and environment data in the advertising business scenario, and pre-process the user behavior and environment data.
[0041] In this step, user behavior data refers to the various behavioral information generated by users in advertising-related activities, such as their historical ad click history, browsing time, purchase behavior, sharing operations, etc. Environmental data refers to the various external environmental information when the ad is delivered, including the time and location of the ad, the type of page it is on, and the current network environment.
[0042] Preprocessing is the process of cleaning, transforming, and integrating raw data to make it suitable for subsequent model training. Common preprocessing operations include data cleaning (removing duplicate, erroneous, or missing data), data normalization (consistently fitting data into a specific range), and data encoding (converting non-numeric data into numeric data).
[0043] In this step, data is collected through various channels, such as recording users' clicks and browsing behaviors on the advertising platform, obtaining basic information from user registration information, and using the log system to record environmental information such as the time and location of advertising. Check whether there are duplicate records, erroneous values, or missing values in the data. Duplicate records can be deleted directly; erroneous values can be corrected or deleted according to business rules; missing values can be processed by filling (such as mean filling, median filling) or deleting related records. Normalize the data, for example, convert numerical data such as the user's age and income to the range of [0,1] to eliminate the dimensional influence between different data features. For non-numeric data, such as the user's gender, occupation, etc., use encoding to convert them into numerical data, such as using one-hot encoding.
[0044] This step provides a high-quality data foundation for subsequent model training. By collecting comprehensive user behavior and environmental data, the model can learn more about users and advertising scenarios, thereby improving model accuracy. Preprocessing operations can improve data quality and usability, avoiding poor model training results due to data issues.
[0045] Step 12: Use the pre-processed user behavior and environment data to train the pre-built A-GNN model, reinforcement learning-based policy network, and time series analysis model respectively to obtain the trained A-GNN model, policy network, and time series analysis model.
[0046] In this step, the A-GNN model incorporates a graph neural network model with an attention mechanism. Graph neural networks are designed to process graph-structured data. In this scenario, they construct a user-advertising-scenario relationship graph, where nodes represent users, ads, and ad scenarios, and edges represent the interactions between them. The attention mechanism allows the model to automatically focus on the information most critical to user value estimation during learning.
[0047] Reinforcement Learning-Based Policy Network: Reinforcement learning is a machine learning method that uses reward signals from an intelligent agent interacting with its environment to learn optimal policies. The policy network is a component of reinforcement learning. It outputs an action policy, namely, whether to display a specific ad to the user, based on historical user behavior data, current ad features, and ad context.
[0048] Time series analysis model: used to process data with a time sequence. In this scenario, it can capture the temporal sequence characteristics of user click behavior and the differences in the contribution of clicks at different time points to the final monetization value.
[0049] In this step, the preprocessed user behavior and environment data is divided into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to evaluate the model's performance during training and adjust the model's hyperparameters, and the test set is used to ultimately evaluate the model's generalization ability.
[0050] After that, model training is carried out. The A-GNN model training is as follows: the constructed user-advertising-scenario association graph and related data are input into the A-GNN model, and the model parameters are continuously adjusted through the back-propagation algorithm to minimize the model's estimation error on the training data.
[0051] Reinforcement learning-based policy network: In an advertising simulation environment, ads are displayed according to the action strategy output by the policy network. Based on the user's actual click feedback, reinforcement learning algorithms (such as Q-learning and deep deterministic policy gradient) are used to update the policy network parameters to continuously optimize the decision of whether to display ads.
[0052] Time series analysis model: User click data containing time series information is input into the time series analysis model. By optimizing the model parameters, the model can accurately learn the trends and periodic characteristics of user click behavior.
[0053] The primary purpose of this step is to train a model that can accurately estimate user value. By training different types of models, we can fully utilize the diverse information in the data and model and learn user value from different perspectives. The trained model will be used for subsequent real-time user value estimation.
[0054] Step 13: Input the user behavior and environment data to be estimated into the trained A-GNN model, policy network, and time series analysis model respectively to calculate the overall monetization benefits after the user enters the internal product, thereby achieving real-time estimation of user value.
[0055] In this step, the user behavior and environment data to be estimated refers to the behavior and environment-related data of new users for which user value estimation is required. Its data type is the same as the training data, including user basic information, current behavior records, advertising features, and advertising scenario information.
[0056] Overall monetization revenue refers to the cumulative value of monetization revenue generated by multiple clicks on different code positions in different advertising scenarios after a user enters an internal product through a single advertisement.
[0057] In this step, the estimated user behavior and environment data are subjected to the same preprocessing operations as in step 11 to ensure that the format and range of the data are consistent with the training data. This will not be repeated here.
[0058] The preprocessed data to be estimated is then fed into the trained A-GNN model, policy network, and time series analysis model. The policy network first predicts the likelihood of a user clicking on a specific ad. If the prediction is positive, the time series analysis model estimates the value of that click. Combined with the A-GNN model's learning of the user-ad-scenario relationship, the overall monetization revenue after the user enters the internal product is calculated.
[0059] This step enables real-time estimation of user value. By inputting new user data into the trained model, the user's overall monetization revenue can be quickly and accurately calculated, providing an important basis for advertising and monetization decisions, helping companies optimize advertising strategies and improve advertising effectiveness and monetization revenue.
[0060] In this step, the policy network predicts the likelihood of a user clicking on a specific ad by performing the following steps:
[0061] ① Data Input: The user's behavior and environmental data (such as basic user information, historical click behavior, current ad features, and ad delivery scenarios) to be estimated are pre-processed and then fed into the trained policy network. This data is converted into numerical form that the policy network can process, for example, through encoding and normalization, so that different types of data meet the network's input requirements.
[0062] ② Model Operation: The policy network extracts and analyzes features from the input data based on pre-trained parameters and patterns. Through multi-layer neural network calculations, it mines the data for hidden associations between users and ads, such as the correlation between the types of ads a user has previously clicked on and the type of the current ad, and the user's preferences for different ads in specific scenarios.
[0063] ③ Result Output: Ultimately, the policy network outputs a probability value between 0 and 1, indicating the likelihood that the user will click on the ad. For example, an output of 0.7 means the model believes there is a 70% chance that the user will click on the ad. Users can set a threshold (such as 0.5) based on actual business needs. When the output probability value is greater than the threshold, the user is considered likely to click; when it is less than the threshold, the user is considered unlikely to click.
[0064] In this step, if the prediction is a possible click, the value of the click can be estimated using the time series analysis model by following the steps below:
[0065] ① Data screening and preparation: After the policy network determines that a user is likely to click on an ad, it filters out time series data related to the user's click behavior, including information such as the user's historical click time, click frequency, and the amount spent after each click. This data is organized chronologically into a format that can be processed by the time series analysis model, for example, by converting it into sequence data with fixed time intervals.
[0066] ② Model Analysis and Prediction: Input the organized time series data into a trained time series analysis model. The model leverages its internal algorithms and structures (such as recurrent neural networks and long short-term memory networks) to learn the patterns and trends in user click behavior over time, such as user click habits at different times of the day and the changing value of clicks over time. Based on these learned patterns, the model estimates the value of the current click and outputs a numerical value representing the expected monetization value of that click.
[0067] In this step, the overall monetization revenue after users enter the internal product is calculated by combining the A-GNN model's learning results on the user-advertising-scenario relationship. This can be achieved through the following steps:
[0068] ① A-GNN model calculation: Information related to users, ads, and ad scenarios is constructed into a graph-structured data structure, with users, ads, and scenarios as nodes and their interactions (e.g., whether a user clicked on an ad, or whether an ad was displayed in a specific scenario) as edges. This graph-structured data is input into the A-GNN model, which uses operations such as graph convolution to learn the associations and feature transfer between nodes, thereby mining the complex relationships between users, ads, and scenarios. For example, the model may discover that certain users are more likely to generate high-value clicks on certain types of ads in certain scenarios.
[0069] ② Comprehensive Calculation: The click value estimated by the time series analysis model is integrated with the associations learned by the A-GNN model. Weighted summation and fusion calculations can be used to consider the impact of different factors on overall monetization revenue, ultimately calculating the overall monetization revenue after the user enters the internal product. For example, based on the associations learned by the A-GNN model, different weights are assigned to different types of click value. This is then combined with the results of the time series analysis model to produce a more accurate estimate of overall monetization revenue, thereby providing real-time estimates of user value.
[0070] See Figure 2 , Figure 2 is a flowchart of a method for estimating user value through deep learning disclosed in another embodiment of the present invention, the method comprising:
[0071] Step 21: Collect user behavior and environment data in an advertising business scenario, and pre-process the user behavior and environment data.
[0072] Step 22: Use the pre-processed user behavior and environment data to train the pre-built A-GNN model, the reinforcement learning-based policy network, and the time series analysis model to obtain the trained A-GNN model, the policy network, and the time series analysis model.
[0073] Step 23: Predict the possibility of the user clicking on the specific advertisement through the policy network. If the prediction is that the user is likely to click, execute step 24; if the prediction is that the user is unlikely to click, execute step 25.
[0074] Step 24: Use the time series analysis model to estimate the value of the click, combine the learning results of the A-GNN model on the user-advertising-scenario relationship, and comprehensively calculate the overall monetization benefits after the user enters the internal product to complete the real-time estimation of user value.
[0075] Step 25: Select other advertising materials that are more relevant to the current user's interest tags or browsing history from the advertising material library for display.
[0076] In this embodiment, step 23 predicts the possibility of a user clicking on a specific advertisement through a policy network. If the prediction is that a click is possible, the time series analysis model is used to estimate the value of the click. Combined with the learning results of the A-GNN model on the user-advertising-scenario association relationship, the overall monetization revenue after the user enters the internal product is comprehensively calculated to complete the real-time estimation of the user value. If the prediction is that a click is unlikely, other advertising materials that are more relevant to the current user's interest tags or browsing history are selected from the advertising material library for display.
[0077] In this embodiment, step 25 is a response strategy adopted when the policy network predicts that the user is unlikely to click on a specific advertisement. Its core purpose is to improve the attractiveness and click-through rate of the advertisement by replacing the advertisement material with one that better suits the user's interests, thereby more accurately estimating the user value and achieving better advertising monetization results.
[0078] After the reinforcement learning-based policy network analyzes the input user behavior to be estimated and the environmental data, if the output probability of the user clicking on a specific advertisement is lower than a preset threshold (e.g., 0.5), and it is determined that the user is unlikely to click on the advertisement, the system will execute step 25.
[0079] First, the system extracts the current user's interest tags and browsing history information from the collected and pre-processed user behavior and environmental data. Interest tags may be automatically marked by the system during the user registration or use of the product, such as "technology", "sports", "food", etc.; browsing history records the advertisements, product pages, articles and other content that the user has browsed in the past. The advertising creative library is a collection that stores a variety of different types of advertising creatives. Each advertising creative has corresponding attribute tags. These tags describe the theme, content, target audience and other information of the advertisement. For example, the material of a mobile phone advertisement may have tags such as "technology", "smartphone", and "Apple".
[0080] Next, an appropriate algorithm is used to calculate the relevance of each creative in the creative library to the current user's interest tags and browsing history. Common methods include keyword matching and vector similarity calculation. For example, this can be done by counting the overlap between creative tags and user interest tags, or by converting creative and user history data into vector representations and then calculating the cosine similarity between these vectors.
[0081] Based on the calculated relevance score, the creatives with the highest relevance are selected from the creative library. You can set a relevance threshold to select only creatives with scores above that threshold, or you can sort creatives by relevance score and select a certain number of creatives that are ranked at the top.
[0082] Ad creatives that are most relevant to users' interest tags or browsing history are displayed to them. The display method can be selected based on the specific advertising business scenario, such as displaying in a specific ad space on a webpage or as pop-up ads in a mobile app.
[0083] This embodiment, by displaying advertising materials that better match user interests, can increase user attention and interest in ads, thereby improving ad click-through rates and providing more opportunities for subsequent user value estimation and ad monetization. This avoids displaying ads that users are not interested in, reduces user aversion and distraction, enhances the user experience in advertising business scenarios, and strengthens user favorability and loyalty to the platform. At the same time, more ad clicks can provide the model with richer data, helping the model better learn user behavior patterns and preferences, thereby improving the accuracy of user value estimation.
[0084] See Figure 3 , Figure 3 is a flowchart of a method for estimating user value through deep learning disclosed in another embodiment of the present invention, the method comprising:
[0085] Step 31: Collect user behavior and environment data in an advertising business scenario, and pre-process the user behavior and environment data.
[0086] Step 32: Use the pre-processed user behavior and environment data to train the pre-built A-GNN model, the reinforcement learning-based policy network, and the time series analysis model to obtain the trained A-GNN model, the policy network, and the time series analysis model.
[0087] Step 33: Predict the possibility of the user clicking on the specific advertisement through the policy network. If the prediction is that the user is likely to click, execute step 34; if the prediction is that the user is unlikely to click, execute step 35.
[0088] Step 34: Use the time series analysis model to estimate the value of the click, combine the learning results of the A-GNN model on the user-advertising-scenario relationship, and comprehensively calculate the overall monetization benefits after the user enters the internal product to complete the real-time estimation of user value.
[0089] Step 35: Based on the historical click-through rate data of different ad positions and the user's behavioral characteristics on the current page, the ad placement is changed to an area that is more likely to attract the user's attention.
[0090] In this embodiment, step 35 is an adjustment strategy taken to improve the attractiveness and click-through rate of an advertisement when the policy network predicts that the user is unlikely to click on a specific advertisement. The adjustment strategy mainly optimizes the advertisement display position by analyzing historical data and the user's current behavioral characteristics.
[0091] After the reinforcement learning-based policy network analyzes the input user behavior to be estimated and the environmental data, if it determines that the user is less likely to click on a specific advertisement (lower than a pre-set threshold, such as 0.5) and believes that the user is unlikely to click on the advertisement, the system will execute step 35.
[0092] First, the system collects click-through rate data for different ad placements over a period of time. This data records the click-through rates for each ad placement in various advertising scenarios, such as the click-through rates for different ad placements on different pages (such as the top, middle, and bottom of a webpage, the splash screen ads in mobile apps, and the carousel ads on the homepage). By analyzing this historical data, we can understand which ad placements have higher overall click-through rates and which have lower performance.
[0093] At the same time, the system monitors the user's behavioral characteristics on the current page in real time. This includes the user's browsing trajectory (such as the user's scrolling speed, pause position, browsing order, etc.), click behavior (whether other elements on the page are clicked, such as links and buttons), and dwell time (how long the user stays in different areas of the page). By analyzing these behavioral characteristics, it is possible to infer the user's focus and points of interest on the current page.
[0094] The attractiveness of each ad spot on the current page is evaluated by combining historical click-through rate data and the user's current page behavior. Ad spots with high historical click-through rates are given a higher weight in the evaluation; ad spots where users' current behavior indicates a focused attention span will also have their attractiveness scores increased accordingly. For example, if an ad spot not only has a high historical click-through rate, but also users currently spend a long time near it and frequently scroll, then this ad spot will have a relatively high attractiveness score.
[0095] Based on the evaluation results, ads originally displayed in areas less likely to attract user attention are moved to more attractive placements. For example, if users spend a lot of time in the middle of the current page, and the placements in that area have a good historical click-through rate, but the current ad is displayed at the bottom of the page and is predicted to be less likely to be clicked, the ad will be moved from the bottom to the middle of the page.
[0096] This embodiment can increase the exposure and attractiveness of advertisements by displaying advertisements in areas that are more likely to attract users' attention, thereby improving the click-through rate of advertisements. More click behaviors help to achieve better advertising monetization and improve the effectiveness and revenue of advertising. At the same time, a reasonable ad placement can reduce users' aversion to advertisements, because advertisements appear in areas that users may be interested in and will not cause too much interference with users' normal browsing. This can improve the user experience in the advertising business scenario and enhance user satisfaction and loyalty to the platform. This embodiment makes decisions based on the analysis of historical data and real-time user behavior, reflecting a data-driven optimization approach. By continuously collecting and analyzing data, the system can continuously improve the adjustment strategy of ad placement to make it more accurate and effective, thereby further improving the overall performance of the advertising business.
[0097] See Figure 4 , Figure 4 is a flowchart of a method for estimating user value through deep learning disclosed in another embodiment of the present invention, the method comprising:
[0098] Step 41: Collect user behavior and environment data in an advertising business scenario, and pre-process the user behavior and environment data.
[0099] Step 42: Use the pre-processed user behavior and environment data to train the pre-built A-GNN model, the reinforcement learning-based policy network, and the time series analysis model to obtain the trained A-GNN model, the policy network, and the time series analysis model.
[0100] Step 43: Predict the possibility of the user clicking on the specific advertisement through the policy network. If the prediction is that the user is likely to click, execute step 44; if the prediction is that the user is unlikely to click, execute step 45.
[0101] Step 44: Use the time series analysis model to estimate the value of the click, combine the learning results of the A-GNN model on the user-advertising-scenario relationship, and comprehensively calculate the overall monetization benefits after the user enters the internal product to complete the real-time estimation of user value.
[0102] Step 45: Collect the user's real-time behavior data within the product, combine it with the machine learning algorithm to recalculate and calibrate the user's interest tags, and based on the new interest tags, filter out advertisements that better match the user's interests from the ad library.
[0103] After the reinforcement learning-based policy network analyzes the input user behavior to be estimated and the environmental data, if it determines that the user is less likely to click on a specific advertisement (lower than a preset threshold, such as 0.5) and believes that the user is unlikely to click on the advertisement, the system will execute step 45.
[0104] First, the system monitors and collects various user behavior data within the product in real time. This data includes, but is not limited to, the pages users browse (such as product details pages and article information pages), their actions on the pages (such as clicks, scrolling, zooming in and out), their dwell time, and their interactions with other users (such as comments and sharing). By collecting this real-time behavioral data, we can more accurately understand users' current interests and needs.
[0105] Secondly, the collected real-time behavioral data is combined with the user's original behavioral data (such as historical browsing records, purchase records, etc.) and input into a pre-set machine learning algorithm. The machine learning algorithm will analyze and process this data to extract the user's interest characteristics. For example, a clustering algorithm is used to cluster the user's behavioral data, and the user's interest category is determined based on the clustering results; or a classification algorithm is used to classify the user's behavior to determine the user's interest in different types of content. Then, based on the analysis results of the algorithm, the user's interest tags are recalculated and calibrated, and the user's interest profile is updated. For example, the original user's interest tags were "sports" and "travel", but through real-time behavioral data, it was found that the user has recently paid more attention to "technology and digital" content, then "technology and digital" will be added to the user's interest tags, and the weights of each interest tag will be adjusted accordingly.
[0106] Again, the ad library stores a large number of ads of different types and themes, and each ad has a corresponding attribute tag that describes the ad's content, target audience, and other information. Based on the recalculated and calibrated user interest tags, the system will search and filter in the ad library. By matching user interest tags and ad attribute tags, ads that are more in line with user interests can be found. For example, if the user's new interest tags are "Technology and Digital," "Sports," and "Travel," the system will filter out ads with attribute tags such as "Technology and Digital Products," "Sports Equipment," and "Recommended Tourist Attractions" from the ad library. The screening process can use methods such as keyword matching and vector similarity calculation to determine the degree of match between ads and user interests, and ultimately select ads with higher matching degrees for display.
[0107] This embodiment adjusts the advertising content according to the user's real-time interests and displays advertisements that better match the user's interests, which can increase the user's attention and interest in the advertisement, thereby improving the click-through rate of the advertisement. When users see an advertisement that interests them, they are more likely to click on it and generate conversion behaviors such as purchases, thereby improving the monetization effect and revenue of the advertisement. More accurate user interest tags can provide more reliable input information for the user value estimation model. By timely updating user interests, the model can better learn the user's behavior patterns and preferences, thereby improving the accuracy of user value estimation and providing stronger support for advertising and business decision-making.
[0108] See Figure 5 , Figure 5 is a flowchart of a method for estimating user value through deep learning disclosed in another embodiment of the present invention, the method comprising:
[0109] Step 51: Collect user behavior and environment data in an advertising business scenario, and pre-process the user behavior and environment data.
[0110] Step 52: Use the pre-processed user behavior and environment data to train the pre-built A-GNN model, the reinforcement learning-based policy network, and the time series analysis model to obtain the trained A-GNN model, the policy network, and the time series analysis model.
[0111] Step 53: Predict the possibility of the user clicking on the specific advertisement through the policy network. If the prediction is that the user is likely to click, execute step 54; if the prediction is that the user is unlikely to click, execute step 55.
[0112] Step 54: Use the time series analysis model to estimate the value of the click, combine the learning results of the A-GNN model on the user-advertising-scenario relationship, and comprehensively calculate the overall monetization benefits after the user enters the internal product to complete the real-time estimation of user value.
[0113] Step 55: Predict the likelihood of the user clicking on a specific advertisement through the policy network. If the prediction is that the user is unlikely to click, evaluate whether there are new advertising scenarios that have not yet been developed but may be suitable for the user based on the content of the page the user is currently on and the user's behavior pattern.
[0114] When the reinforcement learning-based policy network analyzes the input user behavior and environmental data to be estimated, it determines that the user is less likely to click on a specific advertisement (lower than a pre-set threshold, such as 0.5) and believes that the user is unlikely to click on the advertisement, triggering the evaluation operation of the new advertising scenario.
[0115] First, the system conducts an in-depth analysis of the content of the page the user is currently browsing. This includes information such as the page's subject, the type of information contained, and the product or service description. For example, if a user is browsing a fitness equipment product page, the page might include detailed descriptions of various fitness equipment, instructions for use, and user reviews. By analyzing the page's content, the system can understand the user's current areas of interest and potential points of interest.
[0116] At the same time, the system will study the user's behavior patterns on the current page. This includes the user's browsing trajectory (such as the order of scrolling from the top to the bottom of the page, the length of time spent in different sections), click behavior (which links, buttons or images were clicked), and interactive behavior (whether comments, shares or added to favorites were made). For example, if a user spends a long time on a specific equipment introduction section on a fitness equipment page and clicks on the link to the detailed parameters of the equipment, this indicates that the user may have a high interest in the equipment.
[0117] Based on the analysis of page content and user behavior patterns, the system will try to evaluate whether there are new advertising scenarios that have not yet been developed but may be suitable for the user. This requires considering a variety of factors, such as whether there are some blank areas or available interactive elements on the page that can be used to display ads, and whether the advertising content displayed in these locations is related to the user's current interests and the theme of the page. For example, in the sidebar or related recommendation area of the fitness equipment page, can ads related to fitness courses, sports nutrition products, etc. be displayed? These ads are related to the page theme and may meet the user's potential needs. If such a potential area and a suitable combination of advertising content are found, it can be identified as a possible new advertising scenario.
[0118] Step 56: For the evaluated potential new advertising scenarios, conduct an advertising delivery test among user groups with similar characteristics.
[0119] After identifying potential new advertising scenarios, actual advertising testing is needed to verify the effectiveness and feasibility of the scenario and understand users' reactions to ads displayed in the new scenario in order to further optimize the advertising delivery strategy.
[0120] In this step, the system will filter out user groups with similar characteristics to the current user based on various user characteristics (such as age, gender, interests, hobbies, and historical behavior data). For example, if the current user is a young male fitness enthusiast, the system will find other users with similar age, gender, and fitness interests in the user database to form a user group with similar characteristics.
[0121] Ads prepared for potential new advertising scenarios are delivered to selected user groups with similar characteristics. During the delivery process, the system closely monitors and collects various user feedback data, including ad click-through rates, conversion rates (such as purchases, registrations, downloads, etc.), and subsequent changes in user behavior on the page (such as whether the ad led to increased browsing of other pages).
[0122] Based on the collected test data, analyze and evaluate the effectiveness of potential new advertising scenarios. If the test results show that the ads in the new scenario have high click-through rates and conversion rates, and that users respond positively to the ads, then the new advertising scenario can be considered effective and worthy of further promotion and application. Conversely, if the test results are unsatisfactory, the new advertising scenario needs to be adjusted and optimized, or re-evaluated to find other potential new advertising scenarios.
[0123] Through the operations of step 55 and step 56, it is possible to continuously explore and optimize the scenarios and strategies for advertising delivery, improve the targeting and effectiveness of advertising, and thus better realize the real-time estimation of user value and the realization of advertising business.
[0124] See also Figure 6 , Figure 6 : is a schematic diagram of the structure of a deep learning user value estimation device provided by an embodiment of the present invention, the device comprising:
[0125] A collection module 61 is used to collect user behavior and environment data in advertising business scenarios and pre-process the user behavior and environment data;
[0126] A construction module 62 is used to train the pre-built A-GNN model, the reinforcement learning-based policy network, and the time series analysis model using the pre-processed user behavior and environment data to obtain the trained A-GNN model, the policy network, and the time series analysis model;
[0127] The analysis module 63 is used to input the user behavior and environmental data to be estimated into the trained A-GNN model, policy network and time series analysis model respectively to calculate the overall monetization benefits after the user enters the internal product, thereby realizing real-time estimation of user value.
[0128] In some embodiments of the present invention, the user behavior and environment data includes at least one of the following data: user basic information, historical behavior data, advertising feature data, and advertising scenario data.
[0129] In some embodiments of the present invention, the analysis module 63 may include:
[0130] A preprocessing unit 631 is used to preprocess the user behavior and environment data to be estimated;
[0131] Input unit 632, used to input the pre-processed user behavior and environment data into the trained A-GNN model, policy network and time series analysis model respectively;
[0132] The estimation unit 633 is used to predict the possibility of a user clicking on a specific advertisement through the policy network. If the prediction is a possible click, the time series analysis model is used to estimate the value of the click. Combined with the learning results of the A-GNN model on the user-advertising-scenario association, the overall monetization revenue after the user enters the internal product is comprehensively calculated to complete the real-time estimation of user value.
[0133] In some embodiments of the present invention, the analysis module 63 may further include:
[0134] The first display unit 632 is used to select other advertising materials that are more relevant to the current user's interest tags or browsing history from the advertising material library for display.
[0135] In some embodiments of the present invention, the analysis module 63 may further include:
[0136] The display position changing unit 633 is used to change the advertisement display position to an area that is more likely to attract the user's attention based on the historical click-through rate data of different advertisement positions and the user's behavioral characteristics on the current page.
[0137] In some embodiments of the present invention, the analysis module 63 may further include:
[0138] The screening and matching unit 634 is used to collect real-time behavioral data of users within the product, recalculate and calibrate the user's interest tags in combination with machine learning algorithms, and screen out advertisements that better match the user's interests from the advertisement library based on the new interest tags.
[0139] In some embodiments of the present invention, the analysis module 63 may further include:
[0140] New advertising scenario evaluation unit 635, for evaluating whether there is a new advertising scenario that has not yet been developed but may be suitable for the user based on the content of the page the user is currently on and the user's behavior pattern;
[0141] The testing unit 636 is configured to conduct an advertisement delivery test among user groups with similar characteristics for the evaluated potential new advertisement scenarios.
[0142] It should be noted that the deep learning user value estimation device in the embodiment of the present invention and the deep learning user value estimation method in the above embodiment belong to the same inventive concept. The technical details not described in detail in this device can be found in the previous description of the method and will not be repeated here.
[0143] In addition, an embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the aforementioned method when running.
[0144] Figure 7 1 is a schematic diagram of the structure of an electronic device 10 provided by an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0145] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0146] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0147] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the idle detection method.
[0148] In some embodiments, the idle detection method can be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the idle detection method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the idle detection method in any other suitable manner (e.g., by means of firmware).
[0149] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0150] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0151] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0152] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0153] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0154] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0155] It should be understood that the various forms of the processes shown above can be used to re-order, add, or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0156] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A deep learning method for estimating user value, characterized in that: The method comprises: Collecting user behavior and environment data in advertising business scenarios and preprocessing the user behavior and environment data; The pre-processed user behavior and environment data are used to train the pre-built A-GNN model, the reinforcement learning-based policy network, and the time series analysis model, respectively, to obtain the trained A-GNN model, the policy network, and the time series analysis model; The user behavior and environmental data to be estimated are respectively input into the trained A-GNN model, policy network and time series analysis model to calculate the overall monetization benefits after the user enters the internal product, thereby realizing real-time estimation of user value.
2. The method according to claim 1, characterized in that The user behavior and environment data includes at least one of the following data: user basic information, historical behavior data, advertising feature data, and advertising scenario data.
3. The method according to claim 1, characterized in that The user behavior and environment data to be estimated are input into the trained A-GNN model, policy network, and time series analysis model to calculate the overall monetization benefits after the user enters the internal product, achieving real-time estimation of user value, including: Preprocess the estimated user behavior and environment data; The pre-processed user behavior and environment data are input into the trained A-GNN model, policy network and time series analysis model respectively; The policy network is used to predict the possibility of a user clicking on a specific advertisement. If the prediction is a possible click, the time series analysis model is used to estimate the value of the click. Combined with the learning results of the A-GNN model on the user-advertising-scenario relationship, the overall monetization benefits after the user enters the internal product are comprehensively calculated to complete the real-time estimation of user value.
4. The method according to claim 3, characterized in that The method further comprises: The possibility of a user clicking on a specific advertisement is predicted by the policy network. If the prediction is that the user is unlikely to click, other advertisements with a higher relevance to the current user's interest tags or browsing history are selected from the advertisement library for display.
5. The method according to claim 3, characterized in that The method further comprises: The policy network is used to predict the likelihood of a user clicking on a specific ad. If the prediction is that the user is unlikely to click, the ad placement is changed to an area that is more likely to attract the user's attention based on the historical click-through rate data of different ad positions and the user's behavioral characteristics on the current page.
6. The method according to claim 3, characterized in that The method further comprises: The policy network is used to predict the likelihood of a user clicking on a specific ad. If the prediction is that the user is unlikely to click, the user's real-time behavior data within the product is collected, and the user's interest tags are recalculated and calibrated in combination with a machine learning algorithm. Based on the new interest tags, ads that better match the user's interests are screened from the ad library.
7. The method according to claim 3, characterized in that The method further comprises: Predicting the likelihood of a user clicking on a specific ad using the policy network; if the likelihood of a click is predicted to be unlikely, evaluating whether there are new, undeveloped advertising scenarios that may be suitable for the user based on the content of the page the user is currently on and the user's behavior patterns; For potential new advertising scenarios that have been evaluated, advertising delivery tests are conducted among user groups with similar characteristics.
8. A deep learning user value estimation device, characterized in that: The device comprises: A collection module, used to collect user behavior and environment data in advertising business scenarios and pre-process the user behavior and environment data; The construction module is used to train the pre-built A-GNN model, the reinforcement learning-based policy network, and the time series analysis model using the pre-processed user behavior and environment data to obtain the trained A-GNN model, policy network, and time series analysis model; The analysis module is used to input the user behavior and environmental data to be estimated into the trained A-GNN model, policy network and time series analysis model respectively to calculate the overall monetization benefits after the user enters the internal product, thereby realizing real-time estimation of user value.
9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 7.