Push solution optimization method and system based on user-side response data collection
By analyzing the user-side response data, calculating the similarity between the user and the sample group and evaluating the possible clicks, optimizing push for users with less user response data, solving the problem of declining user experience and achieving a push effect that is more in line with user interests.
Patent Information
- Application Number
- CN202411835944.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-12-13
AI Technical Summary
In the prior art, when users browse less, it is difficult to analyze user behavior, resulting in the inability to perform fitting push, resulting in a decline in user experience.
By obtaining multiple users and their user response data on the user side, the similarity between the target user and each sample in the sample group is calculated, the target user's possible clicks on each content type, and the content type whose possible degree is greater than the set threshold is selected for pushing.
Effectively push content of target users, avoid random push, and improve users' experience of using the user side.
Smart Images

Figure CN119293346B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data push technology, and more specifically, to a push solution optimization method and system based on user end response data collection. Background Art
[0002] With the development of informatization, not only users actively look for information and products, but information and products also actively look for users, thus big data push services come into being.
[0003] Data push refers to predicting users' interests and hobbies based on their behavior information, and then pushing relevant content to them to increase information reading or product purchases. When data push meets user needs, it can not only reconcile the contradiction between information overload and personalized needs, reduce invalid interruptions, and improve user experience, but also enhance the effectiveness of information transmission and increase interaction rates, which plays a key role in improving the profitability of commercial platforms and the effectiveness of media communication. Nowadays, this kind of data push exists in all aspects of life, such as entertainment, news, and shopping.
[0004] Nowadays, most software with data push services have huge amounts of data and retrieval functions. For example, when you open a news website, you can get some of the latest and important news-related information, but it may not be the information that users need to know. Users can search for the content they need on the website. At this time, the website collects the user's search target and pushes the search target and related information to the customer.
[0005] For example, the patent document with authorization announcement number CN110753248B discloses a data collection system and method for IPTV user terminals, which reads the user behavior data in the stored user operation log and classifies it to determine the user's preference tags, classify the users, and finally push programs according to the user's preference tags.
[0006] However, the above solution does not take into account that when the user's operation log and physical data are less (the user's browsing behavior is less), the user's behavior is more random when the user operates the client, and it is difficult to analyze the user's behavior to make more appropriate push notifications, resulting in a decrease in user experience. Summary of the invention
[0007] The purpose of the present invention is to propose a push scheme optimization method and system based on user-side response data collection, so as to solve the problem in the prior art that when there are fewer users' browsing behaviors, it is difficult to perform user behavior analysis and cannot perform tailored push, resulting in a decrease in user experience; to this end, the present invention provides solutions in the following two aspects.
[0008] In a first aspect, the present invention provides a method for optimizing a push solution based on user terminal response data collection, comprising:
[0009] Acquire multiple users who are browsing the user terminal and their user response data; the user response data includes multiple response data, each response data includes the push content type, push time and click status;
[0010] The users whose total number of response data in the user response data is less than a set value are regarded as target users;
[0011] Obtain the probability of target users clicking on various types of content; probability for: , is the similarity between the target user and the i-th sample in the sample group; , They are the click status and push time of the dth response data when the content type of the i-th sample is c; Index for the moment; , are the total number of content types of all response data of the i-th sample and the total number of response data with content type c, respectively; is the total number of samples, is a normalization function; the similarity is the similarity between the target user and each sample in the sample group; each sample in the sample group is a user whose total number of response data in the user response data is greater than or equal to a set value;
[0012] Select content types with a probability greater than a set threshold and push them to target users.
[0013] In the above scheme, for users with less user response data, by obtaining the similarity between the target user and each sample, and then obtaining the possibility of the target user clicking on each content type, the target user's preferred content can be effectively pushed, random push can be avoided, and the user experience of using the client can be improved.
[0014] Optionally, the similarity for:
[0015] ;
[0016] In the formula, is the mean of the maximum correlation between the target user’s user response data and the i-th sample; The mth basic information of the target user; is the mth basic information in the basic information of the i-th sample; is the total number of basic information items; It is a normalization function, and the basic information includes but is not limited to age, gender, and region.
[0017] The similarity in the above scheme can measure the similarity between the user response data of the target user and the user response data of the sample.
[0018] Optionally, the maximum correlation value is the maximum correlation value between any response data in the user response data of the target user and each response data in the user response data of the sample;
[0019] The relevance for:
[0020] ;
[0021] In the formula, The content type of the target user's response data in item a; is the content type of the bth response data of the i-th sample; ; The click status of the target user's response data in item a; is the click status of the bth response data of the i-th sample; The time when the response data of item a is pushed; is the push time of the bth response data of the i-th sample; is the normalization function.
[0022] Optionally, it also includes:
[0023] Input the target user's basic information, as well as the push time and push content type in the user response data into the network model, and output the first click rate;
[0024] The likelihood level is modified using the first click rate.
[0025] The above scheme is based on the degree of possibility and further combines the first click rate predicted by the network model to obtain the final degree of possibility, which can further improve the accuracy of the pushed content type.
[0026] Optionally, the modifying the possibility degree by using the click rate includes: performing a weighted summation of the click rate and the possibility degree.
[0027] Optionally, it also includes: in response to the total number of response data in the user response data of each user being greater than or equal to a set value, inputting the basic information of each user and the push time and push content type in the user response data into the network model, outputting the second click-through rate, selecting the content type whose second click-through rate is greater than the set threshold, and pushing to the corresponding user.
[0028] The above solution can timely and efficiently push content to a group of users with more user response data.
[0029] Optionally, the method further includes a process of training the network model, specifically, obtaining a training set; the training set includes basic information of the user, push time, push content type and click status in the user response data; the click status is divided into two categories, click is 1, and no click is 0;
[0030] Input all users’ basic information, as well as push time and push content type in user response data into the network model, and use the corresponding click situations as labels to train the network model;
[0031] Use gradient descent to adjust model parameters to minimize prediction error;
[0032] Iteratively adjust the parameters of the network model until the loss is less than a certain value or the set number of training times is reached to obtain a trained network model.
[0033] In a second aspect, a push solution optimization system based on user terminal response data collection includes:
[0034] processor;
[0035] The memory stores computer instructions for optimizing a push solution based on user-side response data collection. When the computer instructions are executed by the processor, the system executes the above-mentioned method for optimizing a push solution based on user-side response data collection.
[0036] The beneficial effects of the present invention are:
[0037] The solution of the present invention is aimed at users with less user response data. By analyzing the similar behaviors of the target users and users with more user response data, the optimal push solution for the target users is determined. Compared with the existing solution of randomly pushing to users with less user response data, it can better fit the interest preferences of the target users and improve the browsing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0039] Figure 1 The following is a schematic flowchart showing the steps of the push solution optimization method based on user terminal response data collection in this embodiment;
[0040] Figure 2The structural block diagram of the push solution optimization system based on user terminal response data collection in this embodiment is schematically shown. DETAILED DESCRIPTION
[0041] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0042] When users use the client (such as agricultural web pages, news web pages, shopping platforms, social media platforms, etc.), some content is often pushed after the user refreshes the page. However, when the user has a lot of historical user data, the client can push content accurately, but when the user has less historical user data (historical browsing behavior), some content will be pushed randomly. These contents may not be the user's interest preference content, which obviously reduces the user experience.
[0043] Based on the above problems, the present invention provides a push solution optimization method based on user terminal response data collection to solve the push problem of users with less browsing behavior (users who do not frequently use the user terminal, have few browsing times or have just registered), that is, to provide users with the best push solution when the amount of user data is limited.
[0044] Specifically, taking a video webpage as an example, the push scheme optimization method based on user terminal response data collection in this embodiment is introduced. Figure 1 As shown, the following steps are included:
[0045] Step S1, obtaining multiple users who are browsing the user terminal and corresponding user data.
[0046] The user data in this embodiment includes basic information of the user and user response data.
[0047] The basic information includes but is not limited to the user's age, gender, and region. Specifically, in this embodiment, the user management system in the user terminal can obtain the age, gender, and region of each user as the user's basic information, where gender includes male and female, represented by the numbers 0 (male) and 1 (female) respectively; the region is the latitude and longitude of the current user's location, that is, the user's region is represented in digital form.
[0048] Different users have different acceptance levels for different pushed content. For example, users of different age groups often have different interests and needs; men and women have differences in interests and hobbies; the user's geographical location will determine their living environment, cultural background, consumption habits, etc., so it is necessary to collect the user's basic information.
[0049] Furthermore, basic information may also include education background, occupation, marital status, etc.
[0050] Among them, the user response data includes multiple response data, each of which includes the push time, push content type, and click status (whether the user clicks). Specifically, the user end collects the push time, push content type, and click status of any content received by the user, wherein the push time is the time node when the user receives the push, which can be specific to the hour, such as 16:00 on January 1, 2024 can be recorded as [24,1,1,16]; the push content type is a label of the pre-marked content type. When the staff edits the push, they will select the label of the content type, and number the push content type labels in the order of the existing push editing time. When a new category appears, the number is backward, so as to represent the push content type in digital form. For example, when the push content type is literary and artistic, variety show, urban, legal system, etc., the corresponding number can be 1, 2, 3, 4, and when a new category is added, such as ancient costume, the push content type is numbered 5; the click status includes click and no click, which are represented by 1 and 0 respectively.
[0051] Since push time, push content type, and click status directly reflect user interests and preferences, it is necessary to collect the above user response data.
[0052] Step S2: taking users with insufficient user response data as target users.
[0053] In this embodiment, the target user needs to be screened out from multiple users who are browsing the user terminal. The target user can be regarded as a "new user" (relative to all users), that is, a user with less user response data. Since the historical user data of the target user is insufficient, it is impossible to accurately evaluate the corresponding interest preference content, so it is necessary to screen out the user and push the most suitable content as possible.
[0054] In this embodiment, conditions for screening target users are set, that is, when a user meets at least one of the following conditions, the user can be used as a target user.
[0055] Among them, condition 1 is: the total usage time of the user is less than the time threshold; the time threshold can be the average usage time of all users. Specifically, it is necessary to first obtain the usage time of each user (the user management system obtains the first registration time of each user on the user terminal, and the total time of each user from the first registration time to the present on the user terminal is used as the usage time of each user), obtain the average usage time of all users, and use the response data of users whose total usage time is less than the average usage time as the target user data set.
[0056] As another implementation, the duration threshold may also be the average usage duration of all users since the registration time of the user, and of course may also be a manually set value, such as 24 hours.
[0057] The second condition is: the total number of response data in the user response data is less than a set value. Specifically, whether the user is a target user is evaluated based on the total number of contents browsed by the user in the user terminal.
[0058] Condition three is: the number of login times of the user is less than the set number of login times. Specifically, the number of login times of the user on the user terminal can be used to indicate whether the user frequently uses the user terminal, so as to evaluate whether the user is a target user.
[0059] Step S3, calculating the likelihood of the target user clicking on each type of content.
[0060] The process of obtaining the possible degree in this embodiment is as follows:
[0061] First, a sample group is obtained, where the sample group includes users who do not meet the three conditions in the above step S2.
[0062] Since the user response data of the target user cannot provide the target user with useful push content types, that is, it is impossible to accurately evaluate the corresponding interest preference content, it is necessary to select users with more user response data to help simulate the user response data of the target user.
[0063] Secondly, determine the similarity between the target user and each sample in the sample group.
[0064] Samples whose user response data at the initial registration stage are significantly different from those of the target user are not suitable for simulating the target user's future user response data. Therefore, it is necessary to first obtain the similarity between the target user and each sample for preliminary screening to find samples whose initial registration behavior is similar to that of the target user.
[0065] Among them, the similarity for: ;
[0066] In the formula, is the mean of the maximum correlation between the target user’s user response data and the i-th sample; The mth basic information of the target user; is the mth basic information in the basic information of the i-th sample; is the total number of basic information items; is the normalization function.
[0067] The above basic information includes but is not limited to age, gender, region, etc.
[0068] in, It indicates the degree of consistency between the basic information of the target user and the i-th sample. When the difference between the basic information of the target user and the i-th sample is small as a whole, the degree of consistency between the basic information of the target user and the i-th sample is higher; when the degree of consistency between the basic information of the target user and the i-th sample is higher, the similarity between the target user and the i-th sample in the target user sample group is higher.
[0069] Among them, the mean correlation The response data and the The mean of the maximum correlation values of samples.
[0070] Since there may be multiple response data in the user response data of the target user, and there are also many response data in the user response data of the sample, each response data in the user response data of the target user has a correlation with each response data in the user response data of the sample. Thus, the maximum value of the correlation between any response data in the user response data and each response data in the sample can be obtained. for:
[0071] ;
[0072] In the formula, The content type of the target user's response data in item a; is the content type of the bth response data of the i-th sample; ; The click status of the target user's response data in item a; is the click status of the bth response data of the i-th sample; The time when the response data of item a is pushed; is the push time of the bth response data of the i-th sample; is the normalization function.
[0073] in, It is used to determine whether the content type and click status of the a-th response data of the target user are consistent with the b-th response data of the i-th sample in the sample group; the smaller the push time difference between the a-th response data of the target user and the response data in the i-th sample, the higher the correlation between the a-th response data and the b-th response data of the i-th sample.
[0074] The above push time is to select data that is closer in time to the target user's response data, that is, the closer the time when the same content is clicked, the more similar the user behavior habits are, so as to avoid the influence of the user's response data at a distant time.
[0075] Then, determine the degree to which the target user is likely to be pushed the type of content after refreshing the node.
[0076] Specifically, the likelihood that the target user will click on content in category c The calculation formula is as follows:
[0077] ;
[0078] In the formula, is the similarity between the target user and the i-th sample in the sample group; , They are the click status and push time of the dth response data when the content type of the i-th sample is c; Index for the moment; , are the total number of content types of all response data of the i-th sample and the total number of response data with content type c, respectively; is the total number of samples, is a normalized function; the similarity is the similarity between the target user and each sample in the sample group.
[0079] in, is the time interval weight, and the closer the push time interval is to each response data of the i-th sample, the higher the influence weight is; that is, when the overall value of the influence of each response data in the sample with high similarity weighted by the time interval weight is high, the target user is more likely to click on the c-th content type.
[0080] The above normalization function is a linear normalization function, such as a maximum and minimum value normalization function.
[0081] For example, if the current time K is 14:00 on November 29, 2024, then k=[24, 11, 29, 14]. The probability obtained at this time is the probability of clicking on a certain type of content after refreshing at the current time.
[0082] Step S4, selecting content types with a probability greater than a set threshold and pushing them to target users.
[0083] In this embodiment, content types whose click probability of each target user is greater than a set threshold are selected for push.
[0084] The threshold is set as , of course, it can be determined according to actual conditions.
[0085] When the probability is less than or equal to the set threshold, the content can be pushed according to the content type before the target user refreshes the node.
[0086] Furthermore, in order to obtain the optimal push solution for the target user, the degree of possibility is also modified in this embodiment, specifically:
[0087] Input the target user's basic information, as well as the push time and push content type in the user response data into the network model, and output the first click rate;
[0088] The first click rate is used to modify the likelihood. The modification is to perform a weighted summation of the first click rate and the likelihood. The weights of the first click rate and the likelihood can be set equal or determined according to actual conditions.
[0089] The network model mentioned above can be a random forest model or a support vector machine model. Taking the random forest model as an example, the random forest is an integrated learning method composed of multiple decision trees. It improves the overall prediction performance and robustness by combining multiple weak classifiers. The advantage is that the random forest model has a strong generalization ability, can avoid overfitting problems, and the parameter adjustment is relatively simple.
[0090] Specifically, the number of trees in the random forest model in this embodiment is 100, the maximum depth of each tree is 10, and the loss function is the mean square error loss function.
[0091] Among them, the process of training the network model is:
[0092] First, get the training dataset.
[0093] The training data set in this embodiment can be the user data of all samples in the sample group; the basic information in the user data and the push time and push content type in the user response data are used as the input of the model, and the click situation is used as the label.
[0094] It should be noted that the label is not limited to 0 or 1. That's it.
[0095] Of course, as another implementation, the training data set may also include users with insufficient user response data, and these users may be part of the target users (these part of the target users are users with insufficient user response data other than the browsing user). In this case, the user response data of these users also need to be expanded to obtain the simulated training data of the corresponding users.
[0096] Specifically, the likelihood of some target users clicking on various push content types at various times can be obtained, and simulation training data including basic information of some target users, push time in user response data, push content type, and likelihood can be obtained.
[0097] It should be noted that the method for obtaining the possibility degree in the above-mentioned simulation training data is the same as the method for obtaining the possibility degree of the target user, which will not be described in detail here.
[0098] The purpose of obtaining the simulated training data is that in actual applications, new users do not have sufficient historical data and account for a small proportion of the total users. The data of such users is too sparse, which will reduce the accuracy of the model when training the random forest model. Other users have more sample data and are also new users at the beginning of registration. In order to make the push plan obtained by the model more accurate, some target users are matched with old users who have similar behaviors to them at the beginning of registration. The historical data of old users can be effectively used to simulate some target user data, thereby approximating their future behavior trends, and then the data sets of some target users and old users can be obtained. This can fully utilize the existing data and ensure the accuracy of the network model in obtaining the push plan.
[0099] Secondly, the network model is trained using the training data set to obtain a trained network model.
[0100] When the training data set is a combination of users with insufficient user response data and users with sufficient user response data, the input of the model in this embodiment is other data except click status. Users with sufficient user response data use click status as the corresponding label, and some target users use the possibility of each type of content at each time as the label to train the random forest model. The output of the model is output in the form of probability value, which is taken in within the range.
[0101] Furthermore, in response to the total number of response data in the user response data of each user being greater than or equal to a set value, the basic information of the user and the push time and push content type in the user response data can be directly input into the network model to obtain the second click-through rate, and the content type with a second click-through rate greater than the set threshold can be selected for priority push.
[0102] The solution of the present invention can specify the optimal push solution for users with less user response data.
[0103] The present invention also provides a push solution optimization system based on user-side response data collection. Figure 2 As shown, the system includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the push scheme optimization method based on user terminal response data collection according to the present invention is implemented.
[0104] The management system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, and the configuration and functions of these components are known in the art, so they will not be described in detail here.
[0105] In the present invention, the aforementioned memory may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus or device. For example, a computer-readable storage medium may be any appropriate magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory RRAM (Resistive Random Access Memory), a dynamic random access memory DRAM (Dynamic Random Access Memory), a static random access memory SRAM (Static Random-Access Memory), an enhanced dynamic random access memory EDRAM (Enhanced Dynamic Random Access Memory), a high-bandwidth memory HBM (High-Bandwidth Memory), a hybrid memory cube HMC (Hybrid Memory Cube), etc., or any other medium that can be used to store the required information and can be accessed by an application, a module, or both. Any such computer storage medium may be part of a device or accessible or connectable to a device. Any application or module described in the present invention may be implemented using computer-readable / executable instructions that may be stored or otherwise maintained by such a computer-readable medium.
[0106] In the description of this specification, “plurality” means at least two, such as two, three or more, etc., unless otherwise clearly and specifically defined.
[0107] Although this specification has shown and described a number of embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will conceive of many modifications, changes and alternatives without departing from the ideas and spirit of the present invention. It should be understood that in the practice of the present invention, various alternatives to the embodiments of the present invention described herein may be employed.
Claims
1. A push scheme optimization method based on user-side response data collection, characterized in that: include: Get multiple users browsing the client and their user response data; The user response data includes multiple pieces of response data, each piece of response data includes the push content type, push time, and click status; The users whose total number of response data in the user response data is less than a set value are regarded as target users; Get the probability of target users clicking on various types of content; Possibility for: , is the similarity between the target user and the i-th sample in the sample group; , They are the click status and push time of the dth response data when the content type of the i-th sample is c; Moment index for target users; , are the total number of content types of all response data of the i-th sample and the total number of response data with content type c, respectively; is the total number of samples, is a normalization function; each sample in the sample group is a user whose total number of response data in the user response data is greater than or equal to a set value; Select the content type with a probability greater than the set threshold and push it to the target user; the similarity for: ; In the formula, is the mean of the maximum correlation between the target user’s user response data and the i-th sample; The mth basic information of the target user; is the mth basic information in the basic information of the i-th sample; is the total number of basic information items; is a normalization function, and the basic information includes but is not limited to age, gender, and region; The maximum correlation value is the maximum correlation value between any response data in the user response data of the target user and each response data in the user response data of the sample; The relevance for: ; In the formula, The content type of the target user's response data in item a; is the content type of the bth response data of the i-th sample; ; The click status of the target user's response data in item a; is the click status of the bth response data of the i-th sample; The time when the response data of item a is pushed; is the push time of the bth response data of the i-th sample; is the normalization function.
2. The push scheme optimization method based on user terminal response data collection according to claim 1 is characterized in that: Also includes: Input the target user's basic information, as well as the push time and push content type in the user response data into the network model, and output the first click rate; The likelihood level is modified using the first click rate.
3. The push scheme optimization method based on user terminal response data collection according to claim 2 is characterized in that: The modifying the possibility degree by using the first click rate includes: performing a weighted summation of the first click rate and the possibility degree.
4. The push scheme optimization method based on user terminal response data collection according to claim 1 is characterized in that: Also includes: In response to the fact that the total number of response data in the user response data of each user is greater than or equal to a set value, the basic information of each user and the push time and push content type in the user response data are input into the network model, the second click-through rate is output, and the content type whose second click-through rate is greater than the set threshold is selected and pushed to the corresponding user.
5. The push scheme optimization method based on user terminal response data collection according to claim 2 or 4 is characterized in that: The process of training the network model is also included, specifically: obtaining a training set; the training set includes basic information of the user, push time, push content type and click status in the user response data; The click situation is divided into two categories, click is 1, and no click is 0; Input all users’ basic information, as well as push time and push content type in user response data into the network model, and use the corresponding click situations as labels to train the network model; Use gradient descent to adjust model parameters to minimize prediction error; Iteratively adjust the parameters of the network model until the loss is less than a certain value or the set number of training times is reached to obtain a trained network model.
6. A push solution optimization system based on user-side response data collection, characterized in that: include: processor; A memory storing computer instructions for optimizing a push solution based on user-side response data collection. When the computer instructions are executed by the processor, the system executes the push solution optimization method based on user-side response data collection according to any one of claims 1-5.
Citation Information
Patent Citations
A system and method for collecting data from IPTV user terminals
CN110753248B
Pushing method and system based on big data
CN110134870A