User screening method and device

By screening the user groups that have not triggered the preset operation behavior and combining it with the probability prediction model of users in the target category and channel in the future, the problem of mismatch of user screening logic in the existing technology is solved, and a more efficient user screening effect is achieved.

CN113378043BActive Publication Date: 2025-09-19BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110620154.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-03
Publication Date
2025-09-19
Estimated Expiration
2041-06-03

AI Technical Summary

Technical Problem

Existing technologies do not fully consider the intrinsic characteristics of users in user screening, resulting in a mismatch between screening logic and screening targets, affecting the screening effect.

Method used

By screening candidate user groups that have not triggered preset operational behaviors, combined with the user's future probability prediction model in the target category and channel, the first and second probabilities of the user are determined, and the third probability is obtained by fusion to screen out the target user group.

Benefits of technology

It improves the matching and effectiveness of user screening, ensures that the screened users have a high preference for target categories and channels, and improves the accuracy and efficiency of user screening.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113378043B_ABST
    Figure CN113378043B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for user screening, which relates to the field of computer technology. A specific implementation of the method includes: screening a candidate user group from all users who have not triggered a preset operation behavior in a target category; determining a first probability that any user in the candidate user group will trigger a preset operation behavior in the target category in the future, and a second probability that any user in the candidate user group will trigger a preset operation behavior in the target category through a target channel in the future, and determining a third probability that any user will trigger a preset operation behavior in the target category through a target channel in the future based on the first probability and the second probability; screening users whose third probability meets the preset conditions from the candidate user group to obtain a target user group corresponding to the target category. This implementation can screen users based on the user's preference for the target category and the user's preference for the target channel, ensure the match between the screening logic and the screening target, and improve the user screening effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method and device for user screening. Background Art

[0002] Existing technologies typically screen users based on user profiles. This screening logic doesn't fully consider the user's inherent characteristics, isn't highly digitized, and can lead to mismatches between the screening logic and the screening target, making screening effectiveness unreliable. Summary of the Invention

[0003] In view of this, an embodiment of the present invention provides a method and device for user screening, which can screen users according to their preferences for target categories and their preferences for target channels, ensure the match between screening logic and screening targets, and improve user screening effects.

[0004] To achieve the above object, according to one aspect of an embodiment of the present invention, a method for user screening is provided, comprising:

[0005] Filter candidate user groups from all users who have not triggered preset operational behaviors in the target category;

[0006] Determining a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future, and a second probability that the preset operation behavior will be triggered in the target category through a target channel in the future, and determining a third probability that any user will trigger the preset operation behavior in the target category through a target channel in the future based on the first probability and the second probability;

[0007] Users whose third probability meets a preset condition are screened from the candidate user group to obtain a target user group corresponding to the target category.

[0008] Optionally, filter candidate user groups that do not have preset operating behaviors in the target category, including:

[0009] The historical behavior data of all users are obtained, and users who have not triggered the preset operation behavior in the target category are filtered according to the historical behavior data to obtain the candidate user group.

[0010] Optionally, filter candidate user groups that do not have preset operating behaviors in the target category, including:

[0011] Divide all users into multiple target groups and obtain historical behavior data of each user in the target group; determine the TGI index of the target group based on the historical behavior data of each user in the target group, using the triggering of the preset operation behavior in the target category as a common feature, and select target groups with a TGI index greater than a preset TGI threshold as candidate target groups;

[0012] According to the historical behavior data of each user in the candidate target group, users who have not triggered the preset operation behavior in the target category are screened from the candidate target group to obtain the candidate user group.

[0013] Optionally, filter candidate user groups that do not have preset operating behaviors in the target category, including:

[0014] Obtaining historical behavior data of all users in multiple historical time periods, and determining the user's preference value for the target category in the corresponding historical time period based on the historical behavior data in each historical time period;

[0015] Accumulating the user's preference values ​​for the target category in each historical period in a time-decay manner to obtain a user preference index for the target category, and filtering users whose preference index is greater than a set threshold to obtain a candidate user set;

[0016] According to the historical behavior data of each user in the candidate user set, users who have not triggered the preset operation behavior in the target category are filtered from the candidate user set to obtain the candidate user group.

[0017] Optionally, filter candidate user groups that do not have preset operating behaviors in the target category, including:

[0018] Obtaining a user profile of each user, determining a fourth probability of the user triggering the preset operation behavior in the target category based on the user profile, screening users whose fourth probability is greater than a set threshold or several users with the highest fourth probability to obtain the candidate user set;

[0019] The historical behavior data of each user in the candidate user set is obtained, and users who have not triggered the preset operation behavior in the target category are filtered from the candidate user set according to the historical behavior data of each user in the candidate user set to obtain the candidate user group.

[0020] Optionally, determining a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future includes:

[0021] Obtaining historical behavior data of any user, determining the probability of any user triggering the preset operation behavior in each category based on the historical behavior data of any user, and obtaining a first feature vector of any user;

[0022] Obtaining a user profile of the any user, determining a probability of the any user triggering the preset operation behavior in each category based on the user profile of the any user, and obtaining a second feature vector of the any user;

[0023] Determine a category association feature matrix corresponding to any one of the users based on the historical behavior data of the user;

[0024] Obtaining a category feature vector of the target category;

[0025] The first feature vector, the second feature vector, the category association feature matrix, and the category feature vector are input into a pre-trained first model to obtain a first probability that any user will trigger the preset operation behavior in the target category in the future.

[0026] Optionally, determining a second probability that any user in the candidate user group will trigger the preset operation behavior in the target category through a target channel in the future includes:

[0027] Acquire historical behavior data of the any user, determine a triggering channel for each of the preset operation behaviors triggered by the any user in the target category based on the historical behavior data of the any user, and obtain a channel preference vector of the any user;

[0028] All channel preference vectors of any one of the users are input into a pre-trained second model to obtain a second probability that any one of the users will trigger the preset operation behavior through the target channel in the target category in the future.

[0029] Optionally, the third probability that any one of the users will trigger the preset operation behavior in the target category through the target channel in the future is determined based on the first probability and the second probability, including: taking the product of the first probability and the second probability as the third probability that any one of the users will trigger the preset operation behavior in the target category through the target channel in the future.

[0030] Optionally, before screening candidate user groups that have not triggered preset operation behaviors in the target category from all users, the following steps are further included:

[0031] Receive a user filtering request input by a user, and parse a target category from the user filtering request; confirm that a target user group corresponding to the target category does not exist in the cache; if a target user group corresponding to the target category exists in the cache, obtain the target user group corresponding to the target category from the cache.

[0032] According to a second aspect of an embodiment of the present invention, there is provided an apparatus for user screening, comprising:

[0033] The recall module screens candidate user groups from all users who have not triggered preset operation behaviors in the target category;

[0034] an estimation module, determining a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future, and a second probability that the preset operation behavior will be triggered in the target category through a target channel in the future;

[0035] The fusion module determines, based on the first probability and the second probability, a third probability that any one of the users will trigger the preset operation behavior in the target category through the target channel in the future; and filters users whose third probability meets the preset conditions from the candidate user group to obtain a target user group corresponding to the target category.

[0036] Optionally, the recall module screens candidate user groups that do not have preset operating behaviors in the target category, including:

[0037] The historical behavior data of all users are obtained, and users who have not triggered the preset operation behavior in the target category are filtered according to the historical behavior data to obtain the candidate user group.

[0038] Optionally, the recall module screens candidate user groups that do not have preset operating behaviors in the target category, including:

[0039] Divide all users into multiple target groups and obtain historical behavior data of each user in the target group; determine the TGI index of the target group based on the historical behavior data of each user in the target group, using the triggering of the preset operation behavior in the target category as a common feature, and select target groups with a TGI index greater than a preset TGI threshold as candidate target groups;

[0040] According to the historical behavior data of each user in the candidate target group, users who have not triggered the preset operation behavior in the target category are screened from the candidate target group to obtain the candidate user group.

[0041] Optionally, the recall module screens candidate user groups that do not have preset operating behaviors in the target category, including:

[0042] Obtaining historical behavior data of all users in multiple historical time periods, and determining the user's preference value for the target category in the corresponding historical time period based on the historical behavior data in each historical time period;

[0043] Accumulating the user's preference values ​​for the target category in each historical period in a time-decay manner to obtain a user preference index for the target category, and filtering users whose preference index is greater than a set threshold to obtain a candidate user set;

[0044] According to the historical behavior data of each user in the candidate user set, users who have not triggered the preset operation behavior in the target category are filtered from the candidate user set to obtain the candidate user group.

[0045] Optionally, the recall module screens candidate user groups that do not have preset operating behaviors in the target category, including:

[0046] Obtaining a user profile of each user, determining a fourth probability of the user triggering the preset operation behavior in the target category based on the user profile, screening users whose fourth probability is greater than a set threshold or several users with the highest fourth probability to obtain the candidate user set;

[0047] The historical behavior data of each user in the candidate user set is obtained, and users who have not triggered the preset operation behavior in the target category are filtered from the candidate user set according to the historical behavior data of each user in the candidate user set to obtain the candidate user group.

[0048] Optionally, the estimation module determines a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future, including:

[0049] Obtaining historical behavior data of any user, determining the probability of any user triggering the preset operation behavior in each category based on the historical behavior data of any user, and obtaining a first feature vector of any user;

[0050] Obtaining a user profile of the any user, determining a probability of the any user triggering the preset operation behavior in each category based on the user profile of the any user, and obtaining a second feature vector of the any user;

[0051] Determine a category association feature matrix corresponding to any one of the users based on the historical behavior data of the user;

[0052] Obtaining a category feature vector of the target category;

[0053] The first feature vector, the second feature vector, the category association feature matrix, and the category feature vector are input into a pre-trained first model to obtain a first probability that any user will trigger the preset operation behavior in the target category in the future.

[0054] Optionally, the estimation module determines a second probability that any user in the candidate user group will trigger the preset operation behavior in the target category through the target channel in the future, including:

[0055] Acquire historical behavior data of the any user, determine a triggering channel for each of the preset operation behaviors triggered by the any user in the target category based on the historical behavior data of the any user, and obtain a channel preference vector of the any user;

[0056] All channel preference vectors of any one of the users are input into a pre-trained second model to obtain a second probability that any one of the users will trigger the preset operation behavior through the target channel in the target category in the future.

[0057] Optionally, the fusion module determines a third probability that any one of the users will trigger the preset operation behavior in the target category through the target channel in the future based on the first probability and the second probability, including: taking the product of the first probability and the second probability as the third probability that any one of the users will trigger the preset operation behavior in the target category through the target channel in the future.

[0058] Optionally, the device also includes an input and output module, which is used to: receive a user screening request input by the user before the recall module screens the candidate user group that has not triggered the preset operation behavior in the target category from all users, and parse the target category from the user screening request; confirm that there is no target user group corresponding to the target category in the cache; if there is a target user group corresponding to the target category in the cache, obtain the target user group corresponding to the target category from the cache.

[0059] According to a third aspect of an embodiment of the present invention, there is provided an electronic device for user screening, comprising:

[0060] one or more processors;

[0061] a storage device for storing one or more programs,

[0062] When the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect of the embodiment of the present invention.

[0063] According to a fourth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method provided by the first aspect of the embodiment of the present invention is implemented.

[0064] One embodiment of the above invention has the following advantages or beneficial effects: by screening a candidate user group that has not triggered a preset operation behavior in a target category, new users of the target category can be screened; by determining a first probability that a user will trigger a preset operation behavior in the target category in the future, users with a high preference for the target category can be screened from new users; by determining a second probability that a user will trigger a preset operation behavior in the target category through a target channel in the future, users with a high preference for the target channel can be screened from new users. The present invention can ensure a match between the screening logic and the screening target, so that the target user group obtained by screening has a high preference for the target category and target channel, thereby improving the user screening effect.

[0065] The further effects of the above-mentioned non-conventional optional manner will be described below in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The accompanying drawings are provided for a better understanding of the present invention and are not intended to limit the present invention.

[0067] Figure 1 is a schematic diagram of the main process of the user screening method according to an embodiment of the present invention;

[0068] Figure 2 is a schematic diagram of the architecture of a method for user screening in an optional embodiment of the present invention;

[0069] Figure 3 is a flow chart of a method for user screening in an optional embodiment of the present invention;

[0070] Figure 4 is a schematic diagram of screening a candidate user group in an optional embodiment of the present invention;

[0071] Figure 5 is a schematic diagram of screening candidate user groups by user portraits in an optional embodiment of the present invention;

[0072] Figure 6 is a schematic diagram of category relationship mining in an optional embodiment of the present invention;

[0073] Figure 7 is a schematic diagram of determining a first probability using a DNN in an optional embodiment of the present invention;

[0074] Figure 8 is a schematic diagram of main modules of an apparatus for user screening according to an embodiment of the present invention;

[0075] Figure 9 is an exemplary system architecture diagram in which embodiments of the present invention may be applied;

[0076] Figure 10 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. DETAILED DESCRIPTION

[0077] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, in which various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0078] According to one aspect of an embodiment of the present invention, a method for user screening is provided.

[0079] Figure 1 FIG. 1 is a schematic diagram of the main process of the method for user screening according to an embodiment of the present invention. Figure 1 As shown, the user screening method according to an embodiment of the present invention includes: steps S101 to S105.

[0080] Step S101 : screening candidate user groups from all users who have not triggered a preset operation behavior in the target category.

[0081] The "all users" mentioned here refers to all users within the screening scope. For example, in the e-commerce sector, when selecting advertising promotion targets from all registered users on an e-commerce platform, the screening scope is all registered users on the e-commerce platform. For example, in the brokerage sector, when selecting advertising promotion targets from users who have opened accounts in the past three years, the screening scope includes users who have opened accounts on the brokerage platform in the past three years.

[0082] An action is a user's action that triggers the execution of an action. Preset actions can be set selectively based on actual circumstances, such as browsing, clicking, following, adding to cart, and saving. Failed actions can mean that the action has never been triggered, or that it hasn't been triggered within a set time period.

[0083] The users in the candidate user group are all new users of the target category. This step aims to screen new users of the target category from all users so as to formulate a targeted category acquisition strategy (acquiring new users means absorbing new users). The method of screening the candidate user group can be selectively set according to the actual situation, such as screening users who have not triggered preset operation behaviors in the target category within a specific age or education range, screening users who have not triggered preset operation behaviors in the target category within a specific geographical location range, etc. In an optional embodiment, the candidate user group can be screened by any one or more of the following methods: (1) recall based on the user's category behavior, (2) recall based on the TGI of the target group, (3) recall based on time decay, and (4) recall based on user portraits. The following details the implementation method of each method of screening the candidate user group. When the above multiple methods are adopted, the candidate user groups obtained by different methods can be merged and deduplicated to obtain the final subsequent user group.

[0084] (1) Recall based on user category behavior: Obtain historical behavior data of all users, and filter users who have not triggered the preset operation behavior in the target category based on the historical behavior data to obtain the candidate user group. In actual application, the user's historical behavior data can be obtained by log reporting, and the log data can be analyzed to determine whether the user has triggered the preset operation behavior in the target category. Users who have never triggered the preset operation behavior in the target category or have not triggered the preset operation behavior in the target category within a set time period are filtered as candidate users to obtain the candidate user group.

[0085] (2) Recall based on the TGI of the target group: divide all users into multiple target groups, and obtain the historical behavior data of each user in the target group; take triggering the preset operation behavior in the target category as a common feature, determine the TGI index of the target group based on the historical behavior data of each user in the target group, and screen the target group with a TGI index greater than the preset TGI threshold as the candidate target group; based on the historical behavior data of each user in the candidate target group, screen users who have not triggered the preset operation behavior in the target category from the candidate target group to obtain the candidate user group.

[0086] The TGI (Target Group Index) index is an index that reflects the strength or weakness of the target group within a specific research scope (such as geographical area, demographic field, media audience, product consumer). The higher the TGI index, the closer the relationship between the screened user group and the target category, and the better the recall effect. For example, an e-commerce platform has 10,000 users, of which 5,500 are male. The common feature is triggering preset operation behaviors in digital products. The number of male users with this common feature is 4,000, and the number of users with this common feature among all users is 6,000. Then the TGI index of the male group is: (4,000 / 5,500) / (6,000 / 10,000)=1.21.

[0087] (3) Recall based on time decay: obtain historical behavior data of all users in multiple historical time periods, determine the preference value of the user for the target category in the corresponding historical time period based on the historical behavior data in each historical time period; accumulate the preference value of the user for the target category in each historical time period in a time decay manner to obtain the user's preference index for the target category, filter out users whose preference index is greater than a set threshold to obtain a candidate user set; filter out users who have not triggered the preset operation behavior in the target category from the candidate user set based on the historical behavior data of each user in the candidate user set to obtain the candidate user group.

[0088] The preference value reflects a user's preference for a target category. A higher value indicates a greater interest in the target category and a higher likelihood of purchasing items within it. The preference value can be calibrated based on specific circumstances, such as the percentage of orders in the target category among all orders placed by the user. Another example is predicting the purchase probability of items in the target category using a pre-trained model based on user profiles or historical user behavior. The closer the time to the present moment, the greater the weighting of the preference value; the further the time from the present moment, the less weighting.

[0089] Accumulating in a time-decay manner means determining the weights of the preference values ​​of each historical period in a time-decay manner, and then accumulating the user's preference values ​​in each historical period with the weights determined in the pattern to obtain the user's preference index for the target category. In an optional embodiment, the time-decay formula is:

[0090]

[0091] Where x represents the user; y represents the target category; α is a constant that can be fitted or customized; i represents the historical time from now, and the unit can be customized, such as days; k represents the maximum historical time from now, and the unit is the same as that of i; P{y|x,1,k} represents the preference value weight of user x for target category y when the time from now is k (the unit can be customized, such as days).

[0092] The longer the time it takes for the user to trigger the preset operation behavior in the target category, the larger the corresponding preference value index. In actual application, when analyzing a large amount of historical behavior data, incremental and full comprehensive calculation methods can be used for analysis and processing. For example, if the user's preference index for the target category is determined every day based on all historical behavior data since January 1, 2020, then based on the preference index calculated last time, only the current preference value can be calculated, and the current preference value can be added to the preference index calculated last time as the preference index for the day. By using incremental and full comprehensive calculation methods for analysis and processing, only the data for the day needs to be calculated each time, and all historical behavior data only needs to be calculated once when determining the user's preference index, which can greatly reduce the consumption of computing resources and improve the efficiency of user screening.

[0093] (4) Recall based on user portraits: obtain a user portrait of each user, determine the fourth probability of the user triggering the preset operation behavior in the target category based on the user portrait, screen users with a fourth probability greater than a set threshold or several users with the highest fourth probability to obtain the candidate user set; obtain historical behavior data of each user in the candidate user set, and screen users who have not triggered the preset operation behavior in the target category from the candidate user set based on the historical behavior data of each user in the candidate user set to obtain the candidate user group.

[0094] User portraits, also known as user roles, are virtual representations of real users. They are labels that describe users created through big data.

[0095] In actual application, the probability of users triggering preset operation behaviors in the target category can be directly calculated based on user portraits. Figure 5 For example, first calculate the probability that user A triggers the preset operation behavior in categories 1-5, assuming they are p1, p2, p3, p4 and p5 respectively, and take p1 as the fourth probability that user A triggers the preset operation behavior in category 1.

[0096] Of course, the probability of a user triggering a preset operation in all categories can also be calculated. Based on the probability of a user triggering a preset operation in categories other than the target category and the conversion relationship between other categories and the target category, the probability of a user who has triggered a preset operation in other categories triggering a preset operation in the target category in the future can be determined. This probability can be added to the directly calculated probability of a user triggering a preset operation in the target category to obtain a fourth probability of a user triggering a preset operation in the target category. Figure 5 For example, let's first calculate the probabilities of user A triggering the preset action in categories 1-5, assuming they are p1, p2, p3, p4, and p5, respectively. Assuming the probabilities of a user triggering the preset action in categories 2-5 triggering the preset action in category 1 in the future are p2', p3', p4', and p5', respectively. Then, the fourth probability of user A triggering the preset action in category 1 is: (p1+p2'+p3'+p4'+p5').

[0097] When determining the probability of a user triggering a preset action within a particular category, a prediction model can be trained based on the user's historical behavior data. The trained prediction model can then be used to determine the probability of a user triggering a preset action within a specific category. Of course, for a particular category, the probability of a user triggering a preset action within that category can also be determined based on the ratio between the number of actions that triggered the preset action within that category and the total number of actions included in the historical behavior data. Figure 6 Schematic diagram of category relationship mining in an optional embodiment of the present invention. Figure 6As shown, to determine the probability of a user transitioning from category X to category Y, we can first calculate the probability Sup(X) of a user triggering a preset action in category X. Then, we calculate the probability Sup(X∪Y) of a user triggering a preset action in both categories X and Y. We use Sup(X∪Y) / Sup(X) as the probability of a user transitioning from category X to category Y. This means the probability that a user who triggered a preset action in category X will also trigger a preset action in category Y. By considering the relationships between categories, we can further improve the accuracy of user screening.

[0098] Step S102 : determining a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future.

[0099] For a certain category, the ratio between the number of operation behaviors that triggered the preset operation behaviors in the category and the total number of operation behaviors contained in the historical behavior data can be used as the first probability that the user will trigger the preset operation behavior in the category in the future. Alternatively, an estimation model can be trained based on the user's historical behavior data, and the trained estimation model can be used to determine the first probability that the user will trigger the preset operation behavior in a specific category.

[0100] Optionally, determining the first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future includes: obtaining historical behavior data of any user, determining the probability of any user triggering the preset operation behavior in each category based on the historical behavior data of any user, and obtaining a first feature vector of any user; obtaining a user portrait of any user, determining the probability of any user triggering the preset operation behavior in each category based on the user portrait of any user, and obtaining a second feature vector of any user; determining the category association feature matrix corresponding to any user based on the historical behavior data of any user; obtaining the category feature vector of the target category; inputting the first feature vector, the second feature vector, the category association feature matrix and the category feature vector into a pre-trained first model to obtain a first probability that any user will trigger the preset operation behavior in the target category in the future.

[0101] Each element in the first and second eigenvectors represents the probability of a user triggering a preset action in the corresponding category. The first and second eigenvectors have the same number of dimensions as the number of categories. The difference between the first and second eigenvectors is that the probability of a user triggering a preset action in the corresponding category in the first eigenvector is calculated based on the user's historical behavior data, while the probability of a user triggering a preset action in the corresponding category in the second eigenvector is calculated based on the user's profile.

[0102] Figure 7is a schematic diagram of determining the first probability using DNN (Deep Neural Networks) in an optional embodiment of the present invention, such as Figure 7 As shown, user characteristics (including the first eigenvector and the second eigenvector) and category characteristics (including the category association feature matrix and category feature vector) are input into the DNN deep neural network model to predict the probability of a user placing an order in the target category within the next 7 days. During the user screening process, to fully consider the user's inherent characteristics, features such as item price, brand, and origin from the user's historical behavior data can also be input into the model.

[0103] The first probability is the probability that a candidate user will convert into a new user of the target category, reflecting the user conversion rate. When determining this first probability, the present invention not only considers the user profile but also other, richer user characteristics, such as historical user behavior data, category associations, and user behavior characteristics within each category. Furthermore, by employing a deep neural network model to perform high-level mining of intrinsic features, this allows for an accurate estimation of the user's future conversion rate.

[0104] In some optional embodiments, the probability that the user will trigger a preset operation in the category in the future can be determined first, for example, the ratio between the number of user operation behaviors that trigger the preset operation in the category and the total number of operation behaviors contained in the historical behavior data can be used as the probability, or a prediction model can be trained based on the user's historical behavior data and the trained prediction model can be used to determine the probability; then the probability can be directly used as the first probability that the user will trigger the preset operation in the corresponding category. For example, Figure 5 For example, first calculate the probability of user A triggering the preset operation behavior in categories 1-5. Assume that they are p1, p2, p3, p4 and p5 respectively, and directly use p1 as the first probability of user A triggering the preset operation behavior in category 1.

[0105] In actual application, we can further determine the probability that a user who has triggered a preset operation in other categories will trigger a preset operation in the target category in the future based on the association between other categories and the target category, and add this probability to the directly calculated probability that the user will trigger a preset operation in the target category to obtain a fourth probability that the user will trigger a preset operation in the target category. Figure 5 For example, let's first calculate the probabilities of user A triggering the preset action in categories 1-5, assuming they are p1, p2, p3, p4, and p5, respectively. Assuming the probabilities of a user triggering the preset action in categories 2-5 triggering the preset action in category 1 in the future are p2', p3', p4', and p5', respectively. Then, the fourth probability of user A triggering the preset action in category 1 can also be: (p1+p2'+p3'+p4'+p5').

[0106] The category association feature matrix corresponding to a user refers to the association relationship between categories determined based on the historical behavior data of each user. Figure 6 Schematic diagram of category relationship mining in an optional embodiment of the present invention. Figure 6 As shown, for a given user, to determine the probability of converting from category X to category Y, we can first calculate the probability Sup(X) that the user triggers the preset action in category X. Then, we calculate the probability Sup(X∪Y) that the user triggers the preset action in both categories X and Y. The result Sup(X∪Y) / Sup(X) is used as the conversion probability from category X to category Y. This is the probability that a user who triggers the preset action in category X will also trigger the preset action in category Y. By considering the relationships between categories, the accuracy of user screening can be further improved. When there are many categories associated with the target category, to reduce the computational complexity, we can consider only the conversions of the top few categories with the closest relationships to the target category.

[0107] A category feature vector is a vector formed by the attribute characteristics of a category. The specific content of these attributes can be optionally defined, such as the category's ad impressions and the category's new user conversion rate (i.e., the proportion of new users who trigger a predefined action). To facilitate analysis and processing, category features are encoded using Embedding (a method of converting discrete variables into continuous vector representations), continuous features are normalized, and discrete features are encoded using StringIndex.

[0108] Step S103 , determining a second probability that any user in the candidate user group will trigger the preset operation behavior in the target category through the target channel in the future.

[0109] The second probability reflects the user's preference for the target channel. The higher the user's preference for the target channel, the greater the second probability. Taking advertising as an example, the greater the user's preference for the target channel (advertisement) for the target category (for example, beauty products), the more likely the user is to browse or purchase items in the target category through the ad.

[0110] For a certain category, the ratio between the number of user's operation behaviors that triggered the preset operation behaviors through the target channel in this category and the total number of operation behaviors contained in the historical behavior data can be used as the second probability that the user will trigger the preset operation behavior in this category in the future. Alternatively, an estimation model can be trained based on the user's historical behavior data, and the trained estimation model can be used to determine the second probability that the user will trigger the preset operation behavior through the target channel in a specific category.

[0111] Optionally, determining the second probability that any user in the candidate user group will trigger the preset operation behavior in the target category through the target channel in the future includes: obtaining the historical behavior data of any user, determining the triggering channel of each preset operation behavior triggered by any user in the target category based on the historical behavior data of any user, and obtaining the channel preference vector of any user; inputting all the channel preference vectors of any user into a pre-trained second model to obtain the second probability that any user will trigger the preset operation behavior in the target category through the target channel in the future. The network structure of the second model can be selectively set according to actual conditions. Optionally, XGBoot (an open source software library) is used to train the second model.

[0112] The second model is essentially a CTR (Click-Through-Rate) model, an advertising preference model. By establishing a CTR model for CPM advertising, it's possible to filter or downgrade users with low preference for the target channel, thereby acquiring more new users at a lower CAC (Customer Acquisition Cost). CAC is the total marketing expenditure divided by the total number of new users acquired through that expenditure).

[0113] The channel preference vector reflects the user's preference for the corresponding channel. The output of the second model can adopt the structure of "user identifier + category + brand + touchpoint + second probability". The output of the second model is shown in Table 1 below:

[0114] Table 1 Output of the second model

[0115]

[0116] In Table 1, touchpoints represent channels.

[0117] Step S104 , determining a third probability that any user will trigger the preset operation behavior in the target category through the target channel in the future based on the first probability and the second probability.

[0118] The third probability is positively correlated with the first probability and the second probability. A greater third probability indicates a greater probability that the user will trigger a preset action in the target category through the target channel. Optionally, determining the third probability that any user will trigger the preset action in the target category through the target channel in the future based on the first probability and the second probability includes: multiplying the first probability and the second probability as the third probability that any user will trigger the preset action in the target category through the target channel in the future.

[0119] Step S105 , screening users whose third probability meets a preset condition from the candidate user group to obtain a target user group corresponding to the target category.

[0120] In an optional embodiment, before screening candidate user groups that have not triggered preset operation behaviors in the target category from all users, it also includes: receiving a user screening request input by the user, parsing the target category from the user screening request; confirming that the target user group corresponding to the target category does not exist in the cache; if the target user group corresponding to the target category exists in the cache, then obtaining the target user group corresponding to the target category from the cache. In this embodiment, the target user group corresponding to each category is first determined based on the category as the dimension, and then the target user groups corresponding to each category in the user screening request are merged and returned to the user. By analyzing the user screening request as a screening task based on the category as the dimension, the adaptability and scalability of the user screening method of the embodiment of the present invention can be improved. At the same time, due to the setting of the cache, repeated calculations when the same category is included in multiple user screening requests can be avoided, computing resource consumption can be reduced, and user screening efficiency can be improved.

[0121] Figure 2 is a schematic diagram of the architecture of a method for user screening in an optional embodiment of the present invention, Figure 3 FIG. 1 is a flow chart of a method for user screening in an optional embodiment of the present invention. Figure 2 and 3 As shown, in this embodiment of the present invention, user screening is performed using basic data such as user profiles, category tables (including the inclusion and conversion relationships between categories), behavior tables (i.e., tables showing user searches, browsing, clicks, and follow-up behaviors), and order tables (i.e., tables containing each user's orders). Based on this basic data, category associations, user preferences for each category, user operations within each category, user orders within each category, and TGI indicators for different groups are mined. A user screening request is a user screening task, which is broken down into subtasks based on category dimensions. Candidate user groups for each category are screened using user behavior data, user profiles, time decay, and category relationships. Based on user characteristics, category characteristics, and user behavior characteristics within each category, an estimation model is used to determine the first and second probabilities for each user in the candidate user group. The first and second probabilities are then merged to obtain a third probability that the user will trigger a preset operation in the target category in the future. All candidate users are ranked from highest to lowest according to the third probability, or the top several candidate users whose third probabilities exceed a set threshold are screened to obtain the target user group.

[0122] In actual application, the method of the embodiment of the present invention can be an offline system implemented by Spark (a computing engine) + hive (a data warehouse tool). The entire screening stage, except for user input, is completed in the offline stage. This makes it easier to use larger data sets and more complex algorithms to maximize the effect. The system is input from the outside in the task dimension. For the system inside, there may be repeated categories between multiple tasks. Therefore, the system deduplicates by category, generates category-based input, and the output is also category output. Finally, the categories associated with the task and the category results (target user groups) of each category obtained by internal calculations are used to generate the final task-based population package for external use.

[0123] It should be noted that the user historical behavior data mentioned in the embodiments of the present invention can be all the user's behavior data throughout history, or it can be behavior data within a specific time period. The granularity of historical behavior data can be customized, for example, it can be divided into granularities of 1 day, 2 days, 7 days, 14 days, etc.

[0124] The embodiment of the present invention can use the open source machine learning library Tensorflow to implement a user screening model. Tensorflow is an open source software library for machine learning developed by Google. It provides both low-level and high-level APIs. You can use the high-level API to quickly build a mature deep model, and you can choose to use the low-level API to flexibly build a deep learning network model. For the joint recommendation scenario of coupon products, since there is no ready-made model for use, you can choose to use a series of low-level APIs of Tensorflow to build a deep learning network model. After the user screening model is built, it usually takes a period of time to complete the training of the model. The training time is usually determined by the performance of the model itself, the complexity of the model, the hardware capabilities used to train the model, and the business scenario of the application model. In the scenario of joint recommendation of coupon products, considering the user behavior and the high frequency of coupon updates, the model can be trained once a day, and the data used for each training is historical data from several days before the current time. Model deployment

[0125] When deploying user-filtered models, since the system is discrete, all feature training and processing can rely on the BDP (Business Data Platform) platform. DNN models can be predicted based on Python Spark + TensorFlow batch prediction. XGBoost models can be predicted based on Python Spark + XGBoost packages.

[0126] In an embodiment of the present invention, by screening a candidate user group that has not triggered a preset operation behavior in a target category, new users of the target category can be screened; by determining a first probability that a user will trigger a preset operation behavior in the target category in the future, users with a high preference for the target category can be screened from new users; and by determining a second probability that a user will trigger a preset operation behavior in the target category through a target channel in the future, users with a high preference for the target channel can be screened from new users. The present invention can ensure a match between the screening logic and the screening target, so that the target user group obtained by screening has a high preference for the target category and target channel, thereby improving the user screening effect.

[0127] Category growth is a key area of ​​user growth, and attracting new category users is crucial for businesses. Existing technologies typically filter target category user groups based on user profiles, achieving user reach and exposure through advertising. However, advertising is generally billed using CPM (Cost Per Thousand, a unit of measure for the cost of reaching 1,000 people or "households" through a media or media scheduling schedule. It can be used to calculate any media, any demographic, and any total cost. It facilitates the comparison of the cost of one medium to another, or one media schedule to another. CPM is not the sole metric advertisers use to measure media; it is a relative metric often used to measure the value of a medium. CPC is a pay-per-click advertising method that charges based on the number of times an ad is clicked.) to maximize ROI (Return on Investment) and GMV (Gross Merchandise Volume). It does not directly model CAC (Customer Acquisition Cost, which is the cost of acquiring a new user. CAC is the total market-related expenditure divided by the total number of new users generated by the corresponding total expenditure). Therefore, it is inherently incompatible with marketing objectives.

[0128] By adopting the user screening method of the embodiment of the present invention to attract new category users, a set of category attracting models can be established for the same category or across categories and advertisements, to implement an efficient category attracting algorithm. The algorithm is highly matched with the marketing goal, and as many new category users as possible can be acquired at the same cost to maximize the new user attraction effect.

[0129] According to a second aspect of an embodiment of the present invention, a device for user screening is provided.

[0130] Figure 8 Schematic diagram of main modules of the apparatus for user screening according to an embodiment of the present invention, Figure 8As shown, the user screening device 800 includes:

[0131] The recall module 801 selects candidate user groups from all users who have not triggered a preset operation behavior in the target category;

[0132] The estimation module 802 determines a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future, and a second probability that any user will trigger the preset operation behavior in the target category through a target channel in the future, and determines a third probability that any user will trigger the preset operation behavior in the target category through a target channel in the future based on the first and second probabilities.

[0133] The fusion module 803 selects users whose third probability meets a preset condition from the candidate user group to obtain a target user group corresponding to the target category.

[0134] Optionally, the recall module screens candidate user groups that do not have preset operating behaviors in the target category, including:

[0135] The historical behavior data of all users are obtained, and users who have not triggered the preset operation behavior in the target category are filtered according to the historical behavior data to obtain the candidate user group.

[0136] Optionally, the recall module screens candidate user groups that do not have preset operating behaviors in the target category, including:

[0137] Divide all users into multiple target groups and obtain historical behavior data of each user in the target group; determine the TGI index of the target group based on the historical behavior data of each user in the target group, using the triggering of the preset operation behavior in the target category as a common feature, and select target groups with a TGI index greater than a preset TGI threshold as candidate target groups;

[0138] According to the historical behavior data of each user in the candidate target group, users who have not triggered the preset operation behavior in the target category are screened from the candidate target group to obtain the candidate user group.

[0139] Optionally, the recall module screens candidate user groups that do not have preset operating behaviors in the target category, including:

[0140] Obtaining historical behavior data of all users in multiple historical time periods, and determining the user's preference value for the target category in the corresponding historical time period based on the historical behavior data in each historical time period;

[0141] Accumulating the user's preference values ​​for the target category in each historical period in a time-decay manner to obtain a user preference index for the target category, and filtering users whose preference index is greater than a set threshold to obtain a candidate user set;

[0142] According to the historical behavior data of each user in the candidate user set, users who have not triggered the preset operation behavior in the target category are filtered from the candidate user set to obtain the candidate user group.

[0143] Optionally, the recall module screens candidate user groups that do not have preset operating behaviors in the target category, including:

[0144] Obtaining a user profile of each user, determining a fourth probability of the user triggering the preset operation behavior in the target category based on the user profile, screening users whose fourth probability is greater than a set threshold or several users with the highest fourth probability to obtain the candidate user set;

[0145] The historical behavior data of each user in the candidate user set is obtained, and users who have not triggered the preset operation behavior in the target category are filtered from the candidate user set according to the historical behavior data of each user in the candidate user set to obtain the candidate user group.

[0146] Optionally, the estimation module determines a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future, including:

[0147] Obtaining historical behavior data of any user, determining the probability of any user triggering the preset operation behavior in each category based on the historical behavior data of any user, and obtaining a first feature vector of any user;

[0148] Obtaining a user profile of the any user, determining a probability of the any user triggering the preset operation behavior in each category based on the user profile of the any user, and obtaining a second feature vector of the any user;

[0149] Determine a category association feature matrix corresponding to any one of the users based on the historical behavior data of the user;

[0150] Obtaining a category feature vector of the target category;

[0151] The first feature vector, the second feature vector, the category association feature matrix, and the category feature vector are input into a pre-trained first model to obtain a first probability that any user will trigger the preset operation behavior in the target category in the future.

[0152] Optionally, the estimation module determines a second probability that any user in the candidate user group will trigger the preset operation behavior in the target category through the target channel in the future, including:

[0153] Acquire historical behavior data of the any user, determine a triggering channel for each of the preset operation behaviors triggered by the any user in the target category based on the historical behavior data of the any user, and obtain a channel preference vector of the any user;

[0154] All channel preference vectors of any one of the users are input into a pre-trained second model to obtain a second probability that any one of the users will trigger the preset operation behavior through the target channel in the target category in the future.

[0155] Optionally, the fusion module determines a third probability that any one of the users will trigger the preset operation behavior in the target category through the target channel in the future based on the first probability and the second probability, including: taking the product of the first probability and the second probability as the third probability that any one of the users will trigger the preset operation behavior in the target category through the target channel in the future.

[0156] Optionally, the device also includes an input and output module, which is used to: receive a user screening request input by the user before the recall module screens the candidate user group that has not triggered the preset operation behavior in the target category from all users, and parse the target category from the user screening request; confirm that there is no target user group corresponding to the target category in the cache; if there is a target user group corresponding to the target category in the cache, obtain the target user group corresponding to the target category from the cache.

[0157] According to a third aspect of an embodiment of the present invention, there is provided an electronic device for user screening, comprising:

[0158] one or more processors;

[0159] a storage device for storing one or more programs,

[0160] When the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect of the embodiment of the present invention.

[0161] According to a fourth aspect of an embodiment of the present invention, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method provided by the first aspect of the embodiment of the present invention is implemented.

[0162] Figure 9 An exemplary system architecture 900 is shown to which the method for user screening or the apparatus for user screening according to an embodiment of the present invention may be applied.

[0163] like Figure 9 As shown, system architecture 900 may include terminal devices 901, 902, 903, a network 904, and a server 905. Network 904 is used to provide a medium for communication links between terminal devices 901, 902, 903 and server 905. Network 904 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0164] Users can use terminal devices 901, 902, and 903 to interact with server 905 via network 904 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 901, 902, and 903, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0165] The terminal devices 901 , 902 , and 903 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0166] Server 905 may be a server that provides various services, such as a backend management server (for example only) that supports shopping websites browsed by users using terminal devices 901, 902, and 903. The backend management server may analyze and process received data such as product information query requests, and feed back processing results (for example, target push information and product information—for example only) to the terminal device.

[0167] It should be noted that the user screening method provided in the embodiment of the present invention is generally executed by the server 905 , and accordingly, the user screening device is generally set in the server 905 .

[0168] It should be understood that Figure 9 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0169] Reference below Figure 10 , which shows a schematic structural diagram of a computer system 1000 of a terminal device suitable for implementing an embodiment of the present invention. Figure 10 The terminal device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0170] like Figure 10As shown, the computer system 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the system 1000 are also stored in the RAM 1003. The CPU 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0171] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.

[0172] In particular, according to the embodiments disclosed herein, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed herein include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009, and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the above-mentioned functions defined in the system of the present invention are performed.

[0173] It should be noted that the computer-readable medium described in the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wireline, optical fiber cable, RF, or any suitable combination thereof.

[0174] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0175] The modules involved in the embodiments of the present invention may be implemented in software or in hardware. The modules described may also be provided in a processor. For example, they may be described as follows: a processor including a recall module, an estimation module, and a fusion module. The names of these modules do not, in some cases, constitute limitations on the modules themselves. For example, the estimation module may also be described as a "module for screening candidate user groups from all users who have not triggered preset operational behaviors in the target category."

[0176] As another aspect, the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiment; or it may exist independently and not be assembled into the device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by a device, the device includes: screening a candidate user group that has not triggered a preset operation behavior in the target category from all users; determining a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future, and a second probability that the preset operation behavior will be triggered in the target category through a target channel in the future, and determining a third probability that any user will trigger the preset operation behavior in the target category through a target channel in the future based on the first probability and the second probability; screening users whose third probability meets the preset conditions from the candidate user group to obtain a target user group corresponding to the target category.

[0177] According to the technical solution of the embodiments of the present invention, by screening candidate user groups that have not triggered preset operational behaviors in the target category, it is possible to screen new users for the target category; by determining the first probability that a user will trigger a preset operational behavior in the target category in the future, it is possible to screen new users for users with a high preference for the target category; and by determining the second probability that a user will trigger a preset operational behavior in the target category through a target channel in the future, it is possible to screen new users for users with a high preference for the target channel. The present invention can ensure that the screening logic matches the screening target, so that the target user group obtained by screening has a high preference for the target category and target channel, thereby improving the user screening effect.

[0178] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for user screening, characterized in that: include: Filter candidate user groups from all users who have not triggered preset operational behaviors in the target category; Determining a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future, and a second probability that the preset operation behavior will be triggered in the target category through a target channel in the future; and determining a third probability that any user will trigger the preset operation behavior in the target category through a target channel in the future based on the first probability and the second probability, where the third probability is positively correlated with the first probability and the second probability; Users whose third probability meets a preset condition are screened from the candidate user group to obtain a target user group corresponding to the target category.

2. The method according to claim 1, wherein Screen candidate user groups that do not have preset operating behaviors in the target category, including: The historical behavior data of all users are obtained, and users who have not triggered the preset operation behavior in the target category are filtered according to the historical behavior data to obtain the candidate user group.

3. The method according to claim 1, wherein Screen candidate user groups that do not have preset operating behaviors in the target category, including: Divide all users into multiple target groups and obtain historical behavior data of each user in the target group; determine the TGI index of the target group based on the historical behavior data of each user in the target group, using the triggering of the preset operation behavior in the target category as a common feature, and select target groups with a TGI index greater than a preset TGI threshold as candidate target groups; According to the historical behavior data of each user in the candidate target group, users who have not triggered the preset operation behavior in the target category are screened from the candidate target group to obtain the candidate user group.

4. The method according to claim 1, wherein Screen candidate user groups that do not have preset operating behaviors in the target category, including: Obtaining historical behavior data of all users in multiple historical time periods, and determining the user's preference value for the target category in the corresponding historical time period based on the historical behavior data in each historical time period; Accumulating the user's preference values ​​for the target category in each historical period in a time-decay manner to obtain a user preference index for the target category, and filtering users whose preference index is greater than a set threshold to obtain a candidate user set; According to the historical behavior data of each user in the candidate user set, users who have not triggered the preset operation behavior in the target category are filtered from the candidate user set to obtain the candidate user group.

5. The method according to claim 1, wherein Screen candidate user groups that do not have preset operating behaviors in the target category, including: Obtaining a user profile of each user, determining a fourth probability of the user triggering the preset operation behavior in the target category based on the user profile, screening users whose fourth probability is greater than a set threshold or several users with the highest fourth probability to obtain the candidate user set; The historical behavior data of each user in the candidate user set is obtained, and users who have not triggered the preset operation behavior in the target category are filtered from the candidate user set according to the historical behavior data of each user in the candidate user set to obtain the candidate user group.

6. The method according to claim 1, wherein Determining a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future includes: Obtaining historical behavior data of any user, determining the probability of any user triggering the preset operation behavior in each category based on the historical behavior data of any user, and obtaining a first feature vector of any user; Obtaining a user profile of the any user, determining a probability of the any user triggering the preset operation behavior in each category based on the user profile of the any user, and obtaining a second feature vector of the any user; Determine a category association feature matrix corresponding to any one of the users based on the historical behavior data of the user; Obtaining a category feature vector of the target category; The first feature vector, the second feature vector, the category association feature matrix, and the category feature vector are input into a pre-trained first model to obtain a first probability that any user will trigger the preset operation behavior in the target category in the future.

7. The method according to claim 1, wherein Determining a second probability that any user in the candidate user group will trigger the preset operation behavior in the target category through the target channel in the future includes: Acquire historical behavior data of the any user, determine a triggering channel for each of the preset operation behaviors triggered by the any user in the target category based on the historical behavior data of the any user, and obtain a channel preference vector of the any user; All channel preference vectors of any one of the users are input into a pre-trained second model to obtain a second probability that any one of the users will trigger the preset operation behavior through the target channel in the target category in the future.

8. The method according to any one of claims 1 to 7, wherein: Determining a third probability that any one of the users will trigger the preset operation behavior in the target category through the target channel in the future based on the first probability and the second probability, including: taking the product of the first probability and the second probability as the third probability that any one of the users will trigger the preset operation behavior in the target category through the target channel in the future.

9. The method according to any one of claims 1 to 7, wherein: Before screening candidate user groups that have not triggered preset actions in the target category from all users, it also includes: Receive a user filtering request input by a user, and parse a target category from the user filtering request; confirm that a target user group corresponding to the target category does not exist in the cache; if a target user group corresponding to the target category exists in the cache, obtain the target user group corresponding to the target category from the cache.

10. A user screening device, characterized in that: include: The recall module screens candidate user groups from all users who have not triggered preset operation behaviors in the target category; an estimation module, determining a first probability that any user in the candidate user group will trigger the preset operation behavior in the target category in the future, and a second probability that the preset operation behavior will be triggered in the target category through a target channel in the future; The fusion module determines, based on the first probability and the second probability, a third probability that any of the users will trigger the preset operation behavior in the target category through the target channel in the future, where the third probability is positively correlated with the first probability and the second probability; and screens users whose third probability meets the preset conditions from the candidate user group to obtain a target user group corresponding to the target category.

11. An electronic device for user screening, characterized in that: include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 9.

12. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Target user account acquisition method and apparatus of network product

    CN107016569A

  • User screening method and device, computer readable storage medium and electronic equipment

    CN112085542A