Similar population expansion method and device, electronic equipment and readable storage medium
By determining the filtering algorithm in the target scenario, the original user sample is subjected to anomaly filtering and similar group expansion, which solves the problem that it is difficult to discover the inherent connection between users in the existing technology and improves the accuracy and effectiveness of similar group expansion.
Patent Information
- Application Number
- CN202110278343.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-15
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2041-03-15
AI Technical Summary
In existing technologies, manually selecting tags makes it difficult to uncover the inherent connections between users, and the quality of seed users affects the overall quality of similar user expansion.
By selecting a target scenario and determining the corresponding filtering algorithm, anomaly filtering is performed on the original user samples to obtain seed user samples. Similar groups are then expanded in the target scenario. Pre-defined behavioral features are used to detect the recall quality, and the filtering algorithm corresponding to the optimal recalled user samples is selected.
It improves the effectiveness of expanding similar user groups, reduces the impact of noisy users on the expansion results, and enhances the accuracy of the expanded user groups.
Smart Images

Figure CN115082844B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of Internet technology, and more specifically, to a method, apparatus, electronic device, and readable storage medium for expanding a similar user base. Background Technology
[0002] Look-alike audience expansion refers to the technology of finding people similar to or potentially related to a group of seed users through tagging rules or algorithmic models. Existing technologies mainly fall into two categories: selecting audiences through user profile tags and algorithmic models. The former selects audiences by manually choosing tags such as age and gender, while the latter outputs relevant audiences by inputting user profiles and behavioral characteristics into machine learning or deep learning models.
[0003] In realizing the concept disclosed herein, the inventors discovered at least the following problems in the related technology: Manually selecting tags makes it difficult to uncover the inherent connections between users, and also makes it difficult to utilize the potential characteristics of seed users. Furthermore, the quality of seed users also affects the overall quality of the similar users obtained through expansion. Summary of the Invention
[0004] In view of this, the present disclosure provides a method and apparatus for expanding similar populations.
[0005] One aspect of this disclosure provides a method for expanding a similar user group, comprising: selecting a target scenario; determining a filtering algorithm corresponding to the target scenario; performing anomaly filtering on an original user sample using the filtering algorithm to obtain a seed user sample; and expanding the seed user sample into a similar user group under the target scenario to obtain an expanded user sample.
[0006] According to embodiments of this disclosure, determining the filtering algorithm corresponding to the target scenario includes: obtaining a first user sample from a set of candidate user samples based on preset behavioral features; performing anomaly filtering on the first user sample using multiple preset algorithms to obtain multiple second user samples; in the target scenario, recalling each second user sample in the set of candidate user samples using a recall algorithm to obtain multiple recalled user samples; detecting the recall quality of each recalled user sample based on the preset behavioral features; and determining the filtering algorithm corresponding to the target scenario based on the recall quality, wherein the filtering algorithm is one of the multiple preset algorithms.
[0007] According to an embodiment of this disclosure, the step of detecting the recall quality of each recalled user sample based on the preset behavioral characteristics includes: counting the number of specified users in each recalled user sample, wherein the specified users have preset behavioral characteristics; comparing the number of all specified users to determine the optimal recalled user sample, wherein the optimal recalled user sample is the recalled user sample that contains the most specified users among all the recalled user samples.
[0008] According to embodiments of this disclosure, the step of detecting the recall quality of each recalled user sample based on the preset behavioral features further includes: counting the number of feature users other than the first user sample in the candidate user sample set, wherein the feature users are those with preset behavioral features; calculating the recall rate of each recalled user sample, wherein the recall rate is the ratio of the number of specified users to the number of feature users; comparing all the recall rates; and determining the optimal recalled user sample, wherein the optimal recalled user sample is the recalled user sample with the highest recall rate among all the recalled user samples.
[0009] According to an embodiment of this disclosure, determining the filtering algorithm corresponding to the target scenario based on the recall quality includes: determining a preset algorithm corresponding to the optimal recalled user sample; and determining the preset algorithm as the filtering algorithm corresponding to the target scenario.
[0010] According to embodiments of this disclosure, determining the filtering algorithm corresponding to the target scenario further includes: in the target scenario, recalling a first user sample within the candidate user sample set using a recall algorithm to obtain an original recalled user sample; detecting the original recall quality of the original recalled user sample based on the preset behavioral features; comparing the original recall quality with the recall quality of multiple recalled user samples; if the original recall quality is better than the recall quality of all recalled user samples, then there is no corresponding filtering algorithm for the target scenario.
[0011] According to embodiments of this disclosure, the user sample is a set of user feature data, which includes user profile information and user behavior information.
[0012] Another aspect of this disclosure provides a similar user group expansion device, comprising: a selection module for selecting a target scene; a determination module for determining a filtering algorithm corresponding to the target scene; a filtering module for performing anomaly filtering on the original user sample using the filtering algorithm to obtain a seed user sample; and an expansion module for expanding the seed user sample into a similar user group under the target scene to obtain an expanded user sample.
[0013] Another aspect of this disclosure provides an electronic device including one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method of any of the preceding claims.
[0014] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the method described above.
[0015] According to the embodiments of this disclosure, because a technical means of filtering seed users through a specified filtering algorithm to obtain an expanded population with preset behavioral characteristics in the target scenario is adopted, the technical problem of how to mitigate the impact of noisy users in the seed user sample on the expansion of similar populations is at least partially overcome, thereby achieving the technical effect of increasing the effectiveness of the expanded population. Attached Figure Description
[0016] The above and other objects, features, and advantages of this disclosure will become clearer from the following description of embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0017] Figure 1 An exemplary system architecture for which similar population expansion methods and apparatus of this disclosure can be applied is illustrated schematically;
[0018] Figure 2 A flowchart illustrating a similar population expansion method according to an embodiment of the present disclosure is shown schematically;
[0019] Figure 3A A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically;
[0020] Figure 3B A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically;
[0021] Figure 3C A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically;
[0022] Figure 3D A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically;
[0023] Figure 3E A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically;
[0024] Figure 4 A block diagram schematically illustrates a similar crowd expansion device according to an embodiment of the present disclosure; and
[0025] Figure 5 A block diagram of an electronic device suitable for implementing a similar population expansion device according to an embodiment of the present disclosure is shown schematically. Detailed Implementation
[0026] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0027] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0028] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0029] When using expressions such as "at least one of A, B, and C," the expression should generally be interpreted in accordance with the meaning commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, systems having A alone, having B alone, having C alone, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). Similarly, when using expressions such as "at least one of A, B, or C," the expression should generally be interpreted in accordance with the meaning commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, or C" should include, but is not limited to, systems having A alone, having B alone, having C alone, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0030] Embodiments of this disclosure provide a method and apparatus for expanding a similar user group. The method includes: selecting a target scenario and determining a filtering algorithm corresponding to the target scenario using historical user behavior data; filtering original user samples for anomalies using the filtering algorithm to obtain seed user samples; and expanding the seed user samples into a similar user group within the target scenario to obtain expanded user samples.
[0031] Figure 1 An exemplary system architecture 100, illustrating similar population expansion methods and apparatus according to embodiments of this disclosure, is shown schematically. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0032] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0033] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0034] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0035] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0036] It should be noted that the similar audience expansion method provided in this disclosure embodiment can generally be executed by server 105. Correspondingly, the similar audience expansion device provided in this disclosure embodiment can generally be located in server 105. The similar audience expansion method provided in this disclosure embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the similar audience expansion device provided in this disclosure embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Alternatively, the similar audience expansion method provided in this disclosure embodiment can also be executed by terminal devices 101, 102, or 103, or by other terminal devices different from terminal devices 101, 102, or 103. Accordingly, the similar population expansion device provided in this embodiment of the present disclosure can also be set in terminal device 101, 102, or 103, or in other terminal devices different from terminal device 101, 102, or 103.
[0037] For example, user samples may originally be stored in any one of terminal devices 101, 102, or 103 (e.g., terminal device 101, but not limited thereto), or stored on an external storage device and imported into terminal device 101. Then, terminal device 101 may locally execute the similar population expansion method provided in this disclosure embodiment, or send the user samples to other terminal devices, servers, or server clusters, and have the other terminal devices, servers, or server clusters receiving the user samples execute the similar population expansion method provided in this disclosure embodiment.
[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0039] Figure 2 A flowchart illustrating a similar population expansion method according to an embodiment of the present disclosure is shown schematically.
[0040] like Figure 2 As shown, the method includes operations S201 to S204.
[0041] In operation S201, select the target scene.
[0042] Specifically, a target scenario can be understood as a business marketing scenario for a merchant, that is, a sales scenario for a particular product category. For example, a target scenario could be a merchant conducting promotional activities for daily necessities or for baby and maternity products. Product categories include, but are not limited to, 3C products, jewelry, clothing, and food.
[0043] In operation S202, the filtering algorithm corresponding to the target scene is determined.
[0044] In operation S203, the original user samples are filtered for anomalies using a filtering algorithm to obtain seed user samples.
[0045] During the expansion of the similar user population, the original user sample can be provided by external partners or by an internal database. Using the original user sample as a reference sample during this process yields user samples with the same behavioral characteristics as the original sample. However, the quality of the original user sample cannot be guaranteed; that is, the users included in the original sample may not necessarily possess the behavioral characteristics anticipated by the developers. If the original user sample is not filtered for anomalies and is directly used for similar user expansion, the overall quality of the recalled similar users may be reduced due to the influence of some abnormal user samples.
[0046] Meanwhile, some biases may occur during the acquisition of the original user samples, resulting in poor sample quality. For example, the system may determine that it needs to acquire user samples with the characteristic of having browsed laptops in a target scenario related to the 3C product category, but errors may occur during the acquisition process, resulting in the acquisition of user samples containing users with the characteristic of having only browsed clothing. To ensure the quality of the user samples, a filtering algorithm is needed to filter the original user samples.
[0047] It's important to note that different filtering algorithms have varying degrees of adaptability to different business scenarios. Different filtering algorithms excel in different business domains. Therefore, after developers select a target scenario from among many options, they need to choose the filtering algorithm that best suits that target scenario to ensure that the resulting seed user sample can be expanded to a more accurate target audience.
[0048] Specifically, filtering algorithms include, but are not limited to, outlier detection algorithms, isolated forest algorithms, and autoencoder algorithms.
[0049] Local Outlier Factor (LOF) is a density-based anomaly detection method. It calculates an outlier factor for each data point and uses this outlier factor to determine whether the data point is an outlier. If the LOF is greater than or equal to 1, then the data point is considered an outlier.
[0050] Isolation Forest is a learning algorithm that integrates multiple isolated classification trees. This algorithm has the advantage of not requiring calculations based on density and distance metrics, resulting in low computational complexity. Furthermore, since the computation is based on ensemble learning, its time complexity is linear. Each classification tree is generated independently, facilitating deployment in a distributed system.
[0051] An autoencoder is an unsupervised learning model that generates a low-dimensional representation from a high-dimensional input via a neural network. However, because only the most informative features are retained during training, outliers cannot be well restored when the decoder reconstructs the original data, leading to relatively large errors.
[0052] In operation S204, under the target scenario, the seed user sample is expanded to include similar users to obtain an expanded user sample.
[0053] In this embodiment, the user sample is a collection of user feature data. In actual use, the user feature data includes user profile information and user behavior information. User profile information includes the user's personal attribute information, such as age, gender, occupation, family members, personal hobbies, etc. User behavior information is the user's business behavior information within a certain period of time, such as browsing, clicking, and purchasing of items.
[0054] Typically, e-commerce platforms need to make reasonable predictions about a user's future business behavior based on specific user profile and behavioral information, and provide corresponding business services to improve the user experience, such as providing "recommendations for you" services.
[0055] Specifically, e-commerce platforms obtain a user's mobile phone number, IMEI number, and IFA number, and bind these information to the user's account, using them as unique identifiers. The e-commerce platform can then send relevant messages to the user's account and collect characteristic information related to that account, using this collected information as the user's feature data.
[0056] Through the embodiments of this disclosure, a corresponding target scenario is selected according to different needs, and a filtering algorithm that is most suitable for the target scenario is determined. Anomaly filtering is performed on the original user sample to filter out noisy users that do not meet the conditions in the original user sample, thus obtaining a seed user sample. Under the target scenario, the seed user sample is expanded to a similar group to obtain a similar group with the same behavioral characteristics as the seed user sample, thereby achieving the technical effect of improving the effectiveness of the expanded data.
[0057] The following is for reference. Figures 3A to 3D In conjunction with specific embodiments, Figure 2 The method shown will be further explained.
[0058] Figure 3A A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically.
[0059] like Figure 3A As shown, determining the filtering algorithm corresponding to the target scene also includes operations S301 to S305.
[0060] In operation S301, the first user sample is obtained from the candidate user sample set based on preset behavioral characteristics.
[0061] The candidate user sample set is the entire user set. Specifically, it can be the entire user set consisting of all users in the database, or it can be a user set consisting of a portion of user data. Throughout the process of determining the correspondence between the target scenario and the filtering algorithm, the selection and recall of user samples are both completed within the candidate user sample set. That is, all users, whether in the acquired user samples or the recalled user samples, are users from the candidate user sample set.
[0062] The preset behavioral characteristics can specifically include behaviors such as browsing products, adding products to the shopping cart, and purchasing products. Users in the first user sample possess at least one of these preset behavioral characteristics. For example, in a target scenario for 3C products, the first user sample could consist entirely of users with the characteristic of browsing laptops, entirely of users with the characteristic of purchasing laptops, or a combination of users with both browsing and purchasing laptops characteristics. It could also consist of users with both browsing and adding laptops to their cart, or a combination of users with both browsing and adding laptops to their cart, and users who only possess the characteristic of browsing laptops. This application does not limit the specific composition of the user characteristic types in the first user sample.
[0063] The first user sample consists of a certain number of users with preset behavioral characteristics selected from the candidate sample set. However, due to possible anomalies during the selection process, not all users in the actual first user sample have the preset behavioral characteristics.
[0064] In operation S302, multiple preset algorithms are used to filter out anomalies in the first user sample to obtain multiple second user samples.
[0065] Multiple preset algorithms are included, but are not limited to, outlier detection algorithms, isolated forest algorithms, and autoencoder algorithms. After anomaly filtering, second user samples corresponding to each filtering algorithm are obtained.
[0066] In operation S303, under the target scenario, the recall algorithm is used to recall each second user sample in the candidate user sample set to obtain multiple recalled user samples.
[0067] In this embodiment, the second user sample can be understood as a seed user sample, and the recall algorithm is specifically the Faiss recall algorithm. In the Faiss recall algorithm, the seed user sample is represented by a user embedding vector trained by combining user profiles with recent behavioral features. Specifically, the Faiss recall algorithm determines the N users closest to each seed user in the recall vector space. N is determined by the developers based on the actual situation, and N is the final number of users recalled.
[0068] In operation S304, based on preset behavioral characteristics, the recall quality of each recalled user sample is detected.
[0069] In actual recall results, not every user in the recalled user sample possesses the preset behavioral characteristics. Therefore, it is necessary to test the recall quality of each recalled user sample to ensure the optimal filtering algorithm corresponding to the target scenario is determined.
[0070] In operation S305, based on the recall quality, the filtering algorithm corresponding to the target scenario is determined. The filtering algorithm is one of several preset algorithms.
[0071] Practice shows that the distribution characteristics of embedded representation vectors generated from user profiles and behavior sequences vary depending on the behavior type (e.g., browsing or adding to cart) and the target product (e.g., electronics or daily necessities). Existing technologies only use general outlier detection methods to filter noisy users and cannot filter based on specific behavioral characteristics and target products. Given the heterogeneity of user embedded vectors, simply using general methods to remove anomalous seed users may lead to consequences that deviate from business objectives.
[0072] Therefore, determining the correlation between the target scenario and the preset algorithm reflects the filtering capabilities of different filtering algorithms in different feature domains. This disclosure provides a specific method for determining the correspondence between filtering algorithms and behavioral features. Through periodic offline training, before expanding to similar user groups, the preset algorithm corresponding to the most retrieved user samples with the best recall quality is selected as the filtering algorithm for the target scenario. That is, the filtering algorithm with the best fit is determined from multiple preset algorithms. Determining the correspondence between specific business scenarios and filtering algorithms can improve the accuracy of expansion.
[0073] Figure 3B A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically.
[0074] like Figure 3B As shown, based on preset behavioral characteristics, the recall quality of each recalled user sample is detected, including operations S306 to S307.
[0075] In operation S306, the number of specified users in each recalled user sample is counted, and the specified users have preset behavioral characteristics.
[0076] In operation S307, the number of all specified users is compared to determine the optimal recall user sample. The optimal recall user sample is the recall user sample that contains the most specified users among all recall user samples.
[0077] In this embodiment of the disclosure, the recall quality of the recalled user sample is determined by counting the number of users who meet the preset conditions in the recalled user sample.
[0078] For example, given a target scenario involving 3C products, a first user sample A is obtained from the candidate sample set. This first user sample A consists of users who browse laptops. Different filtering algorithms (outlier detection algorithm, isolated forest algorithm, and autoencoder algorithm) are used to filter the first user sample A, resulting in different second user samples, denoted as S1, S2, and S3. Within the target scenario of the 3C product category, similar user groups are expanded from the second user samples S1, S2, and S3 to obtain recall user samples R1, R2, and R3. The users in the recall user samples R1, R2, and R3 possess at least one of the following characteristics: browsing laptops, adding laptops to their cart, and purchasing laptops.
[0079] Specifically, in the recall vector space, the 100 users closest to the second user sample S1 are identified as the recall user sample R1, the 100 users closest to the second user sample S2 are identified as the recall user sample R2, and the 100 users closest to the second user sample S3 are identified as the recall user sample R3. Ideally, all users in the recall user samples R1, R2, and R3 possess at least one of the characteristics of browsing laptops, adding laptops to their cart, and purchasing laptops. For example, the recall user samples R1, R2, and R3 may include users who only have the characteristic of browsing laptops, users who only have the characteristic of purchasing laptops, users who have both the characteristics of browsing laptops and adding laptops to their cart, or users who have all three characteristics of browsing laptops, adding laptops to their cart, and purchasing laptops.
[0080] In the actual recall results, some users in the recalled user samples R1, R2, and R3 did not possess any of the characteristics of browsing laptops, adding laptops to their cart, or purchasing laptops. Users in the recalled user samples R1, R2, and R3 who possessed at least one of these characteristics were selected as eligible users, i.e., designated users. The number of eligible users was counted, and the recall quality of each recalled user sample was determined based on the number of eligible users.
[0081] Among them, the recall user sample with the most eligible users is determined as the optimal recall user sample, that is, the recall quality of this recall user sample is the best.
[0082] It should be noted that this application does not impose specific limitations on the relationship between the behavioral characteristics of the first user sample and the behavioral characteristics of the recalled user sample, but only needs to ensure that the behavioral characteristics of the first user sample and the behavioral characteristics of the recalled user sample belong to the same target scenario.
[0083] For example, in the above example, the user in the first user sample A only has the characteristic of browsing laptops, while the users in the recalled user samples R1, R2, and R3 can have at least one of the characteristics of browsing laptops, adding laptops to their cart, and purchasing laptops. Alternatively, the user in the first user sample A only has the characteristic of browsing laptops, and the users in the recalled user samples R1, R2, and R3 also only have the characteristic of browsing laptops. Or, the user in the first user sample A has the characteristic of browsing laptops, adding laptops to their cart, and purchasing laptops, and the users in the recalled user samples R1, R2, and R3 can also have at least one of these characteristics. Furthermore, the user in the first user sample A only has the characteristic of browsing laptops, and the users in the recalled user samples R1, R2, and R3 can have at least one of the characteristics of browsing mobile phones, adding mobile phones to their cart, and purchasing mobile phones. Browsing laptops, browsing mobile phones, browsing headphones, browsing speakers, etc., are all preset behavioral characteristics under the 3C product category. Examples of the relationship between other behavioral characteristics of the first user sample and the behavioral characteristics of the recalled user samples are similar to the above examples and will not be repeated here.
[0084] Specifically, Figure 3C A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically.
[0085] like Figure 3C As shown, for each preset feature, the recall quality of each recalled user sample is detected, including operations S308 to S310.
[0086] In operation S308, count the number of characteristic users in the candidate user sample set other than the first user sample. Characteristic users are those with preset behavioral characteristics.
[0087] In operation S309, the recall rate of each recalled user sample is calculated. The recall rate is the ratio of the number of specified users to the number of characteristic users.
[0088] In operation S310, all recall rates are compared to determine the optimal recalled user sample, which is the recalled user sample with the highest recall rate among all recalled user samples.
[0089] In this embodiment of the disclosure, the selection and recall of user samples are both completed within the candidate user sample set during the entire process of determining the correspondence between the target scene and the filtering algorithm. That is, all users in both the acquired user samples and the recalled user samples are users from the candidate user sample set. When the number of users with preset behavioral characteristics in the candidate user sample set is small, the number of qualified specified users in the recalled user samples will decrease accordingly.
[0090] When the number of eligible users in the recalled user sample is too small, the reason may be that there are recall anomalies in the recall process, or that the number of users with the preset behavioral characteristics in the candidate user sample set is too small. To avoid the number of eligible users in the recalled user sample being too small due to recall anomalies, the recall quality can be further evaluated using the recall rate.
[0091] Recall is the ratio of the number of specified users to the number of characteristic users. Characteristic users are those in the candidate user sample set, excluding the first user sample, who possess predefined behavioral characteristics. We can consider the set of users with predefined behavioral characteristics in the candidate user sample set as set M. A predetermined number of users are selected from set M to form the first user sample A, and the remaining users in set M form the test user sample B. Understandably, during the recall process, the users recalled through the second user sample do not include users already included in the second user sample itself, and the recall of user samples is completed within the candidate user sample set. Therefore, the users in the recalled user samples are all users from the candidate user sample set. Thus, the characteristic users recalled in the recalled user samples are the users in the test user sample B.
[0092] Therefore, if the number of users in the first user sample is much larger than the number of users in the test user sample, and users with preset behavioral characteristics in the first user sample are not removed when calculating the number of users in the statistical feature, the calculated recall rate may be extremely small.
[0093] Specifically, the number of users in the test user sample B is N.B The number of characteristic users in the recalled user sample R is N. (R∩B) Finally, the recall rate is calculated.
[0094] in,
[0095] The recalled user sample with the highest recall rate is the optimal recalled user sample, which has the best recall quality.
[0096] It should be noted that this application does not specifically limit the relationship between the behavioral characteristics of the first user sample, the behavioral characteristics of the featured users, and the behavioral characteristics of the recalled user samples in this embodiment. It is sufficient to ensure that the behavioral characteristics of the first user sample, the behavioral characteristics of the featured users, and the behavioral characteristics of the recalled user samples belong to the same target scenario. Those skilled in the art can choose according to actual needs. Specific correspondences have been exemplified above and will not be repeated here. However, it should be emphasized that, to ensure consistency, the types of behavioral characteristics of the featured users are the same as the types of behavioral characteristics of the recalled user samples. For example, if the user in the first user sample A only has the characteristic of browsing laptops, and the users selected for the recalled user samples R1, R2, and R3 also only have the characteristic of browsing laptops, then the statistically analyzed featured users should also be users who only have the characteristic of browsing laptops.
[0097] Figure 3D A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically.
[0098] like Figure 3D As shown, the filtering algorithm for determining the target scenario based on recall quality includes operations S311 to S313.
[0099] In operation S311, the preset algorithm corresponding to the optimal recall user sample is determined.
[0100] In operation S312, the preset algorithm is determined to be the filtering algorithm corresponding to the target scene.
[0101] In operations S307 and S310, optimal recalled user samples with the best recall quality are obtained through different evaluation criteria. If the recall quality can be accurately assessed in operation S307 based on the specified number of recalled user samples, the optimal recalled user samples obtained in operation S307 can determine the corresponding preset algorithm, further determining the filtering algorithm corresponding to the target scenario. If the recall quality cannot be accurately assessed in operation S307 based on the specified number of recalled user samples, the optimal recalled user samples are further determined in operation S310 by calculating the recall rate, further determining the filtering algorithm corresponding to the target scenario, i.e., the most suitable filtering algorithm.
[0102] Figure 3E A flowchart illustrating a similar population expansion method according to another embodiment of this disclosure is shown schematically.
[0103] like Figure 3E As shown, determining the filtering algorithm corresponding to the target scene also includes operations S313 to S316.
[0104] In operation S313, under the target scenario, the first user sample is recalled from the candidate user sample set through the recall algorithm to obtain the original recalled user sample.
[0105] In operation S314, the original recall quality of the original recalled user samples is detected based on preset behavioral characteristics.
[0106] In operation S315, the original recall quality is compared with the recall quality of multiple recalled user samples.
[0107] In operation S316, if the original recall quality is better than the recall quality of all recalled user samples, then there is no corresponding filtering algorithm for the target scenario.
[0108] To further compare the impact of filtering algorithms on recall quality, the first user sample without filtering is also recalled to obtain the original recalled user sample, and its recall quality is tested. This recall quality is used as a reference value to compare with the recall quality of the user samples obtained after filtering. In actual filtering, the selected preset algorithm may not be suitable for the target scenario, and filtering the first user sample might actually lower its quality. Therefore, the original recalled user sample is used as a reference. If the original recall quality is better than the recall quality of all recalled user samples obtained after filtering, it is determined that there is no corresponding filtering algorithm for the target scenario. That is, in actual similar population expansion, filtering the original user sample is not necessary in this target scenario.
[0109] It should also be noted that this disclosure only provides an example of selecting a specific business scenario as the target scenario for similar audience expansion. However, in actual applications, business scenarios are diverse. To ensure the accuracy of similar audience expansion results for each business scenario, it is necessary to determine the same evaluation criteria to identify the optimal filtering algorithm for each business scenario. The most crucial evaluation criterion is to ensure that the relationship between the behavioral characteristics of the first user sample, the behavioral characteristics of the characteristic user, and the behavioral characteristics of the recalled user sample is identical in each business scenario.
[0110] For example, in determining recall quality, the relationship between the behavioral characteristics of the first user sample, the behavioral characteristics of the featured users, and the behavioral characteristics of the recalled user samples is as follows:
[0111] In the business scenario of the 3C product category: the user of the first user sample A only has the characteristic of browsing laptops, and the users of the recalled user samples R1, R2 and R3, as well as the characteristic users, also only have the characteristic of browsing laptops.
[0112] In the business scenario of the apparel category: the user of the first user sample A only has the characteristic of browsing women's clothing, and the users of the recalled user samples R1, R2 and R3, as well as the characteristic users, also only have the characteristic of browsing women's clothing.
[0113] In the business scenario of the maternal and infant product category: the user of the first user sample A only has the characteristic of browsing milk powder, and the users of the recalled user samples R1, R2 and R3, as well as the characteristic users, also only have the characteristic of browsing milk powder.
[0114] In other words, when determining the filtering algorithms corresponding to different business scenarios, the standards for horizontal comparison should be consistent.
[0115] Figure 4 A block diagram of a similar crowd expansion device according to an embodiment of the present disclosure is shown schematically.
[0116] like Figure 4 As shown, the similar population expansion device 400 includes a selection module 410, a determination module 420, a filtering module 430, and an expansion module 440.
[0117] Select module 410 to select the target scene;
[0118] The determination module 420 is used to determine the filtering algorithm corresponding to the target scene;
[0119] The filtering module 430 is used to perform anomaly filtering on the original user samples using a filtering algorithm to obtain seed user samples.
[0120] The extension module 440 is used to expand the seed user sample to a similar user group in the target scenario, so as to obtain an extended user sample.
[0121] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as Field Programmable Gate Arrays (FPGAs), Programmable Logic Arrays (PLAs), Systems-on-Chip, Systems-on-Substrate, Systems-on-Package, Application-Specific Integrated Circuits (ASICs), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0122] For example, any plurality of the selection module 410, the receiving module 420, the filtering module 430, and the expansion module 440 may be combined into one module / unit / subunit, or any one of these modules / units / subunits may be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits may be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of the present disclosure, at least one of the selection module 410, the receiving module 420, the filtering module 430, and the expansion module 440 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in any one of the three methods of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the selected module 410, the receiving module 420, the filtering module 430, and the extension module 440 may be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0123] It should be noted that the similar population expansion device part in the embodiments of this disclosure corresponds to the similar population expansion method part in the embodiments of this disclosure. For a detailed description of the similar population expansion device part, please refer to the similar population expansion method part, which will not be repeated here.
[0124] Figure 5A block diagram of an electronic device suitable for implementing the methods described above, according to embodiments of the present disclosure, is illustrated schematically. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0125] like Figure 5 As shown, an electronic device 500 according to an embodiment of the present disclosure includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0126] RAM 503 stores various programs and data required for the operation of system 500. Processor 501, ROM 502, and RAM 503 are interconnected via bus 504. Processor 501 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 502 and / or RAM 503. It should be noted that the programs may also be stored in one or more memories other than ROM 502 and RAM 503. Processor 501 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0127] According to embodiments of this disclosure, system 500 may further include an input / output (I / O) interface 505, which is also connected to bus 504. System 500 may also include one or more of the following components connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. A removable medium 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 510 as needed so that computer programs read from it can be installed into storage section 508 as needed.
[0128] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by processor 501, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0129] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0130] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0131] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than ROM 502 and RAM 503.
[0132] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0133] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0134] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A method for expanding a similar user base, comprising: Select the target scene; Determine the filtering algorithm corresponding to the target scene; The original user samples are filtered for anomalies using the filtering algorithm to obtain seed user samples. In the target scenario, the seed user sample is expanded to include similar user groups to obtain an expanded user sample. The filtering algorithm for determining the target scene includes: Based on preset behavioral characteristics, the first user sample is obtained from the candidate user sample set; Multiple second user samples are obtained by filtering out anomalies from the first user sample using multiple preset algorithms. In the target scenario, a recall algorithm is used to recall each second user sample in the candidate user sample set to obtain multiple recalled user samples. Based on the preset behavioral characteristics, the recall quality of each recalled user sample is detected respectively; Based on the recall quality, a filtering algorithm corresponding to the target scene is determined, wherein the filtering algorithm is one of the plurality of preset algorithms. The step of detecting the recall quality of each recalled user sample based on the preset behavioral features includes: The number of characteristic users other than the first user sample in the candidate user sample set is counted. The characteristic users are those with preset behavioral characteristics. Calculate the recall rate for each of the recalled user samples, where the recall rate is the ratio of the number of specified users to the number of characteristic users, wherein the specified users are users with preset behavioral characteristics in each of the recalled user samples; Compare all the recall rates to determine the optimal recalled user sample, which is the recalled user sample with the highest recall rate among all the recalled user samples.
2. The method according to claim 1, wherein, The step of detecting the recall quality of each recalled user sample based on the preset behavioral features also includes: Count the number of the specified users in each of the recalled user samples; Compare the number of all specified users to determine the optimal recall user sample, which is the recall user sample that contains the most specified users among all the recall user samples.
3. The method according to claim 1 or 2, wherein, The filtering algorithm for determining the target scenario based on the recall quality includes: Determine the preset algorithm corresponding to the optimal recalled user sample; The preset algorithm is determined to be the filtering algorithm corresponding to the target scene.
4. The method according to claim 1, wherein, The filtering algorithm for determining the target scene further includes: In the target scenario, the first user sample is recalled from the candidate user sample set using a recall algorithm to obtain the original recalled user sample. Based on the preset behavioral characteristics, the original recall quality of the original recalled user samples is detected; Compare the original recall quality with the recall quality of multiple recalled user samples; If the original recall quality is better than the recall quality of all recalled user samples, then there is no corresponding filtering algorithm for the target scenario.
5. The method according to claim 1, wherein, The user sample is a collection of user feature data, which includes user profile information and user behavior information.
6. A similar crowd expansion device, comprising: Select a module to select the target scenario; A determination module is used to determine the filtering algorithm corresponding to the target scene; The filtering module is used to perform anomaly filtering on the original user samples using the filtering algorithm to obtain seed user samples; The extension module is used to expand the seed user sample to include similar user groups in the target scenario, thereby obtaining an expanded user sample. The determining module includes the following sub-modules: The acquisition submodule is used to acquire the first user sample from the candidate user sample set based on preset behavioral characteristics; The filtering submodule is used to perform anomaly filtering on the first user sample using multiple preset algorithms to obtain multiple second user samples. The recall submodule is used to recall each second user sample in the candidate user sample set according to the recall algorithm in the target scenario, so as to obtain multiple recalled user samples. The detection submodule is used to detect the recall quality of each of the recalled user samples based on the preset behavioral characteristics. The determination submodule is used to determine the filtering algorithm corresponding to the target scene based on the recall quality, wherein the filtering algorithm is one of the plurality of preset algorithms. The detection submodule includes the following units: The statistics unit is used to count the number of characteristic users other than the first user sample in the candidate user sample set, wherein the characteristic users are those with preset behavioral characteristics; A calculation unit is used to calculate the recall rate of each of the recalled user samples, wherein the recall rate is the ratio of the number of specified users to the number of characteristic users, and wherein the specified users are users with preset behavioral characteristics in each of the recalled user samples; The comparison unit is used to compare all the recall rates and determine the optimal recalled user sample, which is the recalled user sample with the highest recall rate among all the recalled user samples.
7. An electronic device, comprising: One or more processors; Memory, used to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 5.
8. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Method for detecting abnormal user, electronic equipment and computer readable medium
CN110929799A
Method and device for determining on-call rate of recall strategy, equipment and storage medium
CN111444438A
Similar crowd acquisition method based on browsing behavior optimization
CN112445985A