Label information determination method and device, electronic equipment and medium
By calculating the correlation coefficient between user characteristics and core business metrics, benchmark metrics are selected and binned, solving the problem of poor accuracy of tag information in existing technologies and achieving more accurate user segmentation and business optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2025-12-19
- Publication Date
- 2026-05-08
AI Technical Summary
In existing technologies, user tag information mainly relies on the experience and subjective judgment of business personnel, resulting in poor accuracy of tag information.
By acquiring multiple user characteristics and core business metrics, the correlation coefficient between the characteristics and core business metrics is calculated, a preset number of benchmark metrics are selected, and tag information is determined in combination with a preset binning method.
It improves the rationality and accuracy of tag information, supports precision marketing and personalized recommendations, and maximizes user lifetime value.
Smart Images

Figure CN121998674A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to the fields of data statistical analysis and big data, and particularly to a method, apparatus, electronic device and medium for determining tag information. Background Technology
[0002] In the current e-commerce environment, the user base has reached hundreds of millions, and user behavior data is experiencing explosive growth.
[0003] Traditional extensive operation models can no longer meet the demands of market competition; refined user segmentation has become key to enhancing a platform's core competitiveness. Effective user segmentation enables businesses to deeply understand the characteristics and needs of different user groups, providing a core basis for precision marketing, personalized recommendations, and product optimization, ultimately maximizing user lifetime value. Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, and medium for determining tag information.
[0005] According to one aspect of this disclosure, a method for determining label information is provided, comprising:
[0006] To obtain multiple characteristics and core business metrics of each user across multiple users;
[0007] Based on multiple characteristics and core business indicators of each user among the multiple users, obtain the correlation coefficient between each of the characteristics and the core business indicators;
[0008] Based on the correlation coefficient between each of the aforementioned features and the core business indicators, a preset number of benchmark indicators are obtained from the multiple features of the user.
[0009] Based on the preset number of benchmark indicators and the preset binning method, the label information is determined.
[0010] According to another aspect of this disclosure, a device for determining label information is provided, comprising:
[0011] The feature acquisition module is used to acquire multiple features and core business metrics of each user among multiple users.
[0012] The coefficient acquisition module is used to acquire the correlation coefficient between each feature and the core business indicator based on multiple features and core business indicators of each user among the multiple users.
[0013] The indicator acquisition module is used to acquire a preset number of benchmark indicators from multiple features of the user based on the correlation coefficient between each feature and the core business indicator.
[0014] The scheme determination module is used to determine the tag information based on the preset number of benchmark indicators and the preset binning method.
[0015] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described above and any possible implementations.
[0019] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described above and any possible implementation thereof.
[0020] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the aspects and any possible implementations described above.
[0021] According to the technology disclosed herein, the rationality and accuracy of determined label information can be effectively improved.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Detailed Implementation
[0023] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0024] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;
[0025] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;
[0026] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;
[0027] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0028] Figure 5 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure.
[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0030] Obviously, the described embodiments are only some, not all, of the embodiments disclosed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0031] It should be noted that the terminal devices involved in the embodiments of this disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, personal computers, televisions, and other devices with display functions.
[0032] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0033] The existing user tagging information mainly relies on the experience and subjective judgment of business personnel, resulting in poor accuracy of the tagging information.
[0034] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure; as shown Figure 1 As shown, this embodiment provides a method for determining label information, which may specifically include the following steps:
[0035] S101. Obtain multiple characteristics and core business metrics of each user among multiple users;
[0036] In this embodiment, the core business metric can refer to an important indicator within the business to which the defined tag information is applied. The core business metric can be deeply aligned with the business stakeholders, and should be a North Star indicator capable of measuring business health and strategic objectives.
[0037] In this embodiment, a user's multiple characteristics can cover various aspects such as browsing and transactions. Since in real-world applications, user characteristics and core business metrics are often correlated with time, in this embodiment, obtaining each user's multiple characteristics refers to the values of those characteristics at the current moment—that is, the values of the user's multiple characteristics at the moment the tag information determination begins. The user's core business metrics also refer to the values of the user's core business metrics at the current moment.
[0038] S102. Based on multiple characteristics and core business indicators of each user among multiple users, obtain the correlation coefficient between each characteristic and the core business indicators.
[0039] Since different features represent information from different dimensions, some features will inevitably have a strong correlation with core business metrics, while others will have a weak correlation. To more objectively characterize the relationship between each feature and the core business metrics, in this embodiment, for each feature, the correlation coefficient between the feature and the core business metrics can be obtained based on the values of that feature from multiple users and the values of the core business metrics from multiple users.
[0040] S103. Based on the correlation coefficient between each feature and the core business indicators, obtain a preset number of benchmark indicators from multiple features.
[0041] The preset quantity in this embodiment can be set based on experience or needs, for example, 1, 2, 3 or 4.
[0042] S104. Determine the label information based on the preset number of benchmark indicators and the preset binning method.
[0043] After determining the preset number of baseline indicators in this embodiment, and combining them with the preset binning method, the tag information can be determined. The tag information obtained in this embodiment is used to tag users, thereby enabling user classification. The preset binning method in this embodiment can be based on quantile segmentation or natural breakpoint method. For example, the preset binning method in this embodiment can be divided into two bins, three bins, or four bins, and the specific preset binning method can be set according to needs and experience.
[0044] In this embodiment, the preset bin division method can be to divide the material into three bins using the quantile division method, such as dividing it into three equal parts: 0-33%, 34-66%, and 67-100%. If it is divided into two bins, it can be divided into two equal parts: 0-50% and 51-100%.
[0045] Based on the above, it can be understood that the label information in this embodiment is not a fixed label, but rather a group of labels including labels from multiple intervals. For example, if only one benchmark indicator is included, and the preset binning method divides the data into two bins, the resulting label information may include two labels, such as a label whose benchmark indicator value is located in the first interval and a label whose benchmark indicator value is located in the second interval. If two benchmark indicators are included, and the preset binning method still divides the data into two bins, the resulting label information may include four labels, such as a label whose first benchmark indicator value is located in the first interval and the second benchmark indicator value is located in the third interval, a label whose first benchmark indicator value is located in the first interval and the second benchmark indicator value is located in the fourth interval, a label whose second benchmark indicator value is located in the first interval and the second benchmark indicator value is located in the third interval, and a label whose second benchmark indicator value is located in the first interval and the second benchmark indicator value is located in the fourth interval. Similarly, for other numbers of benchmark indicators and other numbers of bins set by the preset binning method, the number of labels included in the label information can be determined according to a similar combination principle, and each type of label can be clearly identified.
[0046] The method for determining tag information in this embodiment can obtain a preset number of benchmark indicators from multiple features based on the correlation coefficient between each feature and the core business indicators. Then, based on the preset number of benchmark indicators and the preset binning method, the tag information can be determined. By referring to the core business indicators, an accurate and effective preset number of benchmark indicators can be selected. Then, combined with the preset binning method, the tag information can be determined accurately and effectively.
[0047] Moreover, in the process of determining tag information, a preset number of benchmark indicators can be selected based on core business indicators. This can effectively improve the correlation between the preset number of benchmark indicators and core business indicators, thereby ensuring the rationality and accuracy of the obtained preset number of benchmark indicators. This can further improve the rationality and accuracy of the tag information determined based on the preset number of benchmark indicators and the preset binning method, providing an effective foundation for subsequent user segmentation.
[0048] Figure 2 This is a schematic diagram based on the second embodiment of the present disclosure; as shown Figure 2 As shown, the method for determining label information in this embodiment is based on the above... Figure 1 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be described in further detail. For example... Figure 2 As shown, the method for determining label information in this embodiment may specifically include the following steps:
[0049] S201. Obtain multiple characteristics and core business metrics of each user among multiple users;
[0050] For example, in this embodiment, taking the business as an e-commerce platform as an example, the multiple characteristics of each user may include at least one of the following: each user's browsing characteristics, each user's transaction characteristics, and each user's main site characteristics.
[0051] For example, for each user, the user's browsing characteristics may include at least one of the following: the number of days the user watched live streams within a preset time period before the current moment, the duration of the live streams watched, the number of live streams watched, the intended category, the number of product categories viewed, the number of products viewed, and the product card click rate.
[0052] In this embodiment, the number of days a user watched the live stream within a preset time period prior to the current moment represents the number of days the user watched the live stream within that preset time period. The total viewing time within that preset time period represents the total duration the user watched the live stream within that preset time period. The total number of times a user watched the live stream within that preset time period represents the total number of times the user watched the live stream within that preset time period, where each time the user enters and exits the live stream is counted as one time. The user's intended categories within that preset time period represent the number of product categories the user intends to purchase within that preset time period; for example, adding a product to the shopping cart indicates an intention to purchase that product within that category. The number of product categories browsed within that preset time period represents the number of product categories the user browsed within that preset time period. The total number of products browsed within that preset time period represents the total number of products browsed within that preset time period. The click-through rate of product cards within a preset time period before the current moment represents the ratio of the number of product cards clicked by the user within the preset time period before the current moment to the total number of product cards recommended to the user during the corresponding time period.
[0053] In this embodiment, the preset time period length can be set according to actual needs, such as 14 days, 7 days, or 21 days, etc., and can be set according to needs or experience. Optionally, the user browsing features in this embodiment may include multiple features corresponding to multiple preset time period lengths prior to the current moment.
[0054] For example, for each user, the user's transaction characteristics include the total transaction amount paid by the user within a preset time period before the current moment, the total number of orders, the number of days since the order was placed, the average order value, the proportion of orders subsidized by the platform, the refund order rate, the concentration of purchased product categories, and cross-category purchasing behavior characteristics.
[0055] The following metrics are defined as follows: Average Order Value (AOV) within a preset time period prior to the current moment: This represents the ratio of the user's total transaction amount to the total number of orders within that time period. Platform Subsidized Order Share within a Preset Time Period Prior to the Current Moment: This represents the ratio of orders placed using platform subsidy coupons within that time period to the total number of orders within that time period. Refund Order Rate within a Preset Time Period Prior to the Current Moment: This represents the number of refunded orders within that time period, divided by the ratio to the total number of orders within that time period. Concentration of Product Categories Purchased by the User within a Preset Time Period Prior to the Current Moment: This indicates the number of product categories purchased by the user within that time period. Cross-Category Purchase Behavior Characteristics of the User within a Preset Time Period Prior to the Current Moment: This indicates the number of product categories purchased by the user within that time period.
[0056] For example, for each user, the user's main site characteristics include at least one of the following: the user's activity level on the main site at the current moment, the user's usage depth, the number of days logged in within a preset time period before the current moment, the number of consecutive check-in days, the application push open rate, the number of times the user posts product reviews, the number of times the user shares their purchase experience, the number of times the user participates in platform activities, and the number of new users successfully invited.
[0057] A user's current activity level on the main site indicates their activity level on the main site at that moment. User usage depth is a tag configured by the main site based on user usage, which can include first-level, second-level, and third-level users. The app push notification open rate within a preset time period prior to the current moment can be equal to the number of app push notifications received and opened by the user within that preset time period, divided by the total number of all app push notifications received by the user within that preset time period.
[0058] Define the core business metrics for the current phase. These metrics should be the North Star indicator, measuring the health of the business and strategic objectives, for example,
[0059] Taking e-commerce platforms as an example, the core business indicators selected to measure the health of the business and strategic goals could be the total gross merchandise volume (GMV) of users within a future preset time period (such as 30 days), the user retention rate in the following month, or the penetration rate of specific new product categories.
[0060] In this embodiment, the technical solution of this disclosure is described using the aforementioned features as an example. In practical applications, features of other types of users can be obtained based on different business needs. This embodiment provides rich user features, which provides necessary support for obtaining more accurate benchmark indicators in the future, thereby further improving the accuracy of the determined tag information.
[0061] S202. Determine the type of value for each feature;
[0062] Specifically, we analyze whether the values of the features are numerical or textual.
[0063] S203. Determine the type of correlation coefficient based on the type of values of each feature;
[0064] For example, if the value of a feature is numerical, the type of the correlation coefficient is determined to be the Pearson correlation coefficient; the Pearson correlation coefficient is applicable to continuous and approximately normally distributed variables and measures linear correlation.
[0065] If the feature values are text-based, the correlation coefficient should be determined as the Kendall rank correlation coefficient. The Spearman rank correlation coefficient is suitable for ordinal data or data that does not follow a normal distribution; it measures monotonic correlations and is more robust.
[0066] When necessary, certain features can be specified to use Kendall's rank correlation coefficient, which is a non-parametric correlation measurement method suitable for small datasets or situations with a large number of duplicate values.
[0067] S204. Based on multiple characteristics and core business indicators of each user among multiple users, analyze the correlation coefficients of the corresponding types;
[0068] Specifically, for each feature, the correlation coefficient for that feature type can be analyzed based on multiple users' data on that feature and their core business metrics. Therefore, it can be considered that the correlation coefficient for each feature is obtained by analyzing the correlation between the feature variable and the core business metric variable. For the analysis of different types of correlation coefficients, refer to the analysis methods of Pearson correlation coefficient, Spearman's rank correlation coefficient, and Kendall's rank correlation coefficient; these will not be elaborated upon here. Thus, different types of correlation coefficients can be used for different features.
[0069] S205. From multiple user characteristics, obtain a preset number of features whose correlation coefficient with core business indicators is greater than a preset coefficient threshold and whose confidence level also reaches a preset confidence threshold, and use these features as a preset number of benchmark indicators.
[0070] In this embodiment, the confidence level corresponding to each feature is used to indicate the degree of confidence that the correlation coefficient between the feature and the core business indicator is greater than a preset coefficient threshold. The preset coefficient threshold can be set based on experience or requirements, and can be a value greater than 0 and less than 1, such as 0.5 or 0.6.
[0071] In practice, we can first check whether the correlation coefficient between each feature and the core business indicator is greater than a preset threshold. Then, we can filter out the data whose correlation coefficient between the feature and the core business indicator is greater than the preset threshold. Furthermore, we can use statistical methods to calculate the confidence level of the event that the correlation coefficient between the feature and the core business indicator is greater than the preset threshold. In this embodiment, we can set the confidence level to be represented by a p-value, and the smaller the confidence level, the higher the confidence level. Therefore, we can set a small preset confidence threshold, such as 0.01. At this time, the confidence level of the feature reaches the preset confidence threshold, that is, the confidence level of the feature is less than or equal to the threshold confidence level. This means that the confidence level of the feature's correlation coefficient with the core business indicator is high.
[0072] During the specific screening process, you can first select features from all features whose correlation coefficient is greater than a preset threshold and whose confidence level is less than or equal to a preset confidence level threshold. Then, sort them according to their correlation coefficient from largest to smallest to obtain the top N features. For example, N can be 5, 6, or other integer values. Next, select a preset number of features from these as benchmark indicators. The preset number should be less than N, such as the first two. You can temporarily retain N features here for future reference. If the preset number of benchmark indicators proves ineffective, you can return to this point and reselect a preset number of benchmark indicators from the N features. For example, if the first and second ranked features were initially selected as two benchmark indicators, but subsequent validation failed, you can then select the first and third ranked features, or the second and third ranked features, as two benchmark indicators for further validation.
[0073] In this embodiment, the preset number can be 2 or 3. In actual application scenarios, it should not exceed 4, because too many preset numbers will lead to excessive user dispersion, lack of clear hierarchical meaning, and difficulty in effective verification.
[0074] By using this method to obtain a preset number of benchmark indicators, the accuracy of the obtained benchmark indicators can be effectively improved.
[0075] For example, in one application scenario, the analysis results show that the number of items added to cart in the past 7 days (correlation coefficient 0.72) and the number of different product categories viewed in the past 15 days (correlation coefficient 0.61) are highly positively correlated with GMV in the next 30 days; while the refund rate in the past 30 days (correlation coefficient -0.55) is highly negatively correlated. These benchmark indicators can be selected as a preset number of benchmark indicators to determine the tag information for this application.
[0076] S206. Based on each benchmark indicator, dynamically bucket the test user dataset according to the preset bucketing method to obtain the bucketing results;
[0077] The test user dataset in this embodiment is a collection of user data including a preset number of benchmark indicators, collected from the historical behavior data of multiple users; that is, the test user dataset includes multiple user data entries, and each user data entry includes a preset number of benchmark indicators. Each user data entry is collected from the user's historical behavior data.
[0078] The preset bucketing method in this embodiment can achieve two-level or three-level bucketing. Theoretically, it is not recommended to implement more than four levels of bucketing, because the more levels of bucketing there are, the more user layers there are, and too many user layers will defeat the purpose of layering.
[0079] This step does not use fixed, empirical binning rules, but rather dynamically bins user data based on the distribution of selected benchmark metrics. For example, for the number of items added to cart in the past 7 days, a quantile-based partitioning method can be used, such as dividing into three equal parts: 0-33%, 34-66%, and 67-100%. Alternatively, a natural breakpoint method can be used to divide users into three bins: "low add-to-cart intention," "medium add-to-cart intention," and "high add-to-cart intention." This dynamic binning operation is performed independently for each benchmark metric.
[0080] S207. Combine the binning results of each benchmark indicator to obtain label information based on a preset number of benchmark indicators and a preset binning method.
[0081] By combining the binning results of multiple benchmark metrics, a preliminary multi-dimensional user segmentation is formed, enabling the stratification of user data in the test user dataset. Taking three benchmark metrics, each dynamically divided into three bins, as an example, after combining the binning results, the test users in the test dataset can be divided into nine intervals, meaning the resulting tag information includes nine tags. In this embodiment, the determined tag information is used to identify users and achieve user classification.
[0082] For example, the three bins derived from one benchmark indicator could be "Add-to-Cart Intention in the First Zone," "Add-to-Cart Intention in the Second Zone," and "High Add-to-Cart Intention in the Third Zone." Similarly, the three bins derived from another benchmark indicator could be "Refund Rate in the First Zone," "Refund Rate in the Second Zone," and "Refund Rate in the Third Zone." The combined tagging information could include the following nine tags: a tag consisting of "Add-to-Cart Intention in the First Zone" and "Refund Rate in the First Zone," a tag consisting of "Add-to-Cart Intention in the First Zone" and "Refund Rate in the Second Zone," and so on. The labels consist of "Add-to-Cart Intention" and "Refund Rate in the Third Zone", "Add-to-Cart Intention in the Second Zone" and "Refund Rate in the First Zone", "Add-to-Cart Intention in the Second Zone" and "Refund Rate in the Second Zone", "Add-to-Cart Intention in the Second Zone" and "Refund Rate in the Third Zone", "Add-to-Cart Intention in the Third Zone" and "Refund Rate in the First Zone", "Add-to-Cart Intention in the Third Zone" and "Refund Rate in the Second Zone", and "Add-to-Cart Intention in the Third Zone" and "Refund Rate in the Third Zone".
[0083] S208. Based on the benchmark indicators of each user's data, verify and determine the validity of the tag information based on the preset number of benchmark indicators and the preset binning method.
[0084] For example, in this embodiment, this step may specifically include the following steps:
[0085] (1) Perform unsupervised clustering of each benchmark indicator of each user's data;
[0086] To verify the scientific validity and effectiveness of the above-mentioned hierarchical scheme, the benchmark indicators in each user's data can be used as features and input into unsupervised clustering algorithms such as K-Means or Density-Based Spatial Clustering of Applications with Noise to achieve unsupervised clustering of each benchmark indicator in each user's data.
[0087] (2) Obtain the variance of the core business metrics of the test users in each cluster and the overall variance of the test user dataset;
[0088] (3) Based on the variance of the core business indicators of the test users in each cluster and the overall variance of the test user dataset, verify and determine the validity of the label information corresponding to the clusters;
[0089] After clustering, the focus is on examining whether the core business metrics of user data within the same cluster exhibit significant concentration. Specifically, this involves calculating the variance of the core business metrics within each cluster. If the variance within a cluster is significantly less than the overall variance of the test user dataset, the stratification is considered effective, indicating a high degree of homogeneity among user data within the same stratum. Conversely, the selection of benchmark metrics needs to be re-evaluated.
[0090] Specifically, if the relative value of the variance of the core business indicators of the test users in each cluster to the overall variance of the test user dataset is greater than a preset relative difference threshold, such as 5%, or other percentages, then the label information corresponding to that cluster is considered valid.
[0091] For example, for a certain cluster, the variance within the cluster is 0.94, and the overall variance of the test user dataset is 1.00. Then the relative value is 0.94 / 1.00-1=-6%. The relative value is greater than the preset relative difference threshold of 5%, so the variance within the cluster is considered to be significantly less than the overall variance of the test user dataset, which proves that the user stratification is effective.
[0092] (4) If the label information of all clusters is valid, determine that the label information based on the preset number of benchmark indicators and the preset binning method is valid.
[0093] If, after verification, the label information for all clusters is valid, the label information determined based on the preset number of benchmark indicators and the preset binning method can be considered valid. Otherwise, if any label information is invalid, the process returns to reselecting benchmark indicators; if reselecting benchmark indicators is still invalid, the core business indicators need to be redefined, and the benchmark indicators are reselected according to the above process in this embodiment until the label information is verified to be valid.
[0094] In this embodiment, by obtaining the binning results of each benchmark indicator according to a preset binning method, and combining the binning results of each benchmark indicator, label information based on a preset number of benchmark indicators and a preset binning method can be obtained, which can effectively improve the rationality and accuracy of the label information.
[0095] Furthermore, by adding the above verification, this embodiment can further improve the accuracy of the determined label information.
[0096] Based on this, the tag information can be determined. In practical applications, in order to verify whether the tag information is effective, corresponding operational strategies can be designed and implemented based on each user segment through actual online business, and A / B group experiments can be used to verify whether the tag information is effective. If it is effective, it can be implemented. If it is ineffective, the tag information can be re-determined according to the technical solution of this embodiment.
[0097] S209. Obtain the values of each benchmark indicator for the target user;
[0098] S210. Classify the target users based on the values of each benchmark indicator and the determined tag information.
[0099] Steps S209 and S210 are application implementations for classifying users based on the determined tag information. Specifically, users are classified based on the values of various benchmark indicators of the users, referring to the determined tag information. Since this embodiment can obtain more reasonable and accurate tag information, the accuracy of user classification can be effectively improved when classifying users based on this tag information. This, in turn, enables more accurate business recommendations based on user classification, effectively improving the efficiency of business recommendation.
[0100] The method for determining tag information in this embodiment can directly locate features that are strongly related to core business objectives as benchmark indicators, enabling operational strategies to more accurately target user groups and their key behavioral nodes that are most sensitive to business results, thereby greatly improving the return on marketing investment and user satisfaction.
[0101] The method for determining label information in this embodiment can refer to core business indicators, select an accurate and effective preset number of benchmark indicators, and then combine them with a preset binning method to accurately and effectively determine label information.
[0102] In this embodiment, the type of correlation coefficient can be determined based on the value type of each feature, which can more accurately obtain the correlation coefficient between each feature and the core business indicators, and thus obtain more accurate and effective benchmark indicators.
[0103] In this embodiment, the validity of the tag information can be verified based on various benchmark indicators of each user's data, thereby further improving the accuracy and validity of the tag information.
[0104] Figure 3 This is a schematic diagram based on the third embodiment of this disclosure; as shown Figure 3 As shown, this embodiment provides a tag information determination device 300, including:
[0105] The feature acquisition module 301 is used to acquire multiple features and core business indicators of each user among multiple users;
[0106] The coefficient acquisition module 302 is used to acquire the correlation coefficient between each feature and the core business indicator based on multiple features and core business indicators of each user among the multiple users.
[0107] The indicator acquisition module 303 is used to acquire a preset number of benchmark indicators from the multiple features based on the correlation coefficient between each feature and the core business indicator.
[0108] The scheme determination module 304 is used to determine the tag information based on the preset number of benchmark indicators and the preset binning method.
[0109] The label information determination device 300 in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0110] Figure 4 This is a schematic diagram based on the fourth embodiment of the present disclosure; as shown Figure 4 As shown, the label information determining device 400 of this embodiment, in the above-described... Figure 3 Based on the technical solutions shown, the technical solutions of this disclosure will be described in further detail. For example... Figure 4 As shown, the label information determining device 400 of this embodiment includes the above-described... Figure 3 The modules with the same name and function shown are: feature acquisition module 401, coefficient acquisition module 402, index acquisition module 403, and scheme determination module 404.
[0111] like Figure 4 As shown, in this embodiment, the scheme determination module 404 includes:
[0112] Bucketing processing unit 4041 is used to dynamically bucket the test user dataset according to a preset bucketing method based on each of the benchmark indicators to obtain the bucketing result; the test user dataset is a set of user data including a preset number of benchmark indicators collected from the historical behavior data of multiple users.
[0113] User stratification unit 4042 is used to combine the binning results of each benchmark indicator to stratify the user data in the test user dataset;
[0114] Verification unit 4043 is used to verify and determine the validity of tag information based on the preset number of benchmark indicators and the preset binning method based on the benchmark indicators of each user data.
[0115] Further, optionally, in one embodiment of this disclosure, the verification unit 4043 is used for:
[0116] Unsupervised clustering is performed on each of the benchmark metrics of each of the user data.
[0117] Obtain the variance of the core business metrics of each user data in each cluster and the overall variance of the test user dataset;
[0118] Based on the variance of the core business indicators of each user data in each cluster and the overall variance of the test user dataset, the validity of the label information corresponding to the cluster is verified and determined.
[0119] If the label information of all clusters is valid, then the label information based on the preset number of benchmark indicators and the preset binning method is determined to be valid.
[0120] Further optionally, in one embodiment of this disclosure, the multiple characteristics of each user include at least one of the user's browsing characteristics, user's transaction characteristics, and user's main site characteristics;
[0121] The user's browsing characteristics include at least one of the following: the number of days the user watched live streams within a preset time period prior to the current moment, the duration of live stream viewing, the number of live stream viewings, the intended category, the number of product categories viewed, the number of products viewed, and the product card click rate.
[0122] The user's transaction characteristics include the total transaction amount paid by the user within a preset time period before the current moment, the total number of orders, the number of days since the order was placed, the average order value, the proportion of orders subsidized by the platform, the refund order rate, the concentration of purchased product categories, and cross-category purchasing behavior characteristics.
[0123] The user's main site characteristics include at least one of the following: the user's activity level on the main site at the current moment, the user's usage depth, the number of days logged in within a preset time period before the current moment, the number of consecutive check-in days, the application push open rate, the number of times the user posts product reviews, the number of times the user shares their purchase experience, the number of times the user participates in platform activities, and the number of new users successfully invited.
[0124] Further optionally, in one embodiment of this disclosure, the coefficient acquisition module 402 is configured to:
[0125] Determine the type of the value for each of the aforementioned features;
[0126] The type of correlation coefficient is determined based on the type of the values of each of the aforementioned features;
[0127] Based on the multiple characteristics and core business indicators of each user among the multiple users, the correlation coefficients of the corresponding types are analyzed.
[0128] Further optionally, in one embodiment of this disclosure, the coefficient acquisition module 402 is configured to:
[0129] If the value of the feature is of numerical type, the type of the correlation coefficient is determined to be the Pearson correlation coefficient.
[0130] If the value of the feature is of text type, then the type of the correlation coefficient is determined to be Kendall's rank correlation coefficient.
[0131] Further optionally, in one embodiment of this disclosure, the indicator acquisition module 403 is used for:
[0132] From the plurality of features, a preset number of features are selected that have a correlation coefficient greater than a preset coefficient threshold and a confidence level that also reaches a preset confidence level threshold. These features are used as the preset number of benchmark indicators. The confidence level corresponding to each feature is used to identify the degree of confidence that the correlation coefficient between the feature and the core business indicator is greater than the preset coefficient threshold.
[0133] Further optional, such as Figure 5 As shown, in one embodiment of this disclosure, the tag information determining device 400 further includes a layering module 405:
[0134] The feature acquisition module 4041 is also used to acquire the values of each of the benchmark indicators of the target user;
[0135] The stratification module 405 is used to stratify the target user based on the values of each of the target user's benchmark indicators and the preset bucketing method.
[0136] The label information determination device 400 in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0137] The acquisition, storage, and application of any type of information, such as user personal information, involved in the technical solutions disclosed herein comply with relevant laws and regulations and do not violate public order and good morals.
[0138] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0139] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0140] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0141] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0142] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the methods of this disclosure. For example, in some embodiments, the methods of this disclosure may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods of this disclosure described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the methods of this disclosure by any other suitable means (e.g., by means of firmware).
[0143] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0144] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0146] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0147] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0148] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0149] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0150] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for determining tag information, comprising: To obtain multiple characteristics and core business metrics of each user across multiple users; Based on multiple characteristics and core business indicators of each user among the multiple users, obtain the correlation coefficient between each of the characteristics and the core business indicators; Based on the correlation coefficient between each of the aforementioned features and the core business indicators, a preset number of benchmark indicators are obtained from the multiple features. Based on the preset number of benchmark indicators and the preset binning method, the label information is determined.
2. The method according to claim 1, wherein, Based on the preset quantity of benchmark indicators and the preset binning method, the label information is determined, including: Based on the aforementioned benchmark metrics, the test user dataset is dynamically bucketed according to a preset bucketing method to obtain the bucketing results; the test user dataset is a set of user data including a preset number of benchmark metrics collected from the historical behavior data of multiple users. The binning results of each benchmark indicator are combined to obtain tag information based on the preset number of benchmark indicators and the preset binning method; Based on the benchmark indicators of each user data, verify and determine the validity of the tag information based on the preset number of benchmark indicators and the preset binning method.
3. The method according to claim 2, wherein, Based on the benchmark metrics of each user data, verify and determine the validity of the tag information based on the preset number of benchmark metrics and the preset bucketing method, including: Unsupervised clustering is performed on each of the benchmark metrics of each of the user data. Obtain the variance of the core business metrics of each user data in each cluster and the overall variance of the test user dataset; Based on the variance of the core business indicators of each user data in each cluster and the overall variance of the test user dataset, the validity of the label information corresponding to the cluster is verified and determined. If the label information of all clusters is valid, then the label information based on the preset number of benchmark indicators and the preset binning method is determined to be valid.
4. The method according to claim 1, wherein, The multiple characteristics of each user include at least one of the following: user browsing characteristics, user transaction characteristics, and user main site characteristics; The user's browsing characteristics include at least one of the following: the number of days the user watched live streams within a preset time period prior to the current moment, the duration of live stream viewing, the number of live stream viewings, the intended category, the number of product categories viewed, the number of products viewed, and the product card click rate. The user's transaction characteristics include the total transaction amount paid by the user within a preset time period before the current moment, the total number of orders, the number of days since the order was placed, the average order value, the proportion of orders subsidized by the platform, the refund order rate, the concentration of purchased product categories, and cross-category purchasing behavior characteristics. The user's main site characteristics include at least one of the following: the user's activity level on the main site at the current moment, the user's usage depth, the number of days logged in within a preset time period before the current moment, the number of consecutive check-in days, the application push open rate, the number of times the user posts product reviews, the number of times the user shares their purchase experience, the number of times the user participates in platform activities, and the number of new users successfully invited.
5. The method according to claim 1, wherein, Based on multiple characteristics and core business metrics of each user among the multiple users, obtain the correlation coefficient between each characteristic and the core business metrics, including: Determine the type of the value for each of the aforementioned features; The type of correlation coefficient is determined based on the type of the values of each of the aforementioned features; Based on the multiple characteristics and core business indicators of each user among the multiple users, the correlation coefficients of the corresponding types are analyzed.
6. The method according to claim 5, wherein, Based on the value type of each of the aforementioned features, the type of the correlation coefficient is determined, including: If the value of the feature is of numerical type, the type of the correlation coefficient is determined to be the Pearson correlation coefficient. If the value of the feature is of text type, then the type of the correlation coefficient is determined to be Kendall's rank correlation coefficient.
7. The method according to claim 1, wherein, Based on the correlation coefficients between each of the aforementioned features and the core business indicators, a preset number of benchmark indicators are obtained from the plurality of features, including: From the plurality of features, a preset number of features are selected that have a correlation coefficient greater than a preset coefficient threshold and a confidence level that also reaches a preset confidence level threshold. These features are used as the preset number of benchmark indicators. The confidence level corresponding to each feature is used to identify the degree of confidence that the correlation coefficient between the feature and the core business indicator is greater than the preset coefficient threshold.
8. The method according to any one of claims 1-7, wherein, The method further includes: Obtain the values of each of the target user's benchmark metrics; The target users are classified based on the values of each of the benchmark indicators and the determined tag information.
9. A device for determining tag information, comprising: The feature acquisition module is used to acquire multiple features and core business metrics of each user among multiple users. The coefficient acquisition module is used to acquire the correlation coefficient between each feature and the core business indicator based on multiple features and core business indicators of each user among the multiple users. The indicator acquisition module is used to acquire a preset number of benchmark indicators from multiple features of the user based on the correlation coefficient between each feature and the core business indicator. The scheme determination module is used to determine the tag information based on the preset number of benchmark indicators and the preset binning method.
10. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-8.
11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.
12. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.