A big data analysis driven push notification service system
By leveraging big data analytics and knowledge graph construction, the accuracy and timing issues of product discount message pushes in existing technologies have been resolved, enabling personalized and proactive pushes that improve user experience and conversion rates.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- KECHUANGTONG CHENGDU CO LTD
- Filing Date
- 2026-01-26
- Publication Date
- 2026-04-17
AI Technical Summary
Existing product discount message push technology lacks accuracy and cannot personalize pushes based on users' dynamic information and consumption habits, resulting in low click-through rates and conversion rates. Furthermore, inappropriate push timing affects user experience and the efficiency of enterprise resource utilization.
Through big data analysis, multi-dimensional user information is collected to construct a knowledge graph. User clustering is performed using a two-layer update clustering algorithm and an improved butterfly optimization algorithm. Combined with users' shopping and dynamic operation information, target users and push time are accurately located to achieve proactive push.
It improves the accuracy and effectiveness of push notifications, enhances user experience and conversion rates, ensures that push content is highly matched with user interests, and delivers information at the optimal time, thereby increasing user satisfaction and loyalty to the service system.
Smart Images

Figure CN121581942B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data analysis, and in particular to a proactive push notification service system driven by big data analysis. Background Technology
[0002] In today's booming e-commerce industry, proactively pushing product discount messages has become a crucial marketing tool for attracting consumers and boosting sales. However, existing product discount message push technologies have serious shortcomings, significantly impacting push effectiveness and user experience.
[0003] In terms of push notification accuracy, traditional systems mostly rely on users' historical purchase records for simple matching and push notifications. This single-dimensional data utilization method completely ignores dynamic information such as users' current browsing behavior, search keywords, and social interactions. For example, if a user has previously purchased a mobile phone, a traditional system will continuously push various mobile phone offers, unaware that the user may now be more interested in mobile phone accessories or have already upgraded their purchase. Moreover, for newly registered users, due to the lack of historical data, the system often has to push notifications randomly, resulting in content that is far removed from the user's potential needs. Relevant data shows that the click-through rate of product promotion messages using traditional push methods is generally less than 5%, and the conversion rate is even less than 2%. A large number of ineffective pushes not only waste the marketing resources of enterprises but also easily cause user annoyance. In terms of push notification timing, existing technology struggles to accurately determine the timing of push notifications based on users' daily activity patterns and consumption habits. For example, pushing promotional messages during users' busy work hours means users simply don't have time to check them; while pushing them during users' rest periods or when their shopping intentions are strong fails to occur in a timely manner.
[0004] Therefore, there is a need to provide a proactive push notification service system driven by big data analytics to improve the accuracy and effectiveness of information delivery. Summary of the Invention
[0005] This invention provides a big data analytics-driven proactive push notification service system, comprising: a data collection module for collecting multi-dimensional information and historical notification operation information from multiple users, wherein the multi-dimensional information includes at least basic information, shopping information, and dynamic operation information; a graph construction module for clustering multiple sampled users based on the multi-dimensional information and historical notification operation information of multiple users to determine multiple clusters, and for each cluster, constructing a first knowledge graph, a second knowledge graph, and cluster features corresponding to the cluster, wherein the first knowledge graph is used to characterize the correlation between the interest scores and push notification click volumes of multiple target products corresponding to the cluster, and the second knowledge graph is used to characterize the correlation between multiple push time periods corresponding to the cluster; a message acquisition module for acquiring the current push message; and a proactive push module for determining the target user and target push time period of the current push message based on the first knowledge graph, second knowledge graph, and cluster features corresponding to each cluster, and proactively pushing the message based on the target user and target push time period.
[0006] Furthermore, the data collection module is used to: determine multiple background factors; collect background information of multiple users based on the multiple background factors; filter multiple target background factors from the multiple background factors based on the background information, shopping information, dynamic operation information and historical notification operation information of multiple users; and collect basic information of multiple users based on the multiple target background factors.
[0007] Furthermore, the data collection module is used to: for each background factor, based on the background information of multiple users, select multiple sampled users corresponding to the background factor from the multiple users; based on the factor value of the background factor corresponding to each sampled user, divide the multiple sampled users corresponding to the background factor into multiple sampled user groups; based on the historical notification operation information of each sampled user group, calculate the operation difference value within the group and the operation difference value outside the group, and determine the first weight of the background factor; for each background factor, based on the shopping information and dynamic operation information of multiple sampled users, calculate the interest similarity between any two sampled users, and calculate the difference between the background factor corresponding to any two sampled users; based on the interest similarity between any two sampled users and the difference between the corresponding background factor, determine the second weight of the background factor; and based on the first weight and the second weight of each background factor, select multiple target background factors from the multiple background factors.
[0008] Furthermore, the graph construction module is used to: for each user, calculate the user's interest value for each type of product based on the user's shopping information, dynamic operation information, and historical notification operation information; calculate the basic information similarity between any two users based on each user's basic information; calculate the interest similarity between any two users based on each user's interest value for each type of product; and cluster multiple sampled users using a two-layer update clustering algorithm based on the basic information similarity and interest similarity between any two users to determine multiple clusters.
[0009] Further, the graph construction module is used for: S11, calculating the clustering distance between any two users based on the basic information similarity and interest similarity between any two users; S12, selecting multiple users as cluster centers based on the clustering distance between any two users; S13, performing clustering based on the multiple cluster centers to generate clustering results; S14, merging the multiple cluster centers based on the clustering results and updating the clustering results; S15, updating the multiple cluster centers based on the updated clustering results using an improved butterfly optimization algorithm, wherein the switching probability of the improved butterfly optimization algorithm is determined based on the clustering distance between any two clusters corresponding to the updated clustering results, and the fitness function is related to the mean clustering distance corresponding to each cluster; S16, determining whether clustering is complete based on the merged multiple cluster centers and the updated multiple cluster centers; if yes, outputting the clustering results; otherwise, executing S13.
[0010] Furthermore, the graph construction module is used to: calculate the interest value of each product corresponding to the cluster based on the interest values of users for each type of product included in the cluster; determine multiple target products corresponding to the cluster and the interest score of each target product based on the interest of each product corresponding to the cluster, wherein the cluster features include multiple target products corresponding to the cluster; calculate the correlation of push notification click volume between any two target products corresponding to the cluster based on the historical notification operation information of users included in the cluster; and construct a first knowledge graph based on the correlation between the multiple target products corresponding to the cluster, the interest score of each target product, and the push notification click volume between any two target products.
[0011] Furthermore, the graph construction module is used to: determine multiple push time periods; calculate the push reading ratio of the cluster in each push time period corresponding to multiple historical periods based on the historical notification operation information of multiple users included in the cluster, determine the correlation of multiple push time periods corresponding to the cluster, and construct a second knowledge graph.
[0012] Furthermore, the graph construction module is used to: acquire shopping information, dynamic operation information, and notification operation information of multiple users in the current period, calculate the user's interest value for each type of product in the current period, and dynamically update the first knowledge graph and the second knowledge graph based on the user's interest value for each type of product in the current period.
[0013] Furthermore, the proactive push module is used to: perform semantic analysis on the current push message to determine the product and promotion start time corresponding to the current push message; determine the cluster that the current push message matches based on the product corresponding to the current push message and the cluster features of each cluster; for each cluster matched by the current push message, determine the push weight of the cluster corresponding to the current push message based on the interest score of the product corresponding to the current push message and the correlation between the push notification click volume of any two target products; determine the target cluster based on the push weight of each cluster corresponding to the current push message, and take the users included in the target cluster as the target users of the current push message.
[0014] Furthermore, the proactive push module is used to: determine multiple candidate push time periods based on the discount start time corresponding to the current push message; and determine the target push time period for the current push message based on the push reading ratio of the target cluster in the current period corresponding to the candidate push time periods and the push reading ratio of the associated push time periods.
[0015] Compared with existing technologies, the big data analytics-driven proactive push notification service system provided by this invention has at least the following beneficial effects:
[0016] 1. By collecting multi-dimensional user information, including basic, shopping, and dynamic operation information, and constructing a detailed knowledge graph, the proactive push module can accurately target users based on this data, ensuring that the pushed content is highly matched with user interests. For example, based on users' past shopping preferences and operating habits, appropriate product push messages are sent to users who genuinely need them, avoiding disturbing uninterested users. Simultaneously, by considering users' receptiveness to different push times, the optimal push timing is selected, allowing users to obtain the information they need at the appropriate time, greatly improving the user experience and enhancing user satisfaction and loyalty to the service system.
[0017] 2. By employing a two-layer update clustering algorithm and an improved butterfly optimization algorithm to cluster users, we can more accurately group users with similar characteristics and interests into the same cluster, providing a strong basis for formulating personalized push strategies. By constructing first and second knowledge graphs, we clearly present information such as the interest score, relevance, and push time period relevance of target products within each cluster. Based on this information, the proactive push module rationally determines the push weight and target push time period, improving the open rate and read rate of push messages, thereby increasing the conversion rate and making push resources more effectively utilized.
[0018] 3. The system can update the first and second knowledge graphs in real time based on users' shopping, dynamic operations, and notification actions during the current period. As user interests and behaviors continue to change, this dynamic update mechanism ensures that the knowledge graph always reflects the latest information, allowing the push strategy to always keep up with user needs. Attached Figure Description
[0019] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:
[0020] Figure 1 This is a schematic diagram of a big data analytics-driven proactive push notification service system according to some embodiments of this specification;
[0021] Figure 2 This is a flowchart illustrating a two-layer update clustering algorithm according to some embodiments of this specification. Detailed Implementation
[0022] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0023] Figure 1 This is a schematic diagram of a big data analytics-driven proactive push notification service system, as shown in some embodiments of this specification. Figure 1 As shown, a big data analytics-driven proactive push notification service system may include a data collection module, a graph construction module, a message acquisition module, and a proactive push module.
[0024] The data collection module is used to collect multi-dimensional information from multiple users, as well as historical notification and operation information.
[0025] The multi-dimensional information includes at least basic information, shopping information, and dynamic operation information.
[0026] Specifically, basic information encompasses a user's fundamental identity and characteristic data, forming the foundation for building user profiles. This information includes, but is not limited to, the user's age, gender, location, and occupation. Age information helps understand the user's life cycle stage; different age groups exhibit significant differences in product needs and preferences. For example, young people may be more inclined towards trendy electronic products or clothing, while middle-aged and elderly individuals may focus more on health and wellness products. Gender information helps differentiate users' preferences in product selection; for instance, female users may be more interested in cosmetics and jewelry, while male users may have greater needs for sports equipment and digital products. Location information reflects a user's consumption environment and lifestyle habits; users in different regions may have different product needs due to factors such as climate and culture. For example, users in cold regions have a higher demand for warm clothing. Occupation information indirectly reflects a user's income level and spending power, thus influencing their purchasing decisions. For example, high-income professionals may be more willing to pay for high-quality, high-priced goods.
[0027] Shopping information, centered on shopping records, details every purchase a user makes on the platform and is key data reflecting user consumption habits and preferences.
[0028] Dynamic operation information records various user interactions on the platform in real time. These behaviors directly reflect user interests and intentions, serving as crucial evidence for a deeper understanding of user needs. This information may include:
[0029] Adding items to the cart: When a user adds items to their cart, it indicates that they have some interest in and intention to purchase the item, but may not have made a final purchase decision yet. Information such as the type and quantity of items added to the cart, as well as the time of addition, is of significant analytical value. For example, if a user adds multiple similar items to their cart, they may be comparing products to find the one with the best value; if a user adds items to their cart but doesn't make a purchase for a long time, they may be waiting for a suitable discount or are still hesitant about whether to buy.
[0030] Product browsing: Browsing product pages is a process for users to understand product details and evaluate whether the product meets their needs. By recording information such as the types of products browsed, browsing duration, and browsing order, we can gain in-depth analysis of users' interests and concerns. For example, if a user spends a long time browsing a certain type of product, it indicates a high level of interest in that type of product; the browsing order may reflect the user's purchase decision-making process. Browsing higher-priced products first and then lower-priced products may indicate that the user values product quality more, while the reverse may indicate that the user is more concerned about price factors. In addition, the search keywords used by users when browsing products can also provide important clues for analyzing user needs. For example, if a user searches for "summer lightweight and breathable sports shoes," it indicates that the user has a need to buy summer sports shoes and has high requirements for the shoes' lightweight and breathable performance.
[0031] Product Favorites: When a user favorites a product, it indicates that the product holds high value for the user or represents a potential purchase. Information such as the types and quantities of products favorited, as well as the timing of the favorites, helps the platform understand the user's long-term interests and preferences. For example, if a user favorites multiple high-end cosmetics, it suggests they have a sustained interest in the high-end beauty sector; if a user immediately favorites a new product upon its release, it may indicate a high level of anticipation and purchase intent for that new product.
[0032] Product reviews and sharing: Users can review and share the products they've purchased, not only expressing their own consumption experience but also providing a reference for other users. Reviews can reflect user satisfaction with the product, its advantages and disadvantages, etc., and the platform can use this analysis to improve product and service quality. Sharing demonstrates user approval and willingness to recommend the product; sharing products on social media or other platforms may attract more potential users to pay attention to and purchase the product.
[0033] Historical notification activity information mainly records users' feedback behavior on various notifications pushed by the platform. This information helps the platform evaluate the effectiveness of notification pushes, optimize notification push strategies, and improve users' attention to and response rate to notifications. It includes whether users clicked to read the pushed product discount information and the reading time.
[0034] In some embodiments, the data collection module is used for:
[0035] Identify multiple background factors, such as age, gender, region, occupation, education level, and income level;
[0036] Based on multiple background factors, background information of multiple users is collected;
[0037] Based on background information, shopping information, dynamic operation information, and historical notification operation information of multiple users, multiple target background factors are selected from multiple background factors.
[0038] Basic information of multiple users is collected based on multiple target background factors.
[0039] In some embodiments, the data collection module is used for:
[0040] For each background factor, based on the background information of multiple users, multiple sampled users corresponding to the background factor are selected from multiple users. Based on the factor value of the background factor corresponding to each sampled user, the multiple sampled users corresponding to the background factor are divided into multiple sampled user groups. Based on the historical notification operation information of each sampled user group, the operation difference value within the group and the operation difference value outside the group are calculated to determine the first weight of the background factor.
[0041] For each background factor, based on the shopping information and dynamic operation information of multiple sampled users, the interest similarity between any two sampled users is calculated, and the difference between the corresponding background factors of any two sampled users is calculated. Based on the interest similarity between any two sampled users and the difference between the corresponding background factors, the second weight of the background factor is determined.
[0042] Based on the first and second weights of each background factor, multiple target background factors are selected from multiple background factors.
[0043] Specifically, for any two sampled users corresponding to a background factor, the mean similarity of other background factors, excluding that background factor, is greater than a similarity threshold (e.g., 70%).
[0044] Based on the factor value of the background factor corresponding to each sampled user, multiple sampled users corresponding to the background factor are divided into multiple sampled user groups. The K-means clustering algorithm can be used to group multiple sampled users according to the difference between the factor values of the background factor for any two sampled users. The smaller the difference, the higher the probability of them being assigned to the same sampled user group. For example, using "occupation type" as the background factor, sampled users can be divided into different groups according to different occupation types, such as a teacher group, a doctor group, and an engineer group. In this way, users within each group engage in the same occupation, while users between groups engage in different occupations, facilitating subsequent comparative analysis of behavioral differences among users of different occupations.
[0045] The knowledge graph construction module is used to cluster multiple sampled users based on multi-dimensional information and historical notification operation information of multiple users to determine multiple clusters. For each cluster, a first knowledge graph, a second knowledge graph, and cluster features are constructed. The first knowledge graph is used to represent the correlation between the interest scores and push notification clicks of multiple target products corresponding to the cluster. The second knowledge graph is used to represent the correlation between multiple push time periods corresponding to the cluster.
[0046] For each sampled user group, historical notification actions (such as reading, clicking, and purchasing pushed product discount information) are collected, and the degree of difference in these actions among users within the group is calculated. Common calculation methods include calculating statistical indicators such as the standard deviation and variance of user actions within the group. These indicators reflect the dispersion of user behavior within the group, i.e., the intra-group action variance value. A larger variance value indicates greater differences in historical notification actions among users within the group; a smaller variance value indicates more similar user behavior within the group.
[0047] For any two distinct sampled user groups, compare the degree of difference in historical notification behavior between the sampled users of these two groups. For example, this can be achieved by calculating the mean difference in historical notification behavior of the sampled users in the two groups, or by using distance metrics (such as Euclidean distance, Mahalanobis distance, etc.). The mean of the inter-group difference values for any two sampled user groups is then calculated to obtain the out-of-group difference value. This out-of-group difference value reflects the differences in historical notification behavior among the sampled users of multiple sampled user groups corresponding to the background factor.
[0048] The first weight of the background factor can be calculated based on the difference between the intra-group and inter-group operation differences. The larger the difference, the larger the first weight. If the inter-group operation difference is significantly greater than the intra-group operation difference, it indicates that the background factor has a significant impact on the user's historical notification operation behavior and can better distinguish the differences in notification operations among different user groups. Therefore, the background factor can be given a higher first weight. Conversely, if the inter-group operation difference is not much different from the intra-group operation difference, it indicates that the background factor has a smaller impact on the user's historical notification operation behavior and should be given a lower first weight.
[0049] The collected shopping information (such as product category, number of purchases, and purchase amount) and dynamic operation information (such as browsing time, number of times added to cart, and number of times favorited) from multiple sampled users are cleaned to remove duplicate, erroneous, and abnormal data. Next, a feature vector is constructed for each product, with dimensions covering shopping and dynamic operation-related indicators. For example, for shopping information, the number of purchases and purchase amount are used as two dimensions; for dynamic operation information, the standardized value of browsing time, the number of times added to cart, and the number of times favorited are used as three additional dimensions. Then, a clustering algorithm (such as K-Means) is used to cluster these product feature vectors, grouping products with similar characteristics into one category. Finally, the products of interest are determined based on the sampled users' comprehensive behavioral performance across all product categories. Specifically, a weighted behavioral score is calculated for each sampled user's behavior towards products in each cluster. The weights can be set according to the importance of each behavior in reflecting interest, such as placing a higher weight on purchasing behavior than on browsing behavior. The products included in the cluster with the highest score are the products of interest for the sampled users, generating a set of products of interest for each sampled user.
[0050] The overlap of the sets of items of interest of two sampled users can be calculated as the similarity of their interests. The greater the overlap, the higher the similarity.
[0051] The difference between the interest similarity of any two sampled users and the corresponding background factor is used as a key variable. This value is then substituted into the correlation coefficient calculation formula (e.g., Spearman's rank correlation coefficient, Kendall's rank correlation coefficient, etc.) to determine the second weight of the background factor. Taking age as an example, suppose sampled users A and B have a high interest similarity, indicating they share significant commonalities in product preferences. If the difference in their age background factor is small, it suggests that sampled users of similar ages have high similarity in product interests. This means that the age background factor is strongly correlated with user product interest similarity and can effectively explain interest similarity; therefore, a higher second weight can be assigned to the age background factor. Conversely, if the two users have high interest similarity but a large age difference, it indicates that age has a small impact on interest similarity and should be given a lower second weight. This comprehensive approach allows for a more scientific assessment of the influence of each background factor on user product interests.
[0052] For each background factor, the first and second weights of the background factor can be weighted and averaged to obtain the average weight of the background factor.
[0053] The average weight of each background factor is normalized. Based on the normalized average weight, the background factors are sorted, and the top n (e.g., 4, 5, 6, etc.) background factors are selected as target background factors.
[0054] During the data collection phase, target background factors are screened from multiple dimensions. The first weight is determined by comparing historical notification operation differences based on background information groups, reflecting the distinguishing effect of background factors on user notification operation behavior. The second weight is determined by combining shopping and dynamic operation information to calculate interest similarity and correlate background factor differences, reflecting the correlation between background factors and user product interests. This dual evaluation makes the screening of target background factors more scientific and comprehensive, and can accurately focus on the key background factors that affect user behavior.
[0055] In some embodiments, the map building module is used for:
[0056] For each user, based on their shopping information, dynamic operation information, and historical notification operation information, the user's interest score for each type of product is calculated. Specifically, the number of times a user purchases a certain type of product within a certain period (e.g., one month) is counted and normalized; the total amount spent on a certain type of product within that period is counted and normalized; the average time interval between two consecutive purchases of a certain type of product is calculated and normalized; and the three normalized scores are weighted and summed to obtain the user's interest score for a certain type of product under the shopping information dimension. Finally, the total browsing time of a user on a certain type of product page is calculated. Normalization is performed on the following data: the number of times a user adds a certain type of product to their shopping cart is normalized; the number of times a user favorites a certain type of product is normalized; the three scores are then weighted and summed to obtain the user's interest score for a certain type of product under the dynamic operation information dimension; the number of times a user clicks on notifications related to a certain type of product is normalized to obtain the notification click score; and the user's interest score for a certain type of product under the shopping information dimension, the user's interest score for a certain type of product under the dynamic operation information dimension, and the notification click score are weighted and summed to obtain the user's interest value for that type of product.
[0057] Based on each user's basic information, the similarity of basic information between any two users is calculated. Specifically, different preprocessing is required for different types of data in the basic information, such as categorical data (gender, occupation, etc.) and numerical data (age, income level, etc.). For categorical data, one-hot encoding can be used to convert it into numerical form; for numerical data, normalization can be performed to scale the data to the [0, 1] interval to eliminate the influence of different units. For a certain categorical variable of two users, if the values are the same, the similarity is 1; otherwise, it is 0. Then, the similarity of all categorical variables is averaged to obtain the similarity of the categorical data part. The similarity of numerical data can be calculated using methods such as cosine similarity and Euclidean distance. The weighted average of the categorical data similarity and the numerical data similarity is then used to obtain the basic information similarity between any two users.
[0058] Based on each user's interest value for each product category, calculate the interest similarity between any two users. Based on each user's interest value for each product category, calculate the cosine similarity of the interest value vectors of the two users as the interest similarity between the two users.
[0059] By using a two-layer update clustering algorithm, multiple sampled users are clustered based on the similarity of basic information and interest between any two users, thus determining multiple clusters.
[0060] In some embodiments, the map building module is used for:
[0061] S11. Calculate the cluster distance between any two users based on the basic information similarity and interest similarity. The greater the basic information similarity and interest similarity, the shorter the cluster distance.
[0062] S12. Based on the cluster distance between any two users, select multiple users as cluster centers. For example, for each user, calculate the variance of the cluster distance between the user and other users, and select users whose variance is greater than the mean variance as cluster centers. The mean variance is the mean of the variances of all users.
[0063] S13. Based on multiple cluster centers, perform clustering to generate clustering results. Specifically, for each user, calculate its clustering distance to each cluster center, and then assign the user to the cluster represented by the nearest cluster center. In this way, each user will be classified into a certain cluster, thus generating preliminary clustering results.
[0064] S14. Based on the clustering results, merge multiple cluster centers and update the clustering results. Specifically, analyze the situation of each cluster in the preliminary clustering results. If it is found that the distance between some clusters is very close, or the number of users in some clusters is too small, these clusters can be considered for merging. After merging the cluster centers, readjust the user affiliations and update the clustering results.
[0065] S15. Using the improved butterfly optimization algorithm, multiple cluster centers are updated based on the updated clustering results. The switching probability of the improved butterfly optimization algorithm is determined based on the clustering distance between any two clusters corresponding to the updated clustering results. The fitness function is related to the mean clustering distance of each cluster.
[0066] S16. Based on the merged and updated cluster centers, determine if clustering is complete. For example, by setting a threshold for cluster center position changes, when the position changes of each cluster center in both sets are less than this threshold, it indicates that the clustering process has stabilized. For example, setting the threshold to 0.05, if the position change of each cluster center in both the merged and updated sets is less than 0.05, it is considered that the convergence condition has been met. If so, output the clustering results; otherwise, execute S13.
[0067] Specifically, the switching probability is determined based on the clustering distance between any two clusters in the updated clustering result. Specifically, the clustering distance between the cluster centers of any two clusters in the updated clustering result is calculated, and their mean is taken. A larger mean value indicates greater differences between clusters in the current clustering result, suggesting a relatively reliable clustering result. In this case, the switching probability tends to favor local search; that is, the "butterfly" (representing the cluster center combination) performs a more detailed search near the current cluster to find a better local position, rather than performing a large-scale global search. In other words, the larger the mean clustering distance between the cluster centers of any two clusters, the smaller the switching probability, allowing the individual butterflies to perform more local searches, focusing on the vicinity of potential optimal solutions and improving solution accuracy. For example, if there are two clusters with a large mean clustering distance between them, it indicates significant differences in the characteristics of these two clusters. In this case, fine-tuning the cluster centers near the current cluster may be more effective in improving the clustering results.
[0068] The butterfly can represent the updated cluster center in each cluster of the updated clustering results. For each cluster, the clustering distance between each user in the cluster and the updated cluster center of that cluster is calculated, and then the mean of these distances is taken as the mean clustering distance for that cluster. The mean clustering distance for each cluster is then averaged again to take the mean global intra-cluster clustering distance. For any two updated cluster centers, the mean of the cluster centers between any two updated cluster centers is calculated as the mean global inter-cluster clustering distance. The independent variables of the fitness function include the mean global intra-cluster clustering distance and the mean global inter-cluster clustering distance. The smaller the mean global intra-cluster clustering distance and the larger the mean global inter-cluster clustering distance, the larger the fitness function value.
[0069] In the clustering process, cluster distances are first calculated based on similarity, and cluster centers are selected using variance to reasonably determine the initial cluster cores. After preliminary clustering, cluster centers are merged and the clustering results are updated, avoiding unreasonable clustering caused by clusters being too close together or having too few users, thus optimizing the clustering structure. Then, an improved butterfly optimization algorithm is used to update the cluster centers. The switching probability is determined based on the mean cluster distance between clusters; when the mean is large, it tends to search locally, improving the accuracy of the solution. The fitness function is related to the mean intra-cluster and inter-cluster cluster distances globally, guiding the algorithm to find better cluster center positions. Finally, a threshold for changes in cluster center positions is set to determine whether clustering is complete, ensuring the stability of the clustering results.
[0070] In some embodiments, the map building module is used for:
[0071] Based on the interest values of users for each type of product included in the cluster, calculate the interest value of each product corresponding to the cluster;
[0072] Based on the interest in each product corresponding to the cluster, determine the multiple target products corresponding to the cluster and the interest score of each target product. The cluster features include the multiple target products corresponding to the cluster.
[0073] Based on the historical notification operation information of the users included in the cluster, calculate the correlation of push notification click volume between any two target products corresponding to the cluster;
[0074] Based on the correlation between the multiple target products corresponding to the cluster, the interest score of each target product, and the click volume of push notifications for any two target products, a first knowledge graph is constructed.
[0075] Specifically, in calculating the interest value for each product category within a cluster, considering that each cluster consists of multiple users, the interest values of all users within the cluster for each product category are considered together. Specifically, each user within a cluster has their own level of interest in different product categories. These interest values are derived from a series of complex calculations, including normalization and weighted summation, by integrating user shopping information, dynamic operation information, and historical notification operation information. To obtain the overall interest value of the entire cluster for a particular product category, the interest values of all users within the cluster for that product category are aggregated, for example, by averaging, to obtain a value that represents the overall cluster's interest in that product category.
[0076] Based on the calculated interest values for each product corresponding to the cluster, multiple target products for the cluster and an interest score for each target product are determined. This step is similar to selecting the most popular products among users in the cluster from a large pool of products. A threshold is set; when the interest value of a certain type of product exceeds this threshold, it is identified as a target product. Furthermore, the interest values of this type of product exceeding the threshold are further processed (such as by re-normalization) to become the interest score for that target product. This score more intuitively reflects the popularity of the target product among users in the cluster. Moreover, these target products serve as an important component of the cluster features, used to more accurately describe the characteristics of the cluster when constructing the subsequent graph.
[0077] Based on the historical notification activity information of users within a cluster, the correlation between click volumes of push notifications for any two target products within that cluster is calculated. Historical notification activity information records users' click behaviors on various product push notifications, revealing underlying user preferences for different product notifications. For any two target products within a cluster, click volume data for push notifications for these two products from all users within the cluster is collected. Correlation analysis, such as using the Pearson correlation coefficient, is performed on this click volume data to measure the degree of correlation between the click volumes of push notifications for these two target products. If the correlation between the click volumes of push notifications for two products is high, it indicates that users show similar interest in these two types of products and may have similar interests or needs.
[0078] Based on the previously obtained clusters, the various target products corresponding to each cluster, the interest score of each target product, and the correlation between the push notification click volumes of any two target products, a first knowledge graph is constructed. In this first knowledge graph, target product nodes are labeled with their corresponding interest scores, reflecting their importance within the cluster. Simultaneously, any two target product nodes are connected based on the correlation between their push notification click volumes. The strength of this correlation can be represented by the thickness or weight of the edges. Only nodes of two target products whose push notification click volume correlation is greater than a correlation coefficient threshold (e.g., 0.5) can be connected by edges. This constructed first knowledge graph comprehensively and deeply displays the relationship between clusters and target products, as well as the inherent connections between target products.
[0079] In some embodiments, the map building module is used for:
[0080] Define multiple push notification time periods;
[0081] Based on the historical notification operation information of multiple users included in the cluster, the push reading ratio of the cluster in each push time period in multiple historical periods is calculated, the correlation of multiple push time periods corresponding to the cluster is determined, and a second knowledge graph is constructed.
[0082] Specifically, the push notification time slots are not set arbitrarily, but rather comprehensively consider factors such as users' daily device usage habits and the characteristics of business scenarios. For example, different time slots are divided, such as before work in the morning, lunch break, and after get off work in the evening, because users' availability and willingness to receive information may differ at different times. Reasonable segmentation allows for more accurate analysis of user behavior. The push notification reading ratio for each time slot across multiple historical periods is calculated based on the historical notification operation information of multiple users within the cluster. This historical notification operation information records in detail users' viewing and clicking behaviors related to push notifications at various time slots. For each historical period (e.g., one week, one month), the number of push notifications read by all users within the cluster in each push notification time slot is counted, and then divided by the total number of notifications pushed to users within the cluster during that time slot to obtain the push notification reading ratio. This ratio directly reflects the level of attention paid to the push content by users within the cluster during that time slot. Correlation analysis is then performed on the push notification reading ratios of different time slots, for example, by calculating correlation coefficients. If the trends in the read ratios of two push notifications over multiple historical periods are similar and the correlation coefficient is high, it indicates a strong correlation between the two time periods. This suggests that users' behavior patterns in receiving push notifications during these two time periods are similar, and they may have similar willingness to receive notifications and attention spans. In the second knowledge graph, the connections between nodes in the push notification time periods represent the correlation between multiple push notification time periods.
[0083] In some embodiments, the map building module is used for:
[0084] Obtain shopping information, dynamic operation information, and notification operation information of multiple users in the current period, and calculate the user's interest value for each type of product in the current period;
[0085] The first and second knowledge graphs are dynamically updated based on each user's interest in each type of product during the current period.
[0086] Specifically, the first and second knowledge graphs are dynamically updated based on each user's calculated interest value for each product category in the current period. For the first knowledge graph, it re-evaluates the various target products corresponding to each cluster and the interest score for each target product. If some users significantly increase their interest in a category of products that were not originally target products in the current period, exceeding a set threshold, then these products may be included in the target product category, and their interest scores will be updated. Conversely, if the attention to a previously target product decreases significantly in the current period, it may be removed or its interest score may be lowered. Furthermore, the correlation between the click-through rates of push notifications for any two target products is recalculated and adjusted based on the new data. For the second knowledge graph, it recalculates the push reading ratio of clusters in each push time period based on the user's notification actions during different push time periods in the current period, thereby redetermining the correlation between push time periods. Through this dynamic updating, the two knowledge graphs can closely follow changes in user behavior and interests, maintaining a high degree of alignment with reality and providing reliable data support for more accurate subsequent analysis and decision-making.
[0087] The message retrieval module is used to retrieve the current push message.
[0088] Specifically, the current push notifications may include information about new or discounted products.
[0089] The proactive push module is used to determine the target user and target push time period of the current push message based on the first knowledge graph, the second knowledge graph and the cluster features corresponding to each cluster, and to proactively push the message based on the target user and target push time period.
[0090] In some embodiments, the active push module is used for:
[0091] Semantic analysis is performed on the current push notification to determine the corresponding product and the start time of the promotion. Specifically, natural language processing technology is used to analyze keywords, key phrases, and semantic structure in the message text. By identifying specific words describing the product, such as product name, brand, and model, the product corresponding to the push notification is accurately located. Simultaneously, time-related expressions related to the promotion, such as "starting from [Month] [Day]" or "limited-time offer within [Number] days," are carefully searched to determine the start time of the promotion in the push notification.
[0092] Based on the product corresponding to the current push message and the cluster features of each cluster, determine the cluster that the current push message matches. Specifically, multiple target products in the cluster features, including the product corresponding to the current push message, can be used as the cluster that matches the current push message.
[0093] For each cluster matched by the current push message, the push weight of the cluster corresponding to the current push message is determined based on the interest score of the product corresponding to the current push message and the correlation between the push notification click volume of any two target products. Specifically, the average of the interest score of the product corresponding to the current push message and the interest score of the nodes connected to the node of the product corresponding to the current push message in the first knowledge graph can be calculated as the push weight of the cluster corresponding to the current push message.
[0094] Based on the push weight of the current push message corresponding to each cluster, the target cluster is determined, and the users included in the target cluster are used as the target users of the current push message. For example, the cluster with a push weight greater than the push weight threshold (e.g., 0.6) can be used as the target cluster.
[0095] In some embodiments, the active push module is used for:
[0096] Based on the current push message's corresponding discount activation time, determine multiple candidate push time periods;
[0097] Based on the push reading ratio of the target cluster in the current period corresponding to the candidate push time period and the push reading ratio of the associated push time period, the target push time period of the current push message is determined.
[0098] Specifically, considering the start time of the promotional activity in the push message, the proactive push module will determine multiple candidate push time periods based on this. These candidate time periods will be reasonably set around the promotion start time, for example, several different time points will be set as candidate push time periods within a certain period before the promotion starts.
[0099] The proactive push module collects push read ratio data for the target cluster within the current period, targeting various candidate push time slots. It also analyzes the read ratios of related push time slots. Through comprehensive analysis and comparison of this data, it identifies time slots with higher read ratios. If a candidate time slot and its related time slots all have relatively high read ratios, it indicates that pushing messages during that time slot will garner more user attention. After comprehensive consideration, the target push time slot for the current message is determined, ensuring that the message is delivered to the target users at the most appropriate time, further improving push effectiveness and promoting user interaction and conversion with the push content.
[0100] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and are considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. A proactive push notification service system driven by big data analytics, characterized in that, include: The data collection module is used to collect multi-dimensional information and historical notification operation information from multiple users. The multi-dimensional information includes at least basic information, shopping information, and dynamic operation information. The knowledge graph construction module is used to cluster multiple sampled users based on multi-dimensional information and historical notification operation information of multiple users to determine multiple clusters. For each cluster, a first knowledge graph, a second knowledge graph and cluster features are constructed. The first knowledge graph is used to represent the correlation between the interest scores and push notification clicks of multiple target products corresponding to the cluster. The second knowledge graph is used to represent the correlation between multiple push time periods corresponding to the cluster. The message retrieval module is used to retrieve the current push message; The proactive push module is used to determine the target user and target push time period of the current push message based on the first knowledge graph, the second knowledge graph, and cluster features corresponding to each cluster, and to proactively push the message based on the target user and target push time period. Specifically, this includes: Perform semantic analysis on the current push message to determine the corresponding product and promotion start time; Based on the product corresponding to the current push message and the cluster characteristics of each cluster, determine the cluster that the current push message matches; For each cluster matched by the current push message, the push weight of the cluster corresponding to the current push message is determined based on the interest score of the product corresponding to the current push message and the correlation between the push notification click volume of any two target products. Based on the push weight of the current push message corresponding to each cluster, the target cluster is determined, and the users included in the target cluster are taken as the target users of the current push message. Based on the current push message's corresponding discount activation time, determine multiple candidate push time periods; Based on the push reading ratio of the target cluster in the current period corresponding to the candidate push time period and the push reading ratio of the associated push time period, the target push time period of the current push message is determined.
2. A big data analytics driven push notification service system as claimed in claim 1, wherein, The data collection module is used for: Identify multiple background factors; Based on multiple background factors, background information of multiple users is collected; Based on background information, shopping information, dynamic operation information, and historical notification operation information of multiple users, multiple target background factors are selected from multiple background factors. Basic information of multiple users is collected based on multiple target background factors.
3. A big data analytics driven push notification service system as claimed in claim 2, wherein, The data collection module is used for: For each background factor, based on the background information of multiple users, multiple sampled users corresponding to the background factor are selected from multiple users. Based on the factor value of the background factor corresponding to each sampled user, the multiple sampled users corresponding to the background factor are divided into multiple sampled user groups. Based on the historical notification operation information of each sampled user group, the operation difference value within the group and the operation difference value outside the group are calculated to determine the first weight of the background factor. For each background factor, based on the shopping information and dynamic operation information of multiple sampled users, the interest similarity between any two sampled users is calculated, and the difference between the corresponding background factors of any two sampled users is calculated. Based on the interest similarity between any two sampled users and the difference between the corresponding background factors, the second weight of the background factor is determined. Based on the first and second weights of each background factor, multiple target background factors are selected from multiple background factors.
4. The big data analytics driven push notification service system of any of claims 1-3, wherein, The map construction module is used for: For each user, the interest value for each type of product is calculated based on the user's shopping information, dynamic operation information, and historical notification operation information. Calculate the basic information similarity between any two users based on each user's basic information. Calculate the interest similarity between any two users based on each user's interest value for each product category; By using a two-layer update clustering algorithm, multiple sampled users are clustered based on the similarity of basic information and interest between any two users, thus determining multiple clusters.
5. A big data analytics driven push notification service system as claimed in claim 4, wherein, The map construction module is used for: S11. Calculate the cluster distance between any two users based on the similarity of their basic information and their interest similarity. S12. Based on the clustering distance between any two users, select multiple users as cluster centers; S13. Based on multiple cluster centers, perform clustering and generate clustering results; S14. Based on the clustering results, merge multiple cluster centers and update the clustering results; S15. Using the improved butterfly optimization algorithm, multiple cluster centers are updated based on the updated clustering results. The switching probability of the improved butterfly optimization algorithm is determined based on the clustering distance between any two clusters corresponding to the updated clustering results. The fitness function is related to the mean clustering distance of each cluster. S16. Based on the merged cluster centers and the updated cluster centers, determine if clustering is complete. If yes, output the clustering results; otherwise, execute S13.
6. A big data analytics driven push notification service system as claimed in claim 4, wherein, The map construction module is used for: Based on the interest values of users for each type of product included in the cluster, calculate the interest value of each product corresponding to the cluster; Based on the interest in each product corresponding to the cluster, determine the multiple target products corresponding to the cluster and the interest score of each target product. The cluster features include the multiple target products corresponding to the cluster. Based on the historical notification operation information of the users included in the cluster, calculate the correlation of push notification click volume between any two target products corresponding to the cluster; Based on the correlation between the multiple target products corresponding to the cluster, the interest score of each target product, and the click volume of push notifications for any two target products, a first knowledge graph is constructed.
7. A big data analytics driven push notification service system as claimed in claim 4, wherein, The map construction module is used for: Define multiple push notification time periods; Based on the historical notification operation information of multiple users included in the cluster, the push reading ratio of the cluster in each push time period in multiple historical periods is calculated, the correlation of multiple push time periods corresponding to the cluster is determined, and a second knowledge graph is constructed.
8. The big data analytics driven push notification service system of any one of claims 1-3, wherein, The map construction module is used for: Obtain shopping information, dynamic operation information, and notification operation information of multiple users in the current period, and calculate the user's interest value for each type of product in the current period; The first and second knowledge graphs are dynamically updated based on each user's interest in each type of product during the current period.
Citation Information
Patent Citations
Big data collection and analysis system based on computer
CN117557341A
Media big data-based marketing advertisement putting method
CN119887306A