An e-commerce marketing prediction method and system based on big data analysis
Through dynamic window segmentation and improvement of Hoffman coding, multi-stage clustering analysis is carried out in combination with transaction record data, and a time-weighted user behavior prediction map is generated, which solves the limitations of the existing e-commerce marketing prediction methods in capturing the diverse behavior patterns and value levels of users, and achieves efficient and accurate marketing prediction.
Patent Information
- Application Number
- CN202510260743.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing e-commerce marketing prediction methods have limitations in capturing the diverse behavior patterns and value levels of users, and cannot accurately reflect the complexity of users' multi-level needs and behavioral paths, and lack flexibility and adaptability.
By collecting user behavior trajectory data of e-commerce platform, dynamic window segmentation process is performed to generate timing behavior fragments, and the improved Hoffman encoding is used to generate behavior path coding vectors. Combining transaction record data, the value index vector is calculated in real time, and multi-stage coupled clustering analysis is carried out to generate user grouping labels with fusion behavior patterns and value levels. Finally, based on geographical location feature data, a time-weighted user behavior prediction map is generated.
It realizes efficient and accurate marketing prediction results in a large-scale and multi-dimensional data environment, can accurately capture the timing characteristics and dynamic changes of user behavior, integrate user behavior patterns and value characteristics, improve the accuracy of user grouping, and improve the personalization and accuracy of marketing prediction through time-space weighted prediction maps.
Smart Images

Figure CN119762122B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis and prediction, and in particular to an e-commerce marketing prediction method and system based on big data analysis. Background Art
[0002] With the booming development of the e-commerce industry, user behavior data on e-commerce platforms has become an important basis for merchants to conduct precision marketing, optimize product recommendations, and improve conversion rates. In recent years, e-commerce marketing prediction methods based on big data analysis have gradually become a trend in the industry. By collecting and analyzing multi-dimensional data such as user behavior trajectories, transaction data, and geographic locations, e-commerce platforms can deeply explore users' interests, preferences, purchasing habits, and potential needs, thereby providing personalized recommendations and marketing strategies.
[0003] However, existing e-commerce marketing prediction methods still have some limitations, especially in how to accurately capture users' diverse behavior patterns and value levels. Traditional methods usually simplify user behavior data into some basic features, such as purchase frequency, consumption amount, etc. This rough processing method cannot accurately reflect the complexity of users' multi-level needs and behavior paths. Most existing clustering analysis methods rely on similarity calculations in a single dimension, and often cannot comprehensively consider the temporal characteristics and geographic location information of user behavior, resulting in the reliability and accuracy of clustering results being affected. In addition, most methods have not fully utilized the huge user transaction data and social influence data in e-commerce platforms, and have ignored the impact of these factors on user purchasing decisions. Therefore, existing marketing prediction methods lack sufficient flexibility and adaptability in dealing with the complex and dynamic changes in the e-commerce market, and cannot achieve refined user portraits and personalized marketing strategies. Summary of the invention
[0004] In view of the problems existing in the existing e-commerce marketing prediction technology, the present invention is proposed.
[0005] Therefore, the problem to be solved by the present invention is how to provide efficient and accurate marketing prediction results in a large-scale, multi-dimensional data environment.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides an e-commerce marketing prediction method based on big data analysis, which includes collecting user behavior trajectory data of an e-commerce platform, performing dynamic window segmentation processing on the user behavior trajectory data to generate time series behavior segments; inputting the time series behavior segments into an encoder, and using improved Huffman coding at the event type layer to generate a behavior path coding vector; and calculating a value indicator vector in real time based on transaction record data; performing multi-stage coupled clustering analysis on the behavior path coding vector and the value indicator vector to generate a user grouping label that integrates the behavior pattern and the value level; based on geographic location feature data, generating a spatiotemporal weighted user behavior prediction map according to the user grouping label.
[0008] As a preferred solution of the e-commerce marketing prediction method based on big data analysis described in the present invention, wherein: the generation of the behavior path coding vector includes: classifying event types based on time series behavior fragments; counting the user's behavior frequency within a fixed time window, and making timeliness adjustments while calculating the behavior frequency, and dynamically adjusting the weight of each event type; the improved Huffman coding assigns a Huffman code to each event type according to the weight adjustment result; the event codes of each user in the behavior sequence are arranged in order to form a behavior path coding vector, the length of the behavior path coding vector is the total number of events in the user behavior sequence, and each dimension value represents the Huffman coding value of the corresponding event type.
[0009] As a preferred solution of the e-commerce marketing prediction method based on big data analysis described in the present invention, the value indicator vector refers to a set of value indicators calculated based on the user's transaction history data, integrated into a value indicator vector, which is used to describe the user's value; the transaction history data includes RFM value, price sensitivity coefficient and social network influence value.
[0010] As a preferred solution of the e-commerce marketing prediction method based on big data analysis described in the present invention, the multi-stage coupled clustering analysis includes the following operation steps: in the first stage, density clustering is performed using the DBSCAN algorithm based on the behavior path encoding vector; in the second stage, secondary grouping based on hierarchical clustering is performed within the same density cluster according to the value indicator vector.
[0011] As a preferred solution of the e-commerce marketing prediction method based on big data analysis described in the present invention, the first stage includes: calculating the similarity between all user behavior path encoding vectors, and generating a similarity matrix of user behavior paths according to the calculation results; according to the similarity matrix, applying the DBSCAN algorithm to cluster the user behavior path encoding vectors; the DBSCAN algorithm divides the user behavior paths into different clusters according to density areas, and automatically identifies dense areas and noise points.
[0012] As a preferred solution of the e-commerce marketing prediction method based on big data analysis described in the present invention, the second stage includes: selecting a density cluster from the DBSCAN density clustering output of the first stage, which includes a group of users with similar behavior patterns; obtaining the value index vectors of all users in the selected density cluster; calculating the distance between the value index vectors of all users to obtain the similarity matrix of the value index vectors;
[0013] Based on the hierarchical clustering algorithm, the clustering process is initialized based on the similarity matrix of the value indicator vector; the users in the density cluster are secondary clustered by the hierarchical clustering algorithm, and the clusters are gradually merged or split to form a hierarchical structure; the final number of subclusters is determined by pruning the dendrogram or setting a distance threshold, and a subcluster label is assigned to each user according to the hierarchical clustering result to identify the subcluster to which the user belongs; the final cluster label, including the density cluster identifier and the subcluster identifier, is assigned to each user to complete the multi-level clustering of users.
[0014] As a preferred solution of the e-commerce marketing prediction method based on big data analysis described in the present invention, the geographic location feature data includes: dividing the geographic space into multiple grids according to the range of longitude and latitude, each grid represents a geographic area within a certain range; the user's geographic coordinates will be mapped to the corresponding grid number according to its longitude and latitude; each geographic grid area corresponds to a thermal value, and the calculation of the thermal value is determined based on transaction frequency, crowd density and social influence; each user generates geographic location features according to his or her geographic location, including longitude block number, latitude block number and business district thermal value.
[0015] As a preferred solution of the e-commerce marketing prediction method based on big data analysis described in the present invention, wherein: the generation of a spatiotemporal weighted user behavior prediction map includes: dividing the time range into multiple time segments; combining user grouping labels, geographic location feature data and time segment data to generate a spatiotemporal weighted user behavior prediction map; each point of the user behavior prediction map represents a geographic location area and time segment, and combining the weighted information of the user grouping labels and geographic location feature data to calculate the user's behavior prediction results in the geographic location area and time period.
[0016] In a second aspect, the present invention provides an e-commerce marketing prediction system based on big data analysis, which includes:
[0017] The data collection and segmentation module collects user behavior trajectory data from the e-commerce platform, and performs dynamic window segmentation processing on the user behavior trajectory data to generate time series behavior fragments; the vector calculation module inputs the time series behavior fragments into the encoder, and uses improved Huffman coding at the event type layer to generate behavior path coding vectors; and calculates the value indicator vector in real time based on the transaction record data; the multi-stage clustering analysis module performs multi-stage coupled clustering analysis on the behavior path coding vector and the value indicator vector to generate user grouping labels that integrate behavior patterns and value levels; the spatiotemporal weighted prediction module generates a spatiotemporal weighted user behavior prediction map based on the geographic location feature data and the user grouping labels.
[0018] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, the steps of the e-commerce marketing prediction method based on big data analysis as described in the first aspect of the present invention are implemented.
[0019] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, the steps of the e-commerce marketing prediction method based on big data analysis as described in the first aspect of the present invention are implemented.
[0020] The beneficial effects of the present invention are as follows: the present invention overcomes the problem of information loss and inaccuracy caused by fixed windows in traditional technologies by adopting a behavior data processing method based on dynamic window segmentation, and can more accurately capture the temporal characteristics of user behavior and its dynamic changes. The Huffman coding method is improved, which not only improves the coding efficiency, but also effectively retains the hierarchical information of the behavior path. In addition, by performing multi-stage coupled clustering analysis on the behavior path coding vector and the value indicator vector calculated based on transaction record data, the invention effectively integrates the user's behavior pattern and value characteristics, solves the problem of single-dimensional analysis of traditional methods when constructing user portraits, and can achieve more accurate user grouping. When further combining geographic location feature data to generate a spatiotemporal weighted user behavior prediction map, it is possible to better explore regional differences and spatiotemporal laws, greatly improving the personalization and accuracy of marketing predictions. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1Flowchart of the e-commerce marketing prediction method based on big data analysis.
[0023] Figure 2 This is a structural diagram of the e-commerce marketing prediction system based on big data analysis. DETAILED DESCRIPTION
[0024] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0025] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0026] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0027] Embodiment 1:
[0028] Reference Figure 1~Figure 2 , which is the first embodiment of the present invention, and provides an e-commerce marketing prediction method based on big data analysis, such as Figure 1 As shown, including,
[0029] S1: Collecting user behavior trajectory data from the e-commerce platform, and performing dynamic window segmentation processing on the user behavior trajectory data to generate time series behavior segments.
[0030] First, collect user behavior trajectory data from the e-commerce platform, including but not limited to user page visits, clicks, purchases, browsing time, dwell time and other behavior data; at the same time, collect transaction record data related to the user, including order details, transaction time, product category, payment amount and other information.
[0031] Set dynamic window segmentation conditions based on user behavior characteristics. The boundary of window segmentation is affected by two main factors: page dwell time threshold and event type conversion frequency.
[0032] Among them, the page dwell time threshold is the maximum time each user stays on a certain page. If this time is exceeded, it is considered that the user's behavior has changed and the window needs to be re-divided; the event type conversion frequency is the frequency threshold of user behavior conversion (for example, from browsing a page to clicking on a product) within a period of time. When behavior conversion occurs frequently, the window boundary should be automatically adjusted.
[0033] In the embodiment of the present invention, the dynamic window segmentation adaptively adjusts the window boundary according to the page dwell time threshold and the event type conversion frequency.
[0034] Specifically, the initial window size is set to a fixed time interval (such as 30 seconds, 1 minute, etc.), but as user behavior changes, the window will be adaptively adjusted according to the page dwell time and event type conversion;
[0035] If a user stays on a page for longer than the set duration threshold, the current window will end and a new window will start from that time point to recapture the user's new behavior sequence; if frequent event type conversions occur within the window (for example, the user switches from the browsing page to the product details page multiple times in a short period of time), the current window will be automatically divided into multiple smaller sub-windows to accurately capture the user's high-frequency interactive behavior.
[0036] According to the above dynamic window segmentation algorithm, time-series behavior segments are generated, and each segment represents the user's behavior trajectory within a certain period of time. The key data contained in the behavior segment includes: user behavior type (click, browse, purchase, etc.), the time when the behavior occurs, the length of time on the page, the frequency of user interaction with the page, etc.; and the generated time-series behavior segments are standardized to ensure the consistency of each behavior segment in the time dimension to avoid affecting subsequent analysis due to differences in the length of time of different behaviors; each behavior segment will be arranged in chronological order, and the time boundary of the behavior segment will be dynamically adjusted according to the user's behavior history and platform data to ensure that the user's behavior dynamics are accurately captured at the moment of analysis.
[0037] It should be noted that the event-driven dynamic window division method breaks through the limitations of traditional fixed time windows or fixed number of events, and can flexibly adapt to the diversity and irregularity of user behavior; compared with the equal-length window division method, this method can more accurately capture the user's behavioral characteristics and changes in a specific period of time, avoiding data loss or window incompatibility in certain behavioral patterns, and ensuring the accuracy and effectiveness of data analysis.
[0038] S2: Input the time series behavior fragment into the encoder, use improved Huffman coding at the event type layer to generate a behavior path coding vector; and calculate the value indicator vector in real time based on the transaction record data.
[0039] The time-series behavior fragments generated in step S1 are used as input, which contain the behavior records of each user within a certain time window (such as page visits, clicks, browsing time, etc.). These behavior fragments have been processed according to dynamic window division to ensure the integrity and timing of the behavior sequence.
[0040] Classify user behaviors by type. Common event types include page views, product clicks, adding to cart, purchases, etc.
[0041] In traditional Huffman coding, the weight of the coding is usually based on the frequency of occurrence of events. However, in e-commerce data, this method may over-compress frequently occurring behaviors (such as browsing pages, clicking on products, etc.), thereby losing important behavioral differences. Therefore, the present invention improves Huffman coding as follows: by encoding each event type (such as page browsing, clicking on products, adding to shopping carts, etc.) of the time series behavior fragment, the Huffman coding of each event is generated, and finally the behavior path coding is output. The behavior path coding of each user will represent the user's behavior sequence on the platform, and the coding length is adjusted according to the frequency and timeliness of the event type.
[0042] In this embodiment, the event type layer is responsible for encoding according to different event types (such as click, browse, purchase, etc.), and each event type is represented by improved Huffman coding. Improved Huffman coding optimizes the weight calculation method in the original Huffman coding to make the coding length of different event types more efficient, and can better retain the short coding length of frequent events and the long coding length of infrequent events, thereby making the behavior sequence more compact in storage and calculation.
[0043] The encoding of event types follows the principle that behavior types with higher frequency (such as purchases or clicks) have shorter codes, while behavior types with lower frequency (such as browsing product details) have longer codes. This encoding method can improve the compression rate of the code while reducing the consumption of computing resources.
[0044] Specifically, the number of occurrences of each event type is counted based on the frequency of user behavior within a certain time window. For example, if a user frequently clicks on a product page in the past week, the frequency of this event type in this time period is high. When calculating event frequency, a time factor is introduced to give priority to the weight of recent behavior. That is, the user's recent behavior (such as behavior in the past 3 days) will be given a higher weight than historical behavior (such as behavior in the past 30 days), which ensures that the event code can reflect the user's currently active behavior pattern.
[0045] The weight of each event type is dynamically adjusted based on its frequency and timeliness. Frequent and recent event types (such as product clicks) are given higher weights, while less popular or historical events (such as fewer product views) are given lower weights.
[0046] According to the weight adjustment results, a Huffman code is assigned to each event type. Specifically: for high-frequency event types (such as product clicks), shorter Huffman codes are assigned to reduce storage costs and computational complexity; for low-frequency event types (such as adding to a shopping cart), longer Huffman codes are assigned to ensure that there is enough expression space when encoding. When the user's behavior pattern changes, the weight distribution of the Huffman code is automatically adjusted. For example, if the number of clicks on a certain product increases dramatically within a certain period of time, the encoding weight of the product can be dynamically adjusted to ensure that it more accurately reflects the changing user interests.
[0047] Arrange the event codes of each user in the behavior sequence in order to form a behavior path coding vector. The length of the behavior path coding vector is the total number of events in the user behavior sequence, and each dimension value represents the Huffman coding value of the corresponding event type.
[0048] For example, for each user, all of his / her behaviors on the platform are arranged in chronological order to form a complete behavior path. For example, the behavior path of user A may be: browse product A → click product B → add to shopping cart → purchase product B.
[0049] For each event type in the user's behavior path, the improved Huffman coding is applied to encode the user's entire behavior path into a sequence of multiple Huffman codes. Each event type generates an optimal code based on its frequency and timeliness adjusted weight. Finally, a behavior path code representing the user's behavior sequence is output, which contains the code of each event type, and the code length is dynamically adjusted according to the frequency and timeliness of the event type.
[0050] As user behavior changes, the encoding of certain event types may become longer or shorter, ensuring the flexibility and efficiency of the encoding. For each user's behavior path encoding, a dynamic compression mechanism is used to further reduce storage requirements. For example, in a behavior path, if multiple behaviors of the same type occur consecutively (such as multiple product clicks), the storage space can be further reduced by replacing multiple short codes with long codes. Ultimately, the generated behavior path encoding will have a high compression rate and can accurately express the user's behavior pattern on the e-commerce platform. Through this dynamic encoding method, the encoding length can meet the efficient storage requirements of behavioral data while ensuring that important behavioral information is not lost.
[0051] It can be seen that by improving Huffman coding and dynamically calculating the value indicator vector, the coding efficiency of user behavior on the e-commerce platform and the accuracy of user value analysis are effectively improved. In the coding process, traditional Huffman coding has the problem of over-compression of frequent events, which may lead to the loss of important behaviors. In order to overcome this problem, an improved Huffman coding method based on event type and timeliness is proposed to dynamically adjust the event coding length to ensure that the coding of frequent behaviors is short and the coding of rare behaviors is long, thereby improving storage efficiency and reducing the computational burden.
[0052] Furthermore, the value indicator vector refers to a set of value indicators calculated based on the user's transaction history data, integrated into a value indicator vector, which is used to describe the user's value, where the transaction history data includes RFM value, price sensitivity coefficient and social network influence value. These value indicators can be generated into a comprehensive value indicator vector by weighted summation or other methods.
[0053] Among them, the FM value is used to measure the user's most recent transaction time, purchase frequency and total consumption amount; the price sensitivity coefficient is used to calculate an indicator reflecting the user's price sensitivity based on the user's reaction to the price; the social network influence value is based on the user's activity data on the social platform to calculate the size of his or her social influence.
[0054] Exemplarily, the FM value comprehensively considers the time of the user's most recent transaction, purchase frequency and total consumption amount, among which the time of the most recent transaction is calculated by calculating the distance between the user and the current time, usually in days. The closer the distance, the higher the user's activity and the larger the value; the purchase frequency is calculated by calculating the number of purchases made by the user within a certain time range (for example, the past 30 days). The more purchases, the higher the user's activity and loyalty; the total consumption amount is calculated by calculating the total consumption amount of the user within a certain period of time (such as the past 30 days). The higher the amount, the stronger the user's consumption capacity. The calculation of the FM value is to combine the influence of these three items through weighted average to reflect the user's comprehensive activity and consumption capacity.
[0055] The price sensitivity coefficient is used to measure the user's sensitivity to price changes. The calculation process is as follows: Purchase probability change is calculated by calculating the difference in the probability of users purchasing a certain product before and after the price change. The user's purchase behavior at different prices can be obtained through historical data; the price change, that is, the price change of the product, is usually measured by the absolute difference in price. By comparing the price change and the purchase probability change, the user's price sensitivity coefficient is calculated. The larger the coefficient, the more sensitive the user is to price changes.
[0056] The social network influence value reflects the user's interactive ability and influence on the social platform. The calculation method is as follows: the content interaction volume is the total number of likes, comments and shares received by the user's posts. The more interactions, the greater the user's participation and influence on the social platform; the number of fans is the number of fans of the user on the social platform, reflecting the coverage of his social influence. The more fans, the wider the influence; the content sharing volume is the number of content shared by the user on the platform, indicating his ability to spread content. The higher the sharing volume, the stronger the social influence of the user. These indicators are weighted to calculate the social influence value.
[0057] The calculated RFM value, price sensitivity coefficient, social network influence value, etc. are integrated into a value indicator vector, which contains the following dimensions: RFM value, price sensitivity coefficient and social network influence value. The RFM value, price sensitivity coefficient and social network influence value are updated in real time according to each new transaction record, so as to dynamically adjust the value indicator vector. If the user has a new transaction behavior, the transaction frequency, transaction amount and other data are automatically updated, and the corresponding RFM value and price sensitivity coefficient are recalculated and updated.
[0058] Through this process, real-time evaluation of each user's value can be achieved, and the value indicator vector can be dynamically adjusted according to changes in transaction behavior, thereby providing the e-commerce platform with an accurate user value portrait.
[0059] The present invention further calculates the user's value index vector by introducing the RFM value, price sensitivity coefficient and social network influence value based on the user's transaction history, thereby realizing real-time dynamic update of the user's value. The value vector provides strong data support for personalized recommendation and user grouping, and can better reflect the user's activity, consumption ability and sensitivity to price changes.
[0060] S3: Perform multi-stage coupled clustering analysis on the behavior path encoding vector and the value indicator vector to generate user grouping labels that integrate the behavior pattern and value level.
[0061] Specifically, the multi-stage coupled cluster analysis includes the following steps:
[0062] In the first stage, density clustering is performed based on the behavior path encoding vector using the DBSCAN algorithm; in the second stage, secondary clustering based on hierarchical clustering is performed within the same density cluster according to the value indicator vector.
[0063] The first phase of operations specifically includes the following:
[0064] The behavior path encoding vector of each user is used as input data for cluster analysis.
[0065] Determine the two key parameters ε and MinPts (minimum number of neighborhood points) of the DBSCAN algorithm:
[0066] It can be done through visualization methods (such as K distance graph) or empirical adjustment. The K distance graph shows the distances of the k neighbors of each point. The location of the "inflection point" can be used as a candidate value for ε. Generally speaking, MinPts can be selected according to the dimension of the data. Usually MinPts is set to twice the dimension of the data (such as MinPts=4 is usually selected for two-dimensional data), but it can also be adjusted according to the clustering effect.
[0067] In order to effectively cluster user behavior paths, it is first necessary to calculate the similarity between users, which can be achieved by calculating the distance between user behavior path encoding vectors. Common distance measurement methods include Euclidean distance, cosine similarity, and Manhattan distance, etc. However, in the clustering task based on behavior path encoding vectors, the present invention uses cosine similarity to calculate the similarity.
[0068] It should be noted that in clustering tasks based on user behavior paths, the behavior path encoding vector usually contains multiple feature dimensions (such as time, location, activity type, etc.), and the behaviors of different users may have similar patterns, but the behavior characteristics of each user may have a large deviation (such as differences in numerical values). The Euclidean distance focuses more on measuring the absolute difference in numerical values, which is not always effective for measuring the similarity of behavior patterns, especially in high-dimensional sparse data, where the Euclidean distance may underestimate the actual similarity.
[0069] Cosine similarity measures similarity by calculating the angular difference between vectors, which can effectively ignore the absolute numerical differences between different user behavior paths and focus on the relative patterns of user behavior. Therefore, it can better reflect the similarity of user behaviors, especially when the magnitude of behavioral features differs greatly.
[0070] After obtaining the similarity between user behavior paths, the DBSCAN algorithm is used for clustering:
[0071] If a user's behavior path has enough points (≥MinPts) in the ϵ neighborhood, the user is considered a core point, and all density-connected points in the neighborhood are included in the cluster; for each core point, DBSCAN continues to expand the cluster and add all density-connected points to the cluster until there are no more points to add.
[0072] If a point cannot find enough neighbors (less than MinPts) in all ϵ neighborhoods, it will be marked as a noise point and does not belong to any cluster. The noise point will not affect the generation of other clusters. Through this process, the user's behavior path will be divided into different density clusters, the high-density areas are clustered into one cluster, and the sparse areas are treated as noise points.
[0073] After DBSCAN clustering, user behavior paths will be divided into different clusters according to their similarities. Each cluster contains a group of users with similar behavior patterns. Clusters with higher density represent user groups with similar behavior patterns, while areas with lower density may become noise. Output the clustering results and generate multiple density clusters, each of which contains users with similar behavior patterns.
[0074] It should be noted that density clustering can identify different density areas of behavior patterns. High-density areas indicate that user behavior paths have high similarity in the area, while low-density areas (noise points) do not participate in clustering, avoiding explicit matrix construction, reducing computational complexity and improving clustering efficiency.
[0075] The second phase of operations specifically includes the following:
[0076] From the DBSCAN clustering output of the first stage, a density cluster is selected. This cluster contains a group of users with similar behavior patterns. Each cluster represents a group of users with similar behavior paths. Usually, the behavior similarity of each density cluster is high, and the behavior characteristics of users have strong commonalities. For all users in the selected density cluster, the corresponding value indicator vector is extracted.
[0077] Hierarchical clustering is used to perform secondary clustering on the value indicator vectors of users in the selected density clusters. Unlike K-means, hierarchical clustering does not require a preset number of clusters (K value), and can automatically adapt to the number of clusters. It gradually merges or splits clusters by calculating the similarity (distance) between clusters, generating a tree-like clustering hierarchy (dendrogram).
[0078] Furthermore, to calculate the distance between all user value indicator vectors, measurement methods such as Euclidean distance, Manhattan distance or cosine similarity can be used. In order to calculate the distance between users, a weighted distance measurement can be used so that different dimensions of value indicators can reflect their relative importance.
[0079] Starting with each user as an independent cluster, similar clusters are gradually merged until all users are merged into one cluster or a stopping condition is reached (such as the number of clusters reaches a predetermined value).
[0080] It should be noted that the output of hierarchical clustering is a dendrogram, and the final number of clusters can be determined based on the structure of the dendrogram. By pruning the dendrogram and selecting the appropriate distance threshold or number of clusters, multiple sub-clusters are obtained. The number of selected clusters is dynamically adjusted based on factors such as the number of users in the density cluster, the diversity and similarity of value indicators, etc., to avoid artificially set fixed K value restrictions.
[0081] After the hierarchical clustering is completed, the hierarchical clustering algorithm will assign each user to a subcluster. These subclusters represent user groups with similar value characteristics and can reflect the heterogeneity of users in the value dimension. Each subcluster corresponds to a group of users with similar purchasing behavior, consumption ability or social behavior.
[0082] The output result will be the final clustering label of the user in the second stage, and each user will be assigned a new cluster label. The label contains the identifier of the cluster in the first stage and the identifier of the subcluster in the second stage (the subcluster from the hierarchical clustering).
[0083] Through the second stage of hierarchical clustering, the user grouping will be more refined and accurate. The users in the original density cluster will be further divided into multiple more representative subgroups according to their value indicator vectors, thus providing more detailed user portraits for subsequent analysis, recommendation systems, etc.
[0084] It can be seen that compared with K-means, hierarchical clustering avoids the problem of presetting the number of clusters, can automatically adapt to the structure of the data, and determine the final number of clusters through the tree diagram, making the clustering results more natural and consistent with the data distribution. Limitation: Through hierarchical clustering, there is no need to manually set the K value (number of clusters), avoiding the uncertainty caused by the selection of the K value, and the clustering results are more flexible. And by integrating the density clustering results of the first stage with the hierarchical clustering results of the second stage, more accurate and hierarchical user grouping labels can be obtained, providing more valuable user portraits.
[0085] The present invention realizes the deep integration of user behavior and value characteristics through multi-stage coupled cluster analysis, thereby effectively generating more representative and accurate user grouping labels. The first stage adopts the DBSCAN density clustering algorithm to cluster users according to the user behavior path encoding vector. In this stage, the similarity of user behavior is measured by calculating the cosine similarity, which can avoid the shortcomings of the traditional Euclidean distance and ensure that the clustering results can effectively reflect the true similarity of user behavior patterns. DBSCAN clustering identifies user groups with similar behavior patterns through density connection, and can process noise data, avoid meaningless clustering interference, and enhance the stability and accuracy of clustering.
[0086] In the second stage, hierarchical clustering is performed within the density clusters generated in the first stage to further refine the user groups. Hierarchical clustering gradually merges or splits user groups, making users in the same cluster more similar in value characteristics based on their value indicator vectors (such as RFM value, price sensitivity, and social influence). By dynamically adjusting the number of clusters, the fixed number of clusters is avoided, ensuring the flexibility and adaptability of clustering.
[0087] S4: Based on the geographic location feature data and the user grouping labels, a spatiotemporal weighted user behavior prediction map is generated.
[0088] First, extract the clustering labels of each user from the user clustering matrix. These labels contain hierarchical information about the user's behavior patterns, value characteristics, etc. The user clustering labels are generated based on user behavior data and value indicator data, and are used to identify users in different groups. Obtain the geographic location information of each user, including the user's latitude and longitude coordinates (location coordinates), and the thermal value of the corresponding business district or area. The thermal value is calculated based on the user activities in the area (such as traffic, purchase frequency, etc.), reflecting the activity and importance of the area.
[0089] Furthermore, the geographic space is divided into multiple grids according to the range of longitude and latitude, and each grid represents a geographic area within a certain range; the user's geographic coordinates will be mapped to the corresponding grid number according to its longitude and latitude; each geographic grid area corresponds to a thermal value, and the calculation of the thermal value is determined based on transaction frequency, crowd density and social influence; each user generates geographic location features according to his or her geographic location, including longitude block number, latitude block number and business district thermal value.
[0090] Furthermore, the heat value of a business district can be expressed as the comprehensive activity and business potential of a business district through the weighted average of transaction frequency, pedestrian density and social influence. The transaction frequency can be obtained by summing up the behavior frequency of each user in the business district and then dividing it by the number of users in the business district; the pedestrian density is obtained by calculating the visit frequency of each user in the business district; and the social influence can be calculated by summing up the social activities of all users in the business district and dividing it by the number of users engaging in social interactions.
[0091] Furthermore, according to the temporal characteristics of user behavior, the entire time range is divided into multiple time segments, which can be divided into units such as hours, days, and weeks, depending on the prediction needs and data availability. In order to reflect the spatiotemporal characteristics of user behavior, the weighted impact of each user's behavior in different time segments and geographical locations is modeled, including time decay factors and spatial weighting factors.
[0092] For example, the spatial weighting factor reflects the influence of the heat value of the business district. The higher the heat value of the business district, the greater the influence of the behavior of the area on the prediction result. The time decay factor is:
[0093] ;
[0094] in, is the time decay factor. As time goes by, the factor value decreases gradually. The time point when the user behavior occurs. is the time reference point, is the decay coefficient, which controls the decay rate. The larger the value, the faster the decay. The time decay factor reflects that the influence of user behavior on prediction gradually weakens over time, ensuring that the weight of old behavior in the prediction gradually decreases.
[0095] Combining the user's clustering labels, geographic location characteristics, and time segment data, a spatiotemporal weighted user behavior prediction map is generated. The behavior prediction map will show the distribution of user behavior in different geographic locations and time segments, and adjust the prediction results for each region and time period based on the weighting factor. Each point in the map can represent a specific geographic location (identified by the latitude and longitude block code) and a specific time segment. At each location in the map, combined with the information of clustering labels and geographic characteristics, the user's behavior prediction results in the region and time period, such as purchase probability, click-through rate, etc., are calculated.
[0096] Through the time-space weighted graph, the potential behavior of users in specific areas and time periods can be predicted. For example, which business districts may have higher purchasing activity in a certain period of time in the future, or in which time periods, the user behavior of a specific group is more concentrated.
[0097] Output the spatiotemporal weighted user behavior prediction map to visualize the user behavior prediction for each spatiotemporal region. Based on the spatiotemporal weighted prediction map, accurate marketing strategies and personalized recommendations can be implemented for different geographic locations and time periods. For example, in a certain business district or time period, high-potential users can be attracted through targeted promotional strategies.
[0098] It should be noted that by combining behavioral data, value indicators and geographic location characteristics, the temporal and spatial variation characteristics of user behavior can be effectively captured. For example, the activity of certain user groups in different business districts, the laws of behavioral patterns over time, etc., can be dynamically analyzed from multiple dimensions to generate a more predictive user behavior map. By analyzing the potential behavior of users in different geographic areas and time periods, it is possible to accurately identify which areas or time periods have users with higher purchasing activity or interaction frequency. Marketing activities can be targeted and promoted in specific time and space areas based on this information, thereby improving the efficiency and conversion rate of product delivery and avoiding marketing waste in low-activity periods and low-activity areas.
[0099] Further, such as Figure 2 As shown, this embodiment also provides an e-commerce marketing prediction system based on big data analysis, including:
[0100] The data collection and segmentation module 100 collects user behavior trajectory data of the e-commerce platform, and performs dynamic window segmentation processing on the user behavior trajectory data to generate time series behavior segments; the vector calculation module 200 inputs the time series behavior segments into the encoder, and uses improved Huffman coding at the event type layer to generate behavior path coding vectors; and calculates the value index vector in real time based on the transaction record data; the multi-stage clustering analysis module 300 performs multi-stage coupled clustering analysis on the behavior path coding vector and the value index vector to generate user grouping labels that integrate behavior patterns and value levels; the spatiotemporal weighted prediction module 400 generates a spatiotemporal weighted user behavior prediction map based on the geographic location feature data and the user grouping labels.
[0101] In summary, the present invention overcomes the problem of information loss and inaccuracy caused by fixed windows in traditional technologies by adopting a behavioral data processing method based on dynamic window segmentation, and can more accurately capture the temporal characteristics of user behavior and its dynamic changes. The improvement of the Huffman coding method not only improves the coding efficiency, but also effectively retains the hierarchical information of the behavior path. In addition, by performing multi-stage coupled clustering analysis on the behavior path coding vector and the value indicator vector calculated based on the transaction record data, the invention effectively integrates the user's behavior pattern and value characteristics, solves the problem of single-dimensional analysis of traditional methods when constructing user portraits, and can achieve more accurate user grouping. When further combining geographic location feature data to generate a spatiotemporal weighted user behavior prediction map, the invention can better tap into regional differences and spatiotemporal laws, greatly improving the personalization and accuracy of marketing predictions.
[0102] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. An e-commerce marketing prediction method based on big data analysis, characterized by: include: Collecting user behavior trajectory data from the e-commerce platform, and performing dynamic window segmentation processing on the user behavior trajectory data to generate time series behavior segments; The time series behavior fragment is input into the encoder, and the behavior path encoding vector is generated by using improved Huffman coding at the event type layer; and the value indicator vector is calculated in real time based on the transaction record data; Perform multi-stage coupled clustering analysis on the behavior path encoding vector and the value indicator vector to generate user grouping labels that integrate the behavior pattern and the value level; Based on the geographic location feature data and the user grouping labels, a spatiotemporal weighted user behavior prediction graph is generated; The generation of the behavior path encoding vector includes: Classify event types based on temporal behavior fragments; Count the frequency of user behaviors within a fixed time window, and make timeliness adjustments while calculating the behavior frequency, dynamically adjusting the weight of each event type; The improved Huffman coding is to assign a Huffman coding to each event type according to the weight adjustment result; Arrange the event codes of each user in the behavior sequence in order to form a behavior path coding vector, the length of which is the total number of events in the user behavior sequence, and each dimension value represents the Huffman coding value of the corresponding event type; The value indicator vector refers to a set of value indicators calculated based on the user's transaction history data, integrated into a value indicator vector, which is used to describe the user's value; The transaction history data includes RFM value, price sensitivity coefficient and social network influence value.
2. The e-commerce marketing prediction method based on big data analysis according to claim 1, characterized in that: The multi-stage coupled cluster analysis includes the following steps: In the first stage, the DBSCAN algorithm is used to perform density clustering based on the behavior path encoding vector; In the second stage, secondary grouping based on hierarchical clustering is performed within the clusters of the same density according to the value indicator vectors.
3. The e-commerce marketing prediction method based on big data analysis according to claim 2, characterized in that: The first phase includes: Calculate the similarity between all user behavior path encoding vectors, and generate a user behavior path similarity matrix based on the calculation results; According to the similarity matrix, a DBSCAN algorithm is applied to cluster the user's behavior path encoding vector; The DBSCAN algorithm divides user behavior paths into different clusters according to density areas, and automatically identifies dense areas and noise points.
4. The e-commerce marketing prediction method based on big data analysis according to claim 3, characterized in that: The second phase includes: Select a density cluster from the DBSCAN density clustering output of the first stage, which contains a group of users with similar behavior patterns; Obtain the value index vectors of all users in the selected density cluster; calculate the distances between the value index vectors of all users to obtain the similarity matrix of the value index vectors; Initializing a clustering process based on a similarity matrix of the value indicator vectors according to a hierarchical clustering algorithm; The hierarchical clustering algorithm is used to perform secondary grouping of users in the density cluster, and the clusters are gradually merged or split to form a hierarchical structure; The final number of sub-clusters is determined by pruning the dendrogram or setting a distance threshold. A sub-cluster label is assigned to each user based on the hierarchical clustering results to identify the sub-cluster to which the user belongs. Assign a final clustering label to each user, including density cluster identifier and sub-cluster identifier, to complete the multi-level grouping of users.
5. The e-commerce marketing prediction method based on big data analysis according to claim 1, characterized in that: The geographic location feature data includes: The geographic space is divided into multiple grids according to the range of longitude and latitude, and each grid represents a geographic area within a certain range; the user's geographic coordinates will be mapped to the corresponding grid number according to their latitude and longitude; Each geographic grid area corresponds to a heat value, and the calculation of the heat value is determined based on transaction frequency, crowd density and social influence; Each user generates geographic location features based on his / her geographic location, including longitude block number, latitude block number and business district heat value.
6. The e-commerce marketing prediction method based on big data analysis according to claim 5, characterized in that: The generating of the spatiotemporal weighted user behavior prediction graph comprises: Divide the time range into multiple time segments; Combine user grouping labels, geographic location feature data, and time segment data to generate a spatiotemporal weighted user behavior prediction graph; Each point of the user behavior prediction map represents a geographical location area and time segment, and the user grouping label and the weighted information of the geographical location feature data are combined to calculate the user's behavior prediction result in the geographical location area and time segment.
7. An e-commerce marketing prediction system based on big data analysis, based on the e-commerce marketing prediction method based on big data analysis according to any one of claims 1 to 6, characterized in that: Also includes: The data collection and segmentation module collects the user behavior trajectory data of the e-commerce platform, and performs dynamic window segmentation processing on the user behavior trajectory data to generate time series behavior segments; The vector calculation module inputs the time series behavior fragment into the encoder, generates the behavior path encoding vector by using improved Huffman coding at the event type layer; and calculates the value index vector in real time based on the transaction record data; A multi-stage clustering analysis module performs multi-stage coupled clustering analysis on the behavior path encoding vector and the value indicator vector to generate user grouping labels that integrate the behavior pattern and the value level; The spatiotemporal weighted prediction module generates a spatiotemporal weighted user behavior prediction map based on the geographic location feature data and the user grouping labels.
Citation Information
Patent Citations
User portrait construction method and system based on big data
CN118098584A
Customer relationship management method and system for electricity marketing
CN118941298A