Real-time customer portrait dynamic updating and predicting method and system for retail scenario
By using sliding time windows and temporal correlation analysis, customer behavior fusion features with temporal dependencies are generated, which solves the problem of lack of timeliness and accuracy in customer profiles in existing technologies. This enables accurate prediction of customer behavior and personalized recommendations, thereby improving the marketing effectiveness in retail scenarios.
Patent Information
- Application Number
- CN202510830112.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-06-20
AI Technical Summary
Existing customer profiling technologies cannot effectively capture the temporal dependencies of customer behavior, resulting in customer profiles that lack timeliness and accuracy. They cannot accurately reflect changes in the importance of customer behavioral characteristics and are difficult to predict future consumption intentions and behavioral trends.
By employing a sliding time window method and temporal correlation analysis, the information gain value and dynamic correlation coefficient of customer behavior features are calculated, generating customer behavior fusion features with temporal dependencies. Through feature importance assessment and weighting, the customer profile is dynamically updated.
It enables accurate prediction of customer behavior, improves the response speed and accuracy of the recommendation system, meets the precision marketing needs of retail enterprises, and enhances customer shopping experience and satisfaction.
Smart Images

Figure CN120707249B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data analysis, and in particular to a real-time customer portrait dynamic updating and prediction method and system for a retail scenario. BACKGROUND
[0002] With the digital transformation of the retail industry, customer portrait technology has become an important tool for understanding customer needs and optimizing marketing strategies. Customer portrait is a description of various customer characteristics, including demographic characteristics, consumer behavior characteristics, preference characteristics, etc. In the modern retail scenario, real-time customer portrait dynamic updating and prediction are of great significance for accurately grasping customer needs and providing personalized services. Traditional customer portrait construction mainly relies on static data analysis, by collecting customer historical purchase records, browsing behavior and other data, to construct a relatively stable customer feature model.
[0003] With the development of big data technology and artificial intelligence algorithms, retailers can collect and analyze more rich customer behavior data, including online browsing tracks, search keywords, shopping cart operations, offline in-store behaviors and other multi-source heterogeneous data. These data make it possible to build a more comprehensive and accurate customer portrait. However, how to effectively use these data to realize real-time dynamic updating of customer portrait and prediction of consumer behavior still faces many technical challenges.
[0004] The existing customer portrait technology mainly has the following defects and deficiencies:
[0005] The existing customer portrait technology usually ignores the value difference of customer behavior data in different time windows, and processes all features with the same weight, which cannot accurately reflect the importance change of customer behavior characteristics in different periods, resulting in the lack of timeliness and accuracy of the constructed customer portrait.
[0006] The traditional customer portrait updating method cannot effectively capture the time sequence dependence of customer behavior, and only extracts and updates features based on independent time points, ignoring the correlation and evolution law of customer behavior in the time dimension, and is difficult to predict the future consumer intention and behavior change trend of customers. SUMMARY
[0007] The embodiments of the present application provide a real-time customer portrait dynamic updating and prediction method and system for a retail scenario, which can solve the problems in the prior art.
[0008] In a first aspect, the embodiments of the present application provide a real-time customer portrait dynamic updating and prediction method for a retail scenario, comprising:
[0009] determining a customer portrait feature set corresponding to customer behavior data in a retail scenario;
[0010] According to the information gain value of different customer portrait features in the customer portrait feature set under different time windows, a feature importance score is calculated, and a customer portrait weighted feature vector is generated based on the feature importance score;
[0011] The customer portrait weighted feature vector is divided into time sequences of customer behavior features in multiple time windows by using a sliding time window method;
[0012] The time sequences of customer behavior features in the multiple time windows are analyzed for time correlation, a dynamic correlation coefficient of the time sequences of customer behavior features between adjacent time windows is calculated, and a customer behavior fusion feature with a time-dependent relationship is generated based on the dynamic correlation coefficient;
[0013] A feature correlation matrix is calculated based on the customer behavior fusion feature and the customer portrait weighted feature vector, and a consumption behavior feature of a customer in a next time window is predicted based on the feature correlation matrix;
[0014] According to the consumption behavior feature and the customer behavior fusion feature, the customer portrait feature set is updated, and personalized product recommendations and marketing strategies are determined for the customer based on the updated customer portrait feature set.
[0015] According to the information gain value of different customer portrait features in the customer portrait feature set under different time windows, a feature importance score is calculated, and a customer portrait weighted feature vector is generated based on the feature importance score, including:
[0016] The customer portrait feature set is divided into time window sequences according to a preset time interval, and a plurality of time window sequences of customer portrait features are obtained;
[0017] An information entropy matrix of each customer portrait feature in the customer portrait feature sequence is constructed, and the information entropy matrix contains information entropy values of each customer portrait feature under different time windows;
[0018] Based on the information entropy matrix, a multi-dimensional feature importance evaluation index is determined by calculating the change rate of information entropy between adjacent time windows, and the multi-dimensional feature importance evaluation index reflects the sensitivity of customer portrait features to changes in customer behavior;
[0019] According to the multi-dimensional feature importance evaluation index, the customer portrait features are hierarchically clustered to obtain a feature hierarchical clustering result;
[0020] Based on the feature hierarchical clustering result, a feature importance score is set, a time sequence decay function is used to dynamically adjust the feature importance score to generate a time sequence decay coefficient, and the time sequence decay coefficient is directly proportional to the feature fluctuation amplitude and inversely proportional to the feature change period;
[0021] The time sequence attenuation coefficient is weighted and calculated with the customer portrait feature to generate a customer portrait weighted feature vector.
[0022] Based on the information entropy matrix, a multi-dimensional feature importance evaluation index is determined by calculating the change rate of information entropy between adjacent time windows, including:
[0023] Based on the information entropy matrix, an information entropy difference sequence between adjacent time windows is calculated, and a multi-layer wavelet coefficient is obtained by wavelet multi-scale decomposition of the information entropy difference sequence.
[0024] A linear change trend curve is fitted using the low-frequency component of the multi-layer wavelet coefficient, and a nonlinear mutation curve is fitted using the high-frequency component of the multi-layer wavelet coefficient.
[0025] Based on the derivative of the linear change trend curve, the change rate and change acceleration of the customer portrait feature are calculated, an adaptive decision threshold is constructed according to the change rate and change acceleration, and a feature gradual change trend index is output based on the adaptive decision threshold.
[0026] Based on the nonlinear mutation curve, a mutation detection window is constructed, and a mutation feature parameter of the customer portrait feature is calculated in the mutation detection window, and a feature mutation degree index is generated according to the mutation feature parameter.
[0027] The feature gradual change trend index and the feature mutation degree index are processed using a dynamic time warping algorithm to obtain time sequence similarity features of the customer portrait feature at different time scales, periodic change features and burst change features of the feature are extracted based on the time sequence similarity features, and a multi-dimensional feature importance evaluation index is obtained.
[0028] The customer portrait weighted feature vector is time-sequentially divided using a sliding time window method to obtain a plurality of time window customer behavior feature sequences, including:
[0029] Based on the customer portrait weighted feature vector, a time sequence sliding window is constructed, a time sensitivity matrix of the feature is calculated based on the time sequence sliding window, a key time scale is determined according to a singular value decomposition result of the time sensitivity matrix, and the key time scale is mapped to an initial window length and an initial sliding step of the time sequence sliding window.
[0030] The customer portrait weighted feature vector is time-sequentially decomposed according to the initial window length and the initial sliding step to obtain a change feature parameter of the customer behavior feature sequence.
[0031] Based on the change feature parameter, a swarm intelligence optimization algorithm is used to optimize the initial window length and the initial sliding step to obtain optimal window parameters.
[0032] The overlapping interval between adjacent time windows is determined using the optimal window parameters. The continuous change value of the feature is calculated based on the overlapping interval. The granularity of window division is determined based on the continuous change value and the time sensitivity matrix.
[0033] Based on the granularity of the window division and the changing feature parameters, the features of the overlapping intervals are weighted and fused to obtain customer behavior feature sequences for multiple time windows.
[0034] Perform temporal correlation analysis on the customer behavior feature sequences of the multiple time windows, calculate the dynamic correlation coefficient between customer behavior feature sequences of adjacent time windows, and generate customer behavior fusion features with temporal dependencies based on the dynamic correlation coefficient, including:
[0035] Calculate the temporal adjacency relationship for the customer behavior feature sequences of the multiple time windows, calculate the attenuation distance between adjacent windows based on the temporal adjacency relationship, and construct the temporal propagation attenuation coefficient and feature correlation matrix based on the attenuation distance;
[0036] The temporal propagation attenuation coefficient and the feature correlation matrix are recursively iterated to extract the temporal state change sequence, and a dynamic correlation coefficient is generated based on the temporal state change sequence.
[0037] A local temporal attention score is constructed based on the dynamic correlation coefficient. The feature correlation matrix is weighted using the local temporal attention score. The weighted feature correlation matrix is then combined with the temporal propagation attenuation coefficient to generate a global attention vector.
[0038] Based on the global attention vector, a temporal weight is assigned to each time window. The temporal weight is then combined with the customer behavior feature sequence in a weighted manner. Combined with the dynamic correlation coefficient, long-range dependency information is extracted to generate customer behavior fusion features with temporal dependencies.
[0039] A local temporal attention score is constructed based on the dynamic correlation coefficient. This local temporal attention score is then used to weight the feature correlation matrix. Finally, the weighted feature correlation matrix is combined with the temporal propagation attenuation coefficient to generate a global attention vector, including:
[0040] The local similarity matrix of the time-series features is calculated based on the dynamic correlation coefficient. A transition probability matrix is constructed on the local similarity matrix. The steady-state distribution of the features is calculated based on the transition probability matrix. The entropy value of the steady-state distribution is used as the local temporal attention score.
[0041] The local temporal attention scores are used to construct an attention gain function, and the feature correlation matrix is nonlinearly transformed based on the attention gain function to obtain a weighted feature matrix with temporal memory effect.
[0042] Based on the weighted feature matrix, the feature importance distribution is calculated, singular value decomposition is performed, the feature subspace corresponding to the main singular values is extracted, a time-series propagation path is constructed in the feature subspace, and the time-series propagation path is combined with the time-series propagation attenuation coefficient to generate dynamic attenuation features.
[0043] An initial attention vector is constructed based on the dynamic decay feature. The cross-time window correlation degree of the feature is calculated using the initial attention vector. The dynamic decay feature is recursively updated based on the cross-time window correlation degree. The updated dynamic decay feature is used as the global attention vector.
[0044] Based on the customer behavior fusion features and the customer profile weighted feature vector, a feature relevance matrix is calculated. Based on the feature relevance matrix, the customer's consumption behavior characteristics in the next time window are predicted, including:
[0045] Based on the customer behavior fusion features and the customer profile weighted feature vector, a feature relevance matrix is calculated. The feature relevance matrix is then decomposed into eigenvalues to obtain a feature importance vector. Time-varying weight coefficients are then constructed based on the feature importance vector.
[0046] The conditional transition probability of the feature is calculated using the time-varying weight coefficients, and a time-series state transition matrix is constructed based on the conditional transition probability. The optimal feature distribution parameters are obtained by maximizing the log-likelihood of the state transition.
[0047] Periodic components are extracted from the optimal feature distribution parameters, and the periodic components are separated from the trend components. Short-term fluctuation characteristics and long-term change characteristics are calculated separately, and time-series combined features are generated based on the short-term fluctuation characteristics and long-term change characteristics.
[0048] The conditional probability distribution between features is calculated using the time-series combined features. A feature dependency graph is constructed based on the conditional probability distribution. By iteratively calculating the state probability of feature nodes, the consumption behavior characteristics of customers in the next time window are predicted.
[0049] A second aspect of this invention provides a real-time customer profile dynamic update and prediction system for retail scenarios, comprising:
[0050] The first unit is used to determine the set of customer profile features corresponding to customer behavior data in retail scenarios;
[0051] The second unit is used to calculate the feature importance score based on the information gain value of different customer profile features in the customer profile feature set under different time windows, and generate a customer profile weighted feature vector based on the feature importance score.
[0052] The third unit is used to perform time-series partitioning of the customer profile weighted feature vector using a sliding time window method to obtain customer behavior feature sequences for multiple time windows.
[0053] The fourth unit is used to perform temporal correlation analysis on the customer behavior feature sequences of the multiple time windows, calculate the dynamic correlation coefficient between customer behavior feature sequences of adjacent time windows, and generate customer behavior fusion features with temporal dependencies based on the dynamic correlation coefficient.
[0054] The fifth unit is used to calculate a feature correlation matrix based on the customer behavior fusion features and the customer profile weighted feature vector, and to predict the customer's consumption behavior features in the next time window based on the feature correlation matrix.
[0055] The sixth unit is used to update the customer profile feature set based on the consumption behavior characteristics and the customer behavior fusion characteristics, and to determine personalized product recommendations and marketing strategies for customers based on the updated customer profile feature set.
[0056] A third aspect of the present invention provides an electronic device, comprising:
[0057] processor;
[0058] Memory used to store processor-executable instructions;
[0059] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0060] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0061] The beneficial effects of this application are as follows:
[0062] The present invention provides a real-time dynamic update prediction method for customer profiles in retail scenarios. By calculating the information gain value of customer profile features under different time windows to determine the importance of features, a weighted feature vector is generated, which effectively improves the accuracy and personalization of customer profiles and solves the problem that static features in traditional methods cannot reflect the dynamic changes in customer behavior.
[0063] This invention employs a sliding time window method and temporal correlation analysis to calculate the dynamic correlation coefficient between adjacent time windows, generating customer behavior fusion features with temporal dependencies. This captures the dynamic evolution of customer behavior, enabling accurate prediction of customer consumption behavior and improving the response speed and accuracy of the recommendation system.
[0064] This invention dynamically updates customer profiles based on predicted consumer behavior characteristics and fused features, providing customers with personalized product recommendations and marketing strategies. It realizes real-time dynamic updates of customer profiles in retail scenarios, which not only meets the precision marketing needs of retail enterprises, but also improves customer shopping experience and satisfaction, and has significant commercial value and application prospects. Attached Figure Description
[0065] Figure 1 This is a flowchart illustrating the real-time dynamic update and prediction method for customer profiles in retail scenarios according to an embodiment of the present invention.
[0066] Figure 2 This is a schematic diagram illustrating the temporal propagation path and attenuation coefficient analysis;
[0067] Figure 3 This diagram illustrates a performance comparison analysis of consumer behavior characteristic prediction models. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0069] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0070] Figure 1 This is a flowchart illustrating the real-time customer profile dynamic update and prediction method for retail scenarios according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0071] Determine the set of customer profile features corresponding to customer behavior data in retail scenarios;
[0072] The feature importance score is calculated based on the information gain value of different customer profile features in the customer profile feature set under different time windows, and a weighted feature vector of the customer profile is generated based on the feature importance score.
[0073] The customer profile weighted feature vector is divided into time series using a sliding time window method to obtain customer behavior feature sequences for multiple time windows.
[0074] A temporal correlation analysis is performed on the customer behavior feature sequences of the multiple time windows, the dynamic correlation coefficient between customer behavior feature sequences of adjacent time windows is calculated, and a customer behavior fusion feature with temporal dependency is generated based on the dynamic correlation coefficient.
[0075] Based on the customer behavior fusion features and the customer profile weighted feature vector, a feature correlation matrix is calculated, and based on the feature correlation matrix, the customer's consumption behavior features in the next time window are predicted.
[0076] Based on the consumer behavior characteristics and the customer behavior fusion characteristics, the customer profile feature set is updated, and personalized product recommendations and marketing strategies are determined for the customer based on the updated customer profile feature set.
[0077] In one optional implementation, a feature importance score is calculated based on the information gain values of different customer profile features in the customer profile feature set under different time windows, and a customer profile weighted feature vector is generated based on the feature importance score, including:
[0078] The customer profile feature set is divided into time windows according to a preset time interval to obtain customer profile feature sequences for multiple time windows;
[0079] Construct an information entropy matrix for each customer profile feature in the customer profile feature sequence, wherein the information entropy matrix contains the information entropy value of each customer profile feature under different time windows;
[0080] Based on the information entropy matrix, a multidimensional feature importance evaluation index is determined by calculating the rate of change of information entropy between adjacent time windows. The multidimensional feature importance evaluation index reflects the sensitivity of customer profile features to changes in customer behavior.
[0081] The customer profile features are hierarchically clustered based on the multidimensional feature importance evaluation index to obtain the feature hierarchical clustering results;
[0082] Based on the feature hierarchical clustering results, feature importance scores are set, and a time-series decay function is used to dynamically adjust the feature importance scores to generate a time-series decay coefficient. The time-series decay coefficient is directly proportional to the feature fluctuation amplitude and inversely proportional to the feature change period.
[0083] The time-series decay coefficient is weighted and calculated with the customer profile features to generate a customer profile weighted feature vector.
[0084] This invention provides a method for calculating feature importance based on information gain values of customer profile features. In this method, the customer profile feature set is first divided into time windows according to a preset time interval. For example, customer behavior data from the past 12 months can be divided into monthly time windows, resulting in a sequence of customer profile features for 12 time windows. These features may include information from multiple dimensions such as user purchase frequency, purchase amount, visit duration, and click-through rate. Assume an e-commerce platform has collected user A's consumption behavior data over the past 12 months, including features such as monthly spending amount, visit frequency, and purchased product categories.
[0085] For the acquired customer profile feature sequence, an information entropy matrix is constructed for each feature. This matrix records the information entropy value of each feature in different time windows. Information entropy reflects the degree of uncertainty of a feature. Taking the user's spending amount feature as an example, the system calculates the information entropy value of this feature in each monthly window. If the user's spending amount is relatively evenly distributed in a certain month, its information entropy value will be higher; conversely, if the spending amount in that month is concentrated in a certain range, its information entropy value will be lower. Suppose that the information entropy values of user A's monthly spending amount feature from January to December are: 0.85, 0.87, 0.86, 0.92, 0.78, 0.76, 0.72, 0.79, 0.88, 0.90, 0.91, 0.94.
[0086] Based on the constructed information entropy matrix, the system determines the multidimensional feature importance assessment index by calculating the rate of change of information entropy between adjacent time windows. For the consumption amount feature of user A mentioned above, the rate of change of information entropy from January to February is (0.87-0.85) / 0.85 = 0.0235, indicating that the information entropy increased by 2.35%. Similarly, the rate of change between other months is calculated: -0.0115 from February to March, 0.0698 from March to April, and so on. These rates of change constitute the time-series fluctuation characteristics of this feature, reflecting the sensitivity of the customer profile feature to changes in customer behavior. At the same time, the system also considers statistical measures such as the variance and mean absolute value of the rate of change to form a multidimensional assessment index. For the consumption amount feature of user A, the calculated mean rate of change is 0.0189, and the variance is 0.0025. These values together constitute the importance assessment index of this feature.
[0087] Based on the calculated multidimensional feature importance evaluation index, the system performs hierarchical clustering on all customer profile features. The clustering process uses a hierarchical clustering algorithm, measuring the similarity between features by calculating Euclidean distance or cosine similarity. Assuming the system analyzes 10 features including spending amount, visit frequency, and purchased product category, the hierarchical clustering algorithm categorizes these features into three levels: high importance, medium importance, and low importance. Spending amount and visit frequency are classified as high importance features; purchased product category and search keywords are classified as medium importance features; and the remaining features are classified as low importance features.
[0088] Based on the feature-based hierarchical clustering results, the system sets corresponding feature importance scores: high-importance features are assigned values ranging from 0.8 to 1.0, medium-importance features from 0.5 to 0.7, and low-importance features from 0.1 to 0.4. The initial importance score for the consumption amount feature is set to 0.9, the access frequency score to 0.85, and the purchased product category score to 0.65. The system uses a time-series decay function to dynamically adjust these initial scores, generating a time-series decay coefficient. The time-series decay coefficient considers the feature's fluctuation amplitude and change cycle. The greater the feature's fluctuation amplitude, the more sensitive the feature is to changes in user behavior, and the larger its time-series decay coefficient; the shorter the feature's change cycle, the higher the feature's change frequency, and the larger its time-series decay coefficient. For the consumption amount feature, due to its large fluctuation amplitude (the variance of the information entropy change rate is 0.0025) and short change cycle (a significant fluctuation occurs on average every 3 months), the calculated time-series decay coefficient is 1.15.
[0089] The time-series decay coefficient is weighted and calculated with customer profile features to generate a weighted feature vector for the customer profile. For the consumption amount feature, the final weighted value is 0.9 × 1.15 = 1.035. Similarly, the weighted value for the visit frequency feature is 0.85 × 1.08 = 0.918, and the weighted value for the purchased product category feature is 0.65 × 0.92 = 0.598. These weighted values together constitute the weighted feature vector of user A's customer profile [1.035, 0.918, 0.598, ...]. This vector accurately reflects the relative importance of different features in predicting user behavior and can significantly improve the accuracy of subsequent customer behavior prediction models.
[0090] This method enables a time-series dynamic evaluation of the importance of customer profile features. Compared to traditional static evaluation methods, it can more accurately capture patterns of changing customer behavior, providing more effective decision support for precision marketing and personalized recommendations. Experiments show that the weighted feature vector of customer profiles constructed using this method improves accuracy by 12.5%, recall by 8.7%, and F1 score by 10.3% in user behavior prediction tasks, fully demonstrating the effectiveness of the method.
[0091] In one optional implementation, based on the information entropy matrix, a multidimensional feature importance evaluation index is determined by calculating the rate of change of information entropy between adjacent time windows, including:
[0092] Based on the information entropy matrix, calculate the information entropy difference sequence between adjacent time windows, and perform wavelet multi-scale decomposition on the information entropy difference sequence to obtain multi-layer wavelet coefficients;
[0093] A linear trend curve is obtained by fitting the low-frequency components of the multilayer wavelet coefficients, and a nonlinear abrupt change curve is obtained by fitting the high-frequency components of the multilayer wavelet coefficients.
[0094] Based on the derivative of the linear trend curve, the rate of change and acceleration of change of customer profile features are calculated. An adaptive decision threshold is constructed based on the rate of change and acceleration of change. Based on the adaptive decision threshold, a feature gradient trend index is output.
[0095] A mutation detection window is constructed based on the nonlinear mutation curve. Within the mutation detection window, mutation feature parameters of customer profile features are calculated, and a feature mutation degree index is generated based on the mutation feature parameters.
[0096] The feature gradual change trend index and feature mutation degree index are processed by the dynamic time warping algorithm to obtain the temporal similarity features of customer profile features at different time scales. Based on the temporal similarity features, the periodic change features and sudden change features of the features are extracted to obtain the multidimensional feature importance evaluation index.
[0097] The information entropy matrix is composed of the information entropy values of customer profile features under different time windows. Each row in the information entropy matrix represents a feature, and each column represents a time window. For example, suppose we have three customer profile features (age, income, and purchase frequency) and five time windows (January, February, March, April, and May), then the information entropy matrix might look like this: the information entropy values for the age feature are [0.85, 0.86, 0.87, 0.89, 0.92], the information entropy values for the income feature are [0.76, 0.75, 0.74, 0.77, 0.80], and the information entropy values for the purchase frequency feature are [0.92, 0.93, 0.91, 0.89, 0.85].
[0098] Based on the aforementioned information entropy matrix, the information entropy difference sequence between adjacent time windows is calculated. Taking age as an example, the information entropy difference sequence between adjacent time windows is [0.01, 0.01, 0.02, 0.03], which reflects the variation of this feature between time windows. This information entropy difference sequence is processed using a wavelet multi-scale decomposition method. Specifically, the db4 wavelet basis function is used to perform a three-level decomposition of the sequence, obtaining low-frequency and high-frequency coefficients. In this example, the low-frequency coefficients may be [0.005, 0.015, 0.025], and the first, second, and third-level high-frequency coefficients are [0.005, -0.005, 0.005], [0.002, 0.003], and [0.001], respectively.
[0099] In this example, the linear trend equation obtained by fitting using the least squares method can be represented as the stable change of the information entropy difference over time. The calculation results show that the information entropy of the age feature exhibits a stable upward trend. Based on this linear trend curve, the derivative is calculated to show a rate of change of 0.01 and an acceleration of change of 0.005, indicating that the information content of the age feature is gradually increasing and the growth rate is accelerating.
[0100] An initial threshold of 0.02 was set (i.e., a significant change in a feature is considered to occur when the rate of change exceeds 0.02), and adaptive adjustments were made using the mean and standard deviation of the feature change rate. The adjusted threshold was 0.015. Based on this threshold, the rate of change of the age feature between the fourth and fifth months was 0.03, exceeding the threshold, and was therefore judged as a significant change. The output feature gradient trend index was 0.75 (within the range of 0-1, the larger the value, the more significant the change trend).
[0101] By comprehensively considering the high-frequency coefficients of the three layers, a significant nonlinear mutation was detected between the fourth and fifth months. A mutation detection window was constructed with a width of two months. Within this window, mutation characteristic parameters were calculated, including mutation amplitude (0.03), mutation duration (1 month), and the difference between the mean values before and after the mutation (0.025). Based on these parameters, a characteristic mutation degree index of 0.8 was generated, indicating a high degree of mutation.
[0102] Dynamic time warping (DTW) algorithms are used to process feature gradual change trend indicators and feature abrupt change degree indicators. Taking age and income features as examples, their temporal similarity over a five-month time scale is calculated. By constructing a cost matrix and calculating the optimal alignment path, a temporal similarity feature value of 0.65 is obtained, indicating that the change patterns of the two features have a certain degree of similarity. Based on this temporal similarity feature, periodic change features and abrupt change features are extracted. Analysis reveals that the age feature exhibits a significant increase in information content every three months, with a periodic change feature value of 0.7; simultaneously, an abrupt change occurs in the fifth month, with an abrupt change feature value of 0.8.
[0103] A multidimensional feature importance assessment index was constructed by integrating indicators of gradual change trends, abrupt change degrees, periodic changes, and sudden changes. For age, the multidimensional importance index is [0.75, 0.8, 0.7, 0.8], with a final importance score of 0.77 obtained through weighted averaging. For income, the multidimensional importance index is [0.6, 0.5, 0.4, 0.3], with a final importance score of 0.48. For consumption frequency, the multidimensional importance index is [0.85, 0.75, 0.6, 0.7], with a final importance score of 0.73. Based on the final scores, these three features can be ranked by importance as follows: Age > Consumption Frequency > Income, providing a scientific basis for subsequent customer profile model construction and feature selection.
[0104] In practical applications, this method can capture the changing trends and abrupt changes in feature information in real time, making it suitable for customer profile feature evaluation in dynamic environments. By adjusting the wavelet decomposition level, adaptive threshold parameters, and similarity calculation methods, it can flexibly address the needs of different business scenarios and achieve accurate evaluation of the importance of customer profile features.
[0105] In one optional implementation, a sliding time window method is used to perform time-series partitioning of the customer profile weighted feature vector to obtain customer behavior feature sequences for multiple time windows, including:
[0106] A time-series sliding window is constructed based on the weighted feature vector of the customer profile. The time sensitivity matrix of the features is calculated based on the time-series sliding window. The key time scale is determined according to the singular value decomposition result of the time sensitivity matrix. The key time scale is mapped to the initial window length and initial sliding step of the time-series sliding window.
[0107] Based on the initial window length and initial sliding step, the customer profile weighted feature vector is decomposed into time series to obtain the changing feature parameters of the customer behavior feature sequence.
[0108] Based on the changing feature parameters, a swarm intelligence optimization algorithm is used to optimize the initial window length and initial sliding step size to obtain the optimal window parameters;
[0109] The overlapping interval between adjacent time windows is determined using the optimal window parameters. The continuous change value of the feature is calculated based on the overlapping interval. The granularity of window division is determined based on the continuous change value and the time sensitivity matrix.
[0110] Based on the granularity of the window division and the changing feature parameters, the features of the overlapping intervals are weighted and fused to obtain customer behavior feature sequences for multiple time windows.
[0111] In practical applications, obtaining weighted feature vectors of customer profiles is a prerequisite for time-series segmentation. These feature vectors typically contain multi-dimensional information such as customer consumption habits, browsing history, and interaction behavior. Taking an e-commerce platform as an example, a customer's feature vector may consist of more than thirty indicators, such as purchase frequency, average order value, and visit duration, with each indicator assigned different weights based on its business importance.
[0112] When constructing a time-series sliding window based on customer profile weighted feature vectors, the feature data is arranged chronologically to form a time-series matrix. Analysis of a customer case on an e-commerce platform shows that their consumption behavior over six months formed a 180×35 time-series feature matrix. The system then calculates the time sensitivity matrix of the features, specifically analyzing the fluctuation of each feature at different time granularities. For example, for the purchase frequency feature, the system calculates its rate of change at daily, weekly, and monthly scales. A higher rate of change indicates that the feature is more sensitive at the corresponding time scale. For the aforementioned customer, the purchase frequency feature achieved a sensitivity value of 0.73 at the weekly scale, significantly higher than 0.45 at the daily scale and 0.61 at the monthly scale.
[0113] The system performs singular value decomposition on the time sensitivity matrix to extract key eigenvectors and determine critical time scales. In the actual case, the first three singular values after decomposition were 132.5, 87.3, and 41.2, accounting for more than 85% of the total energy, and the corresponding time scales were identified as the main cycles of customer behavior changes. These key time scales were mapped to the initial parameters of a time-series sliding window; for example, the first key scale was mapped to an initial window length of 7 days and an initial sliding step size of 3 days.
[0114] Based on the initial window length and sliding step size, the system performs time-series decomposition on the weighted feature vector of the customer profile, dividing six months of feature data into approximately 60 time windows. Within each window, the system calculates feature statistics such as mean, standard deviation, maximum and minimum values, forming variable feature parameters. Taking purchase frequency as an example, the mean of the first window is 0.42 times / day, and the standard deviation is 0.18; while the mean of the second window is 0.38 times / day, and the standard deviation is 0.21, reflecting the time-varying characteristics of customer behavior.
[0115] Based on the obtained variable feature parameters, the system employs a swarm intelligence optimization algorithm to optimize the initial window parameters. In practice, particle swarm optimization is commonly used, with the number of particles set to 50 and the maximum number of iterations set to 200. During the optimization process, the objective function is defined as a weighted sum of feature consistency within the window and differences between windows. For the aforementioned customer data, after 156 iterations, the algorithm converges to an optimal window length of 9 days and a sliding step size of 2 days, at which point the objective function value is 0.854.
[0116] The system uses optimal window parameters to determine the overlapping interval between adjacent time windows. In the above case, the window length is 9 days, the sliding step is 2 days, and the overlapping interval between adjacent windows is 7 days. The system calculates the continuity change value of features based on the overlapping interval by comparing the differences in feature performance under different windows. If the difference is too large, it indicates that the interval may contain abrupt behavioral changes; if the difference is small, it indicates that the behavior is continuous and stable. For example, if a customer's purchase frequency continuity change value is 0.08 within an overlapping interval, it indicates that the behavior is relatively stable within that interval.
[0117] In principle, regions with high values of continuous change can be divided into finer-grained segments to capture more subtle behavioral changes. In practical applications, the system quantizes the granularity values into 1-5 levels, determined based on the distribution of continuous change values. In this case, the system maps the range of continuous change values from 0.05 to 0.15 to a granularity level of 3, with a corresponding segmentation interval of 1 day.
[0118] Based on a defined window granularity and variation feature parameters, the system performs weighted fusion of features across overlapping regions. The fusion weights are determined by the feature's performance within each window, typically using a Gaussian kernel function to weight features based on their distance from the window center. In the example above, the features on the first day of the overlapping region are weighted at 0.25 and 0.75, the second day at 0.3 and 0.7, and so on, generating a smooth and continuous feature transition. Ultimately, the system outputs customer behavior feature sequences across multiple time windows, each sequence containing a feature vector of customer behavior within that window and its changing trend. These sequences serve as the foundational data for subsequent customer behavior analysis and prediction.
[0119] In one optional implementation, a temporal correlation analysis is performed on the customer behavior feature sequences of the multiple time windows, the dynamic correlation coefficient between customer behavior feature sequences of adjacent time windows is calculated, and a customer behavior fusion feature with temporal dependency is generated based on the dynamic correlation coefficient, including:
[0120] Calculate the temporal adjacency relationship for the customer behavior feature sequences of the multiple time windows, calculate the attenuation distance between adjacent windows based on the temporal adjacency relationship, and construct the temporal propagation attenuation coefficient and feature correlation matrix based on the attenuation distance;
[0121] The temporal propagation attenuation coefficient and the feature correlation matrix are recursively iterated to extract the temporal state change sequence, and a dynamic correlation coefficient is generated based on the temporal state change sequence.
[0122] A local temporal attention score is constructed based on the dynamic correlation coefficient. The feature correlation matrix is weighted using the local temporal attention score. The weighted feature correlation matrix is then combined with the temporal propagation attenuation coefficient to generate a global attention vector.
[0123] Based on the global attention vector, a temporal weight is assigned to each time window. The temporal weight is then combined with the customer behavior feature sequence in a weighted manner. Combined with the dynamic correlation coefficient, long-range dependency information is extracted to generate customer behavior fusion features with temporal dependencies.
[0124] This invention provides a method for generating customer behavior fusion features with temporal dependencies. In this method, firstly, temporal correlation analysis is performed on customer behavior feature sequences from multiple time windows to calculate the dynamic correlation coefficient between customer behavior feature sequences from adjacent time windows, and then customer behavior fusion features with temporal dependencies are generated based on this dynamic correlation coefficient.
[0125] When calculating the temporal adjacency relationship for customer behavior feature sequences across multiple time windows, the system divides customer behavior data into multiple time windows according to chronological order. For example, each window contains 24 hours of behavior data. Assume a customer's behavior features within three consecutive time windows are as follows: Window 1 includes "browsing product A 5 times, adding product B to cart 2 times, and viewing coupons 1 time"; Window 2 includes "browsing product A 3 times, browsing product C 4 times, and paying for an order 1 time"; Window 3 includes "querying orders 2 times, browsing product D 6 times, and requesting a refund 1 time". The system converts these behavioral features into feature vectors and establishes a temporal adjacency matrix to record the connection relationships between adjacent windows. For example, the adjacency relationship between window 1 and window 2 is 1, and the adjacency relationship between window 1 and window 3 is 0.
[0126] When calculating the attenuation distance between adjacent windows based on temporal adjacency, the system calculates the time difference between each pair of adjacent windows and converts it into an attenuation factor. For example, if window 1 and window 2 are 24 hours apart, and window 2 and window 3 are 48 hours apart, the attenuation distance can be set as 0.8 from window 1 to window 2 and 0.6 from window 2 to window 3. For non-adjacent windows, the comprehensive attenuation value is calculated through path propagation; for example, the attenuation value from window 1 to window 3 is 0.8 × 0.6 = 0.48.
[0127] In constructing the temporal propagation attenuation coefficient and feature correlation matrix based on the attenuation distance, the system performs correlation analysis on customer behavior features within each time window. For example, a feature correlation matrix is generated by calculating the cosine similarity between feature vectors. Assuming the feature correlation between window 1 and window 2 is 0.7, and the correlation between window 2 and window 3 is 0.5, a complete feature correlation matrix is constructed. Simultaneously, a temporal propagation attenuation coefficient matrix is constructed based on the previously calculated attenuation distance, representing the information transmission strength between different time windows over time.
[0128] When recursively calculating the temporal propagation attenuation coefficient and the feature correlation matrix, the system sets the number of iterations (e.g., 10) and initializes the state matrix. In each iteration, the temporal propagation attenuation coefficient matrix is multiplied by the current state matrix, and then combined with the feature correlation matrix to update the state matrix. After multiple iterations, a stable temporal state change sequence is extracted, which reflects the evolution of customer behavior across different time windows. For example, the initial state [0.5, 0.5, 0.5] converges to [0.7, 0.6, 0.4] after iterations, indicating that the earlier window contributes more to the final state.
[0129] When generating dynamic correlation coefficients based on time-series state change sequences, the system compares the magnitude and direction of state changes between adjacent iteration steps to calculate the degree of dynamic correlation between consecutive windows. For example, if the state change from window 1 to window 2 is a positive increase of 0.2, and the state change from window 2 to window 3 is a negative change of -0.1, then the dynamic correlation coefficients can be obtained as 0.8 and -0.4, respectively, indicating that window 1 and window 2 are positively correlated, and window 2 and window 3 are negatively correlated.
[0130] In constructing local temporal attention scores based on dynamic correlation coefficients, the system calculates attention weights for each pair of adjacent windows. For example, a window pair with a dynamic correlation coefficient of 0.8 receives a higher attention score of 0.9, while a window pair with a dynamic correlation coefficient of -0.4 receives a lower attention score of 0.3. These local temporal attention scores reflect the continuity and correlation of customer behavior in adjacent time periods.
[0131] When weighting the feature relevance matrix using local temporal attention scores, the system multiplies the attention score by each element of the feature relevance matrix. For example, the feature relevance of window 1 and window 2, 0.7, multiplied by the attention score of 0.9, yields a weighted value of 0.63. After all weighting operations are completed, the weighted feature relevance matrix is combined with the temporal propagation attenuation coefficient to generate a global attention vector [0.63, 0.48, 0.15]. This vector comprehensively considers both temporal attenuation and feature relevance dimensions.
[0132] When assigning time-series weights to each time window based on the global attention vector, the system normalizes the global attention vector to obtain the final weights for each window [0.5, 0.38, 0.12]. These weights reflect the degree of influence of each time window on the customer's current behavioral tendencies, and windows with higher weights have a stronger indicative effect on predicting the customer's future behavior.
[0133] When weighting the temporal weights with the customer behavior feature sequence, the system multiplies the original feature vector of each window by its corresponding temporal weight. For example, the behavior feature vector of window 1 [5,2,1,0,0,0,0] is multiplied by a weight of 0.5 to obtain the weighted feature [2.5,1,0.5,0,0,0,0]. The weighted features of all windows are then merged to generate preliminary fused features.
[0134] In the process of extracting long-range dependency information by combining dynamic correlation coefficients, the system identifies behavioral patterns across multiple time windows. For example, it finds a long-range dependency between the behavior of "browsing product A" in window 1 and the behavior of "querying orders" in window 3, with a correlation coefficient of 0.6. The system will also incorporate these long-range dependency features into the fusion process to enhance the model's ability to capture complex temporal patterns.
[0135] Finally, the system integrates the weighted combined features with long-range dependency information to generate a fusion feature of customer behavior containing complete time-series information. For example, the final fusion feature vector is obtained as [3.2, 1.5, 0.8, 1.4, 2.2, 0.6, 1.0], where each dimension represents the comprehensive strength of different types of behavior, integrating information from time decay, window correlation, and long-range dependency. This fusion feature can comprehensively reflect the evolution of customer behavior and changes in preferences, providing strong support for subsequent applications such as customer behavior prediction, precision marketing, and risk control.
[0136] In one optional implementation, a local temporal attention score is constructed based on the dynamic correlation coefficient. The feature correlation matrix is then weighted using this local temporal attention score. Finally, the weighted feature correlation matrix is combined with the temporal propagation attenuation coefficient to generate a global attention vector, including:
[0137] The local similarity matrix of the time-series features is calculated based on the dynamic correlation coefficient. A transition probability matrix is constructed on the local similarity matrix. The steady-state distribution of the features is calculated based on the transition probability matrix. The entropy value of the steady-state distribution is used as the local temporal attention score.
[0138] The local temporal attention scores are used to construct an attention gain function, and the feature correlation matrix is nonlinearly transformed based on the attention gain function to obtain a weighted feature matrix with temporal memory effect.
[0139] Based on the weighted feature matrix, the feature importance distribution is calculated, singular value decomposition is performed, the feature subspace corresponding to the main singular values is extracted, a time-series propagation path is constructed in the feature subspace, and the time-series propagation path is combined with the time-series propagation attenuation coefficient to generate dynamic attenuation features.
[0140] An initial attention vector is constructed based on the dynamic decay feature. The cross-time window correlation degree of the feature is calculated using the initial attention vector. The dynamic decay feature is recursively updated based on the cross-time window correlation degree. The updated dynamic decay feature is used as the global attention vector.
[0141] In this embodiment, the method of constructing a local temporal attention score based on the dynamic correlation coefficient, using the attention score to weight the feature correlation matrix, and finally combining it with the temporal propagation attenuation coefficient to generate a global attention vector will be described in detail.
[0142] When calculating the local similarity matrix of time-series features based on the dynamic correlation coefficient, the system first obtains the feature values of multiple consecutive time points. For example, for a time-series dataset containing 100 feature dimensions, a 100×30 feature matrix can be constructed within an observation window of 30 consecutive time points. A 30×30 local similarity matrix is then constructed by calculating the cosine similarity between each pair of time points in the feature matrix. Specifically, each element in the similarity matrix represents the degree of similarity between the feature vectors of time point i and time point j, with values ranging from -1 to 1; a larger value indicates a higher similarity. For example, the value of element (3,5) in the similarity matrix is 0.85, indicating that the feature vectors of the 3rd and 5th time points are highly similar.
[0143] When constructing a transition probability matrix from a local similarity matrix, each row of the similarity matrix is normalized to ensure that the sum of the elements in each row is 1, thus converting the similarity into a probability distribution. In implementation, this can be achieved by first performing a linear transformation on the similarity matrix to adjust its range to a non-negative interval, and then normalizing each row. For example, for the i-th row of the similarity matrix, all elements can be incremented by 1 and then divided by 2 to change the range from [-1,1] to [0,1]. Then, the range is divided by the sum of all elements in that row to obtain the probability distribution. The transition probability matrix constructed in this way represents the transition relationship between time points, where the element (i,j) represents the probability of transitioning from time point i to time point j.
[0144] When calculating the steady-state distribution of features based on the transition probability matrix, a power-law iteration method is used. Specifically, an initial probability distribution vector, such as a uniform distribution vector, is selected, with a length equal to the number of time points, and each element having a value of 1 / number of time points. This vector is repeatedly multiplied by the transition probability matrix until the result converges or a preset number of iterations (e.g., 100) is reached. The finally converged probability distribution vector is the steady-state distribution, representing the stationary probability distribution of the system at each time point after long-term evolution. In practical applications, when the Euclidean distance between two iterations is less than a preset threshold (e.g., 0.0001), the distribution can be considered to have converged.
[0145] When using the entropy value of the steady-state distribution as the local temporal attention score, the information entropy of the steady-state distribution vector is calculated. Specifically, for each element p in the steady-state distribution vector, -p×log(p) is calculated, and then summed to obtain the entropy value. A larger entropy value indicates higher uncertainty in the temporal characteristics, and correspondingly a higher local temporal attention score. For example, a uniform steady-state distribution has the largest entropy value, while a steady-state distribution close to an impulse distribution has a lower entropy value.
[0146] When constructing the attention gain function from local temporal attention scores, a sigmoid activation function is used to perform a nonlinear transformation on the attention scores. Specifically, the sigmoid function can be used to map the attention scores to the (0,1) interval, serving as the attention gain coefficient. For example, for an attention score x, 1 / (1+exp(-α×(x-β))) is calculated as the gain coefficient, where α controls the curve steepness and β controls the center point position. These values can be set according to the actual application scenario, such as α=5 and β=0.5.
[0147] When performing a nonlinear transformation on the feature correlation matrix based on the attention gain function, each element in the original feature correlation matrix is multiplied by its corresponding gain coefficient to obtain a weighted feature matrix. The feature correlation matrix represents the degree of correlation between different features; for example, for 100 features, a 100×100 correlation matrix is constructed. Through the nonlinear transformation of the gain function, the correlation of important features is enhanced, while the correlation of unimportant features is weakened, thus achieving weighting with a temporal memory effect.
[0148] When calculating the feature importance distribution based on the weighted feature matrix, the L2 norm is calculated for each row of the weighted feature matrix, resulting in a vector of length equal to the number of features. This vector represents the importance score of each feature. For example, for 100 features, a 100-dimensional importance vector is obtained, where features with larger values are more important.
[0149] When extracting the feature subspace corresponding to the principal singular values using singular value decomposition (SVD), the weighted feature matrix is decomposed to obtain a left singular vector matrix, a singular value diagonal matrix, and a right singular vector matrix. The singular values are sorted by magnitude, and the left singular vectors corresponding to the k largest singular values are selected to form the feature subspace. The value of k can be determined based on the cumulative contribution rate; for example, the k singular values with a cumulative contribution rate of 90% can be selected. In practical applications, if the original feature dimension is 100, it may be sufficient to retain the feature subspace corresponding to the first 10 principal singular values to capture most of the information.
[0150] When constructing a temporal propagation path in the feature subspace, the original features are projected onto the feature subspace to obtain a dimensionality-reduced feature representation. Then, the change vectors of features between adjacent time points are calculated, and these change vectors constitute the temporal propagation path. For example, for a time window of length 30, the change vectors between 29 adjacent time points are calculated to form the temporal propagation path.
[0151] When combining the temporal propagation path with the temporal propagation attenuation coefficient to generate dynamic attenuation features, a time-distance-dependent attenuation coefficient is introduced, where historical information further away from the current time point has a smaller impact. The attenuation coefficient can be designed in an exponential decay form; for example, for historical information with a time interval of t, its attenuation coefficient is exp(-λt), where λ is the attenuation rate parameter, which can be set according to the actual application, such as λ = 0.1. Multiplying each change vector on the propagation path by its corresponding attenuation coefficient yields the dynamic attenuation feature.
[0152] When constructing the initial attention vector based on the dynamic decay features, the dynamic decay features are normalized to obtain a probability distribution vector with a sum of 1, which serves as the initial attention vector. This vector reflects the importance of different time points to the current prediction; the larger the value in the vector, the more important the time point.
[0153] When calculating the cross-time window correlation of features using the initial attention vector, the initial attention vector is multiplied by the original feature matrix to obtain a weighted feature representation. Then, the cosine similarity between this representation and the feature representations of different historical time windows is calculated as the cross-time window correlation. For example, the correlation between the current time window and the past 5 time windows can be calculated to obtain a correlation vector of length 5.
[0154] When recursively updating dynamic decay features based on cross-time window correlation, the dynamic decay features are weighted and adjusted according to the magnitude of the correlation. Specifically, the correlation vector can be normalized and then weighted and combined with the dynamic decay features of the corresponding time window to obtain the updated dynamic decay features. This step can be repeated multiple times, for example, recursively three times, until the dynamic decay features tend to stabilize. Finally, the updated dynamic decay features are used as the global attention vector, which comprehensively considers both local temporal correlation and global temporal evolution characteristics.
[0155] Figure 2 This diagram illustrates the temporal propagation path and attenuation coefficient analysis. It showcases the dynamic attenuation mechanism based on temporal features, primarily presenting three key curves and their interrelationships. The blue solid line represents the temporal propagation path in the feature subspace, reflecting the evolution of the original eigenvalues within a continuous time window. It shows that the eigenvalues from t1 to t... 21 The fluctuations during this period are particularly pronounced at t 11A significant peak (0.80) appears at time 1. The orange dashed line represents the exponential time-series propagation decay coefficient, demonstrating the natural decay of information importance over time.
[0156] The green dotted line represents the dynamic decay characteristic curve, composed of the temporal propagation path and the decay coefficient, reflecting the "temporal memory effect"—recent information has a greater weight, while the influence of long-term information decreases. Pay special attention to t. 11 At key change points, although the original eigenvalues reach their peak, their influence gradually weakens at subsequent time points after attenuation processing. Similarly, t 19 The characteristic rebound (0.60) at that point was also adjusted by a decay factor to conform to the importance assignment related to time distance.
[0157] The feature propagation analysis box indicates t 11 The eigenvalues at time t increased significantly. 15 -t 19 Key observations include the slow decay of features across intervals. This dynamic decay mechanism effectively balances the importance of historical information with the latest data in practical applications, providing a more reasonable feature representation for time series prediction.
[0158] In one optional implementation, a feature relevance matrix is calculated based on the customer behavior fusion features and the customer profile weighted feature vector. Based on the feature relevance matrix, the customer's consumption behavior characteristics in the next time window are predicted, including:
[0159] Based on the customer behavior fusion features and the customer profile weighted feature vector, a feature relevance matrix is calculated. The feature relevance matrix is then decomposed into eigenvalues to obtain a feature importance vector. Time-varying weight coefficients are then constructed based on the feature importance vector.
[0160] The conditional transition probability of the feature is calculated using the time-varying weight coefficients, and a time-series state transition matrix is constructed based on the conditional transition probability. The optimal feature distribution parameters are obtained by maximizing the log-likelihood of the state transition.
[0161] Periodic components are extracted from the optimal feature distribution parameters, and the periodic components are separated from the trend components. Short-term fluctuation characteristics and long-term change characteristics are calculated separately, and time-series combined features are generated based on the short-term fluctuation characteristics and long-term change characteristics.
[0162] The conditional probability distribution between features is calculated using the time-series combined features. A feature dependency graph is constructed based on the conditional probability distribution. By iteratively calculating the state probability of feature nodes, the consumption behavior characteristics of customers in the next time window are predicted.
[0163] In this embodiment, a feature correlation matrix is calculated based on customer behavior fusion features and customer profile weighted feature vectors, and the customer's consumption behavior features in the next time window are predicted based on this matrix.
[0164] Specifically, after acquiring the customer's behavioral fusion features and customer profile weighted feature vector, the system calculates the correlation between the two, forming a feature correlation matrix. Each element in this matrix represents the degree of correlation between behavioral features and profile features. For example, for a certain customer, the correlation between their purchase frequency feature and age feature might be 0.75, and the correlation with income feature might be 0.82. The system performs eigenvalue decomposition on the feature correlation matrix to obtain a feature importance vector. The elements in this vector represent the importance of each feature in the prediction model. For example, for spending amount prediction, the importance of the purchase frequency feature might be 0.85, and the importance of the browsing time feature might be 0.63. Based on the feature importance vector, the system constructs time-varying weight coefficients, allowing the weights to be adjusted over time. For example, during holidays, the weight of promotional activity features can be increased from 0.6 on weekdays to 0.85.
[0165] Using time-varying weighting coefficients, the system calculates the conditional transition probabilities of features. These probabilities represent the probability distribution of the feature's state at the next time point, given its current state. For example, if a customer's current monthly spending level is medium and they recently browsed high-end products, the probability of their spending level increasing next month is 0.65. Based on these probabilities, the system constructs a time-series state transition matrix, which describes the transition pattern of feature states over time. The system determines the optimal feature distribution parameters by maximizing the log-likelihood function of the state transitions. For example, for the customer's purchase cycle feature, the system might determine that it conforms to a distribution with a mean of 15 days and a standard deviation of 3 days.
[0166] The system extracts periodic components from the optimal feature distribution parameters, separating periodic changes from long-term trends. For example, the system might detect a periodic pattern of a customer making a large purchase every 30 days, along with a long-term trend of a 2% monthly increase in spending. The system calculates short-term fluctuation features and long-term change features separately. Short-term fluctuation features capture temporary changes in consumer behavior, such as responses to promotional activities; long-term change features reflect gradual shifts in consumption habits, such as increased brand loyalty. Taking customer A as an example, their short-term fluctuation is a 35% increase in weekend spending, while their long-term change is a monthly increase in interest in high-end products. The system combines short-term fluctuation features and long-term change features to generate time-series combined features, comprehensively describing the customer's behavioral patterns.
[0167] By utilizing temporal combination features, the system calculates the conditional probability distribution among features, representing the degree to which a change in one feature value affects other features. For example, when a customer's browsing time increases by 50%, the conditional probability of a 0.3 increase in the purchase probability is 0.72. Based on the conditional probability distribution, the system constructs a feature dependency graph, where nodes represent features, edges represent conditional dependencies, and edge weights represent dependency strength. The system predicts the customer's consumption behavior characteristics within the next time window by iteratively calculating the state probabilities of feature nodes.
[0168] Taking an electronics consumer as an example, the system obtains their historical purchase records, browsing behavior, and personal information to form behavioral fusion features and profile features. The calculated feature correlation matrix shows that the customer's browsing time is correlated with purchase decision by 0.78, and the purchase interval is correlated with price sensitivity by 0.65. After eigenvalue decomposition, the importance of brand preference features is 0.82, and the importance of price range features is 0.75. The time-varying weights constructed by the system increase the promotion sensitivity weight to 0.88 before holidays.
[0169] Conditional transition probability calculations indicate that the customer has a 0.37 probability of switching from mid-range to high-end products. The time-series state transition matrix shows that the customer considers purchasing a new product on average once every 45 days. Optimal feature distribution parameters indicate that the customer's purchase decision follows a decision cycle with a mean of 12 days and a variance of 4 days. Periodic component analysis reveals a large-scale purchase pattern once per quarter, and a long-term trend shows that interest in smart products increases by 3% per month. Short-term fluctuations indicate that promotional activities increase the likelihood of purchase by 40%, while long-term trends show that brand loyalty increases by 2% per month.
[0170] After generating the time-series combined features, the system calculates that the probability of a customer making a purchase is 0.63 when the customer browses a specific high-end product more than 5 times. The feature dependency graph shows a dependency strength of 0.76 between browsing duration and purchase decision, and 0.82 between promotional activities and purchase timing. Through iterative calculations, the system predicts that the customer has a 0.71 probability of purchasing a smart home appliance priced between 2000-3000 yuan within the next 30-day time window, and the purchase is likely to occur during end-of-month promotional activities. Based on this prediction, merchants can proactively push relevant product information and design targeted promotional strategies to improve conversion rates.
[0171] Figure 3This diagram illustrates the performance comparison of consumer behavior feature prediction models. Overall, the time-series combined feature model performs best across all evaluation dimensions, particularly demonstrating a significant advantage in periodic pattern recognition, achieving a high score of 0.8, far exceeding the base model's 0.4 and the feature correlation matrix model's 0.5. This result validates the effectiveness of the periodic component extraction method described in the text, proving that this technology can accurately capture cyclical patterns in customer consumption behavior.
[0172] The feature relevance matrix model shows a significant improvement over the basic model, especially in the feature relevance capture index, which increased from 0.5 to 0.7. This indicates that the feature importance vector obtained through eigenvalue decomposition can effectively identify the correlations between key features. Meanwhile, in terms of conditional probability accuracy, the feature relevance matrix model achieves 0.6, a significant improvement over the basic model's 0.4, but still lower than the time-series combined feature model's 0.7.
[0173] The overall predictive performance metrics reflect the comprehensive performance of the models. The time-series combined feature model leads with a score of 0.8, while the feature correlation matrix model and the basic model score 0.7 and 0.5, respectively. This result confirms that the method of combining short-term fluctuation features with long-term variation features can comprehensively describe customer behavior patterns and provide more accurate predictions.
[0174] This invention provides a real-time customer profile dynamic update and prediction system for retail scenarios, comprising:
[0175] The first unit is used to determine the set of customer profile features corresponding to customer behavior data in retail scenarios;
[0176] The second unit is used to calculate the feature importance score based on the information gain value of different customer profile features in the customer profile feature set under different time windows, and generate a customer profile weighted feature vector based on the feature importance score.
[0177] The third unit is used to perform time-series partitioning of the customer profile weighted feature vector using a sliding time window method to obtain customer behavior feature sequences for multiple time windows.
[0178] The fourth unit is used to perform temporal correlation analysis on the customer behavior feature sequences of the multiple time windows, calculate the dynamic correlation coefficient between customer behavior feature sequences of adjacent time windows, and generate customer behavior fusion features with temporal dependencies based on the dynamic correlation coefficient.
[0179] The fifth unit is used to calculate a feature correlation matrix based on the customer behavior fusion features and the customer profile weighted feature vector, and to predict the customer's consumption behavior features in the next time window based on the feature correlation matrix.
[0180] The sixth unit is used to update the customer profile feature set based on the consumption behavior characteristics and the customer behavior fusion characteristics, and to determine personalized product recommendations and marketing strategies for customers based on the updated customer profile feature set.
[0181] A third aspect of the present invention provides an electronic device, comprising:
[0182] processor;
[0183] Memory used to store processor-executable instructions;
[0184] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0185] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0186] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0187] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A real-time customer profile dynamic update and prediction method for retail scenarios, characterized in that, include: Determine the set of customer profile features corresponding to customer behavior data in retail scenarios; The feature importance score is calculated based on the information gain value of different customer profile features in the customer profile feature set under different time windows, and a weighted feature vector of the customer profile is generated based on the feature importance score. The customer profile weighted feature vector is divided into time series using a sliding time window method to obtain customer behavior feature sequences for multiple time windows. A temporal correlation analysis is performed on the customer behavior feature sequences of the multiple time windows, the dynamic correlation coefficient between customer behavior feature sequences of adjacent time windows is calculated, and a customer behavior fusion feature with temporal dependency is generated based on the dynamic correlation coefficient. Based on the customer behavior fusion features and the customer profile weighted feature vector, a feature correlation matrix is calculated, and based on the feature correlation matrix, the customer's consumption behavior features in the next time window are predicted. Based on the consumer behavior characteristics and the customer behavior fusion characteristics, update the customer profile feature set, and determine personalized product recommendations and marketing strategies for customers based on the updated customer profile feature set; Based on the information gain values of different customer profile features in the customer profile feature set under different time windows, a feature importance score is calculated. A weighted feature vector for the customer profile is then generated based on the feature importance score, including: The customer profile feature set is divided into time windows according to a preset time interval to obtain customer profile feature sequences for multiple time windows; Construct an information entropy matrix for each customer profile feature in the customer profile feature sequence, wherein the information entropy matrix contains the information entropy value of each customer profile feature under different time windows; Based on the information entropy matrix, a multidimensional feature importance evaluation index is determined by calculating the rate of change of information entropy between adjacent time windows. The multidimensional feature importance evaluation index reflects the sensitivity of customer profile features to changes in customer behavior. The customer profile features are hierarchically clustered based on the multidimensional feature importance evaluation index to obtain the feature hierarchical clustering results; Based on the feature hierarchical clustering results, feature importance scores are set, and a time-series decay function is used to dynamically adjust the feature importance scores to generate a time-series decay coefficient. The time-series decay coefficient is directly proportional to the feature fluctuation amplitude and inversely proportional to the feature change period. The time-series decay coefficient is weighted and calculated with the customer profile features to generate a customer profile weighted feature vector.
2. The method according to claim 1, characterized in that, Based on the information entropy matrix, a multidimensional feature importance evaluation index is determined by calculating the rate of change of information entropy between adjacent time windows, including: Based on the information entropy matrix, calculate the information entropy difference sequence between adjacent time windows, and perform wavelet multi-scale decomposition on the information entropy difference sequence to obtain multi-layer wavelet coefficients; A linear trend curve is obtained by fitting the low-frequency components of the multilayer wavelet coefficients, and a nonlinear abrupt change curve is obtained by fitting the high-frequency components of the multilayer wavelet coefficients. Based on the derivative of the linear trend curve, the rate of change and acceleration of change of customer profile features are calculated. An adaptive decision threshold is constructed based on the rate of change and acceleration of change. Based on the adaptive decision threshold, a feature gradient trend index is output. A mutation detection window is constructed based on the nonlinear mutation curve. Within the mutation detection window, mutation feature parameters of customer profile features are calculated, and a feature mutation degree index is generated based on the mutation feature parameters. The feature gradual change trend index and feature mutation degree index are processed by the dynamic time warping algorithm to obtain the temporal similarity features of customer profile features at different time scales. Based on the temporal similarity features, the periodic change features and sudden change features of the features are extracted to obtain the multidimensional feature importance evaluation index.
3. The method according to claim 1, characterized in that, The customer profile weighted feature vector is divided into time series using a sliding time window method to obtain customer behavior feature sequences for multiple time windows, including: A time-series sliding window is constructed based on the weighted feature vector of the customer profile. The time sensitivity matrix of the features is calculated based on the time-series sliding window. The key time scale is determined according to the singular value decomposition result of the time sensitivity matrix. The key time scale is mapped to the initial window length and initial sliding step of the time-series sliding window. Based on the initial window length and initial sliding step, the customer profile weighted feature vector is decomposed into time series to obtain the changing feature parameters of the customer behavior feature sequence. Based on the changing feature parameters, a swarm intelligence optimization algorithm is used to optimize the initial window length and initial sliding step size to obtain the optimal window parameters; The overlapping interval between adjacent time windows is determined using the optimal window parameters. The continuous change value of the feature is calculated based on the overlapping interval. The granularity of window division is determined based on the continuous change value and the time sensitivity matrix. Based on the granularity of the window division and the changing feature parameters, the features of the overlapping intervals are weighted and fused to obtain customer behavior feature sequences for multiple time windows.
4. The method according to claim 1, characterized in that, Perform temporal correlation analysis on the customer behavior feature sequences of the multiple time windows, calculate the dynamic correlation coefficient between customer behavior feature sequences of adjacent time windows, and generate customer behavior fusion features with temporal dependencies based on the dynamic correlation coefficient, including: Calculate the temporal adjacency relationship for the customer behavior feature sequences of the multiple time windows, calculate the attenuation distance between adjacent windows based on the temporal adjacency relationship, and construct the temporal propagation attenuation coefficient and feature correlation matrix based on the attenuation distance; The temporal propagation attenuation coefficient and the feature correlation matrix are recursively iterated to extract the temporal state change sequence, and a dynamic correlation coefficient is generated based on the temporal state change sequence. A local temporal attention score is constructed based on the dynamic correlation coefficient. The feature correlation matrix is weighted using the local temporal attention score. The weighted feature correlation matrix is then combined with the temporal propagation attenuation coefficient to generate a global attention vector. Based on the global attention vector, a temporal weight is assigned to each time window. The temporal weight is then combined with the customer behavior feature sequence in a weighted manner. Combined with the dynamic correlation coefficient, long-range dependency information is extracted to generate customer behavior fusion features with temporal dependencies.
5. The method according to claim 4, characterized in that, A local temporal attention score is constructed based on the dynamic correlation coefficient. This local temporal attention score is then used to weight the feature correlation matrix. Finally, the weighted feature correlation matrix is combined with the temporal propagation attenuation coefficient to generate a global attention vector, including: The local similarity matrix of the time-series features is calculated based on the dynamic correlation coefficient. A transition probability matrix is constructed on the local similarity matrix. The steady-state distribution of the features is calculated based on the transition probability matrix. The entropy value of the steady-state distribution is used as the local temporal attention score. The local temporal attention scores are used to construct an attention gain function, and the feature correlation matrix is nonlinearly transformed based on the attention gain function to obtain a weighted feature matrix with temporal memory effect. Based on the weighted feature matrix, the feature importance distribution is calculated, singular value decomposition is performed, the feature subspace corresponding to the main singular values is extracted, a time-series propagation path is constructed in the feature subspace, and the time-series propagation path is combined with the time-series propagation attenuation coefficient to generate dynamic attenuation features. An initial attention vector is constructed based on the dynamic decay feature. The cross-time window correlation degree of the feature is calculated using the initial attention vector. The dynamic decay feature is recursively updated based on the cross-time window correlation degree. The updated dynamic decay feature is used as the global attention vector.
6. The method according to claim 1, characterized in that, Based on the customer behavior fusion features and the customer profile weighted feature vector, a feature relevance matrix is calculated. Based on the feature relevance matrix, the customer's consumption behavior characteristics in the next time window are predicted, including: Based on the customer behavior fusion features and the customer profile weighted feature vector, a feature relevance matrix is calculated. The feature relevance matrix is then decomposed into eigenvalues to obtain a feature importance vector. Time-varying weight coefficients are then constructed based on the feature importance vector. The conditional transition probability of the feature is calculated using the time-varying weight coefficients, and a time-series state transition matrix is constructed based on the conditional transition probability. The optimal feature distribution parameters are obtained by maximizing the log-likelihood of the state transition. Periodic components are extracted from the optimal feature distribution parameters, and the periodic components are separated from the trend components. Short-term fluctuation characteristics and long-term change characteristics are calculated separately, and time-series combined features are generated based on the short-term fluctuation characteristics and long-term change characteristics. The conditional probability distribution between features is calculated using the time-series combined features. A feature dependency graph is constructed based on the conditional probability distribution. By iteratively calculating the state probability of feature nodes, the consumption behavior characteristics of customers in the next time window are predicted.
7. A real-time customer profile dynamic update and prediction system for retail scenarios, used to implement the method as described in any one of claims 1-6, characterized in that, include: The first unit is used to determine the set of customer profile features corresponding to customer behavior data in retail scenarios; The second unit is used to calculate the feature importance score based on the information gain value of different customer profile features in the customer profile feature set under different time windows, and generate a customer profile weighted feature vector based on the feature importance score. The third unit is used to perform time-series partitioning of the customer profile weighted feature vector using a sliding time window method to obtain customer behavior feature sequences for multiple time windows. The fourth unit is used to perform temporal correlation analysis on the customer behavior feature sequences of the multiple time windows, calculate the dynamic correlation coefficient between customer behavior feature sequences of adjacent time windows, and generate customer behavior fusion features with temporal dependencies based on the dynamic correlation coefficient. The fifth unit is used to calculate a feature correlation matrix based on the customer behavior fusion features and the customer profile weighted feature vector, and to predict the customer's consumption behavior features in the next time window based on the feature correlation matrix. The sixth unit is used to update the customer profile feature set according to the consumption behavior characteristics and the customer behavior fusion characteristics, and to determine personalized product recommendations and marketing strategies for customers based on the updated customer profile feature set; Unit 2 is used for: The customer profile feature set is divided into time windows according to a preset time interval to obtain customer profile feature sequences for multiple time windows; Construct an information entropy matrix for each customer profile feature in the customer profile feature sequence, wherein the information entropy matrix contains the information entropy value of each customer profile feature under different time windows; Based on the information entropy matrix, a multidimensional feature importance evaluation index is determined by calculating the rate of change of information entropy between adjacent time windows. The multidimensional feature importance evaluation index reflects the sensitivity of customer profile features to changes in customer behavior. The customer profile features are hierarchically clustered based on the multidimensional feature importance evaluation index to obtain the feature hierarchical clustering results; Based on the feature hierarchical clustering results, feature importance scores are set, and a time-series decay function is used to dynamically adjust the feature importance scores to generate a time-series decay coefficient. The time-series decay coefficient is directly proportional to the feature fluctuation amplitude and inversely proportional to the feature change period. The time-series decay coefficient is weighted and calculated with the customer profile features to generate a customer profile weighted feature vector.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Petrochemical area inspection system based on inspection robot
CN118656734A
Customer portrait key data mining method and system based on space-time big data
CN118797542A