Customer personalization recommendation system and method based on retail data
Patent Information
- Application Number
- CN202610890201.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-11
AI Technical Summary
对于行为模式变化较快的客户,固定的更新机制会导致推荐数据滞后,无法及时响应其需求变化;而对于行为偏好相对稳定的客户,过于频繁的更新则可能破坏推荐的连贯性,影响客户体验
通过整合客户历史行为画像与实时行为画像,实现了对客户行为特征的全面刻画。历史行为画像能够反映客户长期形成的消费偏好与行为习惯,实时行为画像则能捕捉客户当下的需求动态与行为倾向,两类画像的结合打破了单一数据维度的局限性,让对客户行为的认知更具完整性与立体性。通过生成行为偏差向量,能够直观呈现客户实时行为与历史行为之间的差异,清晰勾勒出客户行为的变化轨迹,避免了对行为变化的模糊判断。
Smart Images

Figure CN122736670A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of personalized recommendation technology for retail customers, specifically to a personalized recommendation system and method based on retail data. Background Technology
[0002] With the deepening of digital transformation in the retail industry, personalized recommendations have become an important means to improve retail efficiency and optimize the consumer experience. Currently, most mainstream personalized recommendation methods in retail scenarios rely on historical customer behavior data as their core basis. By analyzing customers' past purchase records, browsing history, and other information, a fixed customer profile is constructed, and recommended content is matched based on this profile. The core logic of this type of method is based on the assumption that customers' consumption preferences are stable and that historical behavior can continuously reflect customers' true needs.
[0003] However, the retail market's consumption environment is highly dynamic, and customer preferences are often influenced by a variety of factors. For example, seasonal changes, promotional activities, shifts in fashion trends, and adjustments to individual needs can all lead to significant differences in customer behavior patterns at specific times compared to historical patterns. Existing recommendation methods, relying solely on historical behavioral profiles, struggle to capture fluctuations in customer preferences during real-time behavior, easily resulting in recommendations that are out of sync with current customer needs. While some methods attempt to incorporate real-time data, most simply overlay real-time behavioral data with historical data without effectively analyzing the differences between the two types of data, failing to accurately identify the trends and extent of changes in customer behavior.
[0004] Existing recommendation methods generally lack flexible data update mechanisms. In most cases, recommendation data updates either follow a fixed time period or are triggered based on preset behavior frequency thresholds, ignoring the personalized pace of change in different customer behaviors. For customers whose behavior patterns change rapidly, a fixed update mechanism leads to lagging recommendation data, failing to respond promptly to changes in their needs; while for customers with relatively stable behavioral preferences, overly frequent updates may disrupt recommendation consistency and negatively impact customer experience. Furthermore, existing methods do not consider the impact of the dispersion of behavioral biases on recommendation data updates. When customer behavioral biases exhibit random fluctuations rather than trend changes, blindly updating recommendation data may actually reduce recommendation accuracy. Summary of the Invention
[0005] The purpose of this invention is to provide a personalized customer recommendation system and method based on retail data to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides a customer personalized recommendation method based on retail data, the method comprising: Under the condition of meeting the preset recommendation conditions, the historical behavior profile corresponding to the customer ID is obtained. The historical behavior profile includes the historical active duration and historical behavior parameters. Real-time behavior profiles of the same customer ID are collected. The real-time behavior profile includes the real-time active duration and real-time behavior parameters. A behavior deviation vector is generated using the historical behavior profile and the real-time behavior profile. Search for a set of behavior deviation values that match the behavior deviation vector. When the variance of the set of behavior deviation values reaches or exceeds a preset variance threshold, measure the ratio of the time interval between the data storage start time and the profile acquisition time to the time interval between the data storage start time and the current time, and set this ratio as the data update ratio. Modify the customer recommendation data according to the data update ratio.
[0007] Preferably, the step of obtaining the historical behavior profile corresponding to the customer identifier includes: obtaining multiple preset behavior parameters of the customer identifier; for each preset behavior parameter, calculating the abnormal active duration of customer behavior data based on the historical behavior sample library; selecting the minimum active duration from all abnormal active durations, and multiplying the minimum active duration by a preset coefficient as the preset recommendation condition.
[0008] Preferably, when the behavioral data belongs to the underlying behavioral attributes, the step of calculating the alienated active duration for the preset behavioral parameters includes: searching for multiple sets of underlying behavioral record feature values and corresponding active duration record values that satisfy the preset behavioral parameters from the historical behavioral sample library; performing deviation calculation on the multiple sets of underlying behavioral record feature values and active duration record values to obtain a set of record feature value deviation modulus values and a set of active duration record value deviation modulus values; selecting a subset of active duration record value deviation modulus values from the set of active duration record value deviation modulus values where the record feature value deviation modulus value is not less than a preset feature deviation threshold; and performing mode calculation on the subset of active duration record value deviation modulus values to obtain the alienated active duration.
[0009] Preferably, the underlying behavioral attributes include atomic operation records in customer interaction events, wherein the atomic operation record contains an operation timestamp, an operation type identifier, and an operation object identifier; the set of record feature value deviation modulus values is obtained by calculating the frequency of occurrence of the operation type identifier and the distribution entropy value of the operation object identifier in the atomic operation record; The high-level mapping attribute includes customer intent tags generated by a behavior pattern recognition model, which are associated with a specific product category; the behavior data mapper maps the operation records in the low-level behavior attributes to an intent probability distribution through a neural network model, and the feature value of the high-level mapping attribute is the intent tag corresponding to the highest probability in the intent probability distribution.
[0010] Preferably, when the behavioral data belongs to a high-level mapping attribute, the step of calculating the alienated active duration for the preset behavioral parameters includes: calling a behavioral data mapper associated with the high-level mapping attribute and the preset behavioral parameters, wherein the behavioral data mapper receives the low-level behavioral attributes; retrieving multiple sets of low-level behavioral attribute record feature values and corresponding active duration record values that satisfy the preset behavioral parameters from the historical behavioral sample library; processing each set of low-level behavioral attribute record feature values through the behavioral data mapper to obtain multiple sets of high-level mapping attribute feature values; and deriving the alienated active duration by combining the multiple sets of high-level mapping attribute feature values and the multiple sets of active duration record values.
[0011] Preferably, the method further includes the following steps: configuring preset underlying behavioral attributes of preset high-level mapping attributes through a user interface; collecting multiple sets of training samples with preset customer classification and preset behavioral parameters as constraints to train the behavioral data mapper; binding the trained behavioral data mapper with the preset high-level mapping attributes and the preset behavioral parameters; wherein, each set of training samples includes preset underlying behavioral attribute record data and preset high-level mapping attribute record data.
[0012] Preferably, the behavior deviation vector consists of an activity duration deviation vector and a behavior parameter deviation vector; the step of searching for the behavior deviation value set includes: sending a search request to the cloud, obtaining an encryption key, encrypting the behavior dimension, customer classification, the activity duration deviation vector, and the behavior parameter deviation vector, generating ciphertext and uploading it to the cloud, receiving search feedback information returned by the cloud, the search feedback information containing an initial behavior deviation value set; and performing outlier removal processing on the initial behavior deviation value set to obtain the behavior deviation value set.
[0013] Preferably, the step of sending a retrieval request to the cloud includes: decrypting the ciphertext by the cloud to obtain the behavior dimension, the customer category, the active duration deviation vector, and the behavior parameter deviation vector, forming a retrieval task constraint; sending a retrieval request to a distributed node through the cloud to obtain a distributed node key; encrypting the retrieval task constraint and the retrieval key using the distributed node key to generate ciphertext for the retrieval task, sending it to the distributed node, and receiving retrieval feedback information returned by the distributed node.
[0014] Preferably, the step of modifying customer recommendation data according to the data update ratio includes: determining the amount of data to be updated according to the data update ratio, and selecting a corresponding number of data records to perform the update operation from the data storage start time.
[0015] Preferably, the present invention also includes a customer personalized recommendation system based on retail data, the system including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the computer program, implements the steps of the customer personalized recommendation method based on retail data as described above.
[0016] Compared with the prior art, the beneficial effects of the present invention are: By integrating historical and real-time customer behavior profiles, a comprehensive characterization of customer behavior is achieved. Historical behavior profiles reflect long-term consumption preferences and habits, while real-time behavior profiles capture current needs and behavioral tendencies. The combination of these two types of profiles overcomes the limitations of single data dimensions, making the understanding of customer behavior more complete and multi-dimensional. By generating behavioral deviation vectors, the differences between real-time and historical customer behavior can be intuitively presented, clearly outlining the trajectory of customer behavior changes and avoiding vague judgments about behavioral shifts.
[0017] The system searches for a set of behavioral deviation values that match the behavioral deviation vector and uses variance as a criterion to effectively distinguish the nature of customer behavioral deviations. When the variance reaches or exceeds a preset threshold, it indicates that the customer's behavioral deviation is not a random fluctuation but rather exhibits a trend. At this point, a data update process is triggered to ensure that adjustments to the recommendation strategy are targeted. This mechanism, based on the degree of deviation dispersion, avoids blindly updating recommendation data, reduces the interference of ineffective adjustments on recommendation performance, and makes the optimization of the recommendation strategy more rational.
[0018] The calculation method for the data update ratio fully considers the behavioral correlation over time. By calculating the ratio of the data storage start time, the profile collection time, and the current time, it dynamically reflects the temporal rhythm of changes in customer behavior. Different customers exhibit varying rates and cycles of behavioral change, and this ratio adapts to these individual differences. This ensures that the data update ratio is neither too conservative, lagging behind changes in customer needs, nor too aggressive, deviating from the long-term behavioral foundation of customers. This dynamically adjusted update ratio allows modifications to recommendation data to better align with actual changes in customer behavior, enhancing the flexibility and adaptability of the recommendation strategy.
[0019] In retail scenarios, customer purchasing decisions are often closely related to immediate needs and behavioral states. This method, by dynamically modifying recommendation data, enables recommendations to respond promptly to changes in customer behavior, making the recommendations more aligned with customers' current consumption intentions. This not only reduces the time cost for customers in the product selection process, improving the consumer experience, but also allows retailers to more accurately reach potential needs, enhancing the effectiveness of recommendations. Furthermore, the recommendation logic based on behavioral bias vectors and dynamically updated ratios can continuously adapt to the dynamic changes in the retail market and the evolution of customer preferences, making the personalized recommendation mechanism applicable in the long term. This helps retailers better maintain customer relationships and optimize business performance in a highly competitive market environment.
[0020] This method does not rely on complex hardware or massive data collection; it can be implemented simply by analyzing historical and real-time behavioral parameters and activity durations from existing retail data. It is highly practical and operable. Its core logic is clear, and its execution process is concise, allowing for rapid integration into existing retail digital systems. This lowers the barrier to technology implementation and facilitates its application in various retail scenarios, providing a new path for optimizing personalized recommendation technology in the retail industry. Attached Figure Description
[0021] Figure 1 This is a schematic diagram illustrating the working principle of the customer personalized recommendation method based on retail data described in this invention. Figure 2 A flowchart for obtaining historical behavioral profiles and setting preset recommendation criteria; Figure 3 A distribution analysis diagram of customer intent tags based on behavioral pattern recognition; Figure 4 A flowchart for calculating the alienation activity duration under high-level mapping attributes; Figure 5 A chart comparing the mean deviations in customer behavior. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Please see Figure 1This invention provides a personalized customer recommendation method based on retail data. The method includes: acquiring a historical behavior profile corresponding to a customer identifier, the historical behavior profile including historical active duration and historical behavior parameters. The historical active duration reflects the continuity of the customer's activity over a past time period, and the historical behavior parameters cover quantitative indicators of customer interaction behavior. Simultaneously, the system collects a real-time behavior profile for the same customer identifier, the real-time behavior profile including real-time active duration and real-time behavior parameters, generated by monitoring the customer's current session data. The system calculates the difference between the historical and real-time behavior profiles to generate a behavior deviation vector, which represents the trend of customer behavior over time. The system searches for a set of behavior deviation values matching the behavior deviation vector, the set of behavior deviation values originating from a historical data warehouse. When the variance of the behavior deviation value set reaches or exceeds a preset variance threshold, the system measures the ratio of the time interval between the data storage start time and the profile acquisition time to the time interval between the data storage start time and the current time, setting this ratio as the data update ratio. The data update ratio is used to quantify data freshness requirements, and customer recommendation data is modified according to the data update ratio. Modification operations involve adjusting the priority of the recommendation list or updating product association rules.
[0024] Example 1: See Figure 2 The starting point for obtaining historical behavioral profiles corresponding to customer identifiers is acquiring multiple preset behavioral parameters for those identifiers. These preset behavioral parameters are a predefined set of quantitative indicators used to characterize customer behavior. These parameters may include the number of visits per unit of time, the average number of pages viewed per session, and the time interval between two consecutive purchases. The system extracts historical raw data records associated with customer identifiers from the customer master data warehouse. These records include timestamps, operation types, and operation object codes. The extraction process uses index queries based on customer identifiers to ensure data integrity and accuracy. The values of the preset behavioral parameters are obtained through aggregation calculations and historical data smoothing. For example, the visit frequency is calculated by counting the number of sessions within a specific time window and dividing by the window length.
[0025] For each preset behavioral parameter, the differentiated active duration of customer behavior data is calculated based on a historical behavior sample database. This database is a specially constructed collection storing massive amounts of historical customer behavior records and their derived active duration data. The database's data structure is optimized to support efficient querying and analysis. Each sample record includes a customer identifier, a set of behavioral parameters, the calculated active duration value, and the data collection time. Calculating the differentiated active duration requires searching the historical behavior sample database for multiple sets of underlying behavioral record feature values and their corresponding active duration record values that satisfy the current preset behavioral parameters. The search operation uses parameterized query statements, using the preset behavioral parameter values as filtering conditions, and returns a set of all matching records. The underlying behavioral record feature values are numerical features extracted from atomic operation records, and the active duration record value is a pre-calculated customer active duration. Deviation calculations are performed on the multiple sets of underlying behavioral record feature values and active duration record values obtained from the search. This deviation calculation aims to quantify the degree of difference between different records. The set of record feature value deviation moduli is obtained by calculating the Euclidean distance between every two sets of record feature values. The Euclidean distance comprehensively considers the differences across various feature dimensions. The set of deviation modulus values for active duration records is obtained by calculating the absolute difference between any two sets of active duration records. The absolute difference reflects the direct difference in active duration. The calculation process iterates through all record pairs, generating a symmetric deviation matrix. A diagonal element of zero in the matrix indicates that the deviation of a record from itself is zero.
[0026] The process involves selecting a subset of active duration record value deviation moduli from the set of active duration record value deviation moduli where the moduli of the record feature value deviation is not less than a preset feature deviation threshold. This preset feature deviation threshold is a configurable parameter used to control the strictness of the selection. The selection operation iterates through each element of the deviation matrix, checking whether the corresponding record feature value deviation moduli is greater than or equal to the preset feature deviation threshold. Active duration record value deviation moduli that meet the condition are added to a temporary set, which forms the basis for subsequent analysis. The preset feature deviation threshold is set based on the business scenario and the requirements for sensitivity to behavioral changes; a higher threshold means focusing only on record pairs with significant feature differences. The mode is calculated on the subset of active duration record value deviation moduli to obtain the differentiated active duration. The mode calculation identifies the deviation value with the highest frequency in the subset. The calculation process first discretizes all deviation values in the subset into buckets, with each bucket representing a range of deviation values. The frequency of occurrence of deviation values in each bucket is counted, and the range of deviation values corresponding to the bucket with the highest frequency is selected as the mode interval. Taking the median of the mode interval as the final value of the alienated activity duration helps resist the interference of extreme outliers, making the results more robust. Alienated activity duration characterizes the activity duration change pattern that typically accompanies significant changes in customer behavior under given preset behavioral parameters. The minimum activity duration is selected from all alienated activity durations, and multiplied by a preset coefficient to become the preset recommendation condition. All alienated activity durations refer to the set of alienated activity durations calculated for each preset behavioral parameter. Selecting the minimum activity duration is a simple comparison process, iterating through the entire set and recording the minimum value. The minimum activity duration represents the most sensitive behavioral change signal, i.e., the shortest activity duration change that might correspond to a small but significant change in customer behavior. The preset coefficient is a real number greater than zero, used to adjust the leniency of the recommendation condition. The product of the preset coefficient and the minimum activity duration is the threshold of the preset recommendation condition. When the real-time behavioral feature changes detected by the system reach or exceed this threshold, the personalized recommendation process is triggered.
[0027] When behavioral data belongs to the underlying behavioral attributes, the specific steps for calculating the differentiated active duration based on preset behavioral parameters involve in-depth analysis of atomic operation records. Underlying behavioral attributes include atomic operation records in customer interaction events; these records are the smallest granular unit of customer behavior recorded by the system. Each atomic operation record contains an operation timestamp, an operation type identifier, and an operation object identifier. The timestamp is accurate to the millisecond level, the operation type identifier encodes the nature of the operation (e.g., click, swipe, input), and the operation object identifier uniquely identifies the operated object (e.g., product number, page address). The set of deviation moduli for record feature values is obtained by calculating the frequency of occurrence of the operation type identifier and the distribution entropy of the operation object identifier in the atomic operation records. The frequency of occurrence counts the number of times a specific operation type appears in the record set, and the distribution entropy uses the information entropy formula to calculate the dispersion of the operation object identifier. The frequency of occurrence and the distribution entropy together constitute a two-dimensional vector representation of the record feature values. Searching for multiple sets of underlying behavioral record feature values and corresponding active duration record values that satisfy the preset behavioral parameters from the historical behavioral sample database is a data retrieval and matching process. The system parses the preset behavioral parameters and converts them into database query conditions. Query conditions may involve filtering by time range, operation type, and customer segmentation. After executing the query, an initial record set is obtained. Feature extraction is performed on this initial record set to obtain the underlying behavioral record feature values. Active duration record values are directly read from the active duration field of the record set, which has already been preprocessed and calculated during data entry. Calculating the deviation between multiple sets of underlying behavioral record feature values and active duration record values requires constructing record pairs and calculating the difference between each pair. The deviation modulus calculation for record feature values uses a vector distance formula, taking the feature value vectors of two records as input and outputting a scalar distance value. The deviation modulus calculation for active duration record values uses an absolute difference formula, directly calculating the absolute value of the difference between two active duration values. The calculation process covers all possible record pair combinations, generating a complete set of deviation modulus values. This set of deviation modulus values reflects the distribution of differences among records within the historical behavioral sample database.
[0028] Selecting a subset of active duration record value deviation modulus values from the set of active duration record value deviation modulus values, where the modulus of the record feature value deviation is not less than a preset feature deviation threshold, is a conditional filtering operation. The system iterates through each record pair, checking whether the modulus of the feature value deviation of that pair of records meets the condition of being greater than or equal to the preset feature deviation threshold. The active duration deviation modulus values corresponding to the record pairs that meet the condition are retained and aggregated into a new set. This filtering process ensures that subsequent analysis only focuses on record pairs that have significant differences in behavioral characteristics, ignoring record pairs with minor feature differences. The mode calculation of the subset of active duration record value deviation modulus values uses statistical methods to identify the most frequently occurring deviation values. The mode calculation first discretizes all continuous deviation values in the subset into several equally wide intervals, and counts the frequency of values within each interval. The interval with the highest frequency is determined as the mode interval, and the midpoint of the mode interval is selected as the representative value of the alienated active duration. The mode calculation is not sensitive to outliers and can capture the most common patterns in the data. Alienated active duration serves as the basic input for calculating the preset recommendation conditions.
[0029] Example 2: High-level mapping attributes include customer intent tags generated by a behavioral pattern recognition model, which are associated with specific product categories. The behavioral pattern recognition model is a machine learning-based classifier whose input is a sequence of customer behaviors and whose output is a predefined intent category. Customer intent tags include, for example, "short-term promotion response," "long-term brand interest," and "new product exploration." Each tag corresponds to one or more product categories, and the associations are stored in a category-intent mapping table. The behavioral data mapper uses a neural network model to map the operation records in the underlying behavioral attributes to an intent probability distribution. The neural network model adopts a multilayer perceptron architecture, with the number of input layer nodes matching the dimension of the underlying behavioral features, and the number of output layer nodes corresponding to the number of intent tags. The feature value of the high-level mapping attribute is the intent tag corresponding to the highest probability in the intent probability distribution. The intent probability distribution is a probability vector, where each element represents the probability that a customer belongs to a certain intent category. The system selects the tag corresponding to the index with the highest probability value as the final output.
[0030] The neural network model structure of the behavior data mapper consists of an input layer, two hidden layers, and an output layer. The input layer receives low-level behavioral attributes that undergo feature engineering, transforming the original atomic operation records into fixed-dimensional numerical vectors. Transformation methods include one-time encoding of the operation type, embedding representation of the operation object, and extraction of statistical features from the time series. The first hidden layer contains 512 neurons using the ReLU activation function, which introduces a non-linear transformation. The second hidden layer contains 256 neurons, also using the ReLU activation function; the weight matrix of the hidden layers is optimized during training. The output layer contains the same number of neurons as the intent labels, and the Softmax activation function transforms the network output into a probability distribution, ensuring that the sum of all output values is 1. The neural network model is trained based on historical behavioral data; the training set contains a large number of labeled customer behavior sequences and corresponding real intent labels. The association between customer intent labels and specific product categories is established through manually defined rules; business experts formulate mapping relationships based on market strategies and product characteristics. The mapping relationships are stored in an intent-category association table in a relational database, with the table structure including intent label encoding, category encoding, and association weight fields. Association weights represent the degree of relevance between intent tags and product categories. Weight values are derived by analyzing historical purchase data and behavioral patterns. After generating customer intent tags, the system queries the intent-category association table in real time to obtain a list of associated product categories. This category list is used for subsequent personalized recommendation ranking.
[0031] The training process of the behavior pattern recognition model adopts a supervised learning paradigm, with training data derived from historical customer interaction logs. The data preprocessing stage cleans the raw logs, removing invalid records and outliers while retaining complete conversation sequences. Each training sample consists of a customer behavior sequence and manually labeled intent tags. The behavior sequence includes all atomic operation records within a fixed time window. Labeling is performed by the business team, based on factors including the customer's final purchase behavior, search keyword analysis, and page dwell patterns. Training uses the stochastic gradient descent algorithm, optimizing the cross-entropy loss function, which measures the difference between the model's output probability distribution and the true label distribution. An early stopping strategy is employed during training to prevent overfitting, monitoring accuracy changes on the validation set. The calculation of the intent probability distribution involves the forward propagation process of the neural network. Forward propagation transforms the input feature vector through weight matrices and activation functions at each layer. After linear transformation in the input layer, the input feature vector is passed to the first hidden layer, and the output of the first hidden layer serves as the input to the second hidden layer. The output of the second hidden layer undergoes linear transformation in the output layer and is processed by the Softmax function to generate the intent probability distribution. Each element in the probability distribution represents the model's confidence that the customer belongs to the corresponding intent category. The intent label corresponding to the highest probability value is selected as the high-level mapping attribute feature value. A probability threshold setting filters out low-confidence predictions; when the highest probability value is below the threshold, the system marks it as an unknown intent.
[0032] The transformation from low-level behavioral attributes to high-level mapping attributes achieves semantic enhancement of behavioral data, aggregating low-level interactive actions into high-level commercial intents. Clickstream data in atomic operation records is mapped to tags with clear commercial meaning, and the tagging results facilitate the recommendation algorithm's understanding of user needs. Updating and maintaining the behavioral data mapper requires periodic model retraining, incorporating new behavioral patterns and emerging intent categories. Model version management ensures the stability of online services; new models are deployed only after thorough testing. An A / B testing framework compares the recommendation performance of different model versions, with A / B testing metrics including click-through rate, conversion rate, and average order value. The application of high-level mapping attribute features in the recommendation system is manifested in intent-driven product filtering, which adjusts the recommendation strategy based on the customer's current intent preferences. The system maintains a product knowledge graph containing semantic relationships and attribute features between products. Intent tags are connected to category nodes in the knowledge graph, with connection strength determined by association weights. The recommendation engine uses intent tags as query conditions to retrieve related product sets from the knowledge graph; these product sets are then processed by a ranking algorithm and presented to the customer. The sorting algorithm takes into account factors such as product popularity, inventory status, and profit margin.
[0033] The optimization of the neural network model for the behavior data mapper involves hyperparameter tuning, including learning rate, batch size, and the number of hidden layer neurons. Hyperparameter tuning uses a grid search method, which finds the optimal combination within a predefined hyperparameter space. Model performance evaluation employs k-fold cross-validation, which reduces the sensitivity of evaluation results to data partitioning. Training data augmentation techniques generate synthetic samples, improving the model's generalization ability. Synthetic samples are generated by adding noise or time shifts to the original behavior sequences. The hierarchical structure of customer intent labels supports multi-granularity intent representation, allowing the system to simultaneously identify macro and micro user needs. The intent label system is designed as a tree structure, with the root node representing general intents and leaf nodes representing specific intents. The multi-label classification function of the behavior pattern recognition model can predict multiple intent labels, using the sigmoid activation function instead of the softmax function. Prediction of each intent label is performed independently, as customers may have multiple shopping intents simultaneously. In multi-intent scenarios, the recommendation list integrates products corresponding to different intents, with the fusion strategy based on a weighted average of intent probabilities. A real-time synchronization mechanism between high-level mapping attribute feature values and low-level behavioral attributes ensures the timeliness of intent tags. This mechanism monitors the customer behavior event stream. The event stream processing platform receives interaction events from the front-end application, with the event format following a predefined pattern. The behavior data mapper is deployed as a microservice, and the microservice interface receives behavior events and returns intent tags. Tag caching reduces computational latency, storing the results of recently used intent tags. The cache invalidation strategy is based on time thresholds or event quantity thresholds to ensure timely tag updates.
[0034] Interpretive analysis of behavioral pattern recognition models reveals feature importance, helping business personnel understand the basis of model decisions. Feature importance ranking uses a permutation importance method, which measures the impact of shuffling feature values on model performance. Local interpretability techniques, such as SHAP value calculation, display the feature contribution of individual predictions. Model interpretation results are used for business rule validation, ensuring that model behavior conforms to business logic. An anomaly detection module monitors model prediction deviations, triggering a model retraining process. Persistent storage of high-level mapping attribute feature values supports historical intent analysis, recording intent tags along with timestamps. Intent time-series data reveals patterns in customer interest evolution, used for long-term preference modeling. The intent transition probability matrix quantifies the likelihood of conversion between different intents, calculated using a Markov chain. Long-term preference modeling enhances the forward-looking nature of recommendation systems, allowing for proactive recommendations to prepare potentially interesting products in advance. Customer lifetime value prediction, combined with intent evolution patterns, optimizes marketing resource allocation.
[0035] See Figure 3The figure presents the distribution of customer numbers corresponding to five types of customer intent tags generated by the behavioral pattern recognition model, intuitively reflecting the clustering characteristics of customer groups in high-level semantic attributes. The data in the figure comes from the intent probability distribution generated by mapping the underlying behavioral attributes through a neural network. The peak number of "daily needs" customers (420) indicates that basic consumption needs dominate customer behavior. This statistical result will serve as a key input for calculating the behavioral deviation vector and will be used to calibrate the data update ratio. The significant difference between "short-term promotion response" (350) and "long-term brand attention" (280) reflects the heterogeneity of customer decision-making mechanisms. Such quantitative indicators will be used to optimize the weight allocation in the intent-category association table. The relatively low proportion of "new product exploration" (220) and "social recommendation" (180) suggests that the intensity of behavioral data collection for the corresponding scenarios needs to be strengthened to ensure the prediction accuracy of the behavioral data mapper under long-tail distribution. The overall distribution structure provides a topological constraint basis for the adaptive classifier of the recommendation system. The height difference of the pillars corresponding to different intent labels constitutes the topological features of the customer behavior field, which will participate in the calculation of the topological consistency loss function when adjusting the recommendation strategy, ensuring that the customer behavior evolution path and the recommendation results maintain semantic continuity and interpretability.
[0036] Example 3: See Figure 4When behavioral data belongs to a high-level mapping attribute, the starting point for calculating the alienated activity duration for preset behavioral parameters is to call the behavioral data mapper associated with the high-level mapping attribute and the preset behavioral parameters. The behavioral data mapper is a pre-trained machine learning model whose function is to convert low-level raw behavioral data into high-level semantic labels. The calling process is carried out through a mapper registry, which stores the behavioral data mapper model path and version information corresponding to different combinations of high-level mapping attributes and preset behavioral parameters. The system searches the registry according to the preset behavioral parameters to be calculated and the specified high-level mapping attribute type, locates the specific behavioral data mapper instance, and loads the model parameters and computation graph structure of the behavioral data mapper into memory to prepare for the inference task. The behavioral data mapper receives low-level behavioral attributes as input data. The low-level behavioral attributes are a set of atomic operation records generated by the customer during interaction with the retail system. These records contain raw fields such as operation timestamps, operation type identifiers, and operation object identifiers. The input interface of the behavioral data mapper receives pre-processed feature vectors. The preprocessing steps include data cleaning, feature encoding, and sequence alignment. The system retrieves multiple sets of underlying behavioral attribute record feature values and corresponding active duration record values that meet preset behavioral parameters from a historical behavior sample database. This database is a large data warehouse storing massive amounts of customer historical behavior data. The retrieval operation uses parameterized queries, with preset behavioral parameters serving as filtering conditions. The query returns a set of all matching historical records. Each record contains a vector of underlying behavioral attribute record feature values and a numerical value of active duration. The vector of underlying behavioral attribute record feature values is a numerical representation obtained after feature engineering of the original atomic operation record.
[0037] The behavior data mapper processes the feature values of each set of low-level behavioral attribute records to obtain multiple sets of high-level mapped attribute feature values. The process involves the behavior data mapper performing forward propagation calculations on the input feature vector. The behavior data mapper is typically a deep neural network model, with its input layer dimension matching the length of the low-level behavioral attribute record feature value vector, and its output layer dimension corresponding to the dimension of the high-level mapped attribute feature values. For each set of input low-level behavioral attribute record feature values, the behavior data mapper outputs a high-level mapped attribute feature value vector, which represents the mapping result of the input behavioral data in a high-level semantic space. The high-level mapped attribute feature values can be continuous values, discrete labels, or probability distributions, depending on the design and training objectives of the behavior data mapper.
[0038] By combining multiple sets of high-level mapping attribute feature values and multiple sets of active duration records, the alienated active duration is derived. This process requires establishing a mapping relationship from the high-level mapping attribute feature values to the active duration records. One implementation method is to construct a regression model, using the high-level mapping attribute feature values as independent variables and the active duration records as dependent variables, fitting a functional relationship between the two. The regression model can employ algorithms such as linear regression, support vector regression, or neural network regression. Model training uses the retrieved high-level mapping attribute feature values and active duration records as training samples. After training, the regression model can predict the corresponding active duration based on new high-level mapping attribute feature values. The alienated active duration is determined by analyzing the deviation pattern between the predicted and actual values. Specifically, the alienated active duration can be calculated using the following formula: in: This represents the calculated duration of alienation activity. Indicates the number of samples. Indicates the first The actual active duration recorded for each sample. This indicates the first prediction made by the regression model. The formula calculates the average absolute deviation between the predicted and actual active durations for all samples. This average reflects the average error level of the mapping relationship between the high-level mapping attribute feature values and the active duration, and is defined as the alienated active duration.
[0039] The system configures preset underlying behavioral attributes for preset high-level mapping attributes through a user interface. The user interface provides a graphical configuration environment, allowing administrators to establish associations between high-level mapping attributes and underlying behavioral attributes through drag-and-drop. The left side of the configuration interface displays a list of available underlying behavioral attributes, including various atomic operation types and object identifiers; the right side displays a list of defined high-level mapping attributes, each of which can be bound to multiple underlying behavioral attributes. Administrators use drag-and-drop operations to associate underlying behavioral attributes on the left with high-level mapping attributes on the right. Weight values can be set for these associations, representing the contribution of the underlying behavioral attribute to the high-level mapping attribute. Configuration information is saved to the system database for use during the training and inference of the behavioral data mapper. Multiple training samples are collected, constrained by preset customer categories and preset behavioral parameters, to train the behavioral data mapper. The preset customer categories are the result of dividing customer groups according to certain criteria. These criteria can be based on demographic information, spending power, historical behavioral patterns, etc., while the preset behavioral parameters limit the data range and behavioral characteristics of the training samples. When collecting training samples, the system first filters the set of customer identifiers that meet the preset customer classification criteria. Then, for each customer identifier, it extracts records from their historical behavior data that meet the preset behavioral parameter requirements. Each training sample includes two parts of data: preset low-level behavioral attribute record data and preset high-level mapping attribute record data. The preset low-level behavioral attribute record data is the feature representation of the customer's original behavioral data, and the preset high-level mapping attribute record data is the corresponding label or target value. The quantity and quality of training samples directly affect the performance of the behavioral data mapper, and it is necessary to ensure that the samples are representative and diverse.
[0040] Each training sample set includes preset low-level behavioral attribute record data and preset high-level mapping attribute record data. The preset low-level behavioral attribute record data is typically a feature vector. The feature vector consists of multiple feature fields, each representing a quantified value of a low-level behavioral attribute. The feature vector has a fixed dimension to facilitate processing by the neural network model. The preset high-level mapping attribute record data is the sample annotation information, which can come from manual annotation, business rule derivation, or historical behavior analysis. Training samples undergo quality checks to remove samples containing erroneous or inconsistent data, ensuring the accuracy and reliability of the training data. Data augmentation techniques can be applied to training sample generation, expanding the training dataset by adding noise, time shifts, or sampling transformations to improve the generalization ability of the behavioral data mapper. The trained behavioral data mapper is then bound to the preset high-level mapping attributes and preset behavioral parameters. This binding operation is completed on the behavioral data mapper management platform. The management platform records the metadata information of the behavioral data mapper, including model version, training data range, performance metrics, and creation time. The binding process establishes a ternary relationship: the association between the preset high-level mapping attributes, preset behavioral parameters, and the behavioral data mapper model instance. This relationship is stored in the model registry. When the system needs to process high-level mapping attribute calculation tasks that conform to specific preset behavioral parameters, it queries the model registry to obtain the corresponding behavioral data mapper model instance, loads the model, and performs inference calculations. The binding relationship supports version management, allowing the system to maintain multiple versions of the behavioral data mapper model simultaneously, facilitating A / B testing and canary releases.
[0041] Example 4: The behavior deviation vector consists of an active duration deviation vector and a behavior parameter deviation vector. The active duration deviation vector quantifies the difference between historical and real-time active durations. The active duration deviation vector is calculated using the difference method, subtracting the corresponding components of the historical active duration from each dimension of the real-time active duration to obtain a difference vector. The behavior parameter deviation vector measures the degree of deviation between historical and real-time behavior parameters. The calculation of the behavior parameter deviation vector is based on the Euclidean distance formula, treating the two sets of behavior parameters as points in a multi-dimensional space and calculating the straight-line distance between the points. The two deviation vectors together constitute a comprehensive numerical description of customer behavior changes. The system concatenates the active duration deviation vector and the behavior parameter deviation vector into a comprehensive behavior deviation vector, which serves as the input for subsequent search operations. The steps for searching the set of behavior deviation values include sending a search request to the cloud. The search request is a structured data packet containing a query intent identifier and timestamp information. The system obtains a temporary encryption key from the local key management service. This encryption key is specifically used for this search session and has a short validity period. The encryption process covers behavioral dimensions, customer classification, active duration deviation vectors, and behavioral parameter deviation vectors. Behavioral dimension parameters define the perspective levels for data analysis, such as individual customer dimensions, customer group dimensions, and time series dimensions. Customer classification parameters divide customers into different groups, with classification criteria based on consumption behavior, demographic attributes, and value ratings. The system uses an encryption key to perform symmetric encryption on the above data, generating unreadable ciphertext data. The encryption algorithm adopts the AES-256 standard, providing high-strength confidentiality protection.
[0042] The encrypted data is uploaded to the cloud server via a secure transmission channel encrypted with the TLS protocol to prevent data theft or tampering during transmission. Upon receiving the encrypted data, the cloud server uses a pre-shared key to decrypt and reconstruct the original data, obtaining plaintext information including behavioral dimensions, customer classifications, active duration deviation vectors, and behavioral parameter deviation vectors. This plaintext information is assembled into a retrieval task constraint structure, which details the scope, conditions, and objectives of the search. The cloud server parses the retrieval task constraint structure, identifies the data distribution characteristics and matching conditions to be queried, and generates a distributed query execution plan. A retrieval request is sent to distributed nodes via the cloud. These distributed nodes are geographically dispersed data storage and computing units, each responsible for maintaining a portion of the behavioral deviation value data. The cloud server obtains the distributed node key for the corresponding target distributed node from the authentication center. This distributed node key is used to verify the cloud server's identity and establish secure communication. The cloud server uses the distributed node key to encrypt the retrieval task constraints and a newly generated retrieval key. The retrieval key is a randomly generated session key used to protect subsequent data transmission. The encrypted retrieval task ciphertext is sent to the relevant distributed nodes, and the sending process may involve message queues or remote procedure calls.
[0043] Referring to Table 1, after receiving the encrypted retrieval task, the distributed nodes use their private keys to decrypt it and obtain the retrieval task constraints and retrieval key. The distributed nodes execute queries in their local behavior deviation value database, with query conditions based on the behavior dimension, customer category, active duration deviation vector, and behavior parameter deviation vector in the retrieval task constraints. Local database queries utilize index optimization to quickly locate data records similar to the input behavior deviation vector; similarity calculation employs a cosine similarity algorithm. The query returns an initial set of behavior deviation values, containing a set of values representing the deviation metrics corresponding to historical records with similar behavior deviation patterns to the current customer. The distributed nodes encrypt the initial set of behavior deviation values using the retrieval key and return the encrypted retrieval feedback information to the cloud server. The cloud server aggregates the retrieval feedback information from multiple distributed nodes, merging it into a complete initial set of behavior deviation values. The cloud server returns the aggregated initial set of behavior deviation values to the requesting client system, with encryption protection also applied during transmission. The client system receives the retrieval feedback information returned from the cloud and decrypts it using the corresponding key to obtain the initial set of behavior deviation values. Outlier removal is performed on the initial set of behavioral deviation values to obtain the final set of behavioral deviation values. Outlier removal uses statistical methods to identify and filter out outliers that significantly deviate from the main data distribution. A common method is to calculate the quartile range of the data. First, the initial set of behavioral deviation values is sorted in ascending order, and the first quartile Q1 and the third quartile Q3 are identified. The interquartile range (IQR) is calculated as IQR = Q3 - Q1. Any value below Q1 - 1.5 × IQR or above Q3 + 1.5 × IQR is considered an outlier and removed from the set.
[0044] Table 1: Encryption Parameter Table for Behavioral Deviation Vector The variance of the behavioral deviation value set is calculated using the standard variance formula, and the variance value is used to assess the dispersion of behavioral deviations. When the variance reaches or exceeds a preset variance threshold, it indicates that the customer behavior change pattern has significant instability, requiring the initiation of a data update process. The preset variance threshold is a configurable parameter; the threshold setting affects the system's responsiveness to behavioral changes. A higher threshold means that it only reacts to very drastic behavioral fluctuations. The variance calculation result is an important basis for subsequent judgments on whether to modify the recommended data. The system compares the variance value with the preset variance threshold and determines the workflow branch based on the comparison result. The entire search process involves multi-layered encryption and secure communication mechanisms to ensure the privacy protection and secure transmission of customer behavior data. The lifecycle management of encryption keys follows the principle of least privilege; keys are generated and used only when needed and destroyed immediately after use. Authentication of distributed nodes is based on a digital certificate system, with each distributed node holding a unique identity certificate issued by a root certificate authority. Secure communication protocols prevent man-in-the-middle attacks and data leaks, and audit logs record all critical operations for traceability and security analysis.
[0045] See Figure 5 This study presents a comparative analysis of the average deviations of total active duration and total behavioral parameters across five customer categories using bar charts. Specifically, the horizontal axis represents the customer category, including low-frequency buyers, seasonal buyers, newly registered customers, churn-risk customers, and high-value loyal customers; the vertical axis represents the average deviation value, ranging from 0 to 5. This dual-axis design ensures data readability. Visually, dark gray bars represent the total deviation of active duration, while light gray bars represent the total deviation of behavioral parameters, clearly distinguishing the two categories through color contrast. Data analysis shows that churn-risk customers exhibit a peak in total active duration deviation (5.8) and a total behavioral parameter deviation of 3.4, indicating significant fluctuations in their behavioral patterns. In contrast, newly registered customers have relatively lower deviation values (4.5 for active duration and 3.8 for behavioral parameters), reflecting higher consistency in their behavior.
[0046] Example 5: Modifying customer recommendation data based on the data update ratio starts by determining the amount of data to be updated based on the data update ratio. The data update ratio is a value between 0 and 1, representing the proportion of the total data that needs to be updated. The system queries the customer recommendation data storage start time from the metadata management service. The data storage start time is the initial moment when customer recommendation data begins to accumulate. The system obtains the current system time and calculates the time interval T1 between the data storage start time and the profile collection time. The profile collection time is the time point when the real-time behavioral profile is generated. The system calculates the time interval T2 between the data storage start time and the current time. The formula for calculating the data update ratio R is R = T1 / T2. The closer the data update ratio R is to 1, the larger the amount of data that needs to be updated, because the profile collection time is far from the data storage start time, and the current time is even further from the data storage start time, making the historical data relatively outdated. The amount of data to be updated is calculated by multiplying the data update ratio R by the total number of customer recommendation data records. The total number of customer recommendation data records is the number of all recommendation-related data entries accumulated since the data storage start time. The calculation process requires accessing the statistics table of the customer recommendation database, which records the data volume distribution across different time ranges. The system executes a counting query to count the total number of customer recommendation data records from the data storage start time to the current time, obtaining the total number N. The formula for calculating the amount of data U to be updated is U = R × N. The result may be a decimal, so the system rounds it up to ensure that at least one record is updated. The value of U determines the range of customer recommendation data that needs to be modified; the larger the value of U, the wider the impact of the data update operation.
[0047] The system selects a corresponding number of data records for update operations starting from the data storage start time, based on chronological order. The system queries the customer recommendation database, sorting the records in ascending order by their timestamp field, which records the generation time of each customer recommendation data entry. Starting with the earliest record corresponding to the data storage start time, the system sequentially selects the first U records as the dataset to be updated. This selection method ensures that the oldest data is updated first. The data record selection operation uses pagination query technology, reading a batch of records from the database for processing at a time to avoid memory overflow caused by loading a large amount of data at once. Query conditions are limited to ensuring that the customer identifier matches the currently processed customer, guaranteeing that the update operation targets the recommendation data of a specific customer. The update operation involves modifying multiple fields in the customer recommendation data, which affect the generation logic of the recommendation results. Fields that need to be updated include product weight score, association rule confidence, and collaborative filtering similarity value. The product weight score reflects the priority of the product in the recommendation list, and the weight score is dynamically adjusted based on the customer's historical behavior. The association rule confidence represents the strength of the association between products, and the confidence is recalculated based on the customer's purchasing patterns. The collaborative filtering similarity value quantifies the behavioral similarity between the customer and other customers, and the similarity value is updated as customer behavior changes. The update operation is not simply replacing old values with new ones; rather, it recalculates the values of these fields based on the latest behavioral profile. The update process employs a database transaction mechanism, which guarantees the atomicity and consistency of data updates. The system initiates a database transaction, executing a series of update statements within the transaction. The update operation for each customer recommendation data entry includes reading the current value, applying the new algorithm to calculate the updated value, and writing the new value back to the database. Before committing the transaction, the system checks data consistency constraints to ensure that the updated data meets business rules. If an error occurs during the update process, the transaction rolls back to its initial state, maintaining data integrity. The transaction log records all change operations, supporting fault recovery and audit trails.
[0048] The update strategy for customer recommendation data considers the data update ratio, triggering different granularity update operations for different data update ratio ranges. When the data update ratio R is less than 0.3, the system only updates the basic weight fields, which include product click weight and purchase weight. When the data update ratio R is between 0.3 and 0.7, the system updates the basic weight fields and association rule fields, which include product co-occurrence frequency and sequence pattern confidence. When the data update ratio R is greater than 0.7, the system performs a full update, covering all weight fields, association rule fields, and collaborative filtering fields. This tiered update strategy balances computational overhead and recommendation accuracy, using lightweight updates for smaller data update ratios and full updates for larger data update ratios.
[0049] A specific example illustrates the entire update process. Assume a customer with the identifier C1001. Data storage started on January 1, 2024, at 00:00:00. The profile was collected on June 1, 2024, at 10:30:00, and the current system time is July 1, 2024, at 14:20:00. The calculation time interval T1 is the profile collection time minus the data storage start time, which is 6 months (accurate to the second of 15,768,000). The calculation time interval T2 is the current time minus the data storage start time, which is also 6 months (accurate to the second of 15,768,000). Data update ratio. A query of the customer referral database revealed that customer C1001 has a total of 5000 records (N), and the amount of data to be updated is... The system updates the first 5000 records in ascending order of timestamps, starting from the data storage start time of January 1, 2024, at 00:00:00. The update operations include recalculating the product weight score for each record, based on the customer's browsing history and purchase records in the latest customer behavior profile; updating the confidence of association rules, which is recalculated by analyzing the customer's recent shopping cart combinations and purchase sequences; and updating the collaborative filtering similarity value, which is adjusted based on the comparison results between the customer's current behavior pattern and similar customer groups. All update operations are completed within a single database transaction to ensure data consistency. The updated customer recommendation data takes effect immediately, affecting subsequent recommendation results. When generating the recommendation list, the recommendation engine reads the updated weight scores, association rule confidence scores, and similarity values, combining these factors to calculate the product recommendation score. The recommendation results are presented to the customer in sorted order of scores. This data update operation enables the recommendation system to quickly respond to changes in customer behavior, maintaining the real-time nature and relevance of the recommended content. The system records metadata for each data update, including update time, updated data volume, previous version number, and new version number. The update history supports troubleshooting and performance analysis; business personnel can query the data update records of specific customers to understand adjustments to recommendation strategies.
[0050] The customer recommendation data version management maintains data snapshots at multiple time points, with different versions marked by timestamps. When the data update ratio is large, the system may create new data branches for new versions, retaining older versions for rollback and comparison. Difference analysis between versions reveals customer behavior evolution trends, which are used for long-term interest modeling and prediction. An A / B testing framework compares the effectiveness of different update strategies. A / B tests group customers with different update parameters and evaluate key performance indicators (KPIs) for each group. KPIs include recommendation click-through rate, conversion rate, and average order value. Test results guide the optimization of update strategies. The system periodically cleans up outdated data versions to free up storage space; retention strategies are based on a balance between data importance and storage costs. The calculation of the data update ratio relies on accurate time information; the system's time synchronization service ensures clock consistency across all nodes. The time synchronization service uses Network Time Protocol (NTP) to calibrate server time, preventing errors in ratio calculation due to clock deviations. The profile collection timestamp is added by the data acquisition module when generating real-time behavioral profiles, with timestamp accuracy down to the millisecond level. The data storage start time is read from configuration management, which supports dynamic adjustment of the start time to adapt to changes in business needs. The monitoring and alarm mechanism detects abnormal data update ratios. When the data update ratio continuously exceeds a reasonable range, an alarm is triggered, prompting the administrator to check the system status.
[0051] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A customer personalized recommendation method based on retail data, characterized in that, The method includes: Under the condition of meeting the preset recommendation conditions, the historical behavior profile corresponding to the customer ID is obtained. The historical behavior profile includes the historical active duration and historical behavior parameters. Real-time behavior profiles of the same customer ID are collected. The real-time behavior profile includes the real-time active duration and real-time behavior parameters. A behavior deviation vector is generated using the historical behavior profile and the real-time behavior profile. Search for a set of behavior deviation values that match the behavior deviation vector. When the variance of the set of behavior deviation values reaches or exceeds a preset variance threshold, measure the ratio of the time interval between the data storage start time and the profile acquisition time to the time interval between the data storage start time and the current time, and set this ratio as the data update ratio. Modify the customer recommendation data according to the data update ratio.
2. The customer personalized recommendation method based on retail data as described in claim 1, characterized in that, The steps for obtaining the historical behavior profile corresponding to the customer identifier include: obtaining multiple preset behavior parameters of the customer identifier; for each preset behavior parameter, calculating the abnormal active duration of customer behavior data based on the historical behavior sample library; selecting the minimum active duration from all abnormal active durations, and multiplying the minimum active duration by a preset coefficient as the preset recommendation condition.
3. The customer personalized recommendation method based on retail data as described in claim 2, characterized in that, When the behavioral data belongs to the underlying behavioral attributes, the step of calculating the alienated active duration for the preset behavioral parameters includes: searching for multiple sets of underlying behavioral record feature values and corresponding active duration record values that satisfy the preset behavioral parameters from the historical behavioral sample library; performing deviation calculation on the multiple sets of underlying behavioral record feature values and active duration record values to obtain a set of record feature value deviation modulus values and a set of active duration record value deviation modulus values; selecting a subset of active duration record value deviation modulus values from the set of active duration record value deviation modulus values where the record feature value deviation modulus value is not less than a preset feature deviation threshold; and performing mode calculation on the subset of active duration record value deviation modulus values to obtain the alienated active duration.
4. The customer personalized recommendation method based on retail data as described in claim 3, characterized in that, The underlying behavioral attributes include atomic operation records in customer interaction events. Each atomic operation record contains an operation timestamp, an operation type identifier, and an operation object identifier. The set of record feature value deviation modulus values is obtained by calculating the frequency of occurrence of the operation type identifier and the distribution entropy value of the operation object identifier in the atomic operation record. The high-level mapping attribute includes customer intent tags generated by a behavior pattern recognition model, which are associated with a specific product category; the behavior data mapper maps the operation records in the low-level behavior attributes to an intent probability distribution through a neural network model, and the feature value of the high-level mapping attribute is the intent tag corresponding to the highest probability in the intent probability distribution.
5. The customer personalized recommendation method based on retail data as described in claim 4, characterized in that, When the behavioral data belongs to a high-level mapping attribute, the step of calculating the alienated active duration for the preset behavioral parameters includes: calling a behavioral data mapper associated with the high-level mapping attribute and the preset behavioral parameters, wherein the behavioral data mapper receives the low-level behavioral attributes; retrieving multiple sets of low-level behavioral attribute record feature values and corresponding active duration record values that satisfy the preset behavioral parameters from the historical behavioral sample library; processing each set of low-level behavioral attribute record feature values through the behavioral data mapper to obtain multiple sets of high-level mapping attribute feature values; and combining the multiple sets of high-level mapping attribute feature values and the multiple sets of active duration record values to deduce the alienated active duration.
6. The customer personalized recommendation method based on retail data as described in claim 5, characterized in that, It also includes the following steps: Preset underlying behavior attributes can be configured through the user interface to define preset high-level mapping attributes; Multiple sets of training samples are collected, constrained by preset customer classification and preset behavioral parameters, to train the behavior data mapper; The trained behavior data mapper is bound to the preset high-level mapping attributes and the preset behavior parameters; wherein, each training sample includes preset low-level behavior attribute record data and preset high-level mapping attribute record data.
7. The customer personalized recommendation method based on retail data as described in claim 1, characterized in that, The behavioral deviation vector consists of an active duration deviation vector and a behavioral parameter deviation vector; the steps of searching for the behavioral deviation value set include: sending a search request to the cloud, obtaining an encryption key, encrypting the behavioral dimension, customer classification, the active duration deviation vector, and the behavioral parameter deviation vector, generating ciphertext and uploading it to the cloud, receiving search feedback information returned by the cloud, the search feedback information containing an initial behavioral deviation value set; and performing outlier removal processing on the initial behavioral deviation value set to obtain the behavioral deviation value set.
8. The customer personalized recommendation method based on retail data as described in claim 7, characterized in that, The step of sending a retrieval request to the cloud includes: the cloud decrypting the ciphertext to obtain the behavior dimension, the customer category, the active duration deviation vector, and the behavior parameter deviation vector, forming a retrieval task constraint; sending a retrieval request to a distributed node through the cloud to obtain a distributed node key; using the distributed node key to encrypt the retrieval task constraint and the retrieval key to generate retrieval task ciphertext, sending it to the distributed node, and receiving retrieval feedback information returned by the distributed node.
9. The customer personalized recommendation method based on retail data as described in claim 1, characterized in that, The steps for modifying customer recommendation data based on the data update ratio include: determining the amount of data to be updated based on the data update ratio, and selecting a corresponding number of data records to perform the update operation starting from the data storage start time.
10. A customer personalized recommendation system based on retail data, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the customer personalized recommendation method based on retail data as described in any one of claims 1 to 9.