A client behavior analysis-oriented burying point analysis system
By constructing a digital twin model using an adaptive sampling interval algorithm and deep learning technology, the problems of low data collection efficiency, insufficient resource utilization, and inadequate prediction in traditional customer behavior analysis systems are solved, achieving efficient and accurate customer behavior analysis and decision support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI ZHULIN INFORMATION TECH CO LTD
- Filing Date
- 2025-07-09
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional customer behavior analysis systems suffer from low data collection efficiency and quality, insufficient resource utilization, lack of consideration for individual customer differences, and inadequate predictive capabilities, which limits the accuracy of analysis results and the scientific nature of corporate decision-making.
An adaptive sampling interval algorithm is used to dynamically adjust data collection. A digital twin model is built by combining deep learning and data mining technologies. Through business scenario simulation and behavior prediction and decision support modules, a deep understanding and accurate prediction of customer behavior can be achieved.
It improves the comprehensiveness and accuracy of data collection, optimizes the efficiency of system resource utilization, enhances the ability to understand customer behavior, assists enterprises in making scientific and reasonable decisions, and enhances market competitiveness.
Smart Images

Figure CN120911248B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of customer behavior analysis, specifically to a data tracking analysis system for customer behavior analysis. Background Technology
[0002] With the rapid development of internet technology and the deepening of digital transformation, enterprises have an increasing need for insights into customer behavior. Customer behavior analysis has become a key means for enterprises to optimize product iteration, adjust marketing strategies, improve user experience, and enhance market competitiveness.
[0003] Traditional customer behavior analysis systems often face problems such as low data collection efficiency and quality, and insufficient utilization of system resources when processing data from tracking points. On the one hand, due to the diversity and dynamism of customer behavior, fixed-interval data collection methods are difficult to adapt to different customer behavior patterns, resulting in data redundancy or missing data, which affects the accuracy of analysis results. On the other hand, traditional systems lack effective consideration for individual customer differences during data processing and analysis, often adopting a one-size-fits-all analysis method, which makes it difficult to accurately capture the unique behavioral characteristics and needs of each customer. In addition, when predicting customer behavior, traditional systems rely heavily on simple statistical models or experience-based judgments, lacking the ability to deeply mine and accurately predict complex customer behavior patterns, which limits the scientific and rational nature of enterprise decision-making.
[0004] In view of the problems of low data collection efficiency and quality, insufficient system resource utilization, and insufficient prediction accuracy of traditional customer behavior analysis systems, it is particularly important to develop a data tracking analysis system for customer behavior analysis. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a data tracking analysis system for customer behavior analysis. By introducing an adaptive sampling interval algorithm, it can dynamically adjust the data sampling interval according to the activity level of customer behavior, effectively improving the comprehensiveness and accuracy of data collection, while optimizing the efficiency of system resource utilization. In addition, the system combines deep learning algorithms and data mining technology to construct a digital twin model that accurately reflects the characteristics of customer behavior. Through business scenario simulation and behavior prediction and decision support modules, it achieves a deep understanding and accurate prediction of customer behavior.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a data tracking analysis system for customer behavior analysis, which includes the following components: a data acquisition module, a digital twin model construction module, a business scenario simulation module, and a behavior prediction and decision support module;
[0007] The data acquisition module collects various customer behavior data in real time by setting up tracking points on the product interface and application client, including but not limited to click behavior, browsing path, operation duration, purchase records, and search keywords. At the same time, it collects basic customer information to provide a data foundation for subsequent modeling.
[0008] The digital twin model construction module: Based on the acquired customer historical behavior data and basic information, it uses deep learning algorithms and data mining technology to construct a digital twin model of each customer. Combined with the customer's basic information, through feature fusion algorithms, the model can comprehensively consider the individual differences of customers, thereby constructing a digital twin model that accurately reflects the customer's behavioral characteristics.
[0009] The business scenario simulation module creates diverse virtual business scenarios based on the actual needs of enterprises. It imports the constructed digital twin models of individual customers into each virtual business scenario and simulates the behavioral decision-making paths of customers in different scenarios through a model-driven approach.
[0010] The behavior prediction and decision support module analyzes and processes the behavior simulation results of the individual customer digital twin models in the business scenario simulation module. Using statistical methods and machine learning algorithms, it predicts how customers will respond to product iteration and marketing strategy adjustment decisions in actual business scenarios. The prediction results are presented to enterprise decision-makers in a visual form, and detailed decision-making plans are generated, including product improvement suggestions, marketing strategy optimization plans, and resource allocation plans, to assist enterprises in making scientific and reasonable decisions.
[0011] Furthermore, in the data acquisition module, when collecting customer behavior data, an adaptive sampling interval algorithm is used to collect the data. This algorithm dynamically adjusts the data sampling interval according to the activity level of customer behavior, thereby ensuring the accuracy of data collection while optimizing the use of system resources.
[0012] Specifically, customer behavior activity is comprehensively evaluated by factors such as the number of behavioral events and the complexity of operations within a unit of time. The algorithm calculates the average value of customer behavior activity within a fixed window by statistically analyzing customer behavior data within that window, and uses this as the basis for adjusting the sampling interval. When a customer triggers a large number of behavioral events within a unit of time and the operations are highly complex, it indicates that the customer behavior is active. At this time, the algorithm will automatically shorten the sampling interval and collect data more frequently to ensure that every detail of the customer behavior can be captured. Conversely, when the number of customer behavioral events is small and the operations are simple, that is, when the behavior is in a flat state, the algorithm will increase the sampling interval to reduce unnecessary data collection, reduce data redundancy, and reduce the pressure on the system to store and process data.
[0013] The window size used for statistical behavioral data was selected as the optimal value after extensive experiments comparing data integrity and computational efficiency under different windows. The key parameter used to adjust the sampling interval in the algorithm is determined according to the characteristics of the business scenario. This adaptive sampling interval algorithm can accurately adapt to the dynamic changes in customer behavior, improve data collection efficiency and quality, and provide a higher quality and more valuable data foundation for subsequent modules such as digital twin model construction.
[0014] Furthermore, in the digital twin model construction module, when extracting features from customer behavior data, a multi-scale attention feature extraction algorithm is used. First, the original behavior data is divided into subsequences of different scales according to the time series. For a length of... The original data sequence Divided into Subsequences of different scales , Then, for each subsequence, the attention weight is calculated using the following formula:
[0015]
[0016] in , For a custom similarity function, The algorithm, defined as the subsequence length, can focus on key features of customer behavior at different scales, effectively extracting feature vectors that reflect customer behavior patterns and decision-making logic. Compared with traditional feature extraction methods, it can capture customer behavior features more comprehensively and accurately, improving the accuracy and effectiveness of digital twin models.
[0017] Furthermore, in the digital twin model construction module, when performing feature fusion by combining customer basic information and behavioral data, a dynamic weight fusion algorithm is adopted. Let the customer basic information feature vector be... The behavioral data feature vector is The fused feature vector The calculation formula is:
[0018]
[0019] in and These are dynamic weighting coefficients, determined based on the correlation between customer behavior and basic information. This correlation is calculated using mutual information. , , They are respectively , The probability distribution of the features in the middle. As a joint probability distribution, it is adjusted using an adaptive fuzzy control algorithm based on the mutual information value. and When the mutual information value is high, increase the weight related to the basic information. Conversely, it increases. This allows the integrated features to reflect both individual customer differences and behavioral characteristics, thereby improving the accuracy of digital twins in simulating customer behavior.
[0020] Furthermore, in the business scenario simulation module, when creating virtual business scenarios, a scenario generation algorithm based on generative adversarial networks is used, and the generator... Receive random noise vector and business requirement condition vector Generate virtual business scenarios Discriminator Used to determine if the input scenario is a real-world scenario sample. Or a generated virtual scene During training, a simplified adversarial training objective function is used:
[0021]
[0022] in Indicates samples from real-world scenarios Expectation calculation, Represents the random noise vector and business requirement condition vector The expected value calculation of the combination is achieved by taking the logarithm of the discrimination probabilities of the real scene and the generated scene to construct the adversarial optimization objective of the generator and the discriminator, so as to prompt the generator to generate more realistic virtual business scenarios and the discriminator to improve its discrimination ability.
[0023] To balance the training process of the generator and discriminator, an improved weight adjustment factor is introduced. The parameter update step size of the generator and discriminator is dynamically adjusted using the following formula:
[0024]
[0025] in The accuracy of the discriminator on the current batch of data, The accuracy at which the scene generated by the generator is identified as a real scene; when the discriminator accuracy is high, Increasing the value allows for a larger step size in the generator parameter updates, accelerating the generator's optimization speed; conversely, decreasing the value indicates that the generator is performing well. Decreasing the value gives the discriminator more opportunities to optimize, thereby achieving a dynamic balance in the training of both.
[0026] Furthermore, in the business scenario simulation module, when importing the digital twin model into a virtual business scenario for behavioral simulation, a reinforcement learning-driven model interaction algorithm is adopted. The state space S is defined as the set of current states of the digital twin model in the scenario, including customer behavioral characteristics and scenario parameters. The action space A is the set of actions the model can take. The reward function R(s, a) is used to evaluate the model's performance in each state. Next action The subsequent benefits are derived from the model's interaction with the virtual scene, based on the state. Select Action Receive rewards And transition to the new state The goal is to maximize long-term cumulative rewards:
[0027]
[0028] in This is a discount factor, and its value ranges from [value range missing]. The model searches for the optimal value within a certain range using the simulated annealing algorithm. The model parameters are updated using an improved algorithm based on trust region policy optimization. By limiting the magnitude of each policy update, the stability and convergence of the algorithm are ensured, enabling the digital twin model to learn more reasonable behavioral decision-making strategies in virtual business scenarios and improving the realism and accuracy of behavioral simulation.
[0029] Furthermore, in the behavior prediction and decision support module, when analyzing and processing the simulation results to predict customer behavior, a spatiotemporal correlation prediction algorithm is used. This algorithm considers the correlation characteristics of customer behavior in the time and space dimensions. For the time dimension, a time series model is constructed, using an improved Transformer architecture and introducing a location encoding mechanism. The location encoding formula is as follows:
[0030]
[0031] in The position in the time series. For dimensional indexing, For the model dimension, and for the spatial dimension, combining customer geographic location information, the influence weights of different spatial locations on customer behavior are calculated using a spatial attention mechanism. The weight calculation formula is as follows:
[0032]
[0033] in , A learnable feature mapping vector. By integrating information from both temporal and spatial dimensions to determine the spatial location, and using a multilayer perceptron for prediction, the probability distribution of customer behavior in different times and spaces in the future can be obtained, thereby achieving more accurate customer behavior prediction and providing a more comprehensive basis for enterprise decision-making.
[0034] Furthermore, in the behavior prediction and decision support module, when generating decision-making plans, a knowledge graph-based intelligent decision recommendation algorithm is adopted. First, a knowledge graph related to customer behavior and business decisions is constructed. Nodes in the graph include customer behavior characteristics, business strategies, and market environment factors, while edges represent the relationships between nodes. For the predicted customer behavior results, retrieval and reasoning are performed in the knowledge graph. A graph neural network is then used to calculate the matching degree between different business decision-making plans and the current customer behavior and market environment. The matching degree calculation formula is as follows:
[0035]
[0036] in Indicates the first The decision-making options and the current situation are in the [number]th [year]. Matching degree in each dimension For activation function, For nodes The set of connected neighbor nodes, This is the weight matrix. The feature vectors of the neighboring nodes, Using a bias vector, decision-making schemes are sorted and filtered based on matching degree. Combined with the company's business objectives and resource constraints, personalized decision-making plans are generated. At the same time, explanations of the decision-making schemes are provided to help corporate decision-makers understand the basis for their decisions and improve the acceptability and effectiveness of their decisions.
[0037] Furthermore, the data acquisition module, digital twin model construction module, business scenario simulation module, and behavior prediction and decision support module interact with each other through a communication mechanism improved based on a message queue telemetry transmission protocol. A data priority labeling and adaptive flow control strategy are introduced. Different priorities are assigned to different types of data, such as collected customer behavior data, model training parameters, and simulation results, based on their importance and timeliness. The priority calculation formula is as follows:
[0038]
[0039] in For data priority, The importance of data is quantified based on its role and impact within the system. This is a quantifiable value for data timeliness, reflecting the urgency at which the data needs to be processed. , The weighting coefficients are determined using the analytic hierarchy process. During communication, the data transmission rate is dynamically adjusted based on network bandwidth and system load. A fuzzy logic-based flow control algorithm is used to adjust the transmission window size according to network latency and packet loss rate indicators, ensuring the efficiency, stability, and real-time performance of data transmission between modules, and guaranteeing the overall normal operation and accuracy of the system analysis.
[0040] Compared with existing technologies, this data tracking analysis system for customer behavior analysis has the following advantages:
[0041] I. This system utilizes a digital twin model building module, combined with deep learning algorithms and data mining techniques, to construct a digital twin model that accurately reflects customer behavioral characteristics. Through a business scenario simulation module, the model is imported into diverse virtual business scenarios for behavioral simulation. Furthermore, a reinforcement learning-driven model interaction algorithm enables the model to learn more reasonable behavioral decision-making strategies. Finally, the behavior prediction and decision support module employs spatiotemporal correlation prediction algorithms and knowledge graph-based intelligent decision recommendation algorithms to accurately predict customer behavior and generate personalized decision-making plans. These functions collectively enhance enterprises' ability to understand customer behavior, assisting them in making more scientific and reasonable decisions and improving market competitiveness.
[0042] Second, this system introduces an adaptive sampling interval algorithm, which dynamically adjusts the data sampling interval in the data acquisition module based on the activity level of customer behavior. When customer behavior is active, the system automatically shortens the sampling interval to capture more details, and when behavior is calm, the sampling interval is increased to reduce redundant data. This mechanism not only ensures the comprehensiveness and accuracy of data acquisition, but also significantly reduces the pressure on the system to store and process data, thereby optimizing the efficiency of system resource utilization.
[0043] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0045] Figure 1 This is an overall flowchart of a data tracking analysis system for customer behavior analysis.
[0046] Figure 2 This is a flowchart framework for key modules of a data tracking analysis system for customer behavior analysis. Detailed Implementation
[0047] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0048] Example 1
[0049] Before a quarterly promotional event, a leading e-commerce platform discovered a significant gap in the user conversion path from browsing to adding to cart to placing an order. Especially during the peak promotional period when traffic surged, traditional marketing strategies struggled to accurately reach different types of users. The platform aimed to use a data tracking system to deeply analyze user behavior, optimize product recommendation logic and promotional activity design, and improve the overall conversion rate during the peak promotional period.
[0050] We deploy tracking points on core interfaces such as the platform homepage, product detail pages, and shopping cart checkout pages to collect real-time user behavior data, including the frequency of user clicks on product categories, the duration of browsing product details, records of adding and deleting products in the shopping cart, and search keywords. We also collect basic information filled in by users during registration, such as female users aged 25-35 who are more interested in beauty and personal care products, and users in North China who have higher requirements for the delivery time of fresh and cold chain products.
[0051] An adaptive sampling interval algorithm is used to optimize data collection efficiency: when the algorithm detects that a user has browsed 20 products continuously within 10 minutes and compared prices multiple times, it determines that the user is in a highly active state and automatically shortens the sampling interval to 5 seconds / time to ensure that every action of adding to cart or favorites is captured. If the user only opens the APP to browse the homepage and then remains idle for 30 minutes, the algorithm extends the sampling interval to 1 minute / time to reduce the storage of invalid data.
[0052] Based on user behavior data and basic information from the past six months, a multi-scale attention feature extraction algorithm is used to process the raw data. For example, the user's browsing history for the past 7 days is divided into different subsequences according to "daily time period" (morning, noon, and evening) and "product category" (clothing / digital / home furnishings). The algorithm automatically focuses on the user's behavior characteristics of frequently browsing digital products at 8 pm on weekdays, assigning higher attention weight to the data in this time period. The formula for calculating attention weight is:
[0053]
[0054] in , For a custom similarity function, is the length of the subsequence.
[0055] Combining basic user information and behavioral characteristics, an individual model is generated using a dynamic weight fusion algorithm. Let the feature vector of basic customer information be... The behavioral data feature vector is The fused feature vector The calculation formula is:
[0056]
[0057] in and The weighting coefficients are dynamic and are adjusted based on mutual information values using an adaptive fuzzy control algorithm. and For a 30-year-old white-collar user with monthly spending of 5,000 yuan, the basic information characteristics of "high spending power" and the behavioral characteristics of "frequent purchase of light luxury accessories" will be weighted by an adaptive fuzzy control algorithm. The resulting digital twin model will be more inclined to recommend brand products with an average order value of 800-1,500 yuan.
[0058] Virtual promotional scenarios are created using a scene generation algorithm based on generative adversarial networks: Generator Receive random noise vector and business requirement condition vector Generate virtual business scenarios Discriminator Used to determine if the input scenario is a real-world scenario sample. Or a generated virtual scene During training, a simplified adversarial training objective function is used:
[0059]
[0060] in Indicates samples from real-world scenarios Expectation calculation, Represents the random noise vector and business requirement condition vector Calculation of the expected value of the combination;
[0061] To balance the training process of the generator and discriminator, an improved weight adjustment factor is introduced. The parameter update step size of the generator and discriminator is dynamically adjusted using the following formula:
[0062]
[0063] in The accuracy of the discriminator on the current batch of data, The accuracy at which the scene generated by the generator is identified as a real scene; when the discriminator accuracy is high, Increasing the value allows for a larger step size in the generator parameter updates, accelerating the generator's optimization speed; conversely, decreasing the value indicates that the generator is performing well. The reduced value allows the discriminator more opportunities to optimize. When inputting business requirements such as "¥50 off for every ¥300 spent", "limited-time flash sale", and "member exclusive discount", the generator will simulate page layout and traffic distribution scenarios under different discount levels. For example, when generating a virtual scenario of "homepage focus image displaying a beauty set flash sale", the discriminator will compare the actual traffic conversion data during historical promotional periods to ensure the authenticity of the virtual scenario. During the training process, if the discriminator accurately identifies the virtual scenario three times in a row, the system will increase the parameter update step size of the generator through an improved weight adjustment factor to accelerate the optimization of the realism of the virtual scenario.
[0064] After importing the user's digital twin model into a virtual scene, the behavior is simulated through a reinforcement learning-driven model interaction algorithm: assuming the model is in the state of "browsing a beauty flash sale page", the action space includes "click to buy now", "add to cart", "view user reviews", etc. The reward function is set according to historical data - the conversion rate of users who click on reviews and then place an order is higher. Therefore, performing the "view reviews" action will obtain a higher reward value. Through continuous interaction, the model simulates the user's decision-making path under different promotional strategies. For example, 60% of users will choose to place an order after viewing more than 3 positive reviews.
[0065] The simulation results were analyzed using a spatiotemporal correlation prediction algorithm: Considering the correlation characteristics of customer behavior in the time and space dimensions, a time series model was constructed for the time dimension. An improved Transformer architecture was adopted, and a location encoding mechanism was introduced. The location encoding formula is as follows:
[0066]
[0067] in The position in the time series. For dimensional indexing, For the model dimension, and for the spatial dimension, combining customer geographic location information, the influence weights of different spatial locations on customer behavior are calculated using a spatial attention mechanism. The weight calculation formula is as follows:
[0068]
[0069] in , A learnable feature mapping vector. To determine the spatial location quantity, information from both temporal and spatial dimensions is fused and predicted using a multilayer perceptron to obtain the probability distribution of customer behavior in different times and spaces in the future. In the temporal dimension, based on the improved Transformer architecture, it is identified that the peak ordering period during major promotions is concentrated between 8 PM and 10 PM, and the location encoding mechanism strengthens the feature weights of time points such as "8 PM" and "9 PM". In the spatial dimension, combined with the user's delivery address, the spatial attention mechanism calculates that users in East China pay 1.8 times more attention to the "next-day delivery" service than those in Northwest China. After fusing spatiotemporal information, the model predicts that "pushing beauty flash sale products with the next-day delivery label at 8 PM in East China" will increase the click-through rate by 25%.
[0070] A knowledge graph-based intelligent decision-making and recommendation algorithm generation scheme is proposed as follows: First, a knowledge graph related to customer behavior and business decisions is constructed. Nodes in the graph include customer behavior characteristics, business strategies, and market environment factors, while edges represent the relationships between nodes. For predicted customer behavior results, retrieval and reasoning are performed within the knowledge graph. A graph neural network is then used to calculate the matching degree between different business decision options and the current customer behavior and market environment. The matching degree calculation formula is as follows:
[0071]
[0072] in Indicates the first The decision-making options and the current situation are in the [number]th [year]. Matching degree in each dimension For activation function, For nodes The set of connected neighbor nodes, This is the weight matrix. The feature vectors of the neighboring nodes, Using the bias vector, a knowledge graph is constructed containing nodes such as "user browsing time > 5 minutes", "previously purchased similar products", and "high price sensitivity during promotional periods". The edge relationship is defined as "positive correlation between browsing time and purchase intention". For the user group with "high intention but no order" in the prediction results, the algorithm recommends a combination strategy of "limited-time free shipping + free trial" after searching the graph. The graph neural network calculates that the matching degree between this strategy and user behavior reaches 82%. The final decision plan includes a specific product recommendation list (such as a certain brand of foundation), discount (30 yuan off) and push time (July 15, 20:00), and a visual explanation of the knowledge graph reasoning process to help the operations team understand "why this strategy is recommended".
[0073] Example 2
[0074] An internet finance platform discovered that middle-aged users aged 35-45 exhibited a phenomenon of "high pageviews but low conversion rates" when purchasing wealth management products. Furthermore, some users suddenly churned after browsing high-risk products. The platform hopes to use a data tracking system to uncover the investment preferences of this group, optimize product recommendation logic, reduce churn rate, and increase the asset allocation scale of high-net-worth users.
[0075] We deploy tracking points on pages such as "Homepage Recommendations," "Wealth Management Product Details," and "Risk Assessment" in financial apps to collect behavioral data such as the number of times users click on products with different risk levels, the duration of viewing historical return curves on product detail pages, operation records of adjusting investment periods, and search keywords. At the same time, we collect basic information such as users' occupation, annual income, and risk assessment results.
[0076] The adaptive sampling interval algorithm dynamically adjusts in this scenario: when a user continuously compares the historical maximum drawdown data of 3 R3-level wealth management products and saves the product details, the algorithm determines that the user is in a critical period of investment decision-making, and the sampling interval is shortened to 10 seconds / time to capture every data filtering action. If the user only opens the APP to check the account balance and then exits, the sampling interval is extended to 2 minutes / time to reduce data redundancy of low-frequency operations.
[0077] The algorithm uses a multi-scale attention feature extraction algorithm to process behavioral data: the user's investment records over the past three months are divided into subsequences based on "weekly transaction frequency" and "product type preference". The algorithm will focus on the behavior pattern of this middle-aged group redeeming financial products at the end of the quarter, and highlight the feature of "funds returning at the end of the quarter" through attention weight.
[0078] The dynamic weight fusion algorithm combines basic user information and behavioral characteristics: For a 40-year-old corporate executive with an annual income of 800,000, the basic information characteristic of "high risk tolerance" and the behavioral characteristic of "70% of funds allocated to equity funds in the past 6 months" will be adjusted by an adaptive fuzzy control algorithm. The resulting digital twin model is more inclined to recommend a mixed asset allocation scheme rather than a single low-risk product.
[0079] A scenario generation algorithm based on generative adversarial networks creates virtual investment scenarios: Inputting business requirements such as "equity-bond ratio 6:4", "linked to gold index", and "guaranteed minimum return of 2%", the generator simulates the return fluctuation curve scenarios of different asset portfolios. For example, when generating a virtual scenario of "mixed fund + gold ETF combination", the discriminator compares it with real historical market data to ensure the rationality of the return curve. When the scenario generated by the generator is misjudged by the discriminator as a real scenario, the system reduces the parameter update step size of the generator through an improved weight adjustment factor to avoid model overfitting.
[0080] The reinforcement learning-driven model interaction algorithm simulates user decision-making: Assuming the model is in the state of "browsing the details page of a mixed fund", the action space includes "viewing holding details", "calculating the return of fixed investment", "consulting investment advisors", etc. The reward function is set according to historical data - the probability of purchasing after viewing the holding details increases by 40%, so performing this action will get a higher reward. The model simulates through interaction that 62% of middle-aged users will further consult investment advisors after viewing the top 10 holding stocks and the fund manager's past performance.
[0081] Analysis of the spatiotemporal correlation prediction algorithm simulation results: In the time dimension, the improved Transformer architecture identifies that middle-aged users have a stronger willingness to invest at the beginning of the quarter, and the location encoding mechanism strengthens the time characteristics of the "first week of the quarter"; In the spatial dimension, combined with the high cost of living pressure in the user's city (such as first-tier cities), the spatial attention mechanism calculates that this group is more sensitive to the combination of "capital-protected products as a base + a small proportion of high-risk products to seek returns". After fusion, the model predicts that "pushing a combination plan of '60% stable wealth management + 30% stock funds + 10% gold ETF' to middle-aged users in first-tier cities in early April" will increase the purchase conversion rate by 35%.
[0082] A knowledge graph-based intelligent decision-making recommendation algorithm is developed: A knowledge graph containing nodes such as "risk assessment result is balanced," "redeemed stock funds in the past six months," and "pays attention to macroeconomic news" is constructed. Edge relationships are defined as the "correlation between redemption behavior and market fluctuations." For users predicted to have "not reinvested after redemption," the algorithm searches the knowledge graph and recommends a combination strategy of "fixed income + fund + market analysis report." The graph neural network calculates a matching degree of 78%. The decision plan includes specific products, allocation ratios, and delivery methods. The knowledge graph is used to visually demonstrate "why this combination can balance returns and risks," helping financial advisors accurately convey the value of the plan to users.
[0083] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A data tracking analysis system for customer behavior analysis, characterized in that, The system comprises the following components: a data acquisition module, a digital twin model construction module, a business scenario simulation module, and a behavior prediction and decision support module; The data acquisition module collects various types of customer behavior data in real time by setting up tracking points on the product interface and application client. These data include, but are not limited to, click behavior, browsing path, operation duration, purchase records, and search keywords. At the same time, it collects basic customer information. When collecting customer behavior data, an adaptive sampling interval algorithm is used to collect the data. This algorithm dynamically adjusts the data sampling interval based on the activity level of customer behavior. Specifically, customer behavior activity is comprehensively evaluated by factors such as the number of behavioral events and the complexity of operations within a unit of time. The algorithm calculates the average customer behavior activity within a fixed window by statistically analyzing customer behavior data. When a customer triggers a large number of behavioral events within a unit of time and the operations are highly complex, it indicates that the customer behavior is active. At this time, the algorithm will automatically shorten the sampling interval and collect data more frequently. Conversely, when the number of customer behavioral events is small and the operations are simple, that is, when the behavior is in a flat state, the algorithm will increase the sampling interval to reduce unnecessary data collection, reduce data redundancy, and alleviate the pressure on the system to store and process data. The window size used for statistical behavioral data was selected as the optimal value after a large number of experiments comparing data integrity and computational efficiency under different windows. The key parameters used to adjust the sampling interval in the algorithm will be determined according to the characteristics of the business scenario. The digital twin model construction module, based on acquired customer historical behavior data and basic information, employs deep learning algorithms and data mining techniques to construct an individual customer digital twin model. Combining this with basic customer information, a feature fusion algorithm enables the model to comprehensively consider individual customer differences. When extracting features from customer behavior data, a multi-scale attention feature extraction algorithm is used. First, the original behavior data is divided into subsequences of different scales according to the time series. For a length of... The original data sequence Divided into Subsequences of different scales , Then, for each subsequence, the attention weight is calculated using the following formula: in , For a custom similarity function, This represents the total number of subsequences at the current scale. The business scenario simulation module creates diverse virtual business scenarios based on the actual needs of enterprises. It imports the constructed digital twin models of individual customers into each virtual business scenario and simulates the behavioral decision-making paths of customers in different scenarios through a model-driven approach. The behavior prediction and decision support module analyzes and processes the behavior simulation results of the individual customer digital twin models in the business scenario simulation module. Using statistical methods and machine learning algorithms, it predicts how customers will respond to product iteration and marketing strategy adjustment decisions in actual business scenarios. The prediction results are presented to enterprise decision-makers in a visual form, and detailed decision-making plans are generated, including product improvement suggestions, marketing strategy optimization plans, and resource allocation plans, to assist enterprises in making scientific and reasonable decisions.
2. The data tracking analysis system for customer behavior analysis according to claim 1, characterized in that, In the digital twin model construction module, a dynamic weighted fusion algorithm is used when combining customer basic information and behavioral data for feature fusion. Let the customer basic information feature vector be... The behavioral data feature vector is The fused feature vector The calculation formula is: in and The weighting coefficients are dynamic and are adjusted based on mutual information values using an adaptive fuzzy control algorithm. and .
3. The data tracking analysis system for customer behavior analysis according to claim 1, characterized in that, In the business scenario simulation module, a scenario generation algorithm based on generative adversarial networks is used when creating virtual business scenarios. Receive random noise vector and business requirement condition vector Generate virtual business scenarios Discriminator Used to determine if the input scenario is a real-world scenario sample. Or a generated virtual scene During training, a simplified adversarial training objective function is used: in Indicates samples from real-world scenarios Expectation calculation, Represents the random noise vector and business requirement condition vector Calculation of the expected value of the combination; To balance the training process of the generator and discriminator, an improved weight adjustment factor is introduced. The parameter update step size of the generator and discriminator is dynamically adjusted using the following formula: in The accuracy of the discriminator on the current batch of data, The accuracy at which the scene generated by the generator is identified as a real scene; when the discriminator accuracy is high, Increasing the value allows for a larger step size in the generator parameter updates, accelerating the generator's optimization speed; conversely, decreasing the value indicates that the generator is performing well. Decreasing the value gives the discriminator more opportunities to optimize.
4. The data tracking analysis system for customer behavior analysis according to claim 1, characterized in that, In the business scenario simulation module, when importing the digital twin model into a virtual business scenario for behavioral simulation, a reinforcement learning-driven model interaction algorithm is used. The state space S is defined as the set of current states of the digital twin model in the scenario, including customer behavioral characteristics and scenario parameters. The action space A is the set of actions the model can take. The reward function R(s, a) is used to evaluate the model's performance in each state. Next action The subsequent benefits are derived from the model's interaction with the virtual scene, based on the state. Select Action Receive rewards And transition to the new state The goal is to maximize long-term cumulative rewards: in This is the discount factor.
5. The data tracking analysis system for customer behavior analysis according to claim 1, characterized in that, In the behavior prediction and decision support module, when analyzing and processing simulation results to predict customer behavior, a spatiotemporal correlation prediction algorithm is used. This algorithm considers the correlation characteristics of customer behavior in the time and space dimensions. For the time dimension, a time series model is constructed, using an improved Transformer architecture and introducing a location encoding mechanism. The location encoding formula is as follows: in The position in the time series. For dimensional indexing, For the model dimension, and for the spatial dimension, the influence weight of different spatial locations on customer behavior is calculated through a spatial attention mechanism, combining the customer's geographical location information.
6. The data tracking analysis system for customer behavior analysis according to claim 1, characterized in that, In the behavior prediction and decision support module, when generating decision-making plans, an intelligent decision recommendation algorithm based on knowledge graphs is used. First, a knowledge graph related to customer behavior and business decisions is constructed. Nodes in the graph include customer behavior characteristics, business strategies, and market environment factors, while edges represent the relationships between nodes. For the predicted customer behavior results, retrieval and reasoning are performed in the knowledge graph. A graph neural network is then used to calculate the matching degree between different business decision-making plans and the current customer behavior and market environment. The matching degree calculation formula is as follows: in Indicates the first The decision-making options and the current situation are in the [number]th [year]. Matching degree in each dimension For activation function, For nodes The set of connected neighbor nodes, This is the weight matrix. The feature vectors of the neighboring nodes, Using a bias vector, decision-making schemes are ranked and filtered based on matching degree. Combined with the company's business objectives and resource constraints, personalized decision-making plans are generated, while explanations of the decision-making schemes are provided to help corporate decision-makers understand the basis for their decisions.
7. The data tracking analysis system for customer behavior analysis according to claim 1, characterized in that, The data acquisition module, digital twin model construction module, business scenario simulation module, and behavior prediction and decision support module interact with each other through a communication mechanism based on an improved message queue telemetry transmission protocol. A data priority labeling and adaptive flow control strategy are introduced. Different priorities are assigned to different types of data—collected customer behavior data, model training parameters, and simulation results—based on their importance and timeliness. The priority calculation formula is as follows: in For data priority, Quantify the importance of data. This is a quantifiable value for data timeliness, reflecting the urgency at which the data needs to be processed. As a weighting coefficient, the data transmission rate is dynamically adjusted during communication based on network bandwidth and system load. A fuzzy logic-based flow control algorithm is used to adjust the transmission window size according to network latency and packet loss rate indicators.