E-commerce platform user behavior analysis and prediction method based on data feedback
By combining Hawkes behavior excitation modeling with an improved multi-armed bandit algorithm, a user behavior analysis method for e-commerce platforms is constructed. This solves the problems of temporal dependence and non-targeted strategy selection in traditional user behavior analysis methods, and realizes personalized and real-time responsive user behavior prediction.
Patent Information
- Application Number
- CN202510822251.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies fail to effectively consider the temporal dependencies and triggering relationships between behaviors in user behavior analysis on e-commerce platforms. The scoring mechanism is single, the strategy selection is not targeted, and the feedback update mechanism is simple, making it difficult to achieve personalization and real-time response.
Combining Hawkes behavior incentive modeling with the improved multi-arm bandit algorithm, a dynamic scoring mechanism is constructed. By integrating historical scores with trend scores and combining the behavior incentive intensity matrix, arm pulling strategy selection is carried out, and reward correction is performed after feedback, forming a closed-loop optimization.
It realizes the characterization of the temporal dependence and dynamic triggering relationship of user behavior, improves the modeling ability of potential user behavior trends, realizes real-time response and adaptive optimization, and enhances the accuracy of prediction and strategy matching.
Smart Images

Figure CN120689081A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data mining and intelligent decision-making technology, and in particular to a method for analyzing and predicting user behavior on an e-commerce platform based on data feedback. Background Art
[0002] With the rapid development of e-commerce platforms, user behavior data has become diverse, dynamic, and time-series. Accurately understanding and predicting user behavior has become a crucial support for e-commerce platforms to optimize recommendation strategies, increase conversion rates, and provide personalized services. However, user behavior is characterized by significant temporal dependencies and multi-factor driven characteristics. Traditional methods based on static statistics or shallow strategies are unable to meet the needs of modeling complex user behavior patterns.
[0003] In existing technologies, common user behavior analysis methods are mostly based on traditional multi-armed bandit algorithms or decision-making mechanisms such as collaborative filtering. These methods have the following major shortcomings in dealing with temporal correlation and dynamic preference modeling:
[0004] 1. Ignoring the behavior triggering process: Traditional methods usually regard user behaviors as isolated events, fail to fully consider the temporal dependency and triggering relationship between behaviors, and lack an effective causal modeling mechanism.
[0005] 2. Single scoring mechanism: Traditional multi-armed bandit algorithms mostly rely solely on historical return information for scoring and decision-making. They lack the ability to model users' current behavioral trends or potential changes in interests, resulting in delayed responses.
[0006] 3. Strategic arm selection is not targeted: During the arm selection stage, existing methods cannot dynamically guide users based on their current status. They often use fixed strategies or static formulas, making it difficult to achieve personalized behavioral incentives.
[0007] 4. Simple feedback update mechanism: After receiving behavioral feedback, existing algorithms often use simple mean correction or confidence update, which cannot effectively integrate prediction error information for deep adjustment and closed-loop optimization of the model.
[0008] Therefore, how to provide an e-commerce platform user behavior analysis and prediction method based on data feedback is an urgent problem that technicians in this field need to solve. Summary of the Invention
[0009] One purpose of the present invention is to propose a user behavior analysis and prediction method for e-commerce platforms based on data feedback. The present invention combines the Hawkes behavior excitation modeling process with the improved multi-arm bandit algorithm to construct a dynamic scoring mechanism that integrates historical scoring and trend scoring. It describes in detail the closed-loop process of realizing personalized arm-pulling strategy selection and feedback correction driven by behavioral data, and has the advantages of timely response, accurate modeling and strong adaptability.
[0010] The method for analyzing and predicting user behavior on an e-commerce platform based on data feedback according to an embodiment of the present invention includes the following steps:
[0011] S1. Collect user behavior data from e-commerce platforms, including click behavior, browsing behavior, add-to-cart behavior, and order behavior, and construct a chronological sequence of user behavior events on the e-commerce platform.
[0012] S2. Based on the user behavior event sequence of the e-commerce platform, the Hawkes process is used to establish the user behavior excitation intensity function of the e-commerce platform, and the behavior excitation intensity matrix is output to describe the time-driven correlation relationship between different behavior events;
[0013] S3. Initialize the improved multi-arm bandit algorithm, establish multiple arm sets for arm pull options, configure the initial reward estimate, arm pull count, and arm pull selection parameters for each arm pull option, and set the excitation intensity input interface to receive the corresponding intensity value in the behavior excitation intensity matrix as the trigger attraction parameter;
[0014] S4. Construct a scoring fusion structure, set a first input channel to receive the reward estimate as the historical scoring input; set a second input channel to receive the behavior stimulation intensity value as the trend scoring input, and set a weight control parameter to adjust the weight ratio of the historical scoring input and the trend scoring input in the fusion score;
[0015] S5. In each round of arm selection, the scoring fusion structure is called to perform weighted calculation on the historical scoring input and the trend scoring input according to the weight control parameters to generate a fusion score value, and the arm selection option of the current round is selected based on the fusion score value;
[0016] S6. Receive behavioral feedback data generated after the arm-pulling option in the current round, jointly compare the behavioral feedback data with the behavioral stimulation intensity matrix, calculate the prediction deviation, and reconstruct the prediction deviation through feedback to correct the reward estimate of the corresponding arm-pulling option;
[0017] S7. Use the updated reward estimate and arm pull count to feed back into the improved multi-arm bandit algorithm to complete a behavioral data-driven self-update cycle. The improved multi-arm bandit algorithm structure is based on a non-deep model framework and combines behavioral causal modeling components to achieve lightweight strategy optimization based on single-step decision units.
[0018] Optionally, constructing an e-commerce platform user behavior event sequence in chronological order specifically includes: after obtaining the e-commerce platform user behavior data, arranging the behavior events in ascending order according to the timestamp corresponding to each behavior event to construct the e-commerce platform user behavior event sequence.
[0019] Optionally, the S2 specifically includes:
[0020] S21. Arrange each e-commerce platform user behavior event in the e-commerce platform user behavior event sequence in ascending time order, extract the behavior type and timestamp information of each e-commerce platform user behavior event, and construct a behavior time index set. The behavior time index set is used to trigger path calculation;
[0021] S22. Setting a basic excitation rate parameter and a time decay parameter based on the Hawkes process, and defining a behavior excitation mapping function based on the user behavior type of the e-commerce platform, wherein the behavior excitation mapping function serves as a structural component of the excitation intensity calculation model;
[0022] S23. For each target e-commerce platform user behavior event in the behavior time index set, retrieve historical e-commerce platform user behavior events that occurred before the corresponding target e-commerce platform user behavior event, and calculate the excitation contribution value of each historical e-commerce platform user behavior event to the target e-commerce platform user behavior event based on the behavior excitation mapping function and the time decay parameter;
[0023] S24. Perform a weighted summation of the excitation contribution values of all historical e-commerce platform user behavior events, and combine this with a basic excitation rate parameter to generate an excitation intensity value for the target e-commerce platform user behavior event. The excitation intensity value constitutes a single-point output of the e-commerce platform user behavior excitation intensity function.
[0024] S25. Organize the excitation intensity values output by the e-commerce platform user behavior excitation intensity function for all e-commerce platform user behavior events into a behavior excitation intensity matrix. Map the behavior excitation intensity matrix to the excitation intensity input interface and input it into the scoring fusion structure as a trigger gravity parameter to guide the fusion score calculation of subsequent arm options.
[0025] Optionally, the S3 specifically includes:
[0026] S31. Construct an improved multi-armed bandit algorithm and initialize a pull-arm option set in the improved multi-armed bandit algorithm. The pull-arm option set includes multiple e-commerce platform behavior strategy arms, and each e-commerce platform behavior strategy arm is set with a unique pull-arm identifier;
[0027] S32. Configuring an initial reward estimate for each e-commerce platform behavior strategy arm in the improved multi-armed bandit algorithm. The initial reward estimate is set to a predefined constant and is used to initialize the reward of the e-commerce platform behavior strategy arm before effective feedback is available.
[0028] S33: Configure an arm pull count counter, the initial value of which is zero, and is used to record the cumulative number of times the e-commerce platform behavior strategy arm is selected;
[0029] S34. Configure arm selection parameters. The arm selection parameters include confidence estimation items and exploration factor items. They are set using a unified initialization strategy and are used for probability calculation in the subsequent scoring fusion structure.
[0030] S35. Setting an excitation intensity input interface in the improved multi-armed bandit algorithm. The excitation intensity input interface establishes an independent input channel for each e-commerce platform behavior strategy arm, for receiving the excitation intensity value corresponding to the corresponding e-commerce platform behavior strategy arm in the e-commerce platform user behavior excitation intensity matrix;
[0031] S36. Data-bind the e-commerce platform user behavior excitation intensity value received in the excitation intensity input interface with the arm selection parameter of the corresponding e-commerce platform behavior strategy arm to form a trigger gravity parameter input group of the e-commerce platform behavior strategy arm. The trigger gravity parameter input group serves as one of the input variables in the scoring fusion structure and is used to participate in the fusion score calculation of the e-commerce platform user behavior strategy arm.
[0032] Optionally, the improved multi-armed bandit algorithm specifically includes:
[0033] Initialize the arm pull option set, set the initial reward estimate for each e-commerce platform behavior strategy arm, and set the arm pull count counter;
[0034] In each round of arm selection, the traditional score of each e-commerce platform behavior strategy arm is calculated based on the reward estimate and arm pull count counter of each e-commerce platform behavior strategy arm;
[0035] Based on the traditional scoring values of all e-commerce platform behavior strategy arms, the e-commerce platform behavior strategy arm with the largest scoring value is selected as the arm pull option for the current round and the arm pull operation is executed;
[0036] Receive user behavior feedback data after the arm pull operation, and update the reward estimate and arm pull count counter of the corresponding e-commerce platform behavior strategy arm;
[0037] An e-commerce platform user behavior excitation intensity input channel is configured for each e-commerce platform behavior strategy arm, where the e-commerce platform user behavior excitation intensity input channel is used to receive the excitation intensity value output by the e-commerce platform user behavior excitation intensity function;
[0038] Constructing a trigger gravity parameter input group, wherein the trigger gravity parameter input group is composed of the excitation intensity value received by the e-commerce platform user behavior excitation intensity input channel and the reward estimation value, confidence estimation term and exploration factor term of the e-commerce platform behavior strategy arm;
[0039] A gravity modulation scoring mechanism is introduced into the scoring fusion structure. The scoring fusion structure takes the trigger gravity parameter input group as input variables and outputs the modulated scoring value:
[0040] S i =α·U i +(1-α)·λ i ;
[0041] Among them, S i is the modulated score value of the behavior strategy arm of the i-th e-commerce platform, λ i is the excitation intensity value, α∈[0,1] is the fusion weight, U i is the traditional scoring value;
[0042] Based on the modulated scores of all e-commerce platform behavior strategy arms, replace the traditional scores, select the e-commerce platform behavior strategy arm with the largest score as the arm pull option for the current round and execute the arm pull operation;
[0043] After the arm pulling operation is completed, the reward estimation value and arm pulling count counter of the selected e-commerce platform behavior strategy arm are updated, and the parameters in the trigger gravity parameter input group are recalculated to complete the closed-loop update of the gravity modulation score of the current round.
[0044] Optionally, the S4 specifically includes:
[0045] S41. Construct a scoring fusion structure. The scoring fusion structure is used to perform weighted fusion calculation on the scoring results of the e-commerce platform behavior strategy arm. The scoring fusion structure includes a historical scoring channel, a trend scoring channel, and a weight control module.
[0046] S42. Set a historical scoring channel as a first input channel, wherein the first input channel is used to receive the estimated reward value of each e-commerce platform behavior strategy arm from the improved multi-armed bandit algorithm. The estimated reward value represents the historical return performance of each e-commerce platform behavior strategy arm and is used as the historical scoring input in the scoring fusion structure.
[0047] S43. Setting a trend scoring channel as a second input channel, the second input channel is used to receive the e-commerce platform user behavior stimulation intensity value of each e-commerce platform behavior strategy arm, the e-commerce platform user behavior stimulation intensity value is derived from the e-commerce platform user behavior stimulation intensity matrix, and is used as a trend scoring input in the scoring fusion structure;
[0048] S44. Setting a weight control module. The weight control module is used to set the weighted ratio of the historical score input and the trend score input in the fusion score calculation. The weight control module uses a linear combination model to set the fusion formula.
[0049] Optionally, the S5 specifically includes:
[0050] S51. At the beginning of each round of arm selection, the scoring fusion structure is called. The scoring fusion structure extracts the estimated reward value of each e-commerce platform behavior strategy arm from the historical scoring channel as the historical scoring input, and simultaneously extracts the e-commerce platform user behavior stimulation intensity value of each e-commerce platform behavior strategy arm from the trend scoring channel as the trend scoring input;
[0051] S52. The scoring fusion structure maps the historical scoring input and the trend scoring input to the same e-commerce platform behavior strategy arm, and performs weighted scoring calculation according to the weighting ratio set by the weight control module. The weighted scoring calculation is fused according to the linear combination model.
[0052] S53: The scoring fusion structure aggregates the fusion scoring values of all e-commerce platform behavior strategy arms and outputs a list of fusion scoring values of each e-commerce platform behavior strategy arm to the arm selection module;
[0053] S54. The arm selection module selects the e-commerce platform behavior strategy arm with the largest score from the fusion score list as the arm selection option for the current round, and performs the arm selection operation.
[0054] Optionally, the S6 specifically includes:
[0055] S61. In the current round arm selection option, receive e-commerce platform user behavior feedback data corresponding to the e-commerce platform behavior strategy arm selected in the current round, wherein the e-commerce platform user behavior feedback data includes the actual occurrence status of the behavior event corresponding to the e-commerce platform behavior strategy arm and is represented in the form of a label value;
[0056] S62. Extracting the e-commerce platform user behavior stimulation intensity value corresponding to the e-commerce platform behavior strategy arm selected in the current round from the e-commerce platform user behavior stimulation intensity matrix, and performing normalization processing on the e-commerce platform user behavior stimulation intensity value to normalize it to the numerical range of [0, 1] to represent the behavior stimulation probability of the behavior event corresponding to the e-commerce platform behavior strategy arm;
[0057] S63: jointly compare the behavior label value in the e-commerce platform user behavior feedback data with the normalized e-commerce platform user behavior stimulation intensity value, and use the difference to calculate the prediction deviation, where the prediction deviation is based on the granularity of the current round of arm pulling events;
[0058] S64: Input the prediction deviation to the feedback reconstruction module. The feedback reconstruction module constructs a correction term based on the prediction deviation and a preset feedback learning rate to update the reward estimate of the e-commerce platform behavior strategy arm selected in the current round.
[0059] S65. Write the updated reward estimate into the score fusion structure and overwrite the historical score input corresponding to the e-commerce platform behavior strategy arm selected in the current round.
[0060] The beneficial effects of the present invention are:
[0061] (1) The present invention introduces the Hawkes behavior excitation modeling process into the e-commerce user behavior analysis, characterizing the temporal dependency and dynamic triggering relationship between user behaviors, and improving the modeling ability of users' potential behavior trends; at the same time, the improved multi-arm bandit algorithm is combined to design the arm selection structure, introduce the excitation intensity input interface, the scoring fusion mechanism and the feedback correction logic, and realize the real-time response and adaptive optimization of the arm selection decision, effectively alleviating the cold start and feedback delay problems.
[0062] (2) The present invention improves the ability to perceive the user's current interests by constructing a scoring structure that integrates historical scores and trend scores, setting weight control parameters and performing weighted calculations; further combining normalized prediction deviation calculation and feedback reconstruction mechanism, the reward estimation is dynamically corrected, so that the model can maintain prediction stability and strategy sensitivity in scenarios with behavioral data fluctuations.
[0063] (3) The overall method of the present invention is based on non-deep modeling, focusing on the refined improvement of the single-step arm-pulling decision structure, taking into account lightweight deployment and interpretability. At the same time, it establishes a strategy adjustment system including multiple closed-loop modules such as excitation modeling, score fusion, and feedback update. It is suitable for various e-commerce behavior modeling needs and achieves a dual improvement in prediction accuracy and execution efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0065] Figure 1 This is a flowchart of the e-commerce platform user behavior analysis and prediction method based on data feedback proposed by the present invention. DETAILED DESCRIPTION
[0066] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.
[0067] refer to Figure 1 ,The e-commerce platform user behavior analysis and prediction method based on data feedback includes the following steps:
[0068] S1. Collect user behavior data from e-commerce platforms, including click behavior, browsing behavior, add-to-cart behavior, and order behavior, and construct a chronological sequence of user behavior events on the e-commerce platform.
[0069] S2. Based on the user behavior event sequence of the e-commerce platform, the Hawkes process is used to establish the user behavior excitation intensity function of the e-commerce platform, and the behavior excitation intensity matrix is output to describe the time-driven correlation relationship between different behavior events;
[0070] S3. Initialize the improved multi-arm bandit algorithm, establish multiple arm sets for arm pull options, configure the initial reward estimate, arm pull count, and arm pull selection parameters for each arm pull option, and set the excitation intensity input interface to receive the corresponding intensity value in the behavior excitation intensity matrix as the trigger attraction parameter;
[0071] S4. Construct a scoring fusion structure, set a first input channel to receive the reward estimate as the historical scoring input; set a second input channel to receive the behavior stimulation intensity value as the trend scoring input, and set a weight control parameter to adjust the weight ratio of the historical scoring input and the trend scoring input in the fusion score;
[0072] S5. In each round of arm selection, the scoring fusion structure is called to perform weighted calculation on the historical scoring input and the trend scoring input according to the weight control parameters to generate a fusion score value, and the arm selection option of the current round is selected based on the fusion score value;
[0073] S6. Receive behavioral feedback data generated after the arm-pulling option in the current round, jointly compare the behavioral feedback data with the behavioral stimulation intensity matrix, calculate the prediction deviation, and reconstruct the prediction deviation through feedback to correct the reward estimate of the corresponding arm-pulling option;
[0074] S7. Use the updated reward estimate and arm pull count to feed back into the improved multi-arm bandit algorithm to complete a behavioral data-driven self-update cycle. The improved multi-arm bandit algorithm structure is based on a non-deep model framework and combines behavioral causal modeling components to achieve lightweight strategy optimization based on single-step decision units.
[0075] In this embodiment, the reward estimate value and the updated number of arm pulls generated based on the current round of behavioral feedback are fed back to the improved multi-arm bandit algorithm in real time to overwrite and replace the historical parameters of the corresponding e-commerce platform behavior strategy arm. The feedback process writes the reward update module and the counting module through a unified parameter interface, and triggers the strategy scoring component to recalculate the trigger gravity parameter input group, thereby completing the adaptive closed-loop update driven by the user behavior data of the e-commerce platform. This mechanism retains the original arm pull record and historical reward accumulation logic, combined with the introduced behavioral causal modeling component, to achieve single-step strategy optimization without the need for complex deep network conditions, allowing the e-commerce platform system to dynamically respond to changes in user behavior, realize continuous learning and lightweight deployment of strategies, and significantly enhance the real-time performance and deployment versatility of the system.
[0076] In this embodiment, constructing the e-commerce platform user behavior event sequence in chronological order specifically includes: after obtaining the e-commerce platform user behavior data, arranging the behavior events in ascending order according to the timestamp corresponding to each behavior event to construct the e-commerce platform user behavior event sequence.
[0077] This embodiment constructs a behavioral event sequence that reflects the chronological relationship of user behaviors by arranging the collected e-commerce platform user behavior data in ascending order according to timestamps, effectively preserving the time dependency and behavioral evolution path in the behavioral data. This time series serves as the basic input for the subsequent Hawkes behavior excitation modeling process, ensuring that the excitation intensity function can accurately reflect the triggering effect of historical behaviors on future behaviors, thereby improving the temporal sensitivity and causal rationality of user behavior predictions. The data format of the time series structure can also be directly connected to the excitation intensity input interface in the improved multi-armed bandit algorithm, so that the strategy scoring structure can optimize the arm-pulling strategy in combination with the behavior occurrence trend, realize dynamic linkage from behavior perception to strategy feedback, and enhance the overall algorithm's ability to respond to changes in user preferences.
[0078] In this embodiment, S2 specifically includes:
[0079] S21. Arrange each e-commerce platform user behavior event in the e-commerce platform user behavior event sequence in ascending time order, extract the behavior type and timestamp information of each e-commerce platform user behavior event, and construct a behavior time index set. The behavior time index set is used to trigger path calculation;
[0080] S22. Setting a basic excitation rate parameter and a time decay parameter based on the Hawkes process, and defining a behavior excitation mapping function based on the user behavior type of the e-commerce platform, wherein the behavior excitation mapping function serves as a structural component of the excitation intensity calculation model;
[0081] S23. For each target e-commerce platform user behavior event in the behavior time index set, retrieve historical e-commerce platform user behavior events that occurred before the corresponding target e-commerce platform user behavior event, and calculate the excitation contribution value of each historical e-commerce platform user behavior event to the target e-commerce platform user behavior event based on the behavior excitation mapping function and the time decay parameter;
[0082] S24. Perform a weighted summation of the excitation contribution values of all historical e-commerce platform user behavior events, and combine this with a basic excitation rate parameter to generate an excitation intensity value for the target e-commerce platform user behavior event. The excitation intensity value constitutes a single-point output of the e-commerce platform user behavior excitation intensity function.
[0083] S25. Organize the excitation intensity values output by the e-commerce platform user behavior excitation intensity function for all e-commerce platform user behavior events into a behavior excitation intensity matrix. Map the behavior excitation intensity matrix to the excitation intensity input interface and input it into the scoring fusion structure as a trigger gravity parameter to guide the fusion score calculation of subsequent arm options.
[0084] This embodiment constructs a user behavior excitation modeling system based on the Hawkes process, arranges user behavior events in chronological order, extracts behavior types and timestamp information to construct a time index set, sets the basic excitation rate and time decay parameters, and defines an excitation mapping function associated with the behavior type. For each target behavior event, its predecessor behavior event is retrieved, the excitation contribution value of each historical event to the target event is calculated, and a weighted sum is performed to form an excitation intensity value, and then the behavior excitation intensity function is output. By aggregating the excitation intensity values of all events into a behavior excitation intensity matrix and inputting it into the scoring structure, the temporal correlation between behaviors can be quantified into a computable incentive signal, providing a time-sensitive drive for subsequent decision modules, and effectively improving the granularity and accuracy of user behavior trend modeling.
[0085] In this embodiment, S3 specifically includes:
[0086] S31. Construct an improved multi-armed bandit algorithm and initialize a pull-arm option set in the improved multi-armed bandit algorithm. The pull-arm option set includes multiple e-commerce platform behavior strategy arms, and each e-commerce platform behavior strategy arm is set with a unique pull-arm identifier;
[0087] S32. Configuring an initial reward estimate for each e-commerce platform behavior strategy arm in the improved multi-armed bandit algorithm. The initial reward estimate is set to a predefined constant and is used to initialize the reward of the e-commerce platform behavior strategy arm before effective feedback is available.
[0088] S33: Configure an arm pull count counter, the initial value of which is zero, and is used to record the cumulative number of times the e-commerce platform behavior strategy arm is selected;
[0089] S34. Configure arm selection parameters. The arm selection parameters include confidence estimation items and exploration factor items. They are set using a unified initialization strategy and are used for probability calculation in the subsequent scoring fusion structure.
[0090] S35. Setting an excitation intensity input interface in the improved multi-armed bandit algorithm. The excitation intensity input interface establishes an independent input channel for each e-commerce platform behavior strategy arm, for receiving the excitation intensity value corresponding to the corresponding e-commerce platform behavior strategy arm in the e-commerce platform user behavior excitation intensity matrix;
[0091] S36. Data-bind the e-commerce platform user behavior excitation intensity value received in the excitation intensity input interface with the arm selection parameter of the corresponding e-commerce platform behavior strategy arm to form a trigger gravity parameter input group of the e-commerce platform behavior strategy arm. The trigger gravity parameter input group serves as one of the input variables in the scoring fusion structure and is used to participate in the fusion score calculation of the e-commerce platform user behavior strategy arm.
[0092] This implementation initializes an improved multi-armed bandit algorithm to construct a set of pull options encompassing multiple e-commerce platform behavioral strategy arms. Each behavioral strategy arm is configured with an initial reward estimate, a pull count counter, and pull selection parameters to form the basis for decision-making. Furthermore, an excitation intensity input interface is introduced, establishing a dedicated channel for each strategy arm to receive behavioral excitation intensity data. This data is then data-bound to the corresponding strategy arm parameters to form a trigger gravity parameter input group, which serves as the input source for score fusion. This mechanism tightly integrates user behavior trends with historical feedback parameters, providing a more dynamic and personalized input basis for subsequent pull scoring decisions. This enables dynamic perception of behavioral responses and flexible regulation of strategy arm selection, enhancing the adaptability and accuracy of the e-commerce behavior prediction system.
[0093] In this embodiment, the improved multi-armed bandit algorithm specifically includes:
[0094] Initialize the arm pull option set, set the initial reward estimate for each e-commerce platform behavior strategy arm, and set the arm pull count counter;
[0095] In each round of arm selection, the traditional score of each e-commerce platform behavior strategy arm is calculated based on the reward estimate and arm pull count counter of each e-commerce platform behavior strategy arm;
[0096] Based on the traditional scoring values of all e-commerce platform behavior strategy arms, the e-commerce platform behavior strategy arm with the largest scoring value is selected as the arm pull option for the current round and the arm pull operation is executed;
[0097] Receive user behavior feedback data after the arm pull operation, and update the reward estimate and arm pull count counter of the corresponding e-commerce platform behavior strategy arm;
[0098] An e-commerce platform user behavior excitation intensity input channel is configured for each e-commerce platform behavior strategy arm, where the e-commerce platform user behavior excitation intensity input channel is used to receive the excitation intensity value output by the e-commerce platform user behavior excitation intensity function;
[0099] Constructing a trigger gravity parameter input group, wherein the trigger gravity parameter input group is composed of the excitation intensity value received by the e-commerce platform user behavior excitation intensity input channel and the reward estimation value, confidence estimation term and exploration factor term of the e-commerce platform behavior strategy arm;
[0100] A gravity modulation scoring mechanism is introduced into the scoring fusion structure. The scoring fusion structure takes the trigger gravity parameter input group as input variables and outputs the modulated scoring value:
[0101] S i =α·U i +(1-α)·λ i ;
[0102] Among them, S i is the modulated score value of the behavior strategy arm of the i-th e-commerce platform, λ i is the excitation intensity value, α∈[0,1] is the fusion weight, U i is the traditional scoring value;
[0103] This formula is used to modulate the final score of each behavioral strategy arm within the scoring fusion structure. In principle, by introducing a weighted combination of the excitation intensity value and historical scoring parameters, the scoring result reflects not only the reward level of the user's historical behavior, but also the degree of excitation of the current behavioral trend. This mechanism effectively integrates the causal and immediacy characteristics of user behavior, enabling scoring results to dynamically respond to changes in user preferences and enhancing the sensitivity of arm selection to the user's current interest state, thereby achieving more accurate and personalized strategy selection.
[0104] Based on the modulated scores of all e-commerce platform behavior strategy arms, replace the traditional scores, select the e-commerce platform behavior strategy arm with the largest score as the arm pull option for the current round and execute the arm pull operation;
[0105] After the arm pulling operation is completed, the reward estimation value and arm pulling count counter of the selected e-commerce platform behavior strategy arm are updated, and the parameters in the trigger gravity parameter input group are recalculated to complete the closed-loop update of the gravity modulation score of the current round.
[0106] Based on the traditional multi-armed bandit algorithm, this implementation introduces an e-commerce platform user behavior stimulation intensity input channel, so that each e-commerce platform behavior strategy arm not only receives the stimulation intensity value output based on the Hawkes behavior modeling based on the historical reward estimation value and the number of arm pulling times, but also constructs a trigger gravity parameter input group; in the scoring fusion structure, the gravity modulation scoring mechanism is used to dynamically generate the modulated scoring value, effectively integrating the user's current behavior trend and historical performance, and executing arm pulling selection and feedback update based on this, thereby forming an integrated self-update mechanism integrating stimulation response and decision-making regulation, thereby improving the real-time nature of user interest identification and the accuracy of behavior strategy matching.
[0107] In this embodiment, the S4 specifically includes:
[0108] S41. Construct a scoring fusion structure. The scoring fusion structure is used to perform weighted fusion calculation on the scoring results of the e-commerce platform behavior strategy arm. The scoring fusion structure includes a historical scoring channel, a trend scoring channel, and a weight control module.
[0109] S42. Set a historical scoring channel as a first input channel, wherein the first input channel is used to receive the estimated reward value of each e-commerce platform behavior strategy arm from the improved multi-armed bandit algorithm. The estimated reward value represents the historical return performance of each e-commerce platform behavior strategy arm and is used as the historical scoring input in the scoring fusion structure.
[0110] S43. Setting a trend scoring channel as a second input channel, the second input channel is used to receive the e-commerce platform user behavior stimulation intensity value of each e-commerce platform behavior strategy arm, the e-commerce platform user behavior stimulation intensity value is derived from the e-commerce platform user behavior stimulation intensity matrix, and is used as a trend scoring input in the scoring fusion structure;
[0111] S44. Setting a weight control module. The weight control module is used to set the weighted ratio of the historical score input and the trend score input in the fusion score calculation. The weight control module uses a linear combination model to set the fusion formula.
[0112] This implementation constructs a scoring fusion structure, using the historical scoring channel and the trend scoring channel to receive the reward estimate and behavioral incentive intensity values of the e-commerce platform's behavioral strategy arm, respectively. A weighted ratio is set by the weight control module, and a linear combination model is used for scoring fusion calculations, achieving a dynamic integration of historical behavioral performance and current incentive trends. This allows for comprehensive consideration of past user response data and current behavioral drivers during the arm selection process, effectively improving the timeliness and pertinence of strategy arm scoring. This enhances the accuracy and responsiveness of predictions based on changing user interests, improving the system's personalized recommendation effectiveness and decision-making sensitivity.
[0113] In this embodiment, the S5 specifically includes:
[0114] S51. At the beginning of each round of arm selection, the scoring fusion structure is called. The scoring fusion structure extracts the estimated reward value of each e-commerce platform behavior strategy arm from the historical scoring channel as the historical scoring input, and simultaneously extracts the e-commerce platform user behavior stimulation intensity value of each e-commerce platform behavior strategy arm from the trend scoring channel as the trend scoring input;
[0115] S52. The scoring fusion structure maps the historical scoring input and the trend scoring input to the same e-commerce platform behavior strategy arm, and performs weighted scoring calculation according to the weighting ratio set by the weight control module. The weighted scoring calculation is fused according to the linear combination model.
[0116] S53: The scoring fusion structure aggregates the fusion scoring values of all e-commerce platform behavior strategy arms and outputs a list of fusion scoring values of each e-commerce platform behavior strategy arm to the arm selection module;
[0117] S54. The arm selection module selects the e-commerce platform behavior strategy arm with the largest score from the fusion score list as the arm selection option for the current round, and performs the arm selection operation.
[0118] In each round of arm selection decisions, this implementation uses a scoring fusion structure to extract the historical reward estimates and user behavior incentive intensity values of the e-commerce platform's behavioral strategy arm. These are then weighted and fused based on the ratios set by the weight control module to form a fused scoring result. This structure ensures that each strategy arm considers both the user's past behavioral performance and current behavioral trends. The fused scores are then uniformly output to the arm selection module, ultimately selecting the highest-scoring behavioral strategy arm for the current round. This approach implements a collaborative decision-making mechanism that combines historical rewards and behavioral trends, effectively enhancing the e-commerce platform's ability to respond in real time to changes in user interests, and improving the sensitivity of behavior prediction and the accuracy of strategy selection.
[0119] In this embodiment, S6 specifically includes:
[0120] S61. In the current round arm selection option, receive e-commerce platform user behavior feedback data corresponding to the e-commerce platform behavior strategy arm selected in the current round, wherein the e-commerce platform user behavior feedback data includes the actual occurrence status of the behavior event corresponding to the e-commerce platform behavior strategy arm and is represented in the form of a label value;
[0121] S62. Extracting the e-commerce platform user behavior stimulation intensity value corresponding to the e-commerce platform behavior strategy arm selected in the current round from the e-commerce platform user behavior stimulation intensity matrix, and performing normalization processing on the e-commerce platform user behavior stimulation intensity value to normalize it to the numerical range of [0, 1] to represent the behavior stimulation probability of the behavior event corresponding to the e-commerce platform behavior strategy arm;
[0122] S63: jointly compare the behavior label value in the e-commerce platform user behavior feedback data with the normalized e-commerce platform user behavior stimulation intensity value, and use the difference to calculate the prediction deviation, where the prediction deviation is based on the granularity of the current round of arm pulling events;
[0123] S64: Input the prediction deviation to the feedback reconstruction module. The feedback reconstruction module constructs a correction term based on the prediction deviation and a preset feedback learning rate to update the reward estimate of the e-commerce platform behavior strategy arm selected in the current round.
[0124] S65. Write the updated reward estimate into the score fusion structure and overwrite the historical score input corresponding to the e-commerce platform behavior strategy arm selected in the current round.
[0125] This implementation constructs a granular prediction bias by comparing the normalized difference between actual user behavior feedback data and the corresponding behavioral stimulation intensity value after each round of arm pull. This bias is then fed into the feedback reconstruction module to dynamically correct the reward estimate, thereby forming an adaptive adjustment process for the e-commerce platform's behavioral strategy arm scoring mechanism. This method not only quantitatively analyzes the difference between users' short-term actual behavior and predicted trends, but also improves the model's stability and sensitivity through an update strategy controlled by feedback learning rates. This significantly enhances the e-commerce platform's decision-making optimization capabilities under conditions of incomplete feedback and behavioral delays, providing a foundation for refined personalized recommendations and strategy adjustments.
[0126] Example 1:
[0127] To verify the feasibility of this invention, we applied it to the personalized behavior recommendation system of a large, comprehensive e-commerce platform. This platform boasts over 20 million daily active users, and the large number of clicks, browsing, add-to-cart, and ordering behaviors generated by these users provides a rich data foundation for accurate analysis and prediction. In the original system, the platform primarily employed the classic multi-armed bandit algorithm to iteratively update the recommendation strategy. While this method offers a certain degree of real-time responsiveness, it still exhibits significant limitations in modeling complex behavioral causal relationships and accurately predicting user behavior.
[0128] During the implementation of the present invention, the platform back-end system first collects all user behavior data and constructs a sequence of user behavior events. To this end, the Hawkes process is introduced to model the behavior excitation intensity, and the time dependency between different behavior events is explicitly expressed through the excitation path function, and a user behavior excitation intensity matrix is generated. The platform initializes the improved multi-arm bandit algorithm, sets multiple behavior strategy arms and corresponding reward estimates and arm selection parameters, and sets an excitation intensity input channel for each strategy arm to receive the behavior excitation data generated by the Hawkes process.
[0129] The system's scoring module constructs a fused scoring structure consisting of two channels: historical reward estimates and behavioral incentive intensity. It also incorporates a weighting mechanism, using a linear fusion model to comprehensively consider historical performance and current behavioral trends. During each behavioral decision cycle, the system invokes the scoring fusion structure to perform modulated scoring calculations, dynamically select the optimal behavioral strategy arm, and execute the arm pull action. The behavioral feedback data is then compared with the behavioral incentive prediction value using a normalized difference, accurately capturing prediction deviations. The reward estimate for the current strategy arm is then corrected through a feedback reconstruction mechanism, continuously improving prediction accuracy and system adaptability.
[0130] To verify the performance of this method in real business scenarios, the platform selected real operational data over a period of time and established an experimental control group. The control group used the traditional multi-armed bandit algorithm to process behavioral strategy selection, while the experimental group fully implemented the improved mechanism proposed in this invention. The experiment used multiple indicators such as coverage, prediction accuracy, feedback response speed, strategy switching frequency, and fusion stability for comparison:
[0131] Table 1 Experimental data comparing the performance of the method of the present invention and the traditional algorithm
[0132]
[0133]
[0134] From the tabular data, we can see that the prediction accuracy increased from 62% to 83%, the feedback response time was shortened by more than 50%, the prediction deviation of the excitation intensity dropped to one-third of the original, and the user behavior coverage rate increased to 92%. At the same time, the weight stability in the fusion scoring process was significantly enhanced, and the strategy arm switching was more efficient and the rhythm was reasonable, highlighting the robustness and sensitivity of this method in dealing with behavior prediction problems.
[0135] In summary, the method proposed in the present invention not only outperforms existing technologies in terms of computing efficiency and real-time feedback capabilities, but also significantly improves the e-commerce platform's dynamic adaptability to changes in user behavior and the accuracy of personalized strategy output, and has good industrial application prospects and promotion value.
[0136] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A method for analyzing and predicting user behavior on an e-commerce platform based on data feedback, characterized in that: The steps include: S1. Collect user behavior data from e-commerce platforms, including click behavior, browsing behavior, add-to-cart behavior, and order behavior, and construct a chronological sequence of user behavior events on the e-commerce platform. S2. Based on the user behavior event sequence of the e-commerce platform, the Hawkes process is used to establish the user behavior excitation intensity function of the e-commerce platform and output the behavior excitation intensity matrix; S3. Initialize the improved multi-arm bandit algorithm, establish multiple arm sets for arm pull options, configure the initial reward estimate, arm pull times, and arm pull selection parameters for each arm pull option, and set the excitation intensity input interface; S4. Build a scoring fusion structure, set the first input channel to receive the reward estimate as the historical score input; set the second input channel to receive the behavior stimulation intensity value as the trend score input, and set the weight control parameters; S5. In each round of arm selection, the scoring fusion structure is called to perform weighted calculation on the historical scoring input and the trend scoring input according to the weight control parameters to generate a fusion score value, and the arm selection option of the current round is selected based on the fusion score value; S6. Receive behavioral feedback data generated after the arm-pulling option in the current round, jointly compare the behavioral feedback data with the behavioral stimulation intensity matrix, calculate the prediction deviation, and reconstruct the prediction deviation through feedback to correct the reward estimate of the corresponding arm-pulling option; S7. Use the updated reward estimate and arm pull count to feed back to the improved multi-arm bandit algorithm, completing a round of behavioral data-driven self-update cycle.
2. The e-commerce platform user behavior analysis and prediction method based on data feedback according to claim 1 is characterized in that: Constructing the e-commerce platform user behavior event sequence in chronological order specifically includes: after obtaining the e-commerce platform user behavior data, arranging the behavior events in ascending order according to the timestamp corresponding to each behavior event, and constructing the e-commerce platform user behavior event sequence.
3. The e-commerce platform user behavior analysis and prediction method based on data feedback according to claim 1 is characterized in that: The S2 specifically includes: S21. Arrange each e-commerce platform user behavior event in the e-commerce platform user behavior event sequence in ascending time order, extract the behavior type and timestamp information of each e-commerce platform user behavior event, and construct a behavior time index set; S22. Setting a basic excitation rate parameter and a time decay parameter based on the Hawkes process, and defining a behavior excitation mapping function based on the user behavior type of the e-commerce platform; S23. For each target e-commerce platform user behavior event in the behavior time index set, retrieve historical e-commerce platform user behavior events that occurred before the corresponding target e-commerce platform user behavior event, and calculate the excitation contribution value of each historical e-commerce platform user behavior event to the target e-commerce platform user behavior event based on the behavior excitation mapping function and the time decay parameter; S24. Perform a weighted summation of the excitation contribution values of all historical e-commerce platform user behavior events, and combine this with a basic excitation rate parameter to generate an excitation intensity value for the target e-commerce platform user behavior event. S25. Organize the excitation intensity values output by the e-commerce platform user behavior excitation intensity function for all e-commerce platform user behavior events into a behavior excitation intensity matrix, and map the behavior excitation intensity matrix to the excitation intensity input interface.
4. The e-commerce platform user behavior analysis and prediction method based on data feedback according to claim 1 is characterized in that: The S3 specifically includes: S31. Construct an improved multi-armed bandit algorithm and initialize a pull-arm option set in the improved multi-armed bandit algorithm. The pull-arm option set includes multiple e-commerce platform behavior strategy arms, and each e-commerce platform behavior strategy arm is set with a unique pull-arm identifier; S32. Configuring an initial reward estimate for each e-commerce platform behavior strategy arm in the improved multi-armed bandit algorithm. The initial reward estimate is set to a predefined constant and is used to initialize the reward of the e-commerce platform behavior strategy arm before effective feedback is available. S33: Configure an arm pull count counter, the initial value of which is zero, and is used to record the cumulative number of times the e-commerce platform behavior strategy arm is selected; S34. Configuring arm selection parameters. The arm selection parameters include a confidence estimation item and an exploration factor item, and are set using a unified initialization strategy. S35. Setting an excitation intensity input interface in the improved multi-armed bandit algorithm. The excitation intensity input interface establishes an independent input channel for each e-commerce platform behavior strategy arm, for receiving the excitation intensity value corresponding to the corresponding e-commerce platform behavior strategy arm in the e-commerce platform user behavior excitation intensity matrix; S36. Bind the e-commerce platform user behavior excitation intensity value received in the excitation intensity input interface with the arm selection parameter of the corresponding e-commerce platform behavior strategy arm to form a trigger gravity parameter input group of the e-commerce platform behavior strategy arm.
5. The e-commerce platform user behavior analysis and prediction method based on data feedback according to claim 4 is characterized in that: The improved multi-armed bandit algorithm specifically includes: Initialize the arm pull option set, set the initial reward estimate for each e-commerce platform behavior strategy arm, and set the arm pull count counter; In each round of arm selection, the traditional score of each e-commerce platform behavior strategy arm is calculated based on the reward estimate and arm pull count counter of each e-commerce platform behavior strategy arm; Based on the traditional scoring values of all e-commerce platform behavior strategy arms, the e-commerce platform behavior strategy arm with the largest scoring value is selected as the arm pull option for the current round and the arm pull operation is executed; Receive user behavior feedback data after the arm pull operation, and update the reward estimate and arm pull count counter of the corresponding e-commerce platform behavior strategy arm; An e-commerce platform user behavior excitation intensity input channel is configured for each e-commerce platform behavior strategy arm, where the e-commerce platform user behavior excitation intensity input channel is used to receive the excitation intensity value output by the e-commerce platform user behavior excitation intensity function; Constructing a trigger gravity parameter input group, wherein the trigger gravity parameter input group is composed of the excitation intensity value received by the e-commerce platform user behavior excitation intensity input channel and the reward estimation value, confidence estimation term and exploration factor term of the e-commerce platform behavior strategy arm; A gravity modulation scoring mechanism is introduced into the scoring fusion structure. The scoring fusion structure takes the trigger gravity parameter input group as input variables and outputs the modulated scoring value: S i =α·U i +(1-a)·l i ; Among them, S i is the modulated score value of the behavior strategy arm of the i-th e-commerce platform, λ i is the excitation intensity value, α∈[0,1] is the fusion weight, U i is the traditional scoring value; Based on the modulated scores of all e-commerce platform behavior strategy arms, replace the traditional scores, select the e-commerce platform behavior strategy arm with the largest score as the arm pull option for the current round and execute the arm pull operation; After the arm pulling operation is completed, the reward estimation value and arm pulling count counter of the selected e-commerce platform behavior strategy arm are updated, and the parameters in the trigger gravity parameter input group are recalculated to complete the closed-loop update of the gravity modulation score of the current round.
6. The e-commerce platform user behavior analysis and prediction method based on data feedback according to claim 1 is characterized in that: The S4 specifically includes: S41. Construct a scoring fusion structure. The scoring fusion structure is used to perform weighted fusion calculation on the scoring results of the e-commerce platform behavior strategy arm. The scoring fusion structure includes a historical scoring channel, a trend scoring channel, and a weight control module. S42. Set a historical scoring channel as a first input channel, wherein the first input channel is used to receive the estimated reward value of each e-commerce platform behavior strategy arm from the improved multi-armed bandit algorithm. The estimated reward value represents the historical return performance of each e-commerce platform behavior strategy arm and is used as the historical scoring input in the scoring fusion structure. S43. Setting a trend scoring channel as a second input channel, the second input channel is used to receive the e-commerce platform user behavior stimulation intensity value of each e-commerce platform behavior strategy arm, the e-commerce platform user behavior stimulation intensity value is derived from the e-commerce platform user behavior stimulation intensity matrix, and is used as a trend scoring input in the scoring fusion structure; S44. Setting a weight control module. The weight control module is used to set the weighted ratio of the historical score input and the trend score input in the fusion score calculation. The weight control module uses a linear combination model to set the fusion formula.
7. The e-commerce platform user behavior analysis and prediction method based on data feedback according to claim 1 is characterized in that: The S5 specifically includes: S51. At the beginning of each round of arm selection, the scoring fusion structure is called. The scoring fusion structure extracts the estimated reward value of each e-commerce platform behavior strategy arm from the historical scoring channel as the historical scoring input, and simultaneously extracts the e-commerce platform user behavior stimulation intensity value of each e-commerce platform behavior strategy arm from the trend scoring channel as the trend scoring input; S52. The scoring fusion structure maps the historical scoring input and the trend scoring input to the same e-commerce platform behavior strategy arm, and performs weighted scoring calculation according to the weighting ratio set by the weight control module. The weighted scoring calculation is fused according to the linear combination model. S53: The scoring fusion structure aggregates the fusion scoring values of all e-commerce platform behavior strategy arms and outputs a list of fusion scoring values of each e-commerce platform behavior strategy arm to the arm selection module; S54. The arm selection module selects the e-commerce platform behavior strategy arm with the largest score from the fusion score list as the arm selection option for the current round, and performs the arm selection operation.
8. The e-commerce platform user behavior analysis and prediction method based on data feedback according to claim 1 is characterized in that: The S6 specifically includes: S61. In the current round arm selection option, receive e-commerce platform user behavior feedback data corresponding to the e-commerce platform behavior strategy arm selected in the current round, wherein the e-commerce platform user behavior feedback data includes the actual occurrence status of the behavior event corresponding to the e-commerce platform behavior strategy arm and is represented in the form of a label value; S62. Extracting the e-commerce platform user behavior stimulation intensity value corresponding to the e-commerce platform behavior strategy arm selected in the current round from the e-commerce platform user behavior stimulation intensity matrix, and performing normalization processing on the e-commerce platform user behavior stimulation intensity value to normalize it to the numerical range of [0, 1] to represent the behavior stimulation probability of the behavior event corresponding to the e-commerce platform behavior strategy arm; S63: jointly compare the behavior label value in the e-commerce platform user behavior feedback data with the normalized e-commerce platform user behavior stimulation intensity value, and use the difference to calculate the prediction deviation, where the prediction deviation is based on the granularity of the current round of arm pulling events; S64: Input the prediction deviation to the feedback reconstruction module. The feedback reconstruction module constructs a correction term based on the prediction deviation and a preset feedback learning rate to update the reward estimate of the e-commerce platform behavior strategy arm selected in the current round. S65. Write the updated reward estimate into the score fusion structure and overwrite the historical score input corresponding to the e-commerce platform behavior strategy arm selected in the current round.