E-commerce promotion method and system based on cloud computing
By constructing a hidden Markov behavior sequence and dynamic time warping algorithm, combined with grey correlation analysis and elastic resource scheduling, the advertising weight and resource allocation are dynamically adjusted, which solves the problem of insufficient user behavior tracking in existing technologies, realizes real-time response and efficient matching of advertising delivery, and improves the accuracy and efficiency of e-commerce promotion.
Patent Information
- Application Number
- CN202510514909.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-04-23
AI Technical Summary
In existing technologies, cloud computing-based e-commerce promotion methods lack the ability to continuously track users' real-time behavior, resulting in delayed strategy updates in interest drift scenarios, misalignment between advertising delivery and demand, insufficient dynamic adaptability of resource allocation and task queues, and inability to respond to high-frequency behavior fluctuations in a timely manner. In addition, the fragmented analysis of positive and negative behavior data makes it difficult to explore the dynamic coupling relationship in the interaction process, reducing the accuracy of advertising reach and increasing operating costs.
By collecting user click stream data, constructing a hidden Markov behavior sequence, using the dynamic time warping algorithm to perform window alignment on the user behavior sequence, extracting the mutation point sequence, and combining the piecewise integral operation of positive and negative behavior data, using the grey correlation analysis method to dynamically reset the advertising weight, and combining the elastic resource scheduling algorithm to optimize advertising resource allocation and realize dynamic adjustment of advertising content.
It improves the temporal correlation of user behavior prediction, enhances the sensitivity of abnormal behavior identification, quantifies the correlation characteristics of implicit preferences and explicit feedback in the interaction process, reduces the risk of misjudgment of single-dimensional indicators, matches advertising delivery strategies with users' real-time needs, alleviates the load imbalance problem in high-concurrency scenarios, and improves delivery efficiency and conversion rate.
Smart Images

Figure CN120387857B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of e-commerce, and in particular to an e-commerce promotion method and system based on cloud computing. Background Art
[0002] The field of e-commerce encompasses the information management and service methods used to implement the entire process of displaying, selling, paying, and logistics goods and services through online platforms. Core aspects of this technology area include the construction and operation of electronic trading platforms, the integration and management of online payment systems, the collection and analysis of consumer data, the automated coordination of supply chains, and the implementation of online marketing strategies. At the application level, e-commerce encompasses a variety of business models, including B2B, B2C, and C2C. Its overall development relies on the security of information network infrastructure, transaction systems, and the continuous optimization of front-end interactive experiences. On this basis, e-commerce is deeply integrated with emerging information technologies such as cloud computing, big data, and artificial intelligence, forming a cross-regional, multi-terminal, and scalable business operation system.
[0003] Among them, the e-commerce promotion method based on cloud computing refers to a technical solution that uses distributed computing resources and network platforms to achieve widespread dissemination and targeted marketing of e-commerce content. The subject of this patent is mainly aimed at the marketing delivery link in e-commerce activities, covering technical matters such as product information collection, user behavior analysis, advertising delivery strategy generation, and multi-channel content distribution. Specifically, it is based on the cloud computing platform, through the collection and aggregation of user access behavior logs, combined with product attribute tags to establish user portraits, and then executes the matching and distribution process of advertising content and target users according to classification logic, while supporting cross-platform synchronous display and status feedback recovery. The above process, supported by the elastic resource scheduling and high concurrent processing capabilities of the cloud platform, completes the data-driven and intelligent matching of the entire e-commerce promotion process.
[0004] Existing technologies rely on static user profiles and preset rules to perform ad matching, lacking the ability to continuously track real-time behavioral sequences. This leads to delayed policy updates in scenarios where interest drifts. For example, the tagging system for user behavior logs is generated based on fixed time windows, making it difficult to capture sudden changes in interest within short periods of time, resulting in misalignment between ad delivery and current demand. Existing ad weighting mechanisms often employ periodic batch updates, which are unable to promptly respond to high-frequency behavioral fluctuations and can easily result in ineffective impressions when user preferences shift rapidly. Traditional resource scheduling strategies scale linearly based on preset capacity, lacking the dynamic adaptability of resource allocation and task queues to sudden traffic spikes and troughs, potentially causing local node overload or idleness. Furthermore, the separate analysis of positive and negative behavioral data tends to overlook the dynamic coupling between the two during interaction. For example, page skipping behavior may have a temporal correlation with dwell time, making it difficult for existing technologies to mine such complex features, leading to biased judgments of user intent. These issues reduce the accuracy of ad reach and increase operating costs. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose an e-commerce promotion method and system based on cloud computing.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a cloud computing-based e-commerce promotion method, comprising the following steps:
[0007] S2: Input the user behavior sequence into the dynamic time warping algorithm, perform sliding window alignment on the Euclidean distance and time interval difference of adjacent jump paths, extract the fluctuation peak points that exceed the threshold three times in a row, and generate a mutation point sequence;
[0008] S3: Call the positive and negative behavior data of the product interaction log, perform piecewise integration calculations on the number of collections, the length of time spent on the detail page, the frequency of page skips, and the number of quick swipes based on the mutation point sequence, and generate positive and negative dynamic parameters;
[0009] S4: Inputting the positive dynamic parameter and the negative dynamic parameter into a grey correlation analysis method, calculating the correlation difference between the two within three consecutive time windows, and when the difference exceeds a preset threshold, performing a priority reset operation on the advertising content weight set to generate a reconstructed weight set;
[0010] S5: Based on the reconstructed weight set, the cloud server elastic resource scheduling algorithm is called to allocate computing resources and storage resources of the advertisement delivery node, map them to the advertisement display queue, and generate advertisement resource distribution instructions.
[0011] As a further solution of the present invention, the mutation point sequence includes the fluctuation peak timestamp, the fluctuation amplitude value, and the interval between adjacent peak points. The positive dynamic parameters are specifically the collection count integral value and the details page stay time integral value. The negative dynamic parameters are specifically the page skip frequency integral value and the quick sliding count integral value. The reconstruction weight set includes the reset weight value, the priority label, and the advertising identifier. The advertising resource distribution instruction specifically refers to the computing node allocation strategy, the storage node allocation strategy, and the advertising display order list.
[0012] As a further embodiment of the present invention, the step of obtaining the mutation point sequence is specifically as follows:
[0013] S201: Acquire adjacent jump paths in the user behavior sequence, calculate spatial difference values of adjacent path coordinate points, extract timestamp differences of corresponding paths, combine Euclidean distances and time interval differences into difference pairs in path order, and generate a path difference pair sequence;
[0014] S202: Based on the path difference pair sequence, a fixed length of a sliding window is set, and an alignment operation is performed on the sum of the discrete degree of the Euclidean distance difference and the absolute value of the time interval difference within the window, and a comprehensive measurement value of the fluctuation intensity within the window is calculated using the formula:
[0015] ;
[0016] The calculation obtains the window fluctuation intensity, compares the measurement value with the preset alignment threshold, and selects the windows exceeding the threshold and marks them as candidate fluctuation peak points;
[0017] in, Representative The fluctuation intensity of the window, Represents the variance of the Euclidean distance between adjacent paths within the window, Represents the sum of the absolute values of the time interval differences between adjacent paths within the window, is the cumulative value of the number of fluctuations in the current window, Fixed length for sliding window;
[0018] S203: calling the candidate fluctuation peak points, counting the number of consecutive appearances of each peak point in the corresponding window, setting a consecutive number threshold, eliminating peak points that do not meet the continuous exceeding threshold condition, and generating a mutation point sequence.
[0019] As a further solution of the present invention, the steps of obtaining the positive dynamic parameters and the negative dynamic parameters are specifically as follows:
[0020] S301: Retrieving the positive and negative behavior data from the product interaction log, based on the timestamps of the mutation point sequence, divide the number of favorites into intervals based on adjacent mutation points, truncate the dwell time on the detail page into discrete segments based on the mutation points, and classify the page skip frequency and swipe count by time window to generate a segmented data set;
[0021] S302: Perform segmented integral calculations on the number of collections, the length of stay on the details page, the page skip frequency, and the number of slides for each segment in the segmented data set, using the formula:
[0022] ;
[0023] The integral value is obtained by operation, and the integral result of the positive behavior data and the integral result of the negative behavior data are calculated to generate an integral vector;
[0024] in, Representative Dynamic integral value of class behavior, Representative The number of collections of the segment, Representative The length of time you stay on the details page for each segment, Representative The time interval of the segmented mutation points, Representative The page skip frequency of the segment, Representative The number of slides in the segment, Representative Segmented user activity coefficient, A balance factor representing interaction density and behavior type, Represents the total number of segments, Representative The start timestamp of the segment;
[0025] S303: Separate the positive integral and the negative integral according to the positive and negative signs of the integral vector, normalize the positive integral using the maximum and minimum method, normalize the negative integral using the standard deviation method, and allocate a time weight factor based on the proportion of segment duration to generate positive dynamic parameters and negative dynamic parameters.
[0026] As a further solution of the present invention, the step of obtaining the reconstruction weight set is specifically as follows:
[0027] S401: Calling the positive dynamic parameter and the negative dynamic parameter, based on the parameter sequence of three consecutive time windows, aligning the positive and negative parameters of each window point by point, and calculating the correlation coefficient. At the same time, taking the mean value within the window as the correlation degree, integrating the three window data, and generating a window correlation degree set;
[0028] S402: extracting adjacent window correlations from the window correlation set, calculating the absolute difference between the subsequent window and the preceding window, obtaining the difference between the first and second windows, and the difference between the second and third windows, comparing the two and taking the maximum value as the current difference, and comparing the current difference with a preset difference threshold. If the difference exceeds the threshold, generating a difference status flag;
[0029] S403: extracting the original priority of the advertising content weight set according to the difference status mark, reallocating the weight values according to a preset rule and arranging them in descending order, overwriting the original set, and generating a reconstructed weight set.
[0030] As a further solution of the present invention, the steps of obtaining the advertising resource distribution instruction are specifically as follows:
[0031] S501: Call the reconstruction weight set to extract the computing resource demand value and storage resource demand value of the advertisement delivery node, and combine the node load value and delay coefficient in the cloud server elastic resource scheduling algorithm to use the formula:
[0032] ;
[0033] Obtain node resource allocation parameter values through calculation and generate a resource allocation parameter set;
[0034] in, Representative Node The resource allocation parameters, To reconstruct the nodes in the weight set Dynamic reconstruction weights, For nodes The computing resource requirements of For nodes The storage resource requirement value, For nodes The current load value, For nodes The delay coefficient, represents the reconstruction weight association term, represents a delayed associated term;
[0035] S502: Based on the resource allocation parameter set, compare the node resource allocation parameter value with the priority threshold of the advertisement display queue, eliminate parameter values below the priority threshold, intercept the parameter value sequence within the queue length upper limit according to the time window constraint, and generate a node resource adjustment queue;
[0036] S503: calling the node resource adjustment queue, binding the computing resource demand value and the storage resource demand value corresponding to each parameter value in the queue one by one according to the time sequence number of the advertisement display queue, and generating an advertisement resource distribution instruction.
[0037] As a further embodiment of the present invention, the method further comprises:
[0038] S1: Obtain user clickstream data through cloud log collection nodes, perform cleaning and alignment operations on page jump paths, dwell time, and product category labels, and construct user behavior sequences based on the Hidden Markov Model;
[0039] The user behavior sequence specifically includes jump timestamp, topic code, and page level.
[0040] As a further solution of the present invention, the steps of obtaining the user behavior sequence are specifically as follows:
[0041] S101: Collecting raw user clickstream data from cloud logs, extracting page jump paths, dwell time, and product category label fields, filtering out missing path nodes and abnormal dwell time values, matching product category labels with page paths, and generating standardized clickstream data;
[0042] S102: Based on the standardized clickstream data, align the timestamps of page jump paths and product category labels, unify the time accuracy of differentiated terminals, accumulate the dwell time by segment according to the path nodes, extract the continuous path, duration, and label combination under the same user ID, and generate a behavior trajectory parameter set;
[0043] S103: Call the path jump frequency, mean dwell time, and label association in the behavior trajectory parameter set, calculate the state transition probability matrix and dwell time distribution density, map the path nodes to the state variables of the hidden Markov model, establish the corresponding relationship between the state transition and the observation sequence, and generate a user behavior sequence model.
[0044] A cloud computing-based e-commerce promotion system, wherein the cloud computing-based e-commerce promotion system is used to execute the cloud computing-based e-commerce promotion method, and the system comprises:
[0045] The behavior sequence construction module is used to obtain user clickstream data through the cloud log collection node, perform data cleaning and timestamp alignment operations on page jump paths, dwell time, and product category labels, input the cleaned data into the hidden Markov model to generate user behavior sequences, and pass the user behavior sequences to the mutation point extraction module;
[0046] A mutation point extraction module is used to call the user behavior sequence, perform a sliding window alignment operation on the Euclidean distance and time interval difference of adjacent jump paths based on the dynamic time warping algorithm, extract the fluctuation peak points that exceed the preset threshold three times in a row, generate a mutation point sequence, and pass the mutation point sequence to the dynamic parameter generation module;
[0047] A dynamic parameter generation module is used to obtain positive and negative behavior data from the product interaction log, perform piecewise integration operations on the number of collections, the length of time spent on the detail page, the frequency of page skips, and the number of quick swipes based on the mutation point sequence, generate positive and negative dynamic parameters, and pass the positive and negative dynamic parameters to the weight reconstruction module;
[0048] a weight reconstruction module, configured to call the positive dynamic parameter and the negative dynamic parameter, calculate the difference in correlation between the two within three consecutive time windows based on a grey correlation analysis method, perform a priority reset operation on the advertising content weight set when the difference exceeds a dynamic threshold, generate a reconstructed weight set, and transmit the reconstructed weight set to the resource scheduling mapping module;
[0049] The resource scheduling mapping module is used to call the cloud server elastic resource scheduling algorithm based on the reconstruction weight set, dynamically adjust the computing resource and storage resource allocation ratio of the advertising delivery node, map the adjusted resource allocation plan to the advertising display queue, generate advertising resource distribution instructions and output them to the cloud execution node.
[0050] Compared with the prior art, the advantages and positive effects of the present invention are:
[0051] In the present invention, by collecting user click stream data and constructing a hidden Markov behavior sequence, continuous dynamic modeling of user browsing trajectories is achieved, thereby improving the temporal correlation of behavior prediction. A dynamic time warping algorithm is used to perform windowed alignment on the peak points of path fluctuations, capture the critical nodes of sudden changes in user interests, and enhance the sensitivity of abnormal behavior identification. Combined with the piecewise integral operation of positive and negative behavior parameters, the correlation characteristics of implicit preferences and explicit feedback in the interaction process are quantified, reducing the risk of misjudgment caused by single-dimensional indicators. The gray correlation difference is used to dynamically reset the advertising weight, establish a matching mechanism between the delivery strategy and the real-time needs of users, and avoid the response delay caused by fixed weight allocation. Based on the elastic resource scheduling algorithm, the reconstructed weights are mapped to distributed advertising nodes to achieve accurate adaptation of computing resources and display needs, and alleviate the load imbalance problem in high-concurrency scenarios. The above steps form a closed-loop optimization link, so that the advertising content can be dynamically adjusted according to the fluctuations of user behavior, improving delivery efficiency and conversion rate while reducing redundant resource consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic diagram of the workflow of the present invention;
[0053] Figure 2 Flowchart of the steps for obtaining the user behavior sequence of the present invention;
[0054] Figure 3 Flow chart of the steps for obtaining the mutation point sequence of the present invention;
[0055] Figure 4 Flowchart of the steps for obtaining positive dynamic parameters and negative dynamic parameters of the present invention;
[0056] Figure 5 A flow chart of the steps for obtaining the reconstructed weight set of the present invention;
[0057] Figure 6 This is a flow chart of the steps for obtaining the advertising resource distribution instruction of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0059] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.
[0060] Example 1
[0061] See also Figure 1 The present invention provides a technical solution: an e-commerce promotion method based on cloud computing, comprising the following steps:
[0062] S1: Obtain user clickstream data through cloud log collection nodes, perform cleaning and alignment operations on page jump paths, dwell time, and product category labels, and construct user behavior sequences based on the Hidden Markov Model;
[0063] S2: Input the user behavior sequence into the dynamic time warping algorithm, perform sliding window alignment on the Euclidean distance and time interval difference of adjacent jump paths, extract the fluctuation peak points that exceed the threshold three times in a row, and generate a mutation point sequence;
[0064] S3: Call the positive and negative behavior data from the product interaction log and perform piecewise integration operations on the number of collections, length of stay on the detail page, page skip frequency, and number of quick swipes based on the mutation point sequence to generate positive and negative dynamic parameters.
[0065] S4: Inputting the positive dynamic parameter and the negative dynamic parameter into the grey correlation analysis method, calculating the correlation difference between the two within three consecutive time windows, and when the difference exceeds a preset threshold, performing a priority reset operation on the advertising content weight set to generate a reconstructed weight set;
[0066] S5: Based on the reconstructed weight set, the cloud server elastic resource scheduling algorithm is called to allocate computing resources and storage resources of the advertising delivery node, map them to the advertising display queue, and generate advertising resource distribution instructions.
[0067] The user behavior sequence specifically includes the jump timestamp, topic code, and page level value. The mutation point sequence includes the fluctuation peak timestamp, fluctuation amplitude value, and adjacent peak interval. The positive dynamic parameters specifically include the number of collection points and the length of stay on the details page. The negative dynamic parameters specifically include the page skip frequency points and the number of quick sliding points. The reconstruction weight set includes resetting the weight value, priority label, and advertising identifier. The advertising resource distribution instructions specifically refer to the computing node allocation strategy, storage node allocation strategy, and advertising display order list.
[0068] See also Figure 2 , the steps for obtaining the user behavior sequence are as follows:
[0069] S101: Collecting raw user clickstream data from cloud logs, extracting page jump paths, dwell time, and product category label fields, filtering out missing path nodes and abnormal dwell time values, matching product category labels with page paths, and generating standardized clickstream data;
[0070] Collect raw user clickstream data from cloud logs. This involves retrieving log files within a specified time range, such as access logs from the past 24 hours, from the server log system that stores user activity records, such as the Object Storage Service (OSS) or database (such as ClickHouse) deployed on Alibaba Cloud or Tencent Cloud. These logs contain timestamps (such as '2025-04-10 10:30:15.123'), user IDs (such as 'User_A1B2'), URLs of visited pages (such as ' / product / detail / item123'), and request types (such as 'GET'). , user agent (such as browser information), IP address and other original fields, then extract key information from these original fields, identify the URL field representing the page jump path, for example, extract the access sequence such as ' / home'->' / category / electronics'->' / product / list / tv'->' / product / detail / tv_brand_X_model_Y', and extract the start and end timestamps corresponding to each page access, and calculate the page stay time, for example, the user visits ' / product / list / tv' The start timestamp of the page 'v' is '10:31:05.200', and the timestamp of jumping to ' / product / detail / tv_brand_X_model_Y' is '10:31:45.700', so the dwell time is calculated to be 40.5 seconds. Furthermore, the product category tag field that may exist in the log is extracted. This field may be recorded directly in the log (such as 'category=TV'), or it needs to be parsed and mapped through the accessed URL path. For example, through the preset URL-category mapping rule, ' / product / detail / tv_brand_X_model_Y' is converted to ' / product / detail / tv_brand_X_model_Y'. nd_X_model_Y' is mapped to the 'TV' category label. The extracted data is then cleaned to filter out records with incomplete path node information, such as URLs missing key page identifiers and abnormal dwell time values. A reasonable dwell time range is set, such as filtering out records with a dwell time of less than 1 second or greater than 1800 seconds (30 minutes), as these are considered invalid clicks or idle behavior. For example, a record showing a dwell time of 3000 seconds on a product details page is filtered out. Finally, the product category labels corresponding to valid path nodes are matched and associated based on the timestamp and user ID, for example, confirming '10:31:45.The product category for user 'User_A1B2' who accesses ' / product / detail / tv_brand_X_model_Y' at '700' is 'TV'. The processed page jump path, dwell time, and product category label are combined into a structured record to form standardized clickstream data. For example, a record set with the format {UserID: 'User_A1B2', Timestamp: '10:31:45.700', PagePath: ' / product / detail / tv_brand_X_model_Y', DwellTime: 40.5, Category: 'TV'} is generated.
[0071] S102: Based on the standardized clickstream data, the timestamps of the page jump paths and the product category labels are aligned to unify the time accuracy of the differentiated terminals. The dwell time is accumulated by segment by path node, and the continuous path, duration, and label combination under the same user ID is extracted to generate a set of behavior trajectory parameters.
[0072] First, we address the differences in log timestamp accuracy that may exist between different sources or terminal devices. For example, we align the millisecond timestamps recorded by the mobile app (such as '10:31:45.715') and the second timestamps recorded by the web (such as '10:31:46') to second-level accuracy, and use a unified rounding or rounding rule. For example, we round down to the nearest second, so that '10:31:45.715' becomes '10:31:45' and '10:31:46' remains unchanged, ensuring that the time base for subsequent processing is consistent. Next, we align the consecutive access records of the same user on the same page path node and accumulate the duration of these records. For example, if user 'User_A1B2' has two consecutive records on the ' / product / list / tv' page with durations of 15.2 seconds and 25.3 seconds respectively, we merge these two records and record the total duration as Seconds, the associated timestamp uses the timestamp of the last record, and the product category labels are also merged or confirmed to be consistent accordingly. Subsequently, all processed records of the user are concatenated in chronological order according to the user ID (such as 'User_A1B2'), and the continuous page jump path sequence, the merged stay time sequence, and the corresponding product category label sequence are extracted. For example, the behavior trajectory of user 'User_A1B2' is formed: [(' / home', 5.0, 'None'), (' / category / electronics', 10.1, 'Electronic products'), (' / product / list / tv', 40.5, 'TV'), (' / product / detail / tv_brand_X_model_Y', 65.8, 'TV'), (' / cart', 30.2, 'None')], where each tuple represents (path node, accumulated stay time, product category label). Such trajectory parameters of all users are combined to generate a behavior trajectory parameter set.
[0073] S103: Call the path jump frequency, mean dwell time, and label association in the behavior trajectory parameter set, calculate the state transition probability matrix and dwell time distribution density, map the path nodes to the state variables of the hidden Markov model, establish the corresponding relationship between the state transition and the observation sequence, and generate a user behavior sequence model.
[0074] First, calculate the state transition probability, traverse the adjacent path jump instances of all users in the behavior trajectory parameter set, and count the number of jumps from the path node Jump to path node Frequency , and from the path node The total frequency of jumps to all other nodes , then the state transition probability Calculated as For example, statistics show that the frequency of jumping from ' / product / list / tv' to ' / product / detail / tv_brand_X_model_Y' is 500 times, and the total frequency of jumping from ' / product / list / tv' to all other pages is 1000 times. The transition probability is , calculate the transition probabilities between all path nodes to form a state transition probability matrix. At the same time, calculate the distribution density of the length of stay of each path node (state), call the length of stay data of each path node in the behavior trajectory parameter set, for example, for the path ' / product / list / tv', the collected length of stay data is {40.5, 35.2, 55.1, ...}, and use methods such as kernel density estimation (KDE) to fit the probability density function of these length of stay data , which describes the state The length of stay is The probability density of the product category label is analyzed, and the correlation between the product category label and the path node is calculated to determine the specific label At the path node Frequency or conditional probability of occurrence For example, calculate the probability that the 'TV' tag appears in the ' / product / list / tv' path. Then, map each unique page path node (such as ' / home', ' / product / list / tv') to a hidden state in the HMM and calculate the distribution density of the dwell time. Correlation with tags As a state The observation probability (emission probability) is part of or based on the observation sequence. The observation sequence can be the actual observed (stay time, product category label) pair, such as (40.5s, 'TV'). The corresponding relationship between the state transition probability matrix and the observation sequence (stay time distribution, label association) is established to clarify the state Transfer to state The probability of Decision, Status Generate observations (such as length of stay ,Label ) is given by and etc. to jointly decide and finally build the user behavior sequence model.
[0075] See also Figure 3 , the specific steps for obtaining the mutation point sequence are:
[0076] S201: Acquire adjacent jump paths in the user behavior sequence, calculate the spatial difference values of the adjacent path coordinate points, extract the timestamp differences of the corresponding paths, combine the Euclidean distances and time interval differences into difference pairs in the order of the paths, and generate a path difference pair sequence;
[0077] Using the path sequence in the behavior trajectory parameter set, extract two temporally adjacent page jump paths. For example, if there is ' / product / list / tv'->' / product / detail / tv_brand_X_model_Y'->' / cart' in the sequence, then the adjacent jump path pairs are (' / product / list / tv',' / product / detail / tv_brand_X_model_Y') and (' / product / detail / tv_brand_X_model_Y',' / cart'). In order to calculate the spatial difference value of the path coordinate points, each page path needs to be mapped to a multidimensional space coordinate in advance. This mapping can be achieved by vectorizing the page content (such as using TF-IDF or BERT models to process page text elements to obtain vectors) or using a graph embedding algorithm based on page link relationships (such as Node2Vec). Assume that the mapping coordinates of ' / product / list / tv' are , the mapping coordinates of ' / product / detail / tv_brand_X_model_Y' are , then the spatial difference between them, that is, the Euclidean distance, is calculated as ;
[0078] At the same time, extract the timestamp difference between the two adjacent paths, that is, the access time of the latter path node minus the access time of the previous path node, which represents the length of stay at the first path node or the jump interval. For example, if the access timestamp (the time of entering the page) of ' / product / list / tv' is '10:31:05' and the access timestamp of ' / product / detail / tv_brand_X_model_Y' is '10:31:45', then the timestamp difference is seconds, the calculated Euclidean distance and time interval difference seconds, combined into a difference pair according to the order of the original path jump ,Repeat this process for all adjacent paths in the user behavior sequence to generate a sequence of path difference pairs, e.g. ,in It is The Euclidean distance of the path of jumps, It is The time interval between jumps.
[0079] S202: Based on the path difference sequence, a fixed length of the sliding window is set, and an alignment operation is performed on the sum of the discrete degree of the Euclidean distance difference and the absolute value of the time interval difference within the window, and a comprehensive measurement value of the fluctuation intensity within the window is calculated using the formula:
[0080] ;
[0081] The calculation obtains the window fluctuation intensity, compares the measurement value with the preset alignment threshold, and selects the windows that exceed the threshold and marks them as candidate fluctuation peak points;
[0082] in, Representative The fluctuation intensity of the window, Represents the variance of the Euclidean distance between adjacent paths within the window, Represents the sum of the absolute values of the time interval differences between adjacent paths within the window, is the cumulative value of the number of fluctuations in the current window, Fixed length for sliding window;
[0083] Set a sliding window of fixed length, and record the window length as , for example, setting , which means that 10 consecutive path difference pairs are analyzed each time, and the window slides across the entire sequence in turn. A window containing to difference pairs, i.e. , calculate the Euclidean distance difference within the window The degree of discreteness is calculated. The variance of the Euclidean distance , the calculation formula is ,in is the average of the Euclidean distances within the window, e.g. The values are {0.7, 0.8, 0.6, 0.9, 0.7, 0.5, 0.8, 0.7, 0.9, 0.6}, calculate their mean ,variance ;
[0084] At the same time, the time interval difference within the window Calculate the sum of the absolute values, denoted as , note that here The formula description refers to the sum of absolute values, not the "sum of absolute values of time interval differences" in the original text. We will execute according to the formula description. Assuming that the window The value is {40,35,50,30,45,60,25,38,32,48} (unit: seconds), then seconds, get the cumulative value of fluctuations in the current window , which is the length of the window , unless otherwise specified (e.g. only Here we set , then, call the fluctuation intensity calculation formula: Here The absolute value symbol is added to the formula, but It is the sum of absolute values and is always non-negative, so the absolute value sign can be omitted. The formula parameters are explained in detail: Representative The fluctuation intensity of a window. A larger value indicates a more drastic change in user behavior within the window. Represents the variance of the Euclidean distance between adjacent paths within the window, which measures the stability of the path space change. The larger the value, the more unstable it is. It represents the sum of the absolute values of the time interval differences between adjacent paths within the window, and measures the total time span or activity within the window; The cumulative value of the number of fluctuations in the current window (here equal to the window length ), through its square root Scale the intensity of fluctuations; The sliding window is fixed in length and is normalized as the denominator. The operation logic is to convert the instability of spatial changes (variance ) and time span ( ) and by the number of data points in the window ( ) is enlarged and then the window length ( ) is normalized to obtain an indicator that comprehensively measures the severity of the behavior fluctuation within the window, and the example values are used for calculation: ;
[0085] The formula is beneficial in that it combines the spatial variation of the path coordinates ( ) and the total amount of time intervals ( ), and considering the number of data points within the window ( ) and the window size ( ), which can more comprehensively capture the comprehensive fluctuations of user behavior within a short time window and identify potential time points when behavior patterns change significantly. Next, set a fluctuation intensity alignment threshold The setting of the threshold can refer to the historical data The distribution of values, such as taking historical The 90th percentile of the value, or set an empirical value based on business needs. Suppose that by analyzing historical data, set , the calculated window fluctuation intensity With threshold For comparison, , so the If a window exceeds the threshold, the window is marked as a candidate fluctuation peak point, and the calculation and comparison process is repeated for all windows to filter out all windows that exceed the threshold and mark them as a candidate fluctuation peak point set.
[0086] This result shows that the The user behavior fluctuation intensity within the window is high, reaching a level that requires attention, which provides a candidate basis for subsequent identification of mutation points. The calculated value It is a candidate signal.
[0087] S203: Call candidate fluctuation peak points, count the number of consecutive appearances of each peak point in the corresponding window, set a consecutive number threshold, eliminate peak points that do not meet the continuous exceeding threshold condition, and generate a mutation point sequence.
[0088] For example, the window numbers marked as candidate peak points are {5, 6, 7, 12, 18, 19, 20, 21, 25, ...}. Check whether these candidate peak points appear continuously in the sequence. Count the number of consecutive appearances of each peak point corresponding window. For example, windows 5, 6, and 7 appear continuously, the number of consecutive appearances is 3. Window 12 appears alone, the number of consecutive appearances is 1. Windows 18, 19, 20, and 21 appear continuously, the number of consecutive appearances is 4. Window 25 appears alone, the number of consecutive appearances is 1. Set a consecutive number threshold. , the threshold is used to filter out accidental, short-duration fluctuations to ensure that the identified mutation points have a certain stability. The threshold can be set based on experience. For example, it is required that at least three consecutive windows appear to be considered a valid mutation. , the number of consecutive occurrences of each candidate fluctuation peak point and For comparison, for windows 5, 6, and 7, the number of consecutive times is 3. , retain these peak points; for window 12, the number of consecutive times 1 does not meet , remove the peak point; for windows 18, 19, 20, 21, the number of consecutive times 4 satisfies , retain these peak points; for window 25, the number of consecutive times 1 does not meet , remove the peak point, and each window (or the time point it represents, such as the center time point or end time point of the window) in the retained continuous peak point sequence (such as windows 5-7, windows 18-21) is confirmed as a behavior mutation point. The timestamps or serial numbers of all confirmed mutation points are collected to generate the final mutation point sequence, such as {Timestamp_peak5, Timestamp_peak6, Timestamp_peak7, Timestamp_peak18, Timestamp_peak19, Timestamp_peak20, Timestamp_peak21,…}.
[0089] See also Figure 4 The specific steps for obtaining positive dynamic parameters and negative dynamic parameters are as follows:
[0090] S301: Retrieving the positive and negative behavior data from the product interaction log, based on the timestamps of the mutation point sequence, divides the number of favorites into intervals based on adjacent mutation points, truncates the dwell time on the detail page into discrete segments based on the mutation points, and categorizes the page skip frequency and swipe count by time window to generate a segmented dataset.
[0091] Product interaction logs are used to record the specific interactions between users and products. It is necessary to distinguish between positive and negative behavior data. Positive behavior refers to behaviors that indicate user interest or intention, such as favorite, add to cart, purchase, and long dwell time on detail page. Negative behavior refers to behaviors that indicate user disinterest or avoidance, such as skip, fast scroll, and quick return from detail page. Based on the timestamps contained in the mutation point sequence {Timestamp_mut1, Timestamp_mut2,…}, these interaction logs are segmented by time. The time range of the first segment is from the beginning of the sequence to the first mutation point. , the second segment is from arrive , and so on, The time range of each segment is (Assume is the sequence start time), for each segment , statistics of various interactive behavior data within the time interval: calculate the number of collections , for example in During the interval, the user collected the product twice; calculate the length of time the user stayed on the details page , accumulate the duration of all visits to the detail page within the interval, or truncate the duration across the mutation point according to the mutation point timestamp and only count the duration that falls within the interval, for example, a detail page visit from Start to End, then it belongs to segment The duration is 10 seconds; the page skip frequency is counted For example, if the user skipped 5 recommended products in this interval, the number of sliding times is counted. For example, if the user performs 30 sliding operations in the interval, all the segmented statistics { }Data and segmented timestamp information are organized into segmented data sets.
[0092] Table 1 Segmented interaction data example table
[0093]
[0094] As shown in Table 1, the table lists the user interaction behavior statistics in three consecutive time segments divided according to the mutation points.
[0095] S302: Perform segmented integral calculations on the number of collections, length of stay on the details page, page skip frequency, and number of slides for each segment in the segmented data set, using the formula:
[0096] ;
[0097] The integral value is obtained by operation, and the integral result of the positive behavior data and the integral result of the negative behavior data are calculated to generate an integral vector;
[0098] in, Representative Dynamic integral value of class behavior, Representative The number of collections of the segment, Representative The length of time you stay on the details page for each segment, Representative The time interval of the segmented mutation points, Representative The page skip frequency of the segment, Representative The number of slides in the segment, Representative Segmented user activity coefficient, A balance factor representing interaction density and behavior type, Represents the total number of segments, Representative The start timestamp of the segment;
[0099] Perform a piecewise integration operation on the interaction data of each segment in the segmented data set to obtain the dynamic integral value of each behavior type (positive / negative). The formula used is: According to the previous analysis, a more reasonable explanation of this formula is: ;
[0100] Detailed description of formula parameters: Representative The cumulative dynamic integral value of the class behavior (for example, s can be "overall" or distinguish between "positive" and "negative"); is the index of the segment, from 1 to the total number of segments ; It is Number of favorites within a segment; It is Total dwell time on detail pages within a segment (in seconds); It is The length of the segment (in seconds), i.e. ; It is Frequency of page skips within a segment; It is Number of slides within a segment; It is The user activity coefficient of each segment is used to normalize the impact of negative behaviors. Its calculation method can be set as the total number of interactions in the segment ( ) and segment duration The ratio of , and normalize it, for example, to the interval [0.1,1], to avoid the denominator being zero or too small. Assume that the segment 1 After calculation, it is normalized to 0.8, and the segment 2 0.5, segment 3 is 0.9; It is a balance factor between interaction density and behavior type, which is used to adjust the weight of negative behaviors (skip, slide) relative to positive behaviors (collect, stay). Its setting can be determined according to the experimental results. For example, , is designed to balance the impact of the two behaviors on the final integral; the summation symbol Indicates that all The final total score is obtained by adding up the values in brackets calculated for each segment. The first term of the formula The contribution of positive behavior is quantified by multiplying the product of the number of collections and the length of stay (representing positive engagement) by the square root of the segment length (giving longer segments higher weight, but the effect is weakened by the square root), and the absolute value is taken to ensure it is non-negative; the second term of the formula By subtracting the number of slides from the number of skips ( may be negative, indicating sliding dominance) to quantify net negative behavior, divided by the activity coefficient Normalize and multiply by the balance factor Adjust its influence weight; add the two together and then sum across segments to obtain an integral value that comprehensively reflects the user's dynamic behavioral tendencies during the entire observation period.
[0101] The benefit of the formula is that it integrates a variety of positive and negative user interaction behaviors and uses segmented calculation and time weighting ( ), as well as activity normalization and balance factor adjustment, can dynamically and quantitatively evaluate the comprehensive participation and interest changes of users in each stage divided by behavioral mutation points. Use the data in Table 1 for example calculation ( , ): Calculate each segment first and hypothetical :
[0102] Segment 1: , ;
[0103] Segment 2: , ;
[0104] Segment 3: , ;
[0105] Calculate the contribution of each segment:
[0106] Segment 1 Contribution ;
[0107] Segment 2 Contribution ;
[0108] Segment 3 Contribution ;
[0109] Total points ;
[0110] The integral value obtained by calculation is This integral value combines the positive and negative behaviors of the user in the three segments. Next, it is necessary to distinguish whether the calculation is based on the positive behavior log or the negative behavior log, or as in this example, the positive and negative indicators are combined to calculate the total integral, and then separated or the difference is calculated. Assuming that a comprehensive integral is calculated here, further processing is required to obtain positive and negative parameters. If the calculation is based on pure positive indicators (such as only ) and purely negative indicators (such as only ) Apply similar logic to calculate and , and then perform difference calculation to generate the integral vector, for example, we get the vector Here This is an example of a comprehensive score. This result shows that the user's overall interaction during this period showed a strong positive trend (because the score is much greater than 0). It will serve as the basis for normalization and weighting in the next step.
[0111] S303: Separate the positive integral and the negative integral according to the positive and negative signs of the integral vector, normalize the positive integral using the maximum and minimum method, normalize the negative integral using the standard deviation method, and allocate a time weight factor based on the proportion of the segment duration to generate positive dynamic parameters and negative dynamic parameters.
[0112] Separate the positive and negative integrals according to the sign (or source) of the integral value. For example, we assume that we can obtain and (These two values add up to approximately the total integral calculated previously ), for the forward integral The Min-Max method is used for normalization, which requires determining the maximum value of the positive integral in a group of users or a time period. and minimum value , assuming and , then the normalized forward integral
[0113] ;
[0114] Negative integral Normalization is done using the standard deviation (Z-score) method, which requires calculating the mean of negative scores for a group of users or within a time period. and standard deviation , assuming and , then the normalized negative integral ;
[0115] Next, the time weight factor is assigned based on the proportion of each segment duration to the total duration, and the total duration is calculated. Second;
[0116] Calculate the duration of each segment ,For example , , ;
[0117] This time weight factor can be used to weight the final normalized parameter, or as the original text suggests, it may be used to generate an overall time weight factor, for example, based on the duration of the segment where positive behavior mainly occurs or the segment where negative behavior mainly occurs. and , assuming that the normalized result is simply used as the final parameter, the final forward dynamic parameter is generated With negative dynamic parameters .
[0118] See also Figure 5 , the specific steps for obtaining the reconstruction weight set are:
[0119] S401: Call the positive dynamic parameters and negative dynamic parameters, align the positive and negative parameters of each window point by point based on the parameter sequence of three consecutive time windows, calculate the correlation coefficient, and take the mean value within the window as the correlation degree. The three window data are integrated to generate a window correlation degree set;
[0120] For example, window , , The parameters are:
[0121] window : ;
[0122] window : ;
[0123] window : (Using the results of the previous step), align the positive and negative parameters in each window point by point and calculate the correlation coefficient between them. The "correlation coefficient" here can be defined as a certain relationship measure between the two, such as the difference , or ratio , or a combination of functions , and "taking the mean value in the window as the correlation" may mean taking the mean value if there are multiple data points (such as multiple users) in the window, but in this case there is only one value in each window. Yes, so the "correlation coefficient" is the "correlation degree" of the window. We use the difference as the correlation degree calculation method: , calculate the correlation between the three windows: , , , integrate the correlation of these three windows to generate a window correlation set .
[0124] S402: Extracting the correlation of adjacent windows from the window correlation set, calculating the absolute difference between the subsequent window and the previous window, obtaining the difference between the first and second windows, and the difference between the second and third windows, comparing the two and taking the maximum value as the current difference, and comparing the current difference with a preset difference threshold. If the difference exceeds the threshold, a difference status flag is generated;
[0125] From the window correlation set Extract the correlation value of the adjacent window, calculate the absolute difference between the correlation of the next window and the previous window, and calculate the absolute difference between the correlation of the first and second windows. ;
[0126] Calculate the absolute difference in correlation between the second and third windows ;
[0127] Compare these two differences and , take the maximum value as the current difference , set a preset difference threshold , the threshold is used to judge whether the change in correlation is significant enough. Its setting can be based on the statistical distribution of historical correlation changes, such as taking the 85th percentile of the distribution, or setting an empirical value. Assume that , the calculated current difference With preset threshold To compare the values, , so the current difference exceeds the threshold, generating a difference status mark, marked as "significant change" (or 1).
[0128] S403: extracting the original priority of the advertising content weight set according to the difference status mark, reallocating the weight values according to a preset rule and arranging them in descending order, overwriting the original set, and generating a reconstructed weight set.
[0129] According to the difference status marking result, check whether the mark is "significant change" (1). In this case, the mark is 1, so the weight reconstruction operation needs to be performed. First, extract the original priority (i.e., the original weight value) of the advertising content weight set. Assume that the original weight set is {Ad A: 0.5, Ad B: 0.3, Ad C: 0.2}. The weight value represents the delivery priority. The larger the value, the higher the priority. Then, reallocate the weight value according to the preset rules. The rules should clearly specify how to adjust the weight when a significant change is detected. For example, the rule can be: "If a significant change is detected, the time period with the largest correlation change (here is arrive changes, The weight of the better-performing ad (it is necessary to associate the ad with the dynamic parameter performance, assuming that Ad A's performance is more strongly correlated with the positive parameter during this period) is increased by 20%, and the weight of the poorly performing ad (assuming Ad C) is reduced by 20%, and then renormalized. Specific implementation: Increase the weight of Ad A , reduce the weight of ad C , the weight of ad B remains unchanged at 0.3, and the new temporary weights are {A: 0.6, B: 0.3, C: 0.16}, and the total is , normalized: new weight A , new weight B , new weight C , arrange the redistributed and normalized weights in descending order (already in descending order), and obtain {Ad A: 0.566, Ad B: 0.283, Ad C: 0.151}. Use this new weight set to overwrite the original weight set. If the difference status mark in S402 is 0 (does not exceed the threshold), reconstruction is not performed, and the original weights remain unchanged. Finally, the reconstructed weight set {Ad A: 0.566, Ad B: 0.283, Ad C: 0.151} is generated.
[0130] See also Figure 6 , the steps for obtaining advertising resource distribution instructions are as follows:
[0131] S501: Call the reconstruction weight set to extract the computing resource demand value and storage resource demand value of the advertising delivery node, and combine the node load value and delay coefficient in the cloud server elastic resource scheduling algorithm to use the formula:
[0132] ;
[0133] Obtain node resource allocation parameter values through calculation and generate a resource allocation parameter set;
[0134] in, Representative Node The resource allocation parameters, To reconstruct the nodes in the weight set Dynamic reconstruction weights, For nodes The computing resource requirement value of For nodes The storage resource requirement value, For nodes The current load value, For nodes The delay coefficient, represents the reconstruction weight association term, represents a delayed associated term;
[0135] Call the reconstruction weight set {Ad A: 0.566, Ad B: 0.283, Ad C: 0.151}, where Representative Advertising (Related items )’s dynamic reconstruction weights, such as , extract each ad delivery node (Assuming each ad corresponds to a delivery node or service) The required computing resource requirements and storage resource requirements , these values are attributes of the ad itself, for example: Ad A needs Unit computing resources, MB storage resources; Ad B needs Unit computing resources, MB storage resources; Ad C requires Unit computing resources, MB storage resources, combined with the real-time status data of each node provided by the cloud server elastic resource scheduling algorithm, obtain the node Current load value (a normalized value between 0 and 1, indicating the comprehensive utilization of CPU, memory, etc.) and network delay coefficient (For example, the normalized latency indicator, the associated items Refers to delay), assuming the current status of each node is: Node A load ,Delay Node B load ,Delay Node C load ,Delay , using the resource allocation parameter calculation formula: Detailed description of formula parameters: Representative Node The resource allocation parameter value of the resource allocation parameter, the higher the value, the higher the allocation priority; is a node (Associated Ads )’s reconstruction weight; is a node The amount of computing resources required; is a node The amount of storage resources required (Note: Directly adding computing units and storage units may require normalization or weighting first. To simplify the calculation, we assume that they are comparable or converted to equivalent units of calculation); is a node The current load; is a node The delay factor (associated delay ); denominator Combines load and delay to form a comprehensive node status penalty term. The higher the load or delay, the larger the denominator, and the distribution parameter The lower the value, the lower the value. The operation logic is: according to the dynamic importance of the advertisement ( ) and its total resource requirements ( ) calculates the basic score, and then adjusts (penalizes) it based on the node's current load and latency. Nodes with low load and latency receive higher allocation parameters. The formula is beneficial in that it not only considers the dynamic priority and resource requirements of the ad content itself, but also combines the running status (load and latency) of the delivery node in real time, making resource allocation decisions more intelligent and efficient, prioritizing the allocation of important and resource-intensive ads to nodes in good condition. Calculate for each ad node (assuming have been converted to equivalent units of calculation, e.g. ):
[0136] Node A: ;
[0137] Node B: ;
[0138] Node C: ;
[0139] The resource allocation parameter values of each node are calculated and the resource allocation parameter set {node A: 4.478, node B: 2.404, node C: 0.396} is generated. The result shows that after considering the reconstruction weight, resource demand and node status, node A has the highest resource allocation priority. These parameter values Will be used to decide which nodes get resource adjustments.
[0140] S502: Based on the resource allocation parameter set, compare the node resource allocation parameter value with the priority threshold of the advertisement display queue, eliminate parameter values below the priority threshold, intercept the parameter value sequence within the queue length upper limit according to the time window constraint, and generate a node resource adjustment queue;
[0141] Based on the resource allocation parameter set {Node A: 4.478, Node B: 2.404, Node C: 0.396}, set a priority threshold for the ad display queue Only nodes whose resource allocation parameter values are higher than this threshold will be considered for resource adjustment. The threshold is set based on the minimum service quality or resource utilization that the system wants to ensure. For example, according to the system capacity and the expected response speed, , for each node in the set Value and Perform numerical comparison: Node A ( ), Node B( ), node C( ), remove the parameter values below the priority threshold, that is, remove the parameter value of node C, and get the candidate queue {node A: 4.478, node B: 2.404}. Next, intercept the parameter value sequence within the queue length upper limit according to the time window constraint. Assuming that the upper limit of the number of resource adjustment operations allowed to be executed in the current time window is Since the candidate queue length is 2, which is less than the upper limit of 5, all of them are retained. If the candidate queue length exceeds , you need to press Sort the values from high to low and only take the top , press Arrange the candidate queue in descending order (if necessary): {Node A: 4.478, Node B: 2.404}, and generate the final node resource adjustment queue [Node A, Node B].
[0142] S503: Call the node resource adjustment queue, bind the computing resource demand value and storage resource demand value corresponding to each parameter value in the queue one by one according to the time sequence number of the advertisement display queue, and generate an advertisement resource distribution instruction.
[0143] Call the node resource adjustment queue [node A, node B], and for each node in the queue, extract its corresponding computing resource demand value and storage resource requirements (Information from S501), for example, node A corresponds to Unit computing resources and MB storage resources, node B corresponds to Unit computing resources and MB storage resources are allocated. These resource demand values and node identifiers are bound item by item according to the sequence number of the ad display queue (or the order of the adjustment queue itself, such as [node A, node B] in this example) to form specific resource scheduling instructions. For example, the generated instruction set is: {Instruction 1: {Target node: 'node A', computing resource adjustment amount: 3, storage resource adjustment amount: 200MB, priority / timing: 1}, Instruction 2: {Target node: 'node B', computing resource adjustment amount: 2, storage resource adjustment amount: 150MB, priority / timing: 2}}. This instruction set is the final advertising resource distribution instruction, which can be executed by the underlying cloud resource management system.
[0144] A cloud computing-based e-commerce promotion system, which is used to execute the above-mentioned cloud computing-based e-commerce promotion method, includes:
[0145] The behavior sequence construction module is used to obtain user clickstream data through cloud log collection nodes, perform data cleaning and timestamp alignment operations on page jump paths, dwell time, and product category labels, input the cleaned data into the hidden Markov model to generate user behavior sequences, and pass the user behavior sequences to the mutation point extraction module;
[0146] The mutation point extraction module is used to call the user behavior sequence, perform a sliding window alignment operation on the Euclidean distance and time interval difference of adjacent jump paths based on the dynamic time warping algorithm, extract the fluctuation peak points that exceed the preset threshold three times in a row, generate a mutation point sequence, and pass the mutation point sequence to the dynamic parameter generation module;
[0147] The dynamic parameter generation module is used to obtain positive and negative behavior data from product interaction logs. Based on the mutation point sequence, it performs piecewise integration operations on the number of collections, the length of time spent on the detail page, the frequency of page skips, and the number of quick swipes to generate positive and negative dynamic parameters. These parameters are then passed to the weight reconstruction module.
[0148] The weight reconstruction module is used to call positive dynamic parameters and negative dynamic parameters, calculate the difference in correlation between the two within three consecutive time windows based on the grey correlation analysis method, perform a priority reset operation on the advertising content weight set when the difference exceeds the dynamic threshold, generate a reconstructed weight set, and pass the reconstructed weight set to the resource scheduling mapping module;
[0149] The resource scheduling mapping module is used to call the cloud server elastic resource scheduling algorithm based on the reconstruction weight set, dynamically adjust the computing resource and storage resource allocation ratio of the advertising delivery node, map the adjusted resource allocation plan to the advertising display queue, generate advertising resource distribution instructions and output them to the cloud execution node.
[0150] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. An e-commerce promotion method based on cloud computing, characterized in that: The following steps are involved: S2: Input the user behavior sequence into the dynamic time warping algorithm, perform sliding window alignment on the Euclidean distance and time interval difference of adjacent jump paths, extract the fluctuation peak points that exceed the threshold three times in a row, and generate a mutation point sequence; The steps for obtaining the mutation point sequence are specifically as follows: S201: Acquire adjacent jump paths in the user behavior sequence, calculate spatial difference values of adjacent path coordinate points, extract timestamp differences of corresponding paths, combine Euclidean distances and time interval differences into difference pairs in path order, and generate a path difference pair sequence; S202: Based on the path difference pair sequence, a fixed length of a sliding window is set, and an alignment operation is performed on the sum of the discrete degree of the Euclidean distance difference and the absolute value of the time interval difference within the window, and a comprehensive measurement value of the fluctuation intensity within the window is calculated using the formula: ; The calculation obtains the window fluctuation intensity, compares the measurement value with the preset alignment threshold, and selects the windows that exceed the threshold and marks them as candidate fluctuation peak points; in, Representative The fluctuation intensity of the window, Represents the variance of the Euclidean distance between adjacent paths within the window, Represents the sum of the absolute values of the time interval differences between adjacent paths within the window, is the cumulative value of the number of fluctuations in the current window, Fixed length for sliding window; S203: Call the candidate fluctuation peak points, count the number of consecutive appearances of each peak point in the corresponding window, set a consecutive number threshold, remove peak points that do not meet the continuous exceeding threshold condition, and generate a mutation point sequence; S3: Call the positive and negative behavior data of the product interaction log, perform piecewise integration calculations on the number of collections, the length of time spent on the detail page, the frequency of page skips, and the number of quick swipes based on the mutation point sequence, and generate positive and negative dynamic parameters; S4: Inputting the positive dynamic parameter and the negative dynamic parameter into a grey correlation analysis method, calculating the correlation difference between the two within three consecutive time windows, and when the difference exceeds a preset threshold, performing a priority reset operation on the advertising content weight set to generate a reconstructed weight set; S5: Based on the reconstructed weight set, the cloud server elastic resource scheduling algorithm is called to allocate computing resources and storage resources of the advertisement delivery node, map them to the advertisement display queue, and generate advertisement resource distribution instructions.
2. The e-commerce promotion method based on cloud computing according to claim 1, characterized in that: The mutation point sequence includes the fluctuation peak timestamp, fluctuation amplitude value, and adjacent peak interval. The positive dynamic parameters are specifically the collection count integral value and the details page stay time integral value. The negative dynamic parameters are specifically the page skip frequency integral value and the quick sliding count integral value. The reconstruction weight set includes the reset weight value, priority label, and advertising identifier. The advertising resource distribution instruction specifically refers to the computing node allocation strategy, storage node allocation strategy, and advertising display order list.
3. The e-commerce promotion method based on cloud computing according to claim 1, characterized in that: The steps for obtaining the positive dynamic parameters and the negative dynamic parameters are specifically as follows: S301: Retrieving the positive and negative behavior data from the product interaction log, based on the timestamps of the mutation point sequence, divide the number of favorites into intervals based on adjacent mutation points, truncate the dwell time on the detail page into discrete segments based on the mutation points, and classify the page skip frequency and swipe count by time window to generate a segmented data set; S302: Perform segmented integral calculations on the number of collections, the length of stay on the details page, the page skip frequency, and the number of slides for each segment in the segmented data set, using the formula: ; The integral value is obtained by operation, and the integral result of the positive behavior data and the integral result of the negative behavior data are calculated to generate an integral vector; in, Representative Dynamic integral value of class behavior, Representative The number of collections of the segment, Representative The length of time you stay on the details page for each segment, Representative The time interval of the segmented mutation points, Representative The page skip frequency of the segment, Representative The number of slides in the segment, Representative Segmented user activity coefficient, A balance factor representing interaction density and behavior type, Represents the total number of segments, Representative The start timestamp of the segment; S303: Separate the positive integral and the negative integral according to the positive and negative signs of the integral vector, normalize the positive integral using the maximum and minimum method, normalize the negative integral using the standard deviation method, and allocate a time weight factor based on the proportion of segment duration to generate positive dynamic parameters and negative dynamic parameters.
4. The e-commerce promotion method based on cloud computing according to claim 3, characterized in that: The steps for obtaining the reconstruction weight set are specifically as follows: S401: Calling the positive dynamic parameter and the negative dynamic parameter, based on the parameter sequence of three consecutive time windows, aligning the positive and negative parameters of each window point by point, and calculating the correlation coefficient. At the same time, taking the mean value within the window as the correlation degree, integrating the three window data, and generating a window correlation degree set; S402: extracting adjacent window correlations from the window correlation set, calculating the absolute difference between the subsequent window and the preceding window, obtaining the difference between the first and second windows, and the difference between the second and third windows, comparing the two and taking the maximum value as the current difference, and comparing the current difference with a preset difference threshold. If the difference exceeds the threshold, generating a difference status flag; S403: extracting the original priority of the advertising content weight set according to the difference status mark, reallocating the weight values according to a preset rule and arranging them in descending order, overwriting the original set, and generating a reconstructed weight set.
5. The e-commerce promotion method based on cloud computing according to claim 4, characterized in that: The steps for obtaining the advertising resource distribution instruction are specifically as follows: S501: Call the reconstruction weight set to extract the computing resource demand value and storage resource demand value of the advertisement delivery node, and combine the node load value and delay coefficient in the cloud server elastic resource scheduling algorithm to use the formula: ; Obtain node resource allocation parameter values through calculation and generate a resource allocation parameter set; in, Representative Node The resource allocation parameters, To reconstruct the nodes in the weight set Dynamic reconstruction weights, For nodes The computing resource requirement value of For nodes The storage resource requirement value, For nodes The current load value, For nodes The delay coefficient, represents the reconstruction weight association term, represents a delayed associated term; S502: Based on the resource allocation parameter set, compare the node resource allocation parameter value with the priority threshold of the advertisement display queue, eliminate parameter values below the priority threshold, intercept the parameter value sequence within the queue length upper limit according to the time window constraint, and generate a node resource adjustment queue; S503: calling the node resource adjustment queue, binding the computing resource demand value and the storage resource demand value corresponding to each parameter value in the queue one by one according to the time sequence number of the advertisement display queue, and generating an advertisement resource distribution instruction.
6. The e-commerce promotion method based on cloud computing according to claim 1, characterized in that: The method further comprises: S1: Obtain user clickstream data through cloud log collection nodes, perform cleaning and alignment operations on page jump paths, dwell time, and product category labels, and construct user behavior sequences based on the Hidden Markov Model; The user behavior sequence specifically includes a jump timestamp, a subject code, and a page level value.
7. The e-commerce promotion method based on cloud computing according to claim 6, characterized in that: The steps for obtaining the user behavior sequence are specifically as follows: S101: Collecting raw user clickstream data from cloud logs, extracting page jump paths, dwell time, and product category label fields, filtering out missing path nodes and abnormal dwell time values, matching product category labels with page paths, and generating standardized clickstream data; S102: Based on the standardized clickstream data, align the timestamps of page jump paths and product category labels, unify the time accuracy of differentiated terminals, accumulate the dwell time by segment according to the path nodes, extract the continuous path, duration, and label combination under the same user ID, and generate a behavior trajectory parameter set; S103: Call the path jump frequency, mean dwell time, and label association in the behavior trajectory parameter set, calculate the state transition probability matrix and dwell time distribution density, map the path nodes to the state variables of the hidden Markov model, establish the corresponding relationship between the state transition and the observation sequence, and generate a user behavior sequence model.
8. An e-commerce promotion system based on cloud computing, characterized in that: The system is used to implement the cloud computing-based e-commerce promotion method according to any one of claims 1 to 7, and the system includes: The behavior sequence construction module is used to obtain user clickstream data through the cloud log collection node, perform data cleaning and timestamp alignment operations on page jump paths, dwell time, and product category labels, input the cleaned data into the hidden Markov model to generate user behavior sequences, and pass the user behavior sequences to the mutation point extraction module; A mutation point extraction module is used to call the user behavior sequence, perform a sliding window alignment operation on the Euclidean distance and time interval difference of adjacent jump paths based on the dynamic time warping algorithm, extract the fluctuation peak points that exceed the preset threshold three times in a row, generate a mutation point sequence, and pass the mutation point sequence to the dynamic parameter generation module; A dynamic parameter generation module is used to obtain positive and negative behavior data from the product interaction log, perform piecewise integration operations on the number of collections, the length of time spent on the detail page, the frequency of page skips, and the number of quick swipes based on the mutation point sequence, generate positive and negative dynamic parameters, and pass the positive and negative dynamic parameters to the weight reconstruction module; a weight reconstruction module, configured to call the positive dynamic parameter and the negative dynamic parameter, calculate the difference in correlation between the two within three consecutive time windows based on a grey correlation analysis method, perform a priority reset operation on the advertising content weight set when the difference exceeds a dynamic threshold, generate a reconstructed weight set, and transmit the reconstructed weight set to the resource scheduling mapping module; The resource scheduling mapping module is used to call the cloud server elastic resource scheduling algorithm based on the reconstruction weight set, dynamically adjust the computing resource and storage resource allocation ratio of the advertising delivery node, map the adjusted resource allocation plan to the advertising display queue, generate advertising resource distribution instructions and output them to the cloud execution node.
Citation Information
Patent Citations
Charging cost analysis method and system for liquid cooling over-charging terminal
CN119692569A