E-commerce data processing method and system
By analyzing the keyword similarity and click sequence in the user behavior path and dynamically adjusting the product partitioning and data classification, the problems of low user behavior prediction accuracy and unbalanced resource allocation in the existing technology are solved, and the personalization and real-time data processing of the e-commerce platform are realized.
Patent Information
- Application Number
- CN202510959918.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-11
AI Technical Summary
Existing e-commerce data processing methods fail to effectively analyze the fine-grained dynamic changes in user behavior paths, resulting in low user behavior prediction accuracy, misalignment between product recommendation results and user intentions, unbalanced resource allocation, and difficulty in dynamically adjusting data processing during hot periods.
By identifying the text similarity between keywords and product titles, analyzing the semantic triggering sequence and behavioral node changes in the user's click path, dynamically adjusting the click volume of product partitions, and classifying data streams based on user active periods, the data processing flow is optimized.
It improves the accuracy of user behavior prediction, optimizes the personalization and time-accurate delivery of product recommendations, and enhances the data processing accuracy and real-time performance of e-commerce platforms.
Smart Images

Figure CN120471650B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of e-commerce, and in particular to an e-commerce data processing method and system. Background Art
[0002] The field of e-commerce encompasses online transactions and related data processing applications, resulting from the integration of information technology and internet services. The core of this technology lies in the information management of the transaction process for goods or services through a computer network environment, encompassing user registration, product display, order generation, payment settlement, logistics tracking, and after-sales service. E-commerce technology relies on efficient data transmission and processing mechanisms to ensure the timeliness and security of transactions. It encompasses network protocol standards, databases, front-end and back-end interaction technologies, data encryption and authentication mechanisms, user behavior analysis, cross-platform compatibility, and the recording and management of transaction data, forming a multi-dimensional, continuous digital transaction support system.
[0003] Among them, the e-commerce data processing method refers to the processing flow of classifying, extracting, transforming, integrating and reorganizing various types of data generated by transactions, browsing, operations, searches and other behaviors in the e-commerce system according to specific logical relationships. The technical matters targeted by this patent subject include user behavior data collection methods, transaction data structured mapping mechanisms, keyword semantic analysis methods, time-series screening rules for multi-source data, product label classification logic and batch splitting methods for order records.
[0004] Existing technologies often rely on static, structured integration of transaction, browsing, and search behavior data, lacking real-time analysis of fine-grained dynamic changes within user behavior paths, resulting in low user behavior prediction accuracy. In terms of search path optimization, conventional mechanisms fail to consider the semantic match between keywords and product titles, easily leading to misalignment between recommendation results and user intent. They primarily analyze behavior node identification time or events independently, ignoring the rate and logic changes between nodes in the behavior chain, making it difficult to accurately capture turning points in user behavior. Product partitioning strategies lack dynamic analysis of changes in click volume for each partition, resulting in a rigid ranking mechanism and prone to resource imbalance. Regional data scheduling fails to classify streaming data based on user activity periods, leading to data congestion during hot periods and delayed resource response. For example, in a region, user activity is concentrated between 8 PM and 10 PM, but existing technologies fail to dynamically classify and process log streams during this period, instead applying a uniform time weight. This causes high-quality product resources to miss their optimal exposure opportunities, limiting the responsiveness and matching accuracy of e-commerce platforms in delivering refined operations and personalized services. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose an e-commerce data processing method and system.
[0006] In order to achieve the above-mentioned object, the present invention adopts the following technical solution: an e-commerce data processing method, comprising the following steps:
[0007] S1: Based on the product search data of e-commerce platform users, including search keywords, search time, and product click items, the text similarity between the keywords and the clicked product titles is identified. By calculating the similarity distribution interval, stable behavior segments are identified and a search behavior matching relationship set is generated.
[0008] S2: Calling the search behavior matching relationship set, extracting keywords from the product details page and search recommendation list clicked by the user, analyzing the order of occurrence and triggering frequency of the keywords on the differentiated pages, and optimizing the keyword ranking by position changes in the click path to obtain a semantic triggering sequence list;
[0009] S3: Based on the semantic trigger sequence list, record the user's behavior chain from browsing to adding to the shopping cart, analyze the change rate between the time nodes and action types between operations, identify the behavior nodes with jump mutations in the continuous path, extract the occurrence location, and obtain the behavior path switching point set;
[0010] S4: Call the behavior path switching point set, analyze the product partition directory and click tracking records, identify the change in the click volume ratio of the product partition before and after the switching point, reconstruct the frequency of the partition whose ranking change exceeds the threshold, and output the product partition collection priority directory.
[0011] As a further solution of the present invention, the search behavior matching relationship set includes keyword association strength, click response ratio, and text content matching level;
[0012] The semantic trigger sequence list includes keyword trigger levels, page associated tags, and trigger frequency weights;
[0013] The behavior path switching point set includes path interruption nodes, behavior mutation labels, and operation timing features;
[0014] The commodity partition collection priority directory includes a partition adjustment factor, a click volume change interval, and a collection priority label.
[0015] As a further solution of the present invention, the steps of obtaining the search behavior matching relationship set are specifically as follows:
[0016] S111: Based on the product search data of users on the e-commerce platform, the search keywords, search time and product click items are extracted, the character edit distance, semantic similarity and word embedding overlap are identified, and the similarity interval value between the keyword and the product title is generated;
[0017] S112: Calling the similarity interval values between the keywords and product titles, combining them with the search time series, extracting the user's continuous click behavior segments, calculating the standard deviation and range of the similarity interval within the segment, screening the segments that meet the interval stability conditions, and obtaining the user search path stable segment interval index;
[0018] S113: Number the keywords and titles in the paragraphs according to the stable segment interval index of the user search path, analyze the relationship between the keywords and titles, identify the numbering density, sequence span and repetition rate, and use the formula:
[0019] ;
[0020] Calculate the search behavior matching offset and cross-search it with the number sequence to generate a search behavior matching relationship set;
[0021] in, Indicates the search behavior matching offset, is the search keyword position of the i-th matching group, is the product title position of the i-th matching group number, is the similarity vector distance of the matching pair of the i-th matching group, is the keyword search frequency of the i-th matching group number, is the click frequency of the title of the i-th matching group, is the total number of matching groups in the stable segment.
[0022] As a further solution of the present invention, the steps of obtaining the semantic trigger sequence list are specifically as follows:
[0023] S211: Calling the search behavior matching relationship set, extracting the keywords clicked by the user on the product details page and the recommendation list, recording the position number of each keyword's first appearance, comparing the first appearance positions of the differentiated pages, and establishing a keyword page offset value set;
[0024] S212: Counting the click frequencies of the keywords in the differentiated pages based on the keyword page offset value set, calculating frequency ratios, and sorting the keywords according to the frequency differences between the pages to obtain a keyword trigger frequency ratio sequence;
[0025] S213: Call the keyword trigger frequency ratio sequence to locate the position change trajectory of the keyword in the recommendation list, identify the position number difference, time position change value and click frequency difference in continuous click behavior, and combine the number of skips to use the formula:
[0026] ;
[0027] Calculate the keyword position change index, and reconstruct the sorting sequence based on the index to obtain a semantic trigger sequence list;
[0028] in, Indicates the keyword position change index, Keywords in the recommendation list The original position number of Keywords Click the location number on the details page. Keywords The time position change value of Keywords Frequency of clicks on the details page, Keywords Frequency of clicks on the recommended list, Keywords The number of skips, is the total number of keywords.
[0029] As a further solution of the present invention, the step of obtaining the behavior path switching point set is specifically as follows:
[0030] S311: Based on the semantic trigger sequence list, analyze the user's behavior records from browsing to adding to the shopping cart, extract node timestamps, action types and sequences, identify time differences between nodes, match action codes, and generate an inter-operation change sequence;
[0031] S312: Calling the inter-operation change sequence, identifying the change value difference and slope and the deviation of the intermediate nodes, determining whether both the change amplitude and the trend mutation threshold are met, screening the intermediate nodes that meet the conditions, and obtaining the mutation behavior node position set;
[0032] S313: Extract the jump index value and trigger sequence value according to the mutation behavior node position set, identify the joint difference structure of the time node and the jump position, and use the formula:
[0033] ;
[0034] Calculate the node jump mutation amplitude value, arrange them in order, and obtain the behavior path switching point set;
[0035] in, Represents the node jump mutation amplitude value, Represents the index difference of the node in the path, Represents the jump trigger sequence value of the node, Indicates the time difference of the first node in the sliding window, Represents the time difference of the tail node in the sliding window, Indicates the time difference between the nodes in the middle segment of the sliding window.
[0036] As a further solution of the present invention, the steps for obtaining the commodity zone collection priority directory are specifically as follows:
[0037] S411: Calling the behavior path switching point set, extracting the product partition click data in the time window before and after the switching point, identifying the change rate of the click volume share in the time period, and obtaining the partition click share change rate value;
[0038] S412: Compare the partition click ranking change with the ranking change threshold based on the partition click ratio change value, filter out abnormal sequences, and obtain a ranking mutation partition number set;
[0039] S413: According to the set of sorted mutation partition numbers, click frequency, path occurrence times and path density are integrated, and the formula is used:
[0040] ;
[0041] Calculate the reconstruction value of click frequency, sort by reconstruction value, identify the top partition numbers, and establish a product partition collection priority directory;
[0042] in, Represents the reconstructed value of click frequency, Represents the frequency of partition clicks, Represents the number of times the partition appears in the path, represents the concentration density of partition paths, Represents the change in click rate. Represents the ranking change threshold.
[0043] As a further embodiment of the present invention, the method further comprises step S5:
[0044] S5: Based on the product partition collection priority directory, obtain product browsing log stream data from the regional node server, classify it by time tag and current time period, aggregate the data streams that match the active time period, and obtain the regional active data stream aggregation set;
[0045] The regional active data flow aggregation set includes high-frequency access records, time period active peaks, and regional traffic intensity indicators.
[0046] As a further solution of the present invention, the steps of obtaining the regional active data flow aggregation set are specifically as follows:
[0047] S511: Based on the product partition collection priority directory, browsing log stream data is called from the regional node server, the log is categorized by label according to the timestamp field, and the corresponding time period classification is matched according to the current time to generate product browsing log time period classification data volume;
[0048] S512: Based on the amount of time-based classification data in the product browsing logs, determine the number of logs in each time period and the active time period parameter interval, retain the log classification data within the parameter interval, and generate an active time period log stream dataset;
[0049] S513: Call the log entries in the active period log stream data set, extract the product partition identification field, merge each type of log entries according to the partition identification, accumulate the number of log entries under the partition, identify the partition-level aggregation mapping relationship, unify the active log stream data clusters of the regional partitions, and obtain the regional active data stream aggregation set.
[0050] The e-commerce data processing system is used to execute the e-commerce data processing method, and the system includes:
[0051] The search behavior extraction module is based on the product search data of e-commerce platform users, including search keywords, search time, and product click items. It performs text matching between keywords and product titles, analyzes the number of repeated characters, word order consistency, and time interval distribution, extracts the stable search-click relationship areas in the similarity continuous segments, and generates a user behavior association dataset.
[0052] The keyword sequence construction module calls the user behavior association data set, extracts the user's keywords on the product details page and the recommendation list, counts the first appearance position and cumulative trigger frequency of the keywords by page type, and generates a keyword trigger sequence table;
[0053] The path node identification module collects the complete path nodes of the user from the browsing page to the shopping cart based on the keyword trigger sequence table, calculates the number of switches between operation types, the jump time interval and the path sequence fluctuation value, marks the node positions where the operation jump is greater than the threshold, summarizes the overlapping nodes and jump patterns in the differentiated paths, and forms a user path change node set;
[0054] The partition weight adjustment module calls the user path change node set, identifies the partitions to which the products belong before and after the node, counts the difference in the percentage of partition clicks before and after the node change, records the partitions whose click percentage changes exceed the threshold and the offset direction, and outputs a dynamic weight table for the product partitions;
[0055] The traffic data aggregation module identifies the product browsing logs from the regional node server based on the product partition dynamic weight table, extracts the time tag and partition number, aggregates the records with click frequency higher than the median value in the same period, and obtains the regional active data flow aggregation set.
[0056] Compared with the prior art, the advantages and positive effects of the present invention are:
[0057] In this invention, a matching relationship set is constructed based on the text similarity between keywords and product titles in user product search data, effectively improving the depth of association between behavioral data and product content, laying a semantic foundation for subsequent analysis. By differentiating the triggering order and click frequency of keywords in the page, the semantic logic analysis of the search path is further strengthened, and the keyword presentation order is optimized, which helps shorten the path for users to achieve their target behavior. The calculation of the change rate of time nodes and action types in the operation behavior chain enables the identification of sudden behavior nodes, improving the accuracy of extracting behavioral turning points. The structured comparison of product partition click records guides the dynamic adjustment of user attention distribution and enhances the adaptive adjustment ability of the product recommendation mechanism at the structural level. The time classification and aggregation strategy of regional active data streams dynamically matches product display resources based on the actual browsing behavior time period distribution, improving data utilization efficiency and time period precision delivery capabilities. The overall logic composed of various processing actions in series runs through the multi-dimensional processes of user behavior data analysis, semantic understanding, path dynamic identification and active flow aggregation, forming an interconnected closed-loop mechanism at the level of refined user behavior mining and data scheduling, promoting the comprehensive improvement of data processing accuracy, real-time performance and personalization level in e-commerce scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic diagram of the workflow of the present invention;
[0059] Figure 2 This is a flowchart for obtaining a search behavior matching relationship set in the present invention;
[0060] Figure 3 This is a flowchart for obtaining the semantic trigger sequence list in the present invention;
[0061] Figure 4 This is a flow chart for obtaining a set of behavior path switching points in the present invention;
[0062] Figure 5 This is a flowchart for obtaining a priority directory for commodity partition collection in the present invention;
[0063] Figure 6 This is a flowchart for obtaining the regional active data flow aggregation set in the present invention. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0065] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.
[0066] Example 1
[0067] See also Figure 1 The present invention provides a technical solution: an e-commerce data processing method, comprising the following steps:
[0068] S1: Based on the product search data of e-commerce platform users, including search keywords, search time, and product click items, the text similarity between the keywords and the clicked product titles is identified. By calculating the similarity distribution interval, stable behavior segments are identified and a search behavior matching relationship set is generated.
[0069] S2: Call the search behavior matching relationship set to extract keywords from the product details page and search recommendation list clicked by users on the e-commerce platform. Analyze the order and trigger frequency of keywords on differentiated pages. Optimize keyword ranking by position changes in the click path to obtain a semantic trigger sequence list.
[0070] S3: Based on the semantic trigger sequence list, the user's behavior chain from browsing to adding to the shopping cart is recorded, and the change rate between the time nodes and action types between operations is analyzed. The behavior nodes with jump mutations in the continuous path are identified, and the occurrence locations are extracted to obtain the set of behavior path switching points.
[0071] S4: Calls the behavioral path switching point set, analyzes the product partition directory and click tracking records, identifies the change in the click volume ratio of the product partition before and after the switching point, reconstructs the frequency of the partitions whose ranking changes exceed the threshold, and outputs the product partition collection priority directory;
[0072] S5: Based on the product partition collection priority directory, obtain product browsing log stream data from the regional node server, classify it by time tag and current time period, aggregate the data streams that match the active time period, and obtain the regional active data stream aggregation set.
[0073] The search behavior matching relationship set includes keyword association strength, click response ratio, and text content matching level. The semantic trigger sequence list includes keyword trigger level, page association label, and trigger frequency weight. The behavior path switching point set includes path interruption nodes, behavior mutation labels, and operation timing characteristics. The product partition collection priority directory includes partition adjustment factors, click volume change range, and collection priority labels. The regional active data flow aggregation set includes high-frequency access records, time period active peaks, and regional traffic intensity indicators.
[0074] See also Figure 2 , the specific steps for obtaining the search behavior matching relationship set are:
[0075] S111: Based on the product search data of users on the e-commerce platform, the search keywords, search time and product click items are extracted, the character edit distance, semantic similarity and word embedding overlap are identified, and the similarity interval value between the keyword and the product title is generated;
[0076] After the search keywords and product title content are retrieved, the character-level edit distance is calculated. This involves comparing the literal differences between each search keyword and the clicked product title. For example, on an actual e-commerce platform, a user searches for "sneakers" and the product title is "Men's Sneakers." Here, the character edit distance helps us understand the degree of direct literal match between the user's intent and the actual product title. Semantic vector similarity and word embedding overlap are calculated. This calculation involves converting the text into mathematical vectors and then evaluating the similarity between the vectors to determine the semantic proximity between the search keyword and the product title. For example, after converting "sneakers" and "Men's Sneakers" into vectors using a word embedding model, the cosine similarity of these two vectors is calculated. The combined indicators form a comprehensive matching value range, which is used to quantitatively describe the overall matching degree between the search keyword and the product title, thereby assisting the platform in optimizing the search engine results for products and generating a similarity range between keywords and product titles.
[0077] S112: Calling the similarity interval values between keywords and product titles, combining them with the search time series, extracting the user's continuous click behavior segments, calculating the standard deviation and range of the similarity interval within the segment, screening the segments that meet the interval stability conditions, and obtaining the user search path stable segment interval index;
[0078] Combined with the retrieval time series, user behavior patterns are further analyzed. First, the user's continuous click behavior segments are extracted, which involves tracking the user's continuous query behavior for products within a certain period of time. For example, after searching for "sports shoes", the user clicks on multiple related products. By calculating the standard deviation of the similarity interval and the difference between the maximum and minimum values in each segment, the user's search stability is quantified. The calculation helps determine which user behaviors show high consistency, that is, repeatedly searching for the same or similar products in a short period of time, thereby indicating the clarity of search intent and the predictability of purchase intention. By comparing the indicators with the preset stability benchmark value, those user search path segments that show high consistency in behavior are screened, thereby optimizing marketing strategies and inventory management, and obtaining user search path stability segment interval indicators.
[0079] S113: Number the keywords and titles in the paragraphs based on the stable segment interval index of the user search path, analyze the relationship between the keywords and titles, identify the numbering density, sequence span, and repetition rate, and use the formula:
[0080] ;
[0081] Calculate the search behavior matching offset and cross-search it with the number sequence to generate a search behavior matching relationship set;
[0082] in, Indicates the search behavior matching offset, is the search keyword position of the i-th matching group, is the product title position of the i-th matching group number, is the similarity vector distance of the matching pair of the i-th matching group, is the keyword search frequency of the i-th matching group number, is the click frequency of the title of the i-th matching group, is the total number of matching groups in the stable segment;
[0083] Number each group of search keywords and product titles within the stable segment one by one, construct each pair into a matching number pair, and record the text position value corresponding to the search keyword and product title in each group in the original search log in the number sequence to obtain the parameters 、 For example, if a user searches for "sports shoes", "men's sports shoes", and "casual sports shoes" in sequence, the system returns and clicks "men's running shoes", "men's sports shoes", and "men's casual shoes" in sequence, then the corresponding number position is , the corresponding product title position is The value of the text position is represented by the record position number in the retrieval sequence, which is a dimensionless value. The similarity vector distance By embedding keywords and product titles into the 300-dimensional word vector space, the Euclidean distance between the two is calculated. For example, the word vector distance between "sports shoes" and "men's running shoes" is , the vector distance unit is also dimensionless, the search frequency of the keyword Click frequency of product titles The total number of times the same keyword or title appears in the log data is obtained. For example, if "men's sports shoes" was searched 15 times and clicked 12 times in the stable segment, then , frequency is measured in times, and the ratio of the two is used as a dimensionless value in the calculation;
[0084] Substituting the above values into the formula: ;
[0085] Substitute the parameter values:
[0086] ;
[0087] ;
[0088] ;
[0089] Calculate each item:
[0090] Item 1: , the internal radical is , the product is ;
[0091] Item 2: , the product is ;
[0092] Item 3: , the internal radical is , the product is ;
[0093] Sum: ;
[0094] The final calculated matching offset is 2.389, which is dimensionless and represents the average relative offset between each group of search keywords and product titles within the current stable segment. If the behavioral consistency offset benchmark value set by the platform is 1.5, the current value is significantly larger, indicating that users have significant intention drift within this stable segment.
[0095] By introducing the retrieval click ratio of keywords and titles, as well as the product of text sequence position offsets, combined with vector semantic similarity, we can more comprehensively characterize the consistency characteristics of user search behavior and achieve precise identification of matching offset behavior. The results show that there is a large diffusion of user behavioral intentions within the stable segment, and similar product groups need to be further classified to adjust the recommendation strategy.
[0096] See also Figure 3 , the steps for obtaining the semantic trigger sequence list are as follows:
[0097] S211: Calling the search behavior matching relationship set, extracting the keywords clicked by the user on the product details page and the recommendation list, recording the position number of each keyword's first appearance, comparing the first appearance positions of different pages, and establishing a keyword page offset value set;
[0098] The search behavior matching relationship set is used to conduct a detailed analysis of the user's click behavior. First, start with the keywords clicked by the user on the product details page and the search recommendation list. The keywords are recorded and the position of their first appearance is identified. For example, after the user searches for "sports shoes", the position number 1 clicked by the user on the details page and the position number 2 in the recommendation list are recorded. This data recording provides a basis for subsequent analysis. Then, the position difference of the first appearance of the keyword on different pages is calculated. If the position number of a keyword on the details page is much smaller than the position number in the recommendation list, it means that the user is more inclined to obtain the product information through the details page, otherwise he prefers to obtain it through the recommendation list. By comparing the position number difference, the e-commerce platform can provide a strategy for how to optimize product display, and finally form a keyword page offset value set. This set can intuitively reflect the first click position offset of each keyword on different pages, thereby helping to analyze user behavior patterns and guide the platform to optimize the display of product information.
[0099] S212: Counting the click frequencies of the keywords in the differentiated pages based on the keyword page offset value set, calculating the frequency ratios, sorting the keywords according to the frequency differences between the pages, and obtaining a keyword trigger frequency ratio sequence;
[0100] Analyze the click frequency of each keyword on different pages. First, perform a statistical analysis on the frequency of the keywords. For example, "sports shoes" is clicked 100 times on the details page and 150 times in the recommendation list. By counting the frequencies, calculate the frequency ratio of each keyword on different pages. This ratio helps to understand which keywords are more popular and on which pages users interact with keywords more frequently. Then, sort the keywords based on the frequency ratio. This sorting takes into account the popularity of the keywords and the users' clicking habits. Finally, obtain the keyword trigger frequency ratio sequence. By adjusting the display order of the keywords, it can better meet the users' search needs and improve the user experience.
[0101] S213: Call the keyword trigger frequency ratio sequence to locate the position change trajectory of the keyword in the recommendation list, identify the position number difference, time position change value and click frequency difference in continuous click behavior, and combine the number of skips to use the formula:
[0102] ;
[0103] Calculate the keyword position change index, and reconstruct the sorting sequence based on the index to obtain a semantic trigger sequence list;
[0104] in, Indicates the keyword position change index, Keywords in the recommendation list The original position number of Keywords Click the location number on the details page. Keywords The time position change value of Keywords Frequency of clicks on the details page, Keywords Frequency of clicks on the recommended list, Keywords The number of skips, is the total number of keywords;
[0105] Analyze the position change trajectory of each keyword in the recommendation list. First, extract the recommended display position and click position of the keyword in the details page at different time nodes, and set them as parameters respectively. and For example, the keyword "sports shoes" is displayed in the 3rd position in the recommendation list. Clicking the position mapping number in the details page is 6th, marked as 、 Then, the time change of the keyword in the click path is collected, and the difference between the start time and the end time of the user's continuous click on the keyword is converted into a position change value. , normalized in hours. If the keyword is clicked continuously for 2 hours, then , and then count the number of clicks on the keyword on the details page , number of clicks on the recommended list and number of skips Taking "sports shoes" as an example, during the analysis period, the number of clicks on the detail page was 80, the number of clicks on the recommendation list was 50, and the number of times the user saw it but did not click on it was 30. 、 、 Similarly, let’s set the other two groups of keywords as “running shoes” and “basketball shoes”, with the following parameters:
[0106] Keyword "running shoes": ;
[0107] Keyword "basketball shoes": ;
[0108] All position numbers are dimensionless values, time changes The unit is hour, and the number of clicks and skips is "times";
[0109] The calculation process is as follows:
[0110] Item 1: , the internal , the product is ;
[0111] Item 2: , the internal , the product is ;
[0112] Item 3: , the internal , the product is ;
[0113] The sum is: ;
[0114] The result is the keyword position change index, a dimensionless value that indicates the average relative position change in the current keyword set. If this value exceeds the platform's benchmark value of 4.5, it is considered a keyword set with large ranking fluctuations and the display order needs to be adjusted;
[0115] By jointly introducing the position number difference, time change value, detail page click frequency and recommendation skip frequency into the calculation, we can effectively measure the stability and change trend of keywords under multi-page paths, and reflect whether there is excessive disturbance in the keyword ranking in the recommendation strategy, thereby establishing a semantic trigger sequence list. The result shows that the current keywords are highly volatile in the ranking, and the ranking mechanism needs to be adjusted to match the user click path behavior.
[0116] See also Figure 4 , the steps for obtaining the behavior path switching point set are as follows:
[0117] S311: Based on the semantic trigger sequence list, analyze the user's behavior records from browsing to adding to the shopping cart, extract node timestamps, action types and sequences, identify time differences between nodes, match action codes, and generate a sequence of changes between operations;
[0118] The user's behavioral data on the path from browsing to adding to the shopping cart on the e-commerce platform is recorded, including the timestamp, action type and corresponding sequence value of each step. For example, in a typical e-commerce platform, the user first browses the recommended products on the homepage, then clicks to enter the detailed page of a specific product, and finally chooses to add to the shopping cart. Each action has a clear time stamp and sequence. The data is then used to calculate the time difference between consecutive actions, and the time difference is associated with the encoding of the action type to form a two-dimensional data sequence. For example, if the user quickly adds to the shopping cart after browsing the products, the time difference is small, otherwise it is large. This associated data can help merchants analyze the smoothness of user behavior and potential operational obstacles. Ultimately, through this analysis, a sequence of changes between operations is generated. This sequence reflects the activity intensity and conversion efficiency of users at different stages, which is of great significance for optimizing user interface design and improving user experience.
[0119] S312: Call the inter-operation change sequence, identify the change value difference and slope and the deviation of the intermediate nodes, determine whether the change amplitude and trend mutation threshold are met at the same time, select the intermediate nodes that meet the conditions, and obtain the mutation behavior node position set;
[0120] Analyze specific mutation points in user behavior and identify areas where users hesitate or are confused. For example, in a typical shopping process, if a user stays too long on a product details page, it means that more product information is needed or there is difficulty in making a decision. By constructing a sliding window, the difference in the change value between the first and last points in the window and the degree of deviation between the middle node and these two points are calculated. If the value exceeds the preset threshold, it is considered that there is a behavioral mutation point. Mutation points can help merchants identify insufficient information in product descriptions or unreasonable user interface designs, so as to carry out targeted optimization. Through this method, eligible behavior nodes are screened, and the node positions are recorded to generate a mutation behavior node position set.
[0121] S313: Extract the jump index value and trigger sequence value based on the mutation behavior node position set, identify the joint difference structure of the time node and the jump position, and use the formula:
[0122] ;
[0123] Calculate the node jump mutation amplitude value, arrange them in order, and obtain the behavior path switching point set;
[0124] in, Represents the node jump mutation amplitude value, Represents the index difference of the node in the path, Represents the jump trigger sequence value of the node, Indicates the time difference of the first node in the sliding window, Represents the time difference of the tail node in the sliding window, Indicates the time difference of the nodes in the middle segment of the sliding window;
[0125] Analyze the specific position of each node in the user operation path. First, extract the index position value of each mutation node in the user path from the platform record. and its jump trigger sequence value For example, a user browses from the home page to the product details page and then to the shopping cart. Assuming that the operation sequence is 1, 2, and 3, if the user has a sudden change in behavior at step 2, the node index position is 2, and the corresponding jump trigger sequence value is 2. Get the time difference corresponding to the first, last, and middle positions in the sliding window before and after this node 、 、 , where the unit of time difference is unified as seconds. The dimension unification method is to divide the number of milliseconds recorded in the original log by 1000 to get the time difference in seconds. For example, the time difference between node 1 and node 2 is 5.2 seconds, and the time difference between node 2 and node 3 is 3.6 seconds. The intermediate node stay time is 7.1 seconds, then 、 、 , construct a joint difference value function to calculate the amplitude index of behavioral mutation;
[0126] The parameter assignments are as follows: (node index difference), (jump trigger sequence value), 、 、 (Time difference, the unit is seconds);
[0127] Substitute the above values into the formula for step-by-step calculation:
[0128] Calculate the square root: ;
[0129] Calculate the average of the three time differences within the sliding window: ;
[0130] Substitute into the formula to find the mutation amplitude index: ;
[0131] This value indicates the degree of jump mutation of the mutation behavior node in the user operation path. The larger the value, the more obvious the inconsistency between the node and the surrounding nodes in terms of operation rhythm and path structure.
[0132] By superimposing the square root of the path position difference and jump order value of the behavior node and performing a difference analysis with the average time difference within the window, not only the discrete degree of the path position information is retained, but also the degree of mutation of the time rhythm can be comprehensively reflected, thereby achieving quantitative identification of the interruption points in the user behavior path. The result shows that the mutation amplitude of the node is 2.472, which has exceeded the set behavior jump mutation threshold of 2.0. Therefore, this node can be included in the behavior path switching point set as a reference for structural adjustment and interaction optimization.
[0133] See also Figure 5 The specific steps for obtaining the product partition collection priority directory are as follows:
[0134] S411: Calling the behavior path switching point set, extracting the product partition click data in the time window before and after the switching point, identifying the change rate of the click volume share in the time period, and obtaining the partition click share change rate value;
[0135] First, relevant data is collected, such as user click records, especially behavioral data within a specific time window. For example, if we consider an online retail platform, we can analyze users' click behavior on different product partitions during a specific promotion period to determine the impact of the promotion. By comparing data in different time periods, such as before and after holidays, we can calculate the change ratio of click volume in each product partition. This involves a large amount of data collection and processing, requiring efficient database query and time series analysis technology to ensure that the required data can be extracted quickly and accurately, and generate a partition click share change rate value, which directly quantifies the changes in user interest in each partition. This value can be used to adjust the product display strategy to increase the user's purchase conversion rate.
[0136] S412: Based on the partition click ratio change value, compare the partition click ranking change with the ranking change threshold, filter out abnormal sequences, and obtain a ranking mutation partition number set;
[0137] Compare the changes in user click rankings of each product partition and screen those product partitions with the most significant changes. For example, on an e-commerce platform, it is found that the click volume of certain products suddenly increases or decreases after a specific promotion. This change suggests a shift in market trends or changes in consumer preferences. By setting a threshold for ranking changes, abnormal changes can be automatically identified. Further statistical analysis of the data can be performed to compare the changes in each product partition before and after the threshold is set, and a set of ranking mutation partition numbers can be obtained. This provides data support for subsequent marketing strategy adjustments, allowing the marketing team to adjust product layout or promotion strategies in a targeted manner.
[0138] S413: Based on the sorted mutation partition number set, click frequency, path occurrence count, and path density are integrated using the formula:
[0139] ;
[0140] Calculate the reconstruction value of click frequency, sort by reconstruction value, identify the top partition numbers, and establish a product partition collection priority directory;
[0141] in, Represents the reconstructed value of click frequency, Represents the frequency of partition clicks, Represents the number of times the partition appears in the path, represents the concentration density of partition paths, Represents the change in click rate. represents the ranking change threshold;
[0142] Integrate the click frequency, occurrence counts in paths, and path concentration density of related product partitions. First, collect click behavior data ψ for the target partition and count the total number of clicks along a specific behavior path. For example, a clothing item received 180 clicks during a promotion.
[0143] Then collect the number of times it appears in all user paths, φ. For example, this category appears 95 times in the path structure;
[0144] For the path concentration density γ, it needs to be defined as the proportion of consecutive clicks of the partition in the path. Assuming the total number of consecutive click paths for this category is 57 times and the total number of user behavior paths is 125, γ = 57 / 125 ≈ 0.456;
[0145] The click share change value Q needs to be calculated by the percentage change in the click share in the previous and next time windows. Assuming the click share before the promotion is 0.15 and after the promotion is 0.27, then Q = 0.27 − 0.15 = 0.12;
[0146] The ranking change threshold ζ is set to 0.08 based on the fluctuation range of the platform's historical click ranking changes. The above parameters need to be unified in dimension. Among them, ψ and φ are the number of clicks or occurrences, which are dimensionless values. γ is the probability value between 0 and 1. Q and ζ are percentage changes, which also need to be converted into decimals and unified into dimensionless units. All participating parameters can be uniformly converted into standard dimensionless form for calculation;
[0147] Substituting into the formula: ;
[0148] The click frequency reconstruction value k = 4643.375 calculated by the formula is a dimensionless numerical indicator representing the reconstruction strength of the current mutation partition in terms of frequency and path behavior. By introducing the path concentration density γ and the difference between the sorting mutation amplitude Q and ζ, the path behavior characteristics and click response changes are dynamically coupled and calculated to reflect the joint influence of the path behavior aggregation and the degree of interest transition. The result shows that the product partition corresponding to the number exhibits strong clustered click and mutation sorting characteristics in the current path behavior, has collection priority, and is in a priority position in the click frequency reconstruction value sorting. Based on this, the product partition collection priority directory can be obtained.
[0149] See also Figure 6 ,The specific steps for obtaining the regional active data flow aggregation set are:
[0150] S511: Based on the product partition collection priority directory, browse log stream data is called from the regional node server, the log is categorized by label according to the timestamp field, and the corresponding time period classification is matched according to the current time to generate the product browsing log time classification data volume;
[0151] Determine the data access path and permissions in the regional node server to ensure the legitimacy and security of data calls, then obtain browsing log data with timestamps from the server, and classify the log data by timestamp field. The corresponding actual application scenario can be that when an e-commerce platform conducts user behavior analysis, it uses timestamps to separate users' browsing behaviors in different time periods in order to analyze the user's activity in a specific time period. For example, if a user frequently browses a certain type of product between 8 and 10 pm, the system classifies this period as a high-activity period, and classifies the log data into the corresponding time period classification according to the current time, such as classifying browsing data in the evening into the "night active period". The final generated data is the summary data of all product browsing in the period. The data is very critical for merchants to adjust their marketing strategies. It can clarify which period needs to push more advertisements or promotional activities, and generate the product browsing log period classification data volume.
[0152] S512: Based on the amount of product browsing logs classified by time period, the number of logs in each time period and the active time period parameter range are determined, and the log classification data within the parameter range are retained to generate an active time period log stream dataset;
[0153] Evaluate whether the number of logs in each time period meets the preset active time period parameters, and compare the matching degree of the number of logs in each time period with the active time period parameters. This can be achieved by setting an activity threshold. For example, if the number of logs in a certain time period exceeds 1,000, it is considered an active time period. This threshold is based on historical data analysis and conforms to the general rules of user behavior on general e-commerce platforms. The data of the time period that meets the conditions is filtered out to further optimize the data processing efficiency. Only the data of the time period is retained for subsequent analysis, thereby saving resources and improving the processing speed. The specific data operations in the screening process include traversing the data of each time period, detecting the number of logs in each data block, and comparing it with the threshold. Through this method, an active time period log stream dataset is finally generated.
[0154] S513: Call the log entries in the active period log stream data set, extract the product partition identification field, merge each type of log entries by partition identification, accumulate the number of log entries under the partition, identify the partition-level aggregation mapping relationship, unify the active log stream data clusters of the regional partitions, and obtain the regional active data stream aggregation set;
[0155] The active period log stream is further aggregated and processed, and the logs are integrated according to the product partition identifier. The specific process includes extracting the partition identifier from each log and then grouping and counting the logs according to this identifier. In actual application scenarios, e-commerce platforms may want to understand the user visits to different product partitions during the active period in order to make targeted adjustments to product layout and promotional activities. For example, if the home appliance area is more active than the clothing area in the evening, then pushing home appliance-related promotional information in the evening will be more effective. The data processing process involves complex data classification and counting operations. After the number of logs in each partition is accumulated, the total number of user visits to each partition during the active period is generated. The statistical data is very suitable for formulating regional marketing strategies and adjusting inventory and human resource allocation. The final result is an aggregated set of regional active data streams, which provides e-commerce platforms with a detailed user behavior analysis tool for different time periods and product partitions.
[0156] The e-commerce data processing system is used to execute the above-mentioned e-commerce data processing method, and the system includes:
[0157] The search behavior extraction module is based on the product search data of e-commerce platform users, including search keywords, search time, and product click items. It performs text matching between keywords and product titles, analyzes the number of repeated characters, word order consistency, and time interval distribution, extracts the stable search-click relationship areas in the similarity continuous segments, and generates a user behavior association dataset.
[0158] The keyword sequence construction module calls the user behavior association dataset to extract the user's keywords on the product detail page and recommendation list, counts the first appearance position and cumulative trigger frequency of the keywords by page type, and generates a keyword trigger sequence table;
[0159] The path node identification module collects the complete path nodes from the user's browsing page to the shopping cart based on the keyword trigger sequence list. It calculates the number of switches between operation types, the jump time interval, and the path sequence fluctuation value. It marks the node locations where the operation jump exceeds the threshold, summarizes the overlapping nodes and jump patterns in the differentiated paths, and forms a user path change node set.
[0160] The partition weight adjustment module calls the user path change node set, identifies the partitions to which the products belong before and after the node, calculates the difference in the percentage of partition clicks before and after the node change, records the partitions whose click percentage changes exceed the threshold and the direction of the deviation, and outputs a dynamic weight table for the product partitions.
[0161] The traffic data aggregation module identifies product browsing logs from regional node servers based on the dynamic weight table of product partitions, extracts time tags and partition numbers, aggregates records with click frequencies higher than the median value within the same period, and obtains the regional active data flow aggregation set.
[0162] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. An e-commerce data processing method, characterized in that: The following steps are involved: S1: Based on the product search data of e-commerce platform users, including search keywords, search time, and product click items, the text similarity between the keywords and the clicked product titles is identified. By calculating the distribution interval of similarity, stable behavior segments are identified and a search behavior matching relationship set is generated. S2: Calling the search behavior matching relationship set, extracting keywords from the user's clicked product details page and search recommendation list, analyzing the order of occurrence and triggering frequency of the keywords on the differentiated pages, and optimizing the keyword ranking by position changes in the click path to obtain a semantic triggering sequence list; S3: Based on the semantic trigger sequence list, record the user's behavior chain from browsing to adding to the shopping cart, analyze the change rate between the operation time node and the action type, identify the behavior nodes with jump mutations in the continuous path, extract the occurrence location, and obtain the behavior path switching point set; S4: Call the behavior path switching point set, analyze the product partition directory and click tracking records, identify the change in the click volume ratio of the product partition before and after the switching point, reconstruct the frequency of the partition whose ranking change exceeds the threshold, and output the product partition collection priority directory.
2. The e-commerce data processing method according to claim 1, characterized in that: The search behavior matching relationship set includes keyword association strength, click response ratio, and text content matching level; The semantic trigger sequence list includes keyword trigger levels, page associated tags, and trigger frequency weights; The behavior path switching point set includes path interruption nodes, behavior mutation labels, and operation timing features; The commodity partition collection priority directory includes a partition adjustment factor, a click volume change interval, and a collection priority label.
3. The e-commerce data processing method according to claim 1, wherein: The steps for obtaining the search behavior matching relationship set are specifically as follows: S111: Based on the product search data of users on the e-commerce platform, the search keywords, search time and product click items are extracted, the character edit distance, semantic similarity and word embedding overlap are identified, and the similarity interval value between the keyword and the product title is generated; S112: Calling the similarity interval values between the keywords and product titles, combining them with the search time series, extracting the user's continuous click behavior segments, calculating the standard deviation and range of the similarity interval within the segment, screening the segments that meet the interval stability conditions, and obtaining the user search path stable segment interval index; S113: Number the keywords and titles in the paragraphs according to the user search path stable segment interval index, analyze the relationship between the keywords and titles, identify the numbering density, sequence span and repetition rate, and use the formula: ; Calculate the search behavior matching offset and cross-search it with the number sequence to generate a search behavior matching relationship set; in, Indicates the search behavior matching offset, is the search keyword position of the i-th matching group, is the product title position of the i-th matching group number, is the similarity vector distance of the matching pair of the i-th matching group, is the keyword search frequency of the i-th matching group number, is the click frequency of the title of the i-th matching group, is the total number of matching groups in the stable segment.
4. The e-commerce data processing method according to claim 3, wherein: The steps for obtaining the semantic trigger sequence list are specifically as follows: S211: Calling the search behavior matching relationship set, extracting the keywords clicked by the user on the product details page and the recommendation list, recording the position number of each keyword's first appearance, comparing the first appearance positions of the differentiated pages, and establishing a keyword page offset value set; S212: Counting the click frequencies of the keywords in the differentiated pages based on the keyword page offset value set, calculating frequency ratios, and sorting the keywords according to the frequency differences between the pages to obtain a keyword trigger frequency ratio sequence; S213: Call the keyword trigger frequency ratio sequence to locate the position change trajectory of the keyword in the recommendation list, identify the position number difference, time position change value and click frequency difference in continuous click behavior, and combine the number of skips to use the formula: ; Calculate the keyword position change index, and reconstruct the sorting sequence based on the index to obtain a semantic trigger sequence list; in, Indicates the keyword position change index, Keywords in the recommendation list The original position number of For keywords Click the location number on the details page. For keywords The time position change value of Keywords Frequency of clicks on the details page, Keywords Frequency of clicks on the recommended list, Keywords The number of skips, is the total number of keywords.
5. The e-commerce data processing method according to claim 4, characterized in that: The steps for obtaining the behavior path switching point set are specifically as follows: S311: Based on the semantic trigger sequence list, analyze the user's behavior records from browsing to adding to the shopping cart, extract node timestamps, action types and sequences, identify time differences between nodes, match action codes, and generate an inter-operation change sequence; S312: Calling the inter-operation change sequence, identifying the change value difference and slope and the deviation of the intermediate nodes, determining whether both the change amplitude and the trend mutation threshold are met, screening the intermediate nodes that meet the conditions, and obtaining the mutation behavior node position set; S313: Extract the jump index value and trigger sequence value according to the mutation behavior node position set, identify the joint difference structure of the time node and the jump position, and use the formula: ; Calculate the node jump mutation amplitude value, arrange them in order, and obtain the behavior path switching point set; in, Represents the node jump mutation amplitude value, Represents the index difference of the node in the path, Represents the jump trigger sequence value of the node, Indicates the time difference of the first node in the sliding window, Represents the time difference of the tail node in the sliding window, Indicates the time difference between the nodes in the middle segment of the sliding window.
6. The electronic commerce data processing method according to claim 5, characterized in that: The steps for obtaining the commodity partition collection priority directory are as follows: S411: Calling the behavior path switching point set, extracting the product partition click data in the time window before and after the switching point, identifying the change rate of the click volume share in the time period, and obtaining the partition click share change rate value; S412: Compare the partition click ranking change with the ranking change threshold based on the partition click ratio change value, filter out abnormal sequences, and obtain a ranking mutation partition number set; S413: According to the set of sorted mutation partition numbers, click frequency, path occurrence times and path density are integrated, and the formula is used: ; Calculate the reconstruction value of click frequency, sort by reconstruction value, identify the top partition numbers, and establish a product partition collection priority directory; in, Represents the reconstructed value of click frequency, Represents the frequency of partition clicks, Represents the number of times the partition appears in the path, represents the concentration density of partition paths, Represents the change in click rate. Represents the ranking change threshold.
7. The electronic commerce data processing method according to claim 1, wherein: The method further comprises step S5: S5: Based on the product partition collection priority directory, obtain product browsing log stream data from the regional node server, classify it by time tag and current time period, aggregate the data streams that match the active time period, and obtain the regional active data stream aggregation set; The regional active data flow aggregation set includes high-frequency access records, time period active peaks, and regional traffic intensity indicators.
8. The electronic commerce data processing method according to claim 7, characterized in that: The steps for obtaining the regional active data flow aggregation set are specifically as follows: S511: Based on the product partition collection priority directory, browsing log stream data is called from the regional node server, the log is categorized by label according to the timestamp field, and the corresponding time period classification is matched according to the current time to generate product browsing log time period classification data volume; S512: Based on the amount of time-based classification data in the product browsing logs, determine the number of logs in each time period and the active time period parameter interval, retain the log classification data within the parameter interval, and generate an active time period log stream dataset; S513: Call the log entries in the active period log stream data set, extract the product partition identification field, merge each type of log entries according to the partition identification, accumulate the number of log entries under the partition, identify the partition-level aggregation mapping relationship, unify the active log stream data clusters of the regional partitions, and obtain the regional active data stream aggregation set.
9. An e-commerce data processing system, characterized in that: The system is used to implement the e-commerce data processing method according to any one of claims 1 to 8, and the system includes: The search behavior extraction module is based on the product search data of e-commerce platform users, including search keywords, search time, and product click items. It performs text matching between keywords and product titles, analyzes the number of repeated characters, word order consistency, and time interval distribution, extracts the stable search-click relationship areas in the similarity continuous segments, and generates a user behavior association dataset. The keyword sequence construction module calls the user behavior association data set, extracts the user's keywords on the product details page and the recommendation list, counts the first appearance position and cumulative trigger frequency of the keywords by page type, and generates a keyword trigger sequence table; The path node identification module collects the complete path nodes of the user from the browsing page to the shopping cart based on the keyword trigger sequence table, calculates the number of switches between operation types, the jump time interval and the path sequence fluctuation value, marks the node positions where the operation jump is greater than the threshold, summarizes the overlapping nodes and jump patterns in the differentiated paths, and forms a user path change node set; The partition weight adjustment module calls the user path change node set, identifies the partitions to which the products belong before and after the node, counts the difference in the percentage of partition clicks before and after the node change, records the partitions whose click percentage changes exceed the threshold and the offset direction, and outputs a dynamic weight table for the product partitions; The traffic data aggregation module identifies the product browsing logs from the regional node server based on the product partition dynamic weight table, extracts the time tag and partition number, aggregates the records with click frequency higher than the median value in the same period, and obtains the regional active data flow aggregation set.
Citation Information
Patent Citations
Personalized search method and system
CN109189904A
E-commerce platform search recommendation and intelligent secretary integrated system
CN119168751A