A user behavior data mining method for e-commerce ecosystem
By constructing an interactive and display feature dataset and combining it with historical transaction records, the weight of influencer contribution is dynamically adjusted, which solves the problem of unfair revenue distribution under changes in influencer content style and achieves accurate evaluation and fair distribution.
Patent Information
- Application Number
- CN202511211656.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing technologies struggle to accurately assess the contribution of content to conversion and ensure fair revenue distribution between platforms and influencers when influencers change their content style. In particular, they lack the ability to uncover the deep connections between content details and user behavior when faced with changes in creator style.
By collecting content from influencers' images and texts, extracting data on fan interaction frequency fluctuations and comments, and constructing an interaction and display feature dataset, combined with historical transaction records, the correlation between interaction patterns and purchasing behavior is determined. This allows for analysis of fan trust and engagement trends, dynamic adjustment of influencer contribution weights, and optimization of platform revenue distribution.
It enables accurate assessment of content contribution in scenarios where influencer content styles change, ensuring fairness in revenue distribution between the platform and influencers and promoting the creation of high-quality content.
Smart Images

Figure CN120705418B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a method for mining user behavior data in an e-commerce ecosystem. Background Technology
[0002] In the fields of digital marketing and content creation, researching the impact of influencer content on product sales conversion is of paramount value. This area not only concerns revenue sharing between platforms and creators but also directly impacts the healthy development of the e-commerce ecosystem and the optimization of user experience. Influencer-posted text and image content can effectively guide user purchasing behavior; however, accurately assessing the contribution of content to conversion and achieving fair revenue sharing has become a pressing challenge for the industry. Currently, while many methods attempt to evaluate the value of influencer content through data analysis, these methods often overlook the dynamic and complex nature of content creation, especially when faced with changes in creator style, lacking in-depth exploration of the deep connections between content details and user behavior. Existing solutions often remain at the level of surface-level data statistics, failing to fully consider the subtle differences in content presentation and the potential impact of changes in user interaction patterns on final conversion, leading to discrepancies between evaluation results and actual contributions. Specifically, when fan interaction shifts from passive browsing to active commenting or sharing, users' trust in and engagement with the content changes significantly. For example, after an influencer posts a recommendation for a skincare product, some users simply browse the content and buy it immediately, while others ask questions in the comments section about its effectiveness, suitable skin types, and even actively share their own experiences. This shift from passive reception to active participation not only increases user trust in the product but also creates a secondary dissemination effect through comments, further amplifying the influencer's content. Furthermore, when users share content on other social media platforms, this cross-platform dissemination generates deeper conversion value, but existing evaluation systems often struggle to accurately capture this complex interaction chain. Therefore, how to dynamically adjust the weighting of creators' contributions in scenarios where influencer content styles change, and how to ensure fairness in revenue distribution between the platform and influencers, have become critical issues that urgently need to be addressed. Summary of the Invention
[0003] This invention provides a method for mining user behavior data in an e-commerce ecosystem, mainly including:
[0004] This process involves collecting influencer content (text and images), extracting fan interaction frequency fluctuations and comment data, recording the proportion of product close-ups and background environment presentation data, and constructing an interaction and display feature dataset. This dataset is then combined with historical transaction records to extract purchase time points and product category data corresponding to fan interaction frequency, establishing a preliminary correlation between interaction patterns and purchasing behavior. Comment data undergoes text processing to extract keyword frequency distribution, identify user browsing time and sharing channel distribution, analyze fan behavior fluctuations across different time periods, and assess the trend fluctuations in fan trust and engagement. Based on these trend fluctuations, the impact of content style changes on fan interaction is evaluated. According to the assessment results of the impact of content style changes on fan interaction, the influencer contribution weight is dynamically adjusted, and a revised weight allocation scheme is determined. Using the revised weight allocation scheme, key areas for mining interactive content themes and sharing channel distribution are analyzed, extracting fan dwell time distribution and forwarding paths related to purchasing behavior, and obtaining interaction frequency fluctuations under different background environments. Based on the interaction frequency fluctuations, platform revenue allocation is optimized, the revenue allocation results are determined, and fan interaction preferences, peak likes, and interaction theme-related data are stored in the database.
[0005] Furthermore, the process of collecting influencer text and image content, extracting fan interaction frequency fluctuations and comment data, recording the proportion of product close-ups and background environment presentation data, and constructing an interaction and display feature dataset includes:
[0006] The system collects text and image content posted by influencers, records the number of likes, favorites, and shares, and uses timestamps to obtain the temporal distribution of interactive behavior. It calculates the difference between posting time and interaction time as the interaction frequency fluctuation value and statistically analyzes the distribution of comment quantity and word count. The system then performs image segmentation on the text and image content, identifying product and background areas. It calculates the proportion of product area pixels to the total image pixels as the product close-up percentage and extracts the background area color histogram and texture features as environmental matching data. Based on the interaction frequency fluctuation value and the product close-up percentage, a corresponding relationship matrix is established. Deep interaction tags are marked for comments exceeding a preset threshold. Combining the dominant hue value of the color histogram and the texture feature entropy value of the environmental matching data, the data is grouped using a clustering algorithm. The grouping results are then integrated with the corresponding relationship matrix to form an interaction and display feature dataset.
[0007] Furthermore, the step of combining the interaction and display feature dataset with historical transaction records to extract purchase time points and product category data corresponding to fan interaction frequency, and determining a preliminary correlation mapping between interaction patterns and purchasing behavior, includes:
[0008] Purchase timestamps and product category identifiers are extracted from historical transaction records. These are matched with fan accounts in the interaction and display feature dataset to obtain interaction frequency data sequences and purchase behavior sequences. Periods with interaction frequencies exceeding a preset threshold are identified as intensive interaction periods, and periods with purchase counts exceeding a preset threshold are identified as peak purchase periods. The overlap ratio is calculated as the proportion of the intersection of the two periods to the union of the two periods. Strongly correlated user groups with overlap exceeding a preset threshold are marked. The proportion of product close-ups within the intensive interaction periods of these strongly correlated user groups is extracted. Based on this product close-up proportion sequence, a category migration matrix is constructed to record product category conversion frequencies. A preliminary association mapping table between interaction patterns and purchase behaviors is generated using an association rule mining algorithm.
[0009] Furthermore, the extraction of the numerical sequence of product close-up proportions during the period of intensive interaction for the strongly associated user group includes:
[0010] Extract product images viewed during periods of high interaction from the strongly associated user group, calculate the proportion of product area pixels to the total number of pixels in each image, and form a product close-up proportion numerical sequence; for the product close-up proportion numerical sequence, count the product categories corresponding to each value in the sequence, construct a category distribution vector, and form the product close-up proportion numerical sequence.
[0011] Furthermore, the text processing of the comment data, extracting keyword frequency distribution, and identifying interactive data such as user browsing time and sharing channel distribution, includes:
[0012] The comment data is segmented into words, and the frequency of each word is counted to construct a keyword frequency distribution table. Based on the user identifier, the page dwell time of the corresponding user is extracted as browsing duration data, and the number of times the content is shared to each platform is obtained to form sharing channel distribution data. The time period is divided by hours, and the total number of likes in each time period is divided by the number of unique users who like it to obtain the like concentration value. The difference in the number of comments between adjacent time periods is counted as the comment activity change value. By matching the high-frequency words in the keyword frequency distribution table with the timestamps corresponding to the comments, and combining the browsing duration data and the sharing channel distribution data, a time-series interaction feature set containing like concentration and comment activity change values is generated.
[0013] Furthermore, the assessment of the impact of content style changes on fan interaction, combining trends in fan trust and engagement, includes:
[0014] The numerical values of fan trust and engagement trends for each user at different time periods are extracted and combined to form a user behavior feature vector, constructing a comprehensive fan behavior profile matrix that includes time-series variation features. Principal component analysis is performed on the comprehensive fan behavior profile matrix, retaining principal components with a cumulative contribution rate exceeding a preset threshold to obtain a dimensionality-reduced feature matrix. Users in the dimensionality-reduced feature matrix are grouped using a hierarchical clustering algorithm to obtain fan group categories and central feature values. Based on the central feature values of the fan group categories, content posting records corresponding to active time periods are extracted, and the time interval from the content posting time to the first interaction time is calculated as the interaction response speed, thus completing the evaluation of the impact of content style changes on fan interaction.
[0015] Furthermore, the dynamic adjustment of influencer contribution weights based on the evaluation results of the impact of content style changes on fan interaction, and the determination of the adjusted weight allocation scheme, includes:
[0016] Data on the impact of content style changes for each influencer on the response speed and engagement depth of fan interactions is extracted to obtain the initial value of the influencer's current contribution weight. The average time interval between the first fan interaction after an influencer publishes content is calculated as a feedback speed indicator, and the proportion of positive words in comments is counted as the emotional positivity value. Based on the feedback speed indicator and the emotional positivity value, if the feedback speed indicator is lower than a preset time threshold and the emotional positivity value exceeds a preset proportion threshold, the initial value of the contribution weight is multiplied by a preset amplification coefficient to obtain the adjusted contribution weight. Based on the adjusted contribution weight, the total weight of all influencers is calculated and normalized to form the adjusted weight allocation scheme.
[0017] Furthermore, the adjusted weighting scheme is used to analyze key areas of interaction content themes and sharing channel distribution, extracting fan dwell time distribution and forwarding propagation paths related to purchasing behavior, including:
[0018] Using the adjusted weight allocation scheme, high-weight influencers with a weight ratio exceeding a preset threshold are selected, and the topic tags and keywords of their published content are extracted. The total interaction volume under each topic is counted, and the distribution of the number of times the content is shared to each platform is obtained. The topic category with the highest interaction volume and the channel with the most sharing are determined. Content records that meet the topic category and channel are selected from the interaction and display feature dataset. The user dwell time data that leads to purchase behavior after browsing content that meets the topic category and channel is extracted and divided into time intervals. The propagation chain of the content records from the original publication to each level of forwarding is tracked, and the user identifier and timestamp of each level of forwarding are recorded to form a forwarding propagation path.
[0019] Furthermore, the step of optimizing platform revenue distribution based on interaction frequency fluctuations and determining the revenue distribution result includes:
[0020] Calculate the purchase conversion rate for each content theme by dividing the number of purchasing users by the total number of interactive users to obtain the conversion ratio. If the conversion ratio exceeds a preset threshold, record the corresponding influencer ID and current revenue ratio. For influencers whose conversion ratio exceeds the preset threshold, multiply the current revenue ratio base by a fixed growth coefficient to obtain a new revenue ratio. Summarize the new revenue ratios and IDs of all influencers to form a revenue distribution result table.
[0021] Furthermore, after forming the profit distribution results table, it includes:
[0022] Extract the influencer identifier and new revenue ratio from the revenue distribution result table, and obtain the peak value of the influencer's fan liking behavior. Extract high-frequency words from the comment keyword distribution and combine them with interactive topic category tags to form an interactive preference feature set. Calculate the correlation coefficient between the peak value of likes and the interactive preference feature set and the purchase conversion rate to generate time-series correlation data containing the correlation coefficient. Store the time-series correlation data, influencer identifier, and new revenue ratio in the database and integrate them to form a training dataset containing time-series correlation and interactive preferences.
[0023] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0024] This invention discloses a user behavior data mining method for e-commerce ecosystems. It collects content from influencers (including images and text), extracts fan interaction frequency fluctuations and comment data, records the proportion of product close-ups and background environment presentation data, and constructs an interaction and display feature dataset. This dataset is combined with historical transaction records to extract purchase time points and product category data, determining the correlation between interaction patterns and purchase behavior. Further analysis of the correspondence between peak periods of likes, peak comment periods, and order generation reveals the impact of product close-up proportions on purchase conversion rates, forming a behavior mapping framework. Keyword extraction through text processing analyzes fan behavior fluctuations and assesses trust and engagement trends. Finally, it constructs fan behavior profiles, evaluates the impact of content style changes on interaction, dynamically adjusts influencer contribution weights, optimizes the platform's revenue distribution mechanism, and provides data support for promoting high-quality content creation. Attached Figure Description
[0025] Figure 1 This is a flowchart of a user behavior data mining method for an e-commerce ecosystem according to the present invention.
[0026] Figure 2 This is a schematic diagram of a user behavior data mining method for an e-commerce ecosystem according to the present invention. Detailed Implementation
[0027] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] like Figure 1-2 This embodiment of a user behavior data mining method for an e-commerce ecosystem may specifically include:
[0029] S101. Collect influencer text and image content, extract fan interaction frequency fluctuations and comment data, record the proportion of product close-ups and background environment presentation data, and construct an interaction and display feature dataset.
[0030] The system collects text and image content posted by influencers and extracts fan interaction data, recording the number of likes, favorites, and shares for each post. It uses timestamps to obtain the temporal distribution of interaction behavior and calculates the interaction frequency fluctuation value based on the difference between the posting time and the interaction time. Simultaneously, it extracts the number of comments and analyzes the distribution of comment word counts. Image segmentation is performed on the collected text and image content. The YOLO algorithm is used to identify product and background regions, calculating the proportion of product pixels to the total image pixels to obtain the product close-up percentage. The color histogram and texture features of the background region are extracted as environmental matching data. A correspondence matrix is established based on the interaction frequency fluctuation value and the product close-up percentage. If the comment word count exceeds a preset threshold, it is marked as a deep interaction tag. The dominant hue value of the color histogram and the entropy value of the texture features in the environmental matching data are used as environmental complexity indicators. The K-means algorithm is used to group the data containing deep interaction tags, the correspondence matrix, and the environmental complexity indicator. By integrating the K-means grouping results with the corresponding relationship matrix, the interaction frequency fluctuation value, the proportion of product close-up, the deep interaction label, and the environmental complexity index are combined into a four-dimensional feature vector. The feature vector is labeled according to the grouping results to form an interaction and display feature dataset containing display style category labels.
[0031] For example, when collecting content from influencers, the extraction of interactive data involves acquiring information from multiple dimensions.
[0032] Specifically, likes, favorites, and shares reflect users' approval of the content, while timestamps record the exact moment each interaction occurred. Calculating the difference between the posting time and the interaction time yields the interaction frequency fluctuation value, reflecting the trend of content popularity over time. The distribution of comment word counts reveals the depth of user engagement; longer comments often indicate deeper user reflection and emotional investment in the content. Image segmentation is a key technology for identifying the proportion of product close-ups. The YOLO algorithm, a classic method in object detection, uses convolutional neural networks to extract features from images, dividing the image into multiple grids, each responsible for predicting whether the region contains a target object. In influencer-generated content, the YOLO algorithm identifies the bounding box of the product, calculates the proportion of pixels in that region relative to the total image pixels, thus obtaining the proportion of product close-ups. The color histogram of the background region is obtained by statistically analyzing the pixel distribution of different color channels, while texture features are extracted using methods such as the gray-level co-occurrence matrix. These data collectively constitute a quantitative representation of the environmental context.
[0033] It's worth noting that the establishment of the correspondence matrix reveals the intrinsic link between interactive behavior and visual presentation. When the number of characters in a comment exceeds a preset threshold, the system marks it as a deep interaction tag; this tagging method can distinguish between shallow browsing and deep engagement. In the calculation of the environmental complexity index, the dominant hue value of the color histogram reflects the color richness of the background, while the entropy value of the texture features quantifies the visual complexity of the background. The K-means algorithm clusters based on these features, grouping similar content into the same group, with each group representing a specific display style.
[0034] In one embodiment, the construction of four-dimensional feature vectors enables a structured representation of the data. Interaction frequency fluctuations serve as the time dimension feature, the proportion of product close-ups as the spatial dimension feature, deep interaction tags as the user participation dimension feature, and environmental complexity indicators as the visual presentation dimension feature. These four dimensions collectively describe the complete characteristics of the influencer's text and image content. The feature vectors are labeled using K-means grouping results, resulting in a dataset that includes not only the original feature values but also display style category labels.
[0035] S102. Combine the interaction and display feature dataset with historical transaction records to extract purchase time points and product category data corresponding to the frequency of fan interaction, and determine the preliminary correlation mapping between interaction patterns and purchasing behavior.
[0036] Purchase timestamps and product category identifiers are extracted from historical transaction records. User identifiers are matched with fan accounts in the interaction and display feature dataset to obtain each user's interaction frequency data sequence and corresponding purchase behavior sequence. The time difference between the interaction time and the purchase time is calculated to form time interval distribution data. For this time interval distribution data, periods with interaction frequencies exceeding a preset threshold are identified as intensive interaction periods, and periods with purchase counts exceeding a preset threshold are identified as peak purchase periods. The overlap is calculated as the ratio of the intersection duration to the union duration of the two periods. If the overlap exceeds a preset threshold, the user group is marked as a strongly associated user group. The percentage of product close-ups viewed by this strongly associated user group during intensive interaction periods is extracted. Based on the product close-up percentage sequence, the product categories corresponding to changes in close-up percentage are recorded, constructing a category transition matrix. Rows in the matrix represent the product category at the previous moment, columns represent the product category at the next moment, and element values are the transition frequency. The category transition matrix is processed using an association rule mining algorithm to generate a rule set containing interaction frequency, close-up percentage changes, and product category transitions. The support and confidence scores of each rule are extracted from the rule set. The support score is the proportion of the rule's occurrence frequency to the total frequency, and the confidence score is the proportion of the rule's occurrence frequency to the occurrence frequency of the preceding item. Based on the support and confidence scores, valid rules are selected, and a preliminary association mapping table is established that includes the relationship between the interaction frequency range, the range of changes in the proportion of close-up shots, and the probability of purchasing product categories.
[0037] For example, obtaining time interval distribution data is fundamental to understanding user behavior patterns.
[0038] Specifically, when users like, comment on, or share content posted by influencers on social media platforms, the timestamps of these interactions are precisely recorded. These timestamps are compared with the user's subsequent purchase timestamps, and the calculated time difference forms a time interval distribution. This distribution reveals the time pattern from when a user pays attention to content to when they make a purchase decision. Short time intervals indicate a strong immediate conversion effect of the content, while long time intervals reflect the content's sustained influence. Identifying periods of intensive interaction and peak purchase periods involves a threshold determination mechanism.
[0039] In one embodiment, the number of interactions within each time window is counted. When the interaction frequency during a consecutive period is consistently higher than 1.5 times the historical average, that period is marked as an interaction-intensive period. Peak purchase periods are determined using a similar method, by analyzing the temporal distribution characteristics of transaction records. The overlap calculation reflects the degree of correlation between two periods. When a user's high-frequency interaction period and concentrated purchase period largely overlap on the timeline, it indicates that content interaction has a direct impact on the purchase decision.
[0040] It's important to note that the sequence of product close-up percentages records changes in users' visual preferences during browsing. When users transition from viewing low-close-up-percentage lifestyle scene images to high-close-up-percentage product detail images, this trajectory suggests an increased purchase intention. The category transition matrix is constructed based on this pattern; each element records the frequency of user eye movement from one product category to another. High-frequency transition paths reveal the correlation between product categories and user selection preferences. When processing the category transition matrix, the association rule mining algorithm scans all possible product combinations to discover frequently occurring patterns.
[0041] For example, when the algorithm finds that under the condition "browsing clothing content and the proportion of close-up shots increases from 30% to 70%", the probability of the consequent "purchasing the clothing" reaching 0.8, a high-confidence rule is formed. Support indicates the prevalence of the rule in all transactions, while confidence measures the reliability of the rule.
[0042] In one possible implementation, the initial association mapping table integrates multi-dimensional behavioral characteristics. Interaction frequency ranges are divided into three levels: low-frequency, medium-frequency, and high-frequency. The range of changes in the proportion of close-up shots is categorized into three modes: slow growth, rapid growth, and stable high level. Each combination corresponds to a different purchase probability value. This mapping relationship enables the system to predict the likelihood of a user purchasing a specific product category based on their current interaction status and browsing patterns, thereby achieving accurate behavioral prediction and personalized recommendations.
[0043] Data on peak periods for likes and comments are extracted from the interaction and display feature dataset. Order generation is matched with the corresponding time windows in historical transaction records. Purchase conversion rate changes are identified when the proportion of product close-ups is higher than the average. The correspondence between background environment matching style and purchase preferences for specific product categories is analyzed to form a behavior mapping framework.
[0044] From the interaction and display feature dataset, consecutive time periods with more than twice the average number of likes are identified as high-like-density periods, and time intervals with more than a preset threshold of comments are identified as high-comment periods. Based on the timestamps of the high-like-density and high-comment periods, order data within the corresponding time windows is retrieved from historical transaction records, and the order generation frequency for each time window is calculated. For time windows with high order generation frequencies, the proportion of close-up product images viewed by users within that window is extracted, the average proportion of close-up images is calculated, and product display records with a proportion of close-up images higher than the average are selected. The number of purchases generated in these records is divided by the total number of display records in that group to obtain the purchase conversion rate under the condition of high close-up proportion. Based on the product display records that generated purchases, the RGB color value distribution of the background area of the corresponding image and scene element labels are extracted. The color distribution data is clustered using the K-means algorithm to obtain the matching style categories. The purchase frequency of each product category under each matching style is counted, and a correspondence matrix is constructed with matching style, product category as column, and purchase frequency as element value. The system uses the start time of the peak period for likes as the timestamp of the interaction trigger point, and the purchase conversion rate under the condition of high close-up ratio as the quantitative indicator of product display effect. It calculates the difference between the timestamp of the interaction trigger point and the timestamp of the actual purchase behavior as the purchase response delay. It integrates the trigger point timestamp, the quantitative indicator of display effect, the response delay duration and the corresponding relationship matrix data to form a behavior mapping framework that includes temporal correlation and visual preference.
[0045] For example, the identification of peak periods for likes and comments is based on statistical principles.
[0046] Specifically, the system statistically analyzes the number of likes for each time unit in historical data, calculating the mean and standard deviation of the likes. When the number of likes within a consecutive period consistently exceeds twice the mean, that period is marked as a period of high like density. Peak comment periods are determined using a similar method, but with a focus on the temporal distribution characteristics of comment numbers. This dual-dimensional identification of interaction periods can capture moments of concentrated user attention, which often indicate a higher purchase intention. The calculation of order generation frequency involves precise matching of time windows.
[0047] In one embodiment, time windows are divided into hourly units, and the number of orders within each window is counted. When the peak period for likes is from 2 PM to 4 PM, the system extracts order data for this time period and the hour before and after it to calculate the order generation frequency. This approach of expanding the time window takes into account the decision-making delay from when a user generates interest to when they complete a purchase, making subsequent conversion rate calculations more accurate.
[0048] It's important to note that calculating conversion rates under high close-up ratio conditions has significant commercial value. The close-up ratio reflects the degree of detail in a product's display; when this ratio exceeds the average, it indicates that users are more focused on the product itself than the usage scenario. The system statistically analyzes the proportion of purchases actually generated from these high close-up ratio displays, and the resulting conversion rate directly reflects the impact of detailed display on purchase decisions. This metric provides content creators with clear guidance on display strategies. The K-means algorithm plays a crucial role in handling background color distribution. The algorithm first converts the RGB values of the background area of each image into feature vectors, with each vector containing distribution information for the red, green, and blue channels. Through iterative calculations, the algorithm groups similar color distributions into the same category, forming different color schemes.
[0049] For example, warm-toned backgrounds are categorized as a cozy style, while cool-toned backgrounds are categorized as a minimalist style. Scene element labels are obtained through image recognition, such as "outdoor," "home," and "office." These labels, together with the color clustering results, define the style category.
[0050] In one possible implementation, the construction of the correspondence matrix embodies the integration of multi-dimensional data. Rows in the matrix represent different style categories, while columns represent product categories such as clothing, accessories, and home furnishings. Each matrix element records the purchase frequency of a specific product category under a particular style. When the purchase frequency of clothing items under a warm and inviting style is significantly higher than other styles, this correspondence is quantified and recorded, providing data support for subsequent personalized recommendations. The behavioral mapping framework integrates features from both temporal and visual dimensions. Interaction trigger timestamps mark the starting point of user interest, purchase response latency quantifies the time interval from interest to action, display effect quantification reflects the effectiveness of visual presentation, and the correspondence matrix reveals the intrinsic link between style preferences and product selection. This framework enables the system to predict purchasing behavior under specific interaction patterns, achieving precise content delivery and product recommendations.
[0051] S103. Perform text processing on the comment data, extract keyword frequency distribution, identify interactive data on user browsing time and sharing channel distribution, analyze fan behavior fluctuations at different time periods, and assess the trend fluctuations of fan trust and engagement.
[0052] Tokenize the comment data, remove stop words, count the occurrences of each word, construct a keyword frequency distribution table. According to the user identifier in the preliminary association mapping, extract the page停留 time of the corresponding user as the browsing duration data, and at the same time obtain the number of times the content is shared to each platform to form the sharing channel distribution data. Use the high-frequency words in the keyword frequency distribution table to match the timestamps corresponding to the comments containing these words, divide the time period by hour, calculate the total number of likes in each time period divided by the number of independent users who liked in that period to obtain the like concentration value, and count the difference in the number of comments in adjacent time periods to determine the comment activity change value. Identify the time periods exceeding the mean value in the like concentration value sequence as high-concentration time periods, mark the positive and negative conversion points of the comment activity change value as activity fluctuation points, match the positive and negative words in the keywords through a sentiment dictionary. If the proportion of the frequency of positive words exceeds a preset threshold and the cumulative value of the browsing duration in the corresponding time period continues to increase, the trust index for that time period is marked as positive. Calculate the number of sharing platform types as the participation breadth value for the sharing channel distribution data, combine the occurrence frequency of high-concentration time periods and the density of activity fluctuation points, and calculate the change rate of the trust index sequence and the dispersion degree of the participation breadth value sequence through a sliding window method to obtain the trend fluctuation evaluation result including the trust change rate and the participation sense fluctuation amplitude for each time period.
[0053] Exemplarily, the tokenization of the comment text involves the basic link of natural language processing.
[0054] Specifically, the system first segments the original comment text, decomposing the continuous character sequence into meaningful lexical units. In the process of removing stop words, the system filters out high-frequency but meaningless words such as "of", "already", "in", etc., and retains words with substantial evaluation significance such as "good quality", "fast logistics", "correct color", etc. The construction of the keyword frequency distribution table is completed by counting the occurrences of each word in all comments. High-frequency words often reflect the product features or service elements that users are most concerned about. The extraction of the browsing duration data is based on the accurate recording of user behavior logs. When a user enters a certain product detail page, the entry timestamp is recorded. When the user leaves the page or performs other operations, the departure timestamp is recorded. The difference between the two is the single browsing duration. The sharing channel distribution data is obtained by tracking the click behavior of the sharing button. The system records the number of times the content is shared to different platforms such as WeChat, Weibo, Xiaohongshu, etc., forming a multi-dimensional propagation path map.
[0055] In one embodiment, the concentration of likes reflects the degree of aggregation of user interaction behavior. If 100 likes come from 100 different users, the concentration of likes is 1, indicating that the interaction is scattered; if 100 likes come from 50 users, with some users liking multiple times, the concentration is 2, indicating that there is repeated interaction among the core fan group. The change in comment activity is obtained by calculating the difference in the number of comments in adjacent time periods. Positive values indicate an increase in activity, and negative values indicate a decrease in activity. The magnitude of the change reflects the dynamic evolution of the content's popularity.
[0056] It's important to note that the sentiment dictionary plays a crucial role in identifying review sentiment. The dictionary includes pre-labeled positive words such as "like," "satisfied," and "recommend," and negative words such as "disappointed," "return," and "poor quality." By matching keywords in reviews with the sentiment dictionary, the ratio of the frequency of positive words to the total frequency of sentiment words is calculated. When this ratio exceeds 0.7, it's considered that positive reviews dominate. Simultaneously, the cumulative browsing time for the corresponding period is observed. If it continuously increases, it indicates that users are willing to invest more time in understanding the product, and trust is developing positively. The participation breadth value is calculated based on the diversity of sharing platforms. When content is shared only on a single platform, the participation breadth value is low; when content spreads across multiple platforms, the participation breadth value increases, reflecting the content's cross-platform influence. The sliding window method plays a smoothing role in calculating trends. The window contains data from multiple consecutive time periods. The rate of change is obtained by calculating the slope of a linear regression of the data within the window; a positive slope indicates an upward trend, and a negative slope indicates a downward trend. The trend fluctuation assessment results integrate dynamic indicators from multiple dimensions. The rate of change in trust reflects the direction and speed of user attitude evolution, while the amplitude of engagement fluctuations is obtained by calculating the standard deviation of the engagement breadth value series; a larger standard deviation indicates more unstable user engagement behavior. This comprehensive evaluation method can fully reflect the behavioral characteristics and psychological state changes of the fan base, providing content creators with accurate user insights.
[0057] S104. By combining the trends and fluctuations in fan trust and engagement, assess the impact of changes in content style on fan interaction.
[0058] By analyzing the trend fluctuations in fan trust and engagement data, we extract the trust value sequences and engagement intensity sequences for each user at different time periods. These two sequences are combined to form a user behavior feature vector. Integrating the feature vectors of all users, we construct a comprehensive fan behavior profile matrix that includes temporal variations. Principal component analysis (PCA) is used to reduce the dimensionality of this comprehensive fan behavior profile matrix, retaining principal components with a cumulative contribution rate exceeding a preset threshold. This yields a dimensionality-reduced feature matrix. Hierarchical clustering is then used to group users within the dimensionality-reduced feature matrix, resulting in fan group categories with different behavioral patterns and their central feature values. Based on the central feature values of each fan group category, we identify the content posting records corresponding to the active periods of each category. We extract the image-to-text ratio, color scheme, and text length from these records as content style parameters. The time interval between content posting and the first interaction is calculated as the interaction response speed. Finally, we weight and sum the comment word count, share count, and browsing time according to preset weighting coefficients to obtain the engagement depth index. For adjacent time periods when content style parameters change, calculate the difference in interaction response speed before and after the change. Obtain the participation depth change rate by subtracting the participation depth index before the change from the participation depth index after the change and then dividing by the index before the change. Establish an evaluation result table containing three columns of data: style parameter change type, response speed difference, and participation depth change rate, to complete the evaluation of the impact of content style changes on fan interaction.
[0059] For example, the construction of a comprehensive fan behavior profile matrix is based on the integration of multi-dimensional time-series data.
[0060] Specifically, the trust score sequence records how users' approval of content creators changes over time, with higher scores indicating stronger trust. The engagement intensity sequence reflects users' willingness to actively interact, quantified by the frequency and depth of behaviors such as liking, commenting, and sharing. Combining these two sequences after aligning them by time results in a feature vector that includes both attitudinal and behavioral dimensions, comprehensively characterizing users' interaction features. Principal component analysis plays a crucial role in dimensionality reduction. The original behavioral feature vector may contain dozens of dimensions of data, and direct processing would lead to excessive computational complexity. Principal component analysis, through linear transformation, projects the original features onto the direction of maximum variance, preserving the main patterns of data variation. When the cumulative contribution rate reaches 85%, it means that the retained principal components contain 85% of the information in the original data, achieving both data compression and preservation of key features.
[0061] In one embodiment, a hierarchical clustering algorithm progressively merges similar user groups by calculating the similarity between users. The algorithm first treats each user as an independent category, then calculates the Euclidean distance between the feature vectors of any two users; the smaller the distance, the more similar the behavioral patterns. By continuously merging the nearest categories, a tree structure is eventually formed, and different fan group categories are obtained based on a preset distance threshold. The central feature value of each category is obtained by calculating the mean of the features of all users within that category.
[0062] It's important to note that the extraction of content style parameters involves quantification across multiple dimensions. The image-to-text ratio is calculated as the percentage of the image area on the entire content page; a high ratio indicates visual dominance, while a low ratio suggests text-based content. Color scheme is determined by extracting the RGB values of the image's primary color tone; warm tones create a cozy atmosphere, while cool tones convey professionalism. Text length is directly calculated based on the number of words; shorter text facilitates quick browsing, while longer text provides detailed information. Interaction response speed reflects the content's immediate appeal. The shorter the time interval between content publication and the first user interaction, the faster the content captures user attention. In the weighted calculation of engagement depth metrics, comment word count reflects depth of thought (weighted at 0.5); share frequency reflects willingness to spread (weighted at 0.3); and browsing time indicates level of attention (weighted at 0.2). This weighting approach balances deep interaction with behavioral diversity. The evaluation results table quantifies the impact relationships. When the content style changes from a high to a low image-to-text ratio, the corresponding response speed difference is recorded; positive values indicate faster response, and negative values indicate slower response. The calculation of the participation depth change rate adopts a relative change method, which eliminates the impact of differences in the base number of different user groups.
[0063] S105. Based on the evaluation results of the impact of content style changes on fan interaction, dynamically adjust the weight of influencer contribution and determine the adjusted weight allocation scheme.
[0064] Based on the assessment results of the impact of content style changes on fan interaction, the interaction response speed difference and engagement depth change rate data for each influencer are extracted to obtain the initial value of the influencer's current contribution weight. The average time interval between the first fan interaction after the influencer publishes content is calculated as the feedback speed indicator. The emotional positivity value is obtained by matching the proportion of positive words in the comments through an emotional dictionary. For the feedback speed indicator and emotional positivity value, combined with the interaction response speed difference and engagement depth change rate, if the feedback speed indicator is lower than a preset time threshold and the emotional positivity value exceeds a preset proportion threshold, the initial value of the contribution weight is multiplied by a preset amplification coefficient to obtain the adjusted contribution weight; otherwise, the initial value of the contribution weight remains unchanged. Using the adjusted contribution weight, the sum of the contribution weights of all influencers is calculated. The contribution weight of each influencer is divided by the sum to obtain the normalized weight proportion, forming an adjusted weight allocation scheme that includes the influencer's identifier and corresponding weight proportion.
[0065] For example, the dynamic adjustment mechanism for influencer contribution weight is based on multi-dimensional behavioral data evaluation.
[0066] Specifically, the initial contribution weight reflects an influencer's historical performance and basic influence on the platform, typically calculated based on a combination of their follower count, historical content quality, and conversion rates. This initial value serves as a benchmark for adjustments, ensuring the continuity and rationality of weight changes. The calculation of the feedback speed metric involves precise measurement of the time dimension. After an influencer publishes new content, the time of the first user interaction is recorded, and the difference between this time and the publication time is calculated. By statistically analyzing the average first interaction time after multiple publications, the influencer's feedback speed metric is obtained. A shorter time interval indicates that the content has strong immediate appeal and can quickly stimulate user interest.
[0067] In one embodiment, the positive sentiment score is obtained using mature sentiment dictionary technology. After segmenting the comment text, it is matched against preset positive and negative sentiment lexicons. Positive words include expressions of positive attitude such as "like," "wonderful," and "recommend," while negative words include expressions of negative attitude such as "disappointed," "average," and "not worth it." The positive sentiment score, ranging from 0 to 1, is obtained by calculating the proportion of positive words appearing out of the total number of sentiment words.
[0068] It's important to note that the introduction of interaction response speed difference and engagement depth change rate enhances the comprehensiveness of the evaluation. The interaction response speed difference reflects the impact of style changes on user attention by comparing response time changes before and after content style adjustments. The engagement depth change rate measures the actual effect of content optimization from the perspective of user engagement. These two indicators, along with feedback speed and emotional positivity, constitute a four-dimensional evaluation system. The conditional judgment for weight adjustment reflects the design philosophy of the incentive mechanism. When the feedback speed is faster than the preset threshold and the emotional positivity is higher than the standard value, it indicates that the influencer's content can quickly attract users and obtain positive feedback; this excellent performance is recognized by increasing the weight. The preset increase coefficient is usually set between 1.1 and 1.5, ensuring the significance of the adjustment while avoiding excessive fluctuations. For influencers who do not meet the improvement conditions, the original weight remains unchanged, maintaining the stability of the system. Normalization ensures the rationality and comparability of weight allocation. By calculating the sum of the adjusted weights of all influencers and then dividing each influencer's weight by this sum, the resulting weight percentage always maintains the characteristic that the sum is 1. This approach ensures that the relative importance of different influencers is reflected, while preventing the weight values from increasing indefinitely. The resulting weight allocation scheme reflects the influencers' real-time performance while maintaining the balance of the overall allocation structure.
[0069] S106. By adjusting the weight allocation scheme, analyze the key mining directions of interactive content themes and sharing channel distribution, extract the fan dwell time distribution and forwarding propagation path related to purchasing behavior, and obtain the interaction frequency fluctuation under different background environments.
[0070] By adjusting the weighting scheme, high-weight influencers with a weight ratio exceeding a preset threshold are selected. The topic tags and keywords of their published content are extracted, and the total interaction volume under each topic is statistically analyzed. Simultaneously, the distribution of content sharing frequency across various platforms is obtained, identifying the topic categories with the highest interaction volume and the channels with the most sharing as key areas for further analysis. Based on the topic categories and channels identified in the key areas for analysis, content records that simultaneously meet the topic category and are disseminated on designated channels are selected from the initial dataset. User dwell time data resulting from viewing this content is extracted and divided into short, medium, and long intervals after being sorted by duration value. The complete propagation chain of content from its original publication to each level of forwarding is tracked, recording the user identifier and timestamp of each level of forwarding to form a forwarding propagation path. For the content and dwell time interval data at each node in the forwarding propagation path, corresponding background environment matching parameters are extracted, including color tone values and scene type tags. The frequency of interactive behavior within each dwell time interval is statistically analyzed according to different background environment matching types. The difference in interaction frequency between adjacent time periods is calculated to obtain interaction frequency fluctuation data, forming a correspondence record between background matching type and interaction frequency fluctuation characteristics.
[0071] For example, the selection mechanism for high-weight influencers reflects the priority principle of resource allocation.
[0072] Specifically, when a KOL's weighting exceeds a preset threshold of 5%, it indicates that the KOL possesses strong content creation capabilities and user influence on the platform. The extraction of hashtags is achieved through natural language processing technology. The system identifies noun phrases and trending keywords in the content, such as "outfit sharing," "makeup tutorials," and "digital product reviews." These tags directly reflect the core attributes of the content. Interaction statistics encompass a comprehensive calculation of various user behaviors.
[0073] In one embodiment, the system assigns 1 point to likes, 3 points to comments, and 5 points to shares, and calculates the total interaction for each topic by weighted summation. The distribution of sharing channels is obtained by tracking click logs of the share buttons, recording the specific number of times content is shared to various platforms such as WeChat, Weibo, Xiaohongshu, and Douyin. When the total interaction for the "Outfit Sharing" topic reaches 100,000 points, and the number of shares on the Xiaohongshu platform accounts for 60% of the total shares, this topic and channel combination becomes a key area for analysis.
[0074] It's important to note that the division of dwell time intervals is based on statistical patterns of user behavior. First, the page dwell time of all users who made a purchase was collected. The data was then sorted from smallest to largest, and the 33rd and 67th percentiles were used as interval boundaries. Short intervals are typically 0-30 seconds, representing impulsive purchases after quick browsing; medium intervals are 30-120 seconds, reflecting rational decisions after moderate consideration; and long intervals exceed 120 seconds, representing cautious purchases after in-depth research. This division method considers both the actual distribution of data and has clear business implications. Tracking the forwarding propagation path reveals the viral spread characteristics of the content. When user A publishes original content, user B forwards it and adds a comment, and user C then forwards it from B, forming a propagation chain A→B→C. The user identifier, forwarding timestamp, and text content added during forwarding are recorded for each node. By analyzing the length and number of branches of the propagation path, content features with high propagation value can be identified. The extraction of background environment parameters involves the application of computer vision technology. Color tone values are obtained by calculating the RGB values of the dominant color in the image; for example, warm tones may appear as orange-red hues with high R values. Scene type labels are determined by an image recognition model, classifying backgrounds into typical scenes such as "home," "outdoor," "office," and "coffee shop." The combination of these parameters forms a unique visual style identifier.
[0075] In one possible implementation, the calculation of interaction frequency fluctuations reflects the temporal changes in user activity. The system counts the number of interactions for each background combination type on an hourly basis, and obtains the fluctuation data by calculating the difference in the number of interactions between adjacent hours. When the interaction frequency of content with warm colors in a home setting surges between 8 and 9 pm, while content with cool colors in an office setting is active on weekday mornings, this correlation provides data support for optimizing the timing of content release.
[0076] S107. Optimize the platform revenue distribution based on the fluctuation of interaction frequency, determine the revenue distribution results, and store the fan interaction preferences, peak likes, and interaction topic association data in the database.
[0077] Based on interaction frequency fluctuation data, the purchase conversion rate for each content theme is calculated. The conversion ratio is obtained by dividing the number of purchasing users by the total number of interacting users. If this ratio exceeds a preset threshold, the interactive content theme is considered to have a high promotional effect on purchasing behavior, and the corresponding influencer identifier and current revenue ratio are recorded. For influencers with a high promotional effect, their current revenue ratio base is obtained. The base is multiplied by a fixed growth factor to obtain the revenue increment. The base is added to the increment to obtain the new revenue ratio. The adjusted revenue ratio, influencer identifier, and adjustment timestamp are stored, and the data of all influencers are summarized to form a revenue allocation result table. Using the revenue allocation result table, the identifier, new revenue ratio, and adjustment timestamp of each influencer are extracted. The times when the corresponding influencer's fan liking behavior is concentrated are obtained from historical data as the peak of likes, the distribution of comment keywords is used as the interaction preference, and the interactive theme category label is obtained. These data are combined with the purchase conversion rate and stored in the database to form the training base dataset. By periodically reading the training dataset, peak times of fan likes and frequency of interaction preference keywords are extracted as feedback indicators. The Pearson correlation coefficient between these indicators and the purchase conversion rate is calculated. By comparing the difference in correlation coefficients between adjacent periods, the trend data of the relationship between fan feedback and purchase behavior is obtained, which is used to optimize the rules for subsequent dynamic adjustment.
[0078] For example, purchase conversion rate is a core metric for evaluating the value of content.
[0079] Specifically, the system calculates the conversion rate by dividing the total number of users who viewed content on a specific topic within a given time window by the number of users who made an actual purchase. For example, if 1000 users viewed a beauty tutorial topic and 80 of them purchased related products within 72 hours, the conversion rate is 8%. If the preset threshold is set at 5%, the topic is considered to have a high promotional effect. This dynamic adjustment mechanism for the revenue ratio reflects an incentive-driven distribution principle.
[0080] In one embodiment, an influencer's initial revenue share is determined based on their follower count and historical performance, typically ranging from 1% to 10% of the platform's total revenue. When an influencer's content is identified as having a high promotional effect, the system uses a fixed increase coefficient of 1.2 for calculation. If the influencer's current revenue share is 3%, then the revenue increment is 3% × 0.2 = 0.6%, resulting in a new revenue share of 3.6%. This progressive adjustment ensures the interests of high-quality creators while maintaining a balanced overall distribution.
[0081] It's important to note that the identification of peak likes is based on time series analysis principles. A 24-hour day is divided into 48 half-hour periods, and the number of likes within each period is counted. After smoothing the data using a moving average, the moments with significantly higher like counts than adjacent periods are identified as peaks. When the number of likes reaches its highest point of the day between 8 PM and 8:30 PM, this period is marked as the peak like time for that influencer's content, reflecting the activity patterns of the fan base. The extraction of interaction preferences involves the application of text mining techniques. All comments under the influencer's posts are segmented, and the frequency of each keyword is counted. High-frequency words such as "recommendation," "purchased," and "link" (purchase intent terms), as well as product feature terms such as "quality," "makes skin look whiter," and "versatile," collectively constitute a profile of fan interaction preferences. This keyword distribution not only reflects user focus but also reveals key factors influencing purchase decisions.
[0082] In one possible implementation, the Pearson correlation coefficient reveals the strength of the linear relationship between variables. The system converts the peak time of likes into a numerical variable, such as 8 PM as 20, and uses the frequency of interaction preference keywords as another variable, calculating its correlation with the purchase conversion rate. The correlation coefficient ranges from -1 to 1, with values close to 1 indicating a strong positive correlation and values close to 0 indicating no correlation. When the correlation coefficient between the peak time of likes and the conversion rate is 0.75 in a certain period and becomes 0.82 in the next period, the difference of 0.07 indicates that the correlation is strengthening. The construction of the training dataset enables the integrated storage of multi-dimensional features. Each record includes fields such as influencer ID, new revenue ratio, adjustment timestamp, peak time of likes, interaction preference keywords and their frequency, interaction topic category, and purchase conversion rate. This structured data organization not only facilitates subsequent query analysis but also provides a standardized input format for training machine learning models. Through regular updates and accumulation, the dataset gradually forms a knowledge base reflecting the evolution of the platform ecosystem.
[0083] Based on interaction frequency fluctuation data, we calculate the fan engagement depth index of influencer content, evaluate the purchase conversion contribution of different interactive content themes, set threshold standards to promote purchase behavior, increase the revenue weight coefficient of influencer content types that exceed the threshold, adjust the platform commission ratio and influencer revenue sharing structure, and increase the income of influencers with high conversion interaction modes to above the benchmark value, thus forming an incentive-based revenue optimization model for high-quality content creation.
[0084] We acquire interaction frequency fluctuation data for influencer content, including the distribution of likes, comments, and shares over time. We calculate the frequency change rate based on the time intervals of these interactions, and obtain the fan engagement depth metric by dividing the change rate by the total number of followers. This metric reflects the sustained level of attention the fan base pays to the influencer's content. For each content theme, we correlate the fan engagement depth metric with historical purchase records under that theme. We use Bayesian inference to calculate the conditional probability of a purchase at a specific engagement depth. The product of the conditional probability and the engagement depth metric determines the purchase conversion contribution of each content theme. Based on the distribution of purchase conversion contribution across all content themes, we rank the contributions from highest to lowest. We set the contribution values in the top 30% percentile as a threshold standard for promoting purchase behavior. If an influencer's contribution for a certain type of content theme exceeds this threshold, the revenue weight coefficient for that theme is set as the ratio of contribution to the threshold. This weight coefficient is used to adjust the proportion received by the influencer in the original revenue sharing structure. The actual earnings of influencers are calculated based on the adjusted revenue sharing structure and compared with the benchmark earnings set by the platform. If the actual earnings do not reach the benchmark, the weighting coefficient is adjusted upward according to the proportion of the difference to the benchmark until the influencer's earnings exceed the benchmark. By periodically executing the above adjustment process, an optimized earnings model that incentivizes the creation of high-quality content is formed.
[0085] In one possible implementation, acquiring interaction frequency fluctuation data involves collecting user behavior data at different time intervals after an influencer posts content.
[0086] Specifically, when an influencer posts a beauty tutorial video, the system records the number of likes: 500 within 0-1 hour, 800 within 1-3 hours, 400 within 3-6 hours, and so on. The frequency change rate is calculated by determining the difference in interaction increments between adjacent time periods. If an influencer has 100,000 followers and the video accumulates 5,000 interactions within 24 hours, the follower engagement depth index is 0.05. This value reflects the activity level of the follower group and the appeal of the content.
[0087] It's important to note that Bayesian inference plays a crucial role in calculating the contribution of purchase conversion. First, purchase behavior data is analyzed at a specific engagement depth. For example, if the engagement depth is 0.05, and 50 out of 1000 users complete a purchase, the conditional probability is 0.05. Multiplying this probability by the engagement depth metric yields a purchase conversion contribution of 0.0025 for that content topic. This calculation method comprehensively considers the correlation between user engagement and actual purchase behavior, avoiding the one-sidedness of judging content value solely based on purchase volume.
[0088] In one embodiment, the threshold standard is set using a dynamic quantile method. The contribution of all content topics is ranked; for example, makeup tutorials have a contribution of 0.0025, fashion sharing 0.0020, and lifestyle vlogs 0.0015. The value of 0.0020, located in the top 30% quantile, is set as the threshold. When the contribution of a makeup tutorial (0.0025) exceeds the threshold, its revenue weighting coefficient is calculated as 0.0025 / 0.0020 = 1.25. In the original revenue sharing structure, the influencer receives 20% of the sales revenue; after adjustment, this becomes 20% × 1.25 = 25%. This adjustment mechanism incentivizes influencers to create more high-converting content.
[0089] Preferably, the revenue optimization mechanism achieves dynamic balance through periodic adjustments. Assuming the platform sets a benchmark revenue of 5000 yuan per month, a creator's actual revenue, calculated using the adjusted revenue-sharing structure, is 4500 yuan, with the difference of 500 yuan representing 10% of the benchmark. The weighting coefficient is then increased from 1.25 to 1.375, bringing the creator's revenue to 5500 yuan. This mechanism is executed monthly, adjusted based on the latest interaction data and purchase conversion rates, ensuring revenue for high-quality content creators while promoting the healthy development of the platform's content ecosystem. By directly linking high-converting content with increased revenue, a positive cycle of creative quality and economic returns is created.
[0090] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of the embodiments of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.
Claims
1. A method for mining user behavior data in an e-commerce ecosystem, characterized in that, The method includes: Collect influencer image and text content, extract fan interaction frequency fluctuations and comment data, record product close-up ratio and background environment matching data, and construct an interaction and display feature dataset. In this process, the image and text content is segmented to identify product area and background area. The proportion of product area pixels to the total number of image pixels is calculated as the product close-up ratio, and the background area color histogram and texture features are extracted as environment matching data. By combining the interaction and display feature dataset with historical transaction records, purchase time points and product category data corresponding to the frequency of fan interaction are extracted to determine the preliminary correlation mapping between interaction patterns and purchasing behavior. Text processing is performed on the comment data to extract keyword frequency distribution, identify interactive data on user browsing time and sharing channel distribution, analyze the behavioral fluctuations of fans at different time periods, and assess the trend fluctuations of fan trust and engagement. By combining trends in fan trust and engagement, we can assess the impact of changes in content style on fan interaction. Based on the assessment results of the impact of content style changes on fan interaction, the weight of influencer contribution is dynamically adjusted to determine the adjusted weight allocation scheme. By adjusting the weighting scheme, we analyze the key areas for mining interactive content themes and sharing channel distribution, extract the fan dwell time distribution and forwarding propagation path related to purchasing behavior, and obtain the interaction frequency fluctuation under different background environments. This includes: using the adjusted weighting scheme, we screen high-weight influencers whose weight ratio exceeds a preset threshold, extract the theme tags and keywords of their published content, count the total interaction volume under each theme, obtain the distribution of the number of times content is shared to each platform, and determine the theme category with the highest interaction volume and the channel with the most sharing times; we screen content records that meet the theme category and channel from the interaction and display feature dataset, extract the user dwell time data of users who made purchases after browsing content that meets the theme category and channel, and divide the dwell time intervals; we track the propagation chain of the content records from the original publication to each level of forwarding, record the user identifier and timestamp of each level of forwarding, and form the forwarding propagation path. The platform's revenue distribution is optimized based on the fluctuations in interaction frequency. The revenue distribution results are determined, and the data related to fan interaction preferences, peak likes, and interaction topics are stored in the database.
2. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that, The process involves collecting influencer content (text and images), extracting fan interaction frequency fluctuations and comment data, recording the proportion of product close-ups and background environment presentation data, and constructing an interaction and display feature dataset, including: Collect the text and image content posted by influencers, record the number of likes, favorites, and shares, and combine the timestamps to obtain the temporal distribution of interactive behavior. Calculate the difference between the posting time and the interaction time as the interaction frequency fluctuation value, and statistically analyze the distribution of the number and word count of comments. Establish a corresponding relationship matrix based on the interaction frequency fluctuation value and the proportion of product close-ups, mark deep interaction tags for comments with word counts exceeding a preset threshold, and combine the dominant color value of the color histogram of the environmental matching data with the texture feature entropy value. Group the data through a clustering algorithm, integrate the grouping results with the corresponding relationship matrix, and form an interaction and display feature dataset.
3. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that, The process of combining interactive and display feature datasets with historical transaction records to extract purchase time points and product category data corresponding to fan interaction frequency, and determining a preliminary correlation mapping between interaction patterns and purchasing behavior, includes: Purchase timestamps and product category identifiers are extracted from historical transaction records. These are matched with fan accounts in the interaction and display feature dataset to obtain interaction frequency data sequences and purchase behavior sequences. Periods with interaction frequencies exceeding a preset threshold are identified as intensive interaction periods, and periods with purchase counts exceeding a preset threshold are identified as peak purchase periods. The overlap ratio is calculated as the proportion of the intersection of the two periods to the union of the two periods. Strongly correlated user groups with overlap exceeding a preset threshold are marked. The proportion of product close-ups within the intensive interaction periods of these strongly correlated user groups is extracted. Based on this product close-up proportion sequence, a category migration matrix is constructed to record product category conversion frequencies. A preliminary association mapping table between interaction patterns and purchase behaviors is generated using an association rule mining algorithm.
4. The user behavior data mining method for an e-commerce ecosystem according to claim 3, characterized in that, The extraction of the numerical sequence of product close-up percentages for the strongly associated user group during periods of intensive interaction includes: Extract product images viewed during periods of high interaction from the strongly associated user group, calculate the proportion of product area pixels to the total number of pixels in each image, and form a product close-up proportion numerical sequence; for the product close-up proportion numerical sequence, count the product categories corresponding to each value in the sequence, construct a category distribution vector, and form the product close-up proportion numerical sequence.
5. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that, The text processing of comment data, extraction of keyword frequency distribution, and identification of user browsing time and sharing channel distribution interaction data include: The comment data is segmented into words, and the frequency of each word is counted to construct a keyword frequency distribution table. Based on the user identifier, the page dwell time of the corresponding user is extracted as browsing duration data, and the number of times the content is shared to each platform is obtained to form sharing channel distribution data. The time period is divided by hours, and the total number of likes in each time period is divided by the number of unique users who like it to obtain the like concentration value. The difference in the number of comments between adjacent time periods is counted as the comment activity change value. By matching the high-frequency words in the keyword frequency distribution table with the timestamps corresponding to the comments, and combining the browsing duration data and the sharing channel distribution data, a time-series interaction feature set containing like concentration and comment activity change values is generated.
6. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that, The assessment of the impact of content style changes on fan interaction, combining trends in fan trust and engagement, includes: The numerical values of fan trust and engagement trends for each user at different time periods are extracted and combined to form a user behavior feature vector, constructing a comprehensive fan behavior profile matrix that includes time-series variation features. Principal component analysis is performed on the comprehensive fan behavior profile matrix, retaining principal components with a cumulative contribution rate exceeding a preset threshold to obtain a dimensionality-reduced feature matrix. Users in the dimensionality-reduced feature matrix are grouped using a hierarchical clustering algorithm to obtain fan group categories and central feature values. Based on the central feature values of the fan group categories, content posting records corresponding to active time periods are extracted, and the time interval from the content posting time to the first interaction time is calculated as the interaction response speed, thus completing the evaluation of the impact of content style changes on fan interaction.
7. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that, The assessment results of the impact of content style changes on fan interaction are used to dynamically adjust the influencer contribution weight and determine the adjusted weight allocation scheme, including: Data on the impact of content style changes for each influencer on the response speed and engagement depth of fan interactions is extracted to obtain the initial value of the influencer's current contribution weight. The average time interval between the first fan interaction after an influencer publishes content is calculated as a feedback speed indicator, and the proportion of positive words in comments is counted as the emotional positivity value. Based on the feedback speed indicator and the emotional positivity value, if the feedback speed indicator is lower than a preset time threshold and the emotional positivity value exceeds a preset proportion threshold, the initial value of the contribution weight is multiplied by a preset amplification coefficient to obtain the adjusted contribution weight. Based on the adjusted contribution weight, the total weight of all influencers is calculated and normalized to form the adjusted weight allocation scheme.
8. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that, The process of optimizing platform revenue distribution based on interaction frequency fluctuations and determining the revenue distribution result includes: Calculate the purchase conversion rate for each content theme by dividing the number of purchasing users by the total number of interactive users to obtain the conversion ratio. If the conversion ratio exceeds a preset threshold, record the corresponding influencer ID and current revenue ratio. For influencers whose conversion ratio exceeds the preset threshold, multiply the current revenue ratio base by a fixed growth coefficient to obtain a new revenue ratio. Summarize the new revenue ratios and IDs of all influencers to form a revenue distribution result table.
9. A method for mining user behavior data in an e-commerce ecosystem according to claim 8, characterized in that, After generating the profit distribution results table, it includes: Extract the influencer identifier and new revenue ratio from the revenue distribution result table, and obtain the peak moment of the corresponding influencer's fan liking behavior as the liking peak; extract high-frequency words from the comment keyword distribution, and combine them with interactive topic category tags to form an interactive preference feature set; Calculate the correlation coefficient between the peak number of likes and the interaction preference feature set and the purchase conversion rate, and generate time-series correlation data containing the correlation coefficient; store the time-series correlation data, influencer identifiers and new revenue ratios in the database, and integrate them to form a training dataset containing time-series correlation and interaction preferences.
Citation Information
Patent Citations
E-commerce precision marketing intelligent system and method fused with AI personalized recommendation
CN120450807A
Method and system for implementing a cloud-based social media marketing method and system
US20140180788A1