User behavior data mining method for e-commerce ecosystem
By constructing a data set of interactive and display features, combining it with historical transaction records, and dynamically adjusting the influencer contribution weights, we can solve the problem of unfair income distribution caused by changes in influencer content styles, and achieve accurate evaluation and fair distribution.
Patent Information
- Application Number
- CN202511211656.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing technologies make it difficult to accurately evaluate the contribution of content to conversions and achieve fair profit distribution when the content style of influencers changes. They ignore the dynamic nature of content creation and the complexity of user behavior, resulting in deviations between evaluation results and actual contributions.
By collecting the graphic and text content of influencers, extracting the fluctuations in fan interaction frequency and comment data, building an interaction and display feature data set, and combining historical transaction records, determining the correlation mapping between interaction patterns and purchasing behaviors, analyzing the trends in fan trust and participation, dynamically adjusting the influencer contribution weight, and optimizing the platform's revenue distribution.
It achieves accurate evaluation in scenarios where the content style of influencers changes, ensures the fairness of revenue distribution between the platform and influencers, and promotes the creation of high-quality content.
Smart Images

Figure CN120705418A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a method for mining user behavior data for an e-commerce ecosystem. Background Art
[0002] In the field of digital marketing and content creation, studying the impact of influencers' content promotion on product sales conversion is crucial. This area not only concerns the distribution of revenue between platforms and creators but also directly impacts the healthy development of the e-commerce ecosystem and the optimization of user experience. The text and image content published by influencers can effectively guide user purchasing behavior. However, accurately evaluating the contribution of content to conversions and ensuring fair revenue sharing based on this contribution remain pressing challenges for the industry. While many current methods attempt to evaluate the value of influencer content through data analysis, these approaches often overlook the dynamic and complex nature of content creation, particularly when creators' styles change, and lack the ability to explore the deeper connections between content details and user behavior. Existing solutions often rely solely on superficial statistical analysis, failing to fully consider the potential impact of subtle differences in content presentation and shifting user interaction patterns on final conversions, resulting in a mismatch between evaluation results and actual contribution. Specifically, when fan interaction shifts from passive browsing to active commenting or sharing, user trust and engagement with the content significantly change. For example, after an influencer posts content promoting a certain skincare product, some users will simply browse the content and then purchase it directly, while others will ask questions in the comment section about the effectiveness and skin compatibility, and even proactively share their own user experiences. This shift from passive reception to active participation not only increases users' trust in the product, but also creates a secondary dissemination effect through interaction in the comment section, further amplifying the influence of the influencer's content. In addition, when users share content on other social platforms, this cross-platform dissemination behavior will generate deeper conversion value, but existing evaluation systems often find it difficult to accurately capture this complex interactive chain. Therefore, how to dynamically adjust the weight calculation of creators' contributions and ensure the fairness of revenue distribution between the platform and influencers in scenarios where the influencer's content style changes has become a key issue that needs to be addressed urgently. Summary of the Invention
[0003] The present invention provides a user behavior data mining method for an e-commerce ecosystem, which mainly includes:
[0004] Collect expert graphic and text content, extract fan interaction frequency fluctuations and comment data, record the proportion of product close-ups and background environment presentation data, and construct an interaction and display feature dataset; combine the interaction and display feature dataset with historical transaction records, extract purchase time points and product category data corresponding to fan interaction frequency, and determine the preliminary correlation mapping between interaction patterns and purchasing behavior; perform text processing on comment data, extract keyword frequency distribution, identify interaction data on user browsing time and sharing channel distribution, analyze fan behavior fluctuations in different time periods, and evaluate trend fluctuations in fan trust and engagement; combine fan trust and engagement trend fluctuations to evaluate the impact of content style changes on fan interaction; based on the evaluation results of the impact of content style changes on fan interaction, dynamically adjust the expert contribution weight and determine the adjusted weight distribution plan; through the adjusted weight distribution plan, analyze the key mining directions of interactive content themes and sharing channel distribution, extract the fan stay time distribution and forwarding propagation path related to purchase behavior, and obtain interaction frequency fluctuations under different background environment combinations; optimize platform revenue distribution based on interaction frequency fluctuations, determine revenue distribution results, and store fan interaction preferences, likes peaks, and interaction theme association data in the database.
[0005] Furthermore, the collection of influencer graphic content, extraction of fan interaction frequency fluctuations and comment data, recording of product close-up ratios and background environment presentation data, and construction of an interaction and display feature dataset include:
[0006] Collect the graphic and text content posted by experts, record the number of likes, collections and shares, combine the timestamp marks to obtain the temporal distribution of interactive behavior, calculate the difference between the publishing time and the interaction time as the interaction frequency fluctuation value, and count the number of comments and the distribution of word count; perform image segmentation on the graphic and text content, identify the product area and background area, calculate the ratio of the number of pixels in the product area to the number of pixels in the overall image as the product close-up ratio, and extract the color histogram and texture features of the background area as the environment matching data; establish a corresponding relationship matrix based on the interaction frequency fluctuation value and the product close-up ratio, mark the deep interaction label with the number of words in the comment exceeding the preset threshold, combine the main color value of the color histogram and the texture feature entropy value of the environment matching data, group the data through the clustering algorithm, integrate the grouping results and the corresponding relationship matrix to form an interaction and display feature data set.
[0007] Furthermore, the interaction and display feature dataset is combined with historical transaction records to extract purchase time points and product category data corresponding to fan interaction frequency, and determine a preliminary correlation mapping between interaction patterns and purchase behaviors, including:
[0008] Extract purchase timestamps and product category identifiers from historical transaction records, match fan accounts in the interaction and display feature dataset, obtain interaction frequency data sequences and purchase behavior sequences, identify time periods when the interaction frequency exceeds a preset threshold as interaction-intensive periods, identify time periods when the number of purchases exceeds a preset threshold as purchase peak periods, calculate the ratio of the intersection duration of the two time periods to the union duration as the overlap, mark strongly associated user groups with an overlap exceeding a preset threshold, and extract a numerical sequence of the product feature ratio of the strongly associated user groups during the interaction-intensive period; construct a category migration matrix based on the product feature ratio numerical sequence, record the frequency of product category conversion, and generate a preliminary association mapping table of interaction patterns and purchase behaviors through an association rule mining algorithm.
[0009] Furthermore, extracting a numerical sequence of product feature ratios of the strongly associated user group during the intensive interaction period includes:
[0010] Product images viewed during the intensive interaction period are extracted from the strongly associated user group, and the ratio of the number of pixels in the product area of each image to the number of pixels in the entire image is calculated to form a numerical sequence of product close-up ratios; for the numerical sequence of product close-up ratios, the product category corresponding to each value in the sequence is counted, and a category distribution vector is constructed to form the numerical sequence of product close-up ratios.
[0011] Furthermore, the text processing of the comment data, extraction of keyword frequency distribution, and identification of interactive data on user browsing time and sharing channel distribution include:
[0012] The comment data is segmented, the number of occurrences of each word is counted, and a keyword frequency distribution table is constructed; based on the user identifier, the corresponding user's page residence time is extracted as browsing time data, and the number of times the content is shared on each platform is obtained to form sharing channel distribution data; the time period is divided into hours, and the total number of likes in each time period is calculated and divided by the number of independent likers as the like concentration value, and the difference in the number of comments in adjacent time periods is counted as the comment activity change value; through the high-frequency words in the keyword frequency distribution table, the timestamps corresponding to the comments are matched, and the browsing time data and the sharing channel distribution data are combined to generate a temporal interaction feature set containing like concentration and comment activity change values.
[0013] Furthermore, the above-mentioned combination of fan trust and engagement trend fluctuations to evaluate the impact of content style changes on fan interaction includes:
[0014] Extract the numerical values of the fan trust and participation trends of each user in different time periods, combine them to form user behavior feature vectors, and construct a fan behavior comprehensive portrait matrix containing time-series change characteristics; perform principal component analysis on the fan behavior comprehensive portrait matrix, retain the principal components whose cumulative contribution rate exceeds a preset threshold, and obtain a reduced-dimensional feature matrix; group the users in the reduced-dimensional feature matrix through a hierarchical clustering algorithm to obtain fan group categories and central eigenvalues; based on the central eigenvalues of the fan group categories, extract the content release records corresponding to the active time period, calculate the time interval from the content release moment to the first interaction moment as the interactive response speed, and complete the evaluation of the impact of content style changes on fan interaction.
[0015] Furthermore, the weight of influencer contribution is dynamically adjusted based on the evaluation results of the impact of content style changes on fan interaction, and the adjusted weight distribution scheme is determined, including:
[0016] Extract data on the impact of each influencer's corresponding content style changes on the response speed and participation depth of fan interactions, and obtain the initial value of the influencer's current contribution weight; calculate the average time interval between the first fan interaction after the influencer releases the content as the feedback speed indicator, and count the proportion of positive words in the comments as the emotional positivity value; based on the feedback speed indicator and the emotional positivity value, if the feedback speed indicator is lower than the preset time threshold and the emotional positivity value exceeds the preset proportion threshold, multiply the initial value of the contribution weight by the preset increase coefficient to obtain the adjusted contribution weight; based on the adjusted contribution weight, calculate the sum of all influencer weights, and normalize them to form an adjusted weight distribution plan.
[0017] Furthermore, the adjusted weight distribution scheme is used to analyze the key mining directions of interactive content themes and sharing channel distribution, and extract the fan stay time distribution and forwarding propagation paths related to purchasing behavior, including:
[0018] Through the adjusted weight distribution scheme, high-weight influencers whose weight ratio exceeds the preset threshold are screened, the subject tags and keywords of their published content are extracted, the total amount of interactions under each topic is counted, the distribution of the number of times the content is shared on each platform is obtained, and the subject category with the highest interaction volume and the channel with the most sharing times are determined; content records that meet the subject categories and channels are screened from the interaction and display feature data set, and the user stay time data of users who browse the content that meets the subject categories and channels and then make purchases are extracted, and the time intervals are divided; the propagation chain of the content record from the original release to the forwarding at each level is tracked, and the user ID and timestamp of each forwarding level are recorded to form a forwarding propagation path.
[0019] Furthermore, the optimization of platform revenue distribution according to interaction frequency fluctuations and determination of revenue distribution results include:
[0020] Calculate the purchase conversion rate under each content theme, and obtain the conversion ratio by dividing the number of purchasing users by the total number of interactive users; if the conversion ratio exceeds the preset threshold, record the corresponding expert ID and current revenue ratio; for experts whose conversion ratio exceeds the preset threshold, multiply the current revenue ratio base by a fixed increase coefficient to obtain a new revenue ratio; summarize the new revenue ratios and IDs of all experts to form a revenue distribution result table.
[0021] Furthermore, after the income distribution result table is formed, it includes:
[0022] Extract the expert ID and the new revenue ratio from the revenue distribution result table, and obtain the concentrated moment of the corresponding expert's fans' likes behavior as the likes peak; extract high-frequency words in the comment keyword distribution, and combine them with the interactive topic category labels to form an interactive preference feature set; calculate the correlation coefficient between the likes peak and the interactive preference feature set and the purchase conversion rate, and generate time series correlation data containing the correlation coefficient; store the time series correlation data, the expert ID and the new revenue ratio in a database, and integrate them to form a training data set containing time series correlation and interactive preference.
[0023] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:
[0024] The present invention discloses a user behavior data mining method for an e-commerce ecosystem. By collecting graphic and text content of influencers, extracting fan interaction frequency fluctuations and comment data, recording the product close-up ratio and background environment matching presentation data, an interaction and display feature data set is constructed. The data set is combined with historical transaction records to extract purchase time points and product category data, and determine the correlation mapping between interaction patterns and purchase behaviors. Further analyze the correspondence between the like-intensive period, comment peak period and order generation situation, identify the impact of the product close-up ratio on the purchase conversion rate, and form a behavior mapping framework. Extract keywords through text processing, analyze fan behavior fluctuations, and evaluate trust and participation trends. Finally, a fan behavior portrait is constructed, the impact of content style changes on interaction is evaluated, the influence of influencer contribution weights is dynamically adjusted, and the platform revenue distribution mechanism is optimized to provide data support for promoting high-quality content creation. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 The present invention is a flowchart of a method for mining user behavior data for an e-commerce ecosystem.
[0026] Figure 2 Schematic diagram of a user behavior data mining method for an e-commerce ecosystem according to the present invention. DETAILED DESCRIPTION
[0027] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0028] like Figure 1-2 In this embodiment, a user behavior data mining method for an e-commerce ecosystem may specifically include:
[0029] S101. Collect the graphic and text content of influencers, extract the fluctuation of fan interaction frequency and comment data, record the proportion of product close-ups and background environment presentation data, and build an interaction and display feature dataset.
[0030] The image and text content posted by influencers is collected and fan interaction data is extracted. The number of likes, favorites, and shares for each piece of content is recorded. The temporal distribution of interactive behavior is obtained through timestamp tags. The interaction frequency fluctuation value is calculated based on the difference between the posting time and the interaction time. The number of comments is also extracted and the word count distribution of the comments is calculated. Image segmentation is performed on the collected image and text content. The product area and background area are identified using the YOLO algorithm. The ratio of the number of pixels in the product area to the total number of pixels in the image is calculated to obtain the proportion of product close-ups. The color histogram and texture features of the background area are extracted as environmental matching data. A correspondence matrix is established based on the interaction frequency fluctuation value and the proportion of product close-ups. If the word count of a comment exceeds a preset threshold, it is marked as a deep interaction label. The main color value of the color histogram and the entropy value of the texture features in the environmental matching data are used as environmental complexity indicators. The data containing deep interaction labels, correspondence matrix, and environmental complexity indicators are grouped using the K-means algorithm. The K-means grouping results and the corresponding relationship matrix were integrated, and the interaction frequency fluctuation value, product close-up ratio, depth of interaction label and environment complexity index were combined into a four-dimensional feature vector. The feature vector was labeled according to the grouping results to form an interaction and display feature dataset that includes display style category labels.
[0031] For example, when collecting graphic and text content from influencers, the extraction of interactive data involves multi-dimensional information acquisition.
[0032] Specifically, the number of likes, favorites, and shares reflects the user's level of engagement with the content, while the timestamp records the exact moment of each interaction. By calculating the difference between the posting time and the interaction time, we can obtain the interaction frequency fluctuation value, which reflects the trend of content popularity over time. The statistical distribution of comment word count reveals the depth of user engagement; longer comments often indicate deeper reflection and emotional investment in the content. Image segmentation is a key technology for identifying the proportion of product close-ups. The YOLO algorithm, a classic method in object detection, uses convolutional neural networks to extract features from images, dividing the image into multiple grids. Each grid is responsible for predicting whether the target object is contained in that area. In influencer image content, the YOLO algorithm identifies the bounding box of the product and calculates the ratio of the number of pixels in this area to the total number of pixels in the image, thereby determining the proportion of product close-ups. The color histogram of the background area is obtained by statistically analyzing the pixel distribution of different color channels, and texture features are extracted using methods such as the gray-level co-occurrence matrix. Together, these data constitute a quantitative representation of the environmental context.
[0033] It should be noted that the establishment of the correspondence matrix reveals the intrinsic connection between interactive behavior and visual presentation. When the number of words in a comment exceeds the preset threshold, the system marks it as a deep interaction label. This labeling method can distinguish between shallow browsing and deep participation. In the calculation of the environmental complexity index, the main color value of the color histogram reflects the color richness of the background, and the entropy value of the texture feature quantifies the visual complexity of the background. The K-means algorithm clusters according to these features, classifying similar content into the same group, each group representing a specific display style.
[0034] In one embodiment, the construction of a four-dimensional feature vector enables structured representation of data. Interaction frequency fluctuations serve as temporal features, product close-up ratios serve as spatial features, depth of interaction tags serve as user engagement features, and environmental complexity indicators serve as visual presentation features. These four dimensions collectively describe the complete characteristics of the influencer's graphic content. The feature vectors are annotated using the K-means grouping results, resulting in a dataset containing not only the original feature values but also display style category labels.
[0035] S102: Combine the interaction and display feature dataset with historical transaction records, extract purchase time points and product category data corresponding to fan interaction frequency, and determine a preliminary correlation mapping between interaction patterns and purchase behaviors.
[0036] Purchase timestamps and product category identifiers are extracted from historical transaction records. User identifiers are then matched with fan accounts in the interaction and display feature dataset to obtain a data series of interaction frequency and corresponding purchase behavior for each user. The time difference between the interaction behavior and the purchase time is calculated to form time interval distribution data. Based on this time interval distribution data, periods where the interaction frequency exceeds a preset threshold are identified as high-interaction periods, and periods where the purchase frequency exceeds a preset threshold are identified as peak purchase periods. The overlap is calculated as the ratio of the intersection duration of these two periods to the union duration. If the overlap exceeds the preset threshold, the user group is labeled as a strongly associated user group. A numerical series of product feature ratios viewed by this strongly associated user group during the high-interaction period is extracted. Based on this numerical series of product feature ratios, the corresponding product categories are recorded when the feature ratio changes. A category transition matrix is constructed, in which rows represent the product category at the previous moment, columns represent the product category at the next moment, and element values represent the frequency of transitions. The category transition matrix is processed using an association rule mining algorithm to generate a set of rules encompassing interaction frequency, feature ratio changes, and product category transitions. The support and confidence of each rule are extracted from the rule set. The support is the ratio of the frequency of rule occurrence to the total frequency, and the confidence is the ratio of the number of times the rule is established to the number of times the previous item occurs. Valid rules are screened based on the support and confidence, and a preliminary association mapping table is established that includes the corresponding relationship between the interaction frequency range, the range of variation of the close-up ratio, and the purchase probability of the product category.
[0037] For example, obtaining time interval distribution data is the basis for understanding user behavior patterns.
[0038] Specifically, when users like, comment, or share content posted by influencers on social platforms, the timestamps of these interactions are accurately recorded. These timestamps are then compared with the timestamps of the user's subsequent purchases, and the calculated time difference forms a time interval distribution. This distribution reveals the temporal patterns from when users engage with content to when they make a purchase decision. Short time intervals indicate strong immediate conversion effects, while long time intervals reflect sustained influence. Identifying periods of intense interaction and peak purchase times involves a threshold determination mechanism.
[0039] In one embodiment, the number of interactions within each time window is counted. When the interaction frequency for a certain continuous period consistently exceeds 1.5 times the historical average, that period is marked as a period of high interaction. Peak purchasing periods are determined using a similar approach, by analyzing the temporal distribution characteristics of transaction records. The calculation of overlap reflects the degree of correlation between the two periods. When a user's high-frequency interaction period and concentrated purchasing period overlap significantly on the timeline, it indicates that content interaction has directly driven purchasing decisions.
[0040] It's important to note that the numerical sequence of product close-up percentages captures changes in visual preferences across user browsing behavior. When users transition from browsing life scenes with a low percentage of close-up images to product detail images with a high percentage of close-up images, this shift suggests a strengthening of purchase intent. The category transition matrix is constructed based on this pattern of change. Each element in the matrix records the frequency with which users shift their gaze from one product category to another. High-frequency transition paths reveal the correlations between product categories and user preferences. When processing the category transition matrix, association rule mining algorithms scan all possible product combinations to discover frequently occurring patterns.
[0041] For example, if the algorithm finds that the probability of "purchasing the clothing item" reaches 0.8 under the antecedent condition "browsing clothing content and the proportion of close-up photos increases from 30% to 70%," a high-confidence rule is formed. Support indicates the prevalence of the rule across all transactions, while confidence measures the reliability of the rule.
[0042] In one possible implementation, the creation of a preliminary association mapping table integrates multi-dimensional behavioral characteristics. Interaction frequency is divided into three levels: low, medium, and high. The range of variation in the proportion of close-up images is divided into three patterns: slow growth, rapid growth, and stable high. Each combination corresponds to a different purchase probability. This mapping relationship enables the system to predict a user's likelihood of purchasing a specific product category based on their current interaction status and browsing patterns, thereby achieving accurate behavioral prediction and personalized recommendations.
[0043] Extract data on periods of high likes and peak comments from the interaction and display feature datasets, match the order generation situation in the corresponding time window in the historical transaction records, identify changes in purchase conversion rates when the proportion of product close-ups is higher than the average, analyze the correspondence between background environment matching style and purchase preferences for specific product categories, and form a behavior mapping framework.
[0044] From the interactive and display feature dataset, continuous periods with more than twice the average number of likes are identified as intensive likes periods, and time intervals with more than a preset threshold number of comments are extracted as comment peak periods. Based on the timestamps of the intensive periods and peak periods, order data within the corresponding time windows are searched in the historical transaction records, and the order generation frequency of each time window is calculated. For time windows with a high frequency of order generation, the close-up percentage data of the product images browsed by users in the window are extracted, the average close-up percentage is calculated, and product display records with a close-up percentage higher than the average are screened out. The number of purchases generated in these records is counted and divided by the total number of display records in the group to obtain the purchase conversion rate value under the condition of high close-up percentage. Based on the product display records that generate purchases, the RGB color value distribution and scene element labels of the background area of the corresponding image are extracted. The color distribution data is clustered using the K-means algorithm to obtain the matching style category. The number of purchases of each product category under each matching style is counted, and a corresponding relationship matrix is constructed with the behavior matching style, the column as the product category, and the element value as the purchase number. The starting time of the period with a high number of likes is used as the interaction trigger point timestamp, and the purchase conversion rate under the condition of high close-up ratio is used as the quantitative indicator of product display effect. The difference between the interaction trigger point timestamp and the actual purchase behavior timestamp is calculated as the purchase response delay. The trigger point timestamp, display effect quantitative indicator, response delay duration and corresponding relationship matrix data are integrated to form a behavior mapping framework that includes temporal association and visual preference.
[0045] For example, the identification of periods of high likes and peak comments is based on statistical principles.
[0046] Specifically, the system counts the number of likes for each time unit in the historical data and calculates the mean and standard deviation of the number of likes. When the number of likes in a certain continuous period exceeds twice the mean, the period is marked as a period of intensive likes. A similar method is used to determine the peak period of comments, but the focus is on the temporal distribution characteristics of the number of comments. This dual-dimensional interaction period identification can capture the moments of concentrated attention of user groups, which often indicate a higher purchase intention. The calculation of order generation frequency involves the precise matching of time windows.
[0047] In one embodiment, time windows are divided into hourly units, and the number of orders within each window is counted. If the peak like-posting period is between 2:00 PM and 4:00 PM, the system extracts order data from this time period and the hour before and after to calculate the order generation frequency. This expanded time window approach accounts for the decision delay between user interest and purchase, making subsequent conversion rate calculations more accurate.
[0048] It should be noted that calculating the purchase conversion rate under conditions of a high close-up ratio has important commercial value. The close-up ratio reflects the degree of product detail displayed. When the ratio exceeds the average, it indicates that users are more concerned with the product itself rather than the usage scenario. The system calculates the proportion of actual purchases generated by these high close-up ratio displays. The resulting conversion rate directly reflects the influence of detailed display on purchasing decisions. This metric provides content creators with clear guidance on display strategies. The K-means algorithm plays a key role in processing background color distribution. The algorithm first converts the RGB values of the background area of each image into feature vectors. Each vector contains the distribution information of the three channels: red, green, and blue. Through iterative calculations, the algorithm classifies similar color distributions into the same category, forming different matching styles.
[0049] For example, a warm-toned background is categorized as cozy, while a cool-toned background is categorized as minimalist. Scene element labels, such as "outdoor," "home," and "office," are obtained through image recognition. These labels, along with color clustering results, define the style categories.
[0050] In one possible implementation, the construction of a correspondence matrix embodies the integration of multidimensional data. The rows of the matrix represent different style categories, and the columns represent product categories such as clothing, accessories, and home furnishings. Each matrix element records the number of purchases of a specific product category within a specific style. When the number of purchases of clothing products within a cozy style context is significantly higher than that of other styles, this correspondence is quantified and recorded, providing data support for subsequent personalized recommendations. The formation of the behavior mapping framework integrates features from both temporal and visual dimensions. The interaction trigger point timestamp marks the onset of user interest, the purchase response delay quantifies the time interval from interest to action, and the display effect quantitative indicator reflects the effectiveness of the visual presentation. The correspondence matrix reveals the inherent connection between style preferences and product selection. This framework enables the system to predict purchasing behavior under specific interaction patterns, enabling precise content delivery and product recommendations.
[0051] S103. Perform text processing on the comment data, extract keyword frequency distribution, identify interactive data on user browsing time and sharing channel distribution, analyze the fluctuations in fans' behavior in different time periods, and evaluate the trend fluctuations in fans' trust and participation.
[0052] Perform word segmentation on the comment data, count the occurrences of each word after removing stop words, construct a keyword frequency distribution table, extract the page停留 time of the corresponding user as the browsing duration data according to the user identifier in the preliminary association mapping, and at the same time obtain the number of times the content is shared on each platform to form the sharing channel distribution data. Use the high-frequency words in the keyword frequency distribution table to match the timestamps corresponding to the comments containing these words, divide the time period by hour, calculate the total number of likes in each time period divided by the number of independent users who liked in that period to obtain the like concentration value, and count the difference in the number of comments in adjacent time periods to determine the comment activity change value. Identify the time periods exceeding the average value in the like concentration value sequence as high-concentration time periods, mark the positive and negative conversion points of the comment activity change value as activity fluctuation points, match the positive and negative words in the keywords through a sentiment dictionary, and if the proportion of the frequency of positive words exceeds the preset threshold and the cumulative value of the browsing duration in the corresponding time period continues to increase, the trust index for that time period is recorded as positive. Calculate the number of sharing platform types as the participation breadth value for the sharing channel distribution data, combine the occurrence frequency of high-concentration time periods and the density of activity fluctuation points, and calculate the change rate of the trust index sequence and the dispersion degree of the participation breadth value sequence through a sliding window method to obtain the trend fluctuation evaluation result including the trust change rate and the participation sense fluctuation amplitude for each time period.
[0053] Exemplarily, the word segmentation of the comment text involves the basic link of natural language processing.
[0054] Specifically, the system first segments the original comment text, decomposing the continuous character sequence into meaningful word units. During the process of removing stop words, the system filters out high-frequency but meaningless words such as "of", "already", "in", etc., and retains words with substantial evaluation significance such as "good quality", "fast logistics", "correct color", etc. The construction of the keyword frequency distribution table is completed by counting the occurrences of each word in all comments. High-frequency words often reflect the product features or service elements that users are most concerned about. The extraction of the browsing duration data is based on the accurate record of the user behavior log. When a user enters a certain product detail page, the entry timestamp is recorded. When the user leaves the page or performs other operations, the departure timestamp is recorded. The difference between the two is the single browsing duration. The sharing channel distribution data is obtained by tracking the click behavior of the sharing button. The system records the number of times the content is shared on different platforms such as WeChat, Weibo, Xiaohongshu, etc., forming a multi-dimensional propagation path map.
[0055] In one embodiment, the calculation of like concentration reflects the degree of aggregation of user interaction behavior. If 100 likes come from 100 different users, the like concentration is 1, indicating dispersed interaction. If 100 likes come from 50 users, and some users like multiple times, the concentration is 2, indicating repeated interaction among a core group of fans. The change in comment activity is calculated by calculating the difference in the number of comments in adjacent time periods. A positive value indicates an increase in activity, while a negative value indicates a decrease in activity. The magnitude of the change reflects the dynamic evolution of content popularity.
[0056] It's important to note that the sentiment lexicon plays a key role in identifying review trends. This lexicon contains pre-annotated positive terms such as "like," "satisfied," and "recommended," as well as negative terms such as "disappointed," "return," and "poor quality." By matching keywords in reviews with the sentiment lexicon, the ratio of positive term frequency to total sentiment term frequency is calculated. When this ratio exceeds 0.7, the review is considered predominantly positive. The cumulative browsing time over the corresponding period is also observed. A sustained increase indicates that users are willing to invest more time in understanding the product, indicating a positive development in trust. The engagement breadth value is calculated based on the diversity of sharing platforms. When content is shared only on a single platform, the engagement breadth value is low; when content spreads across multiple platforms, the engagement breadth value increases, reflecting the content's cross-platform influence. A sliding window method smooths trends. The window contains data from multiple consecutive time periods. The rate of change is calculated by calculating the linear regression slope of the data within the window. A positive slope indicates an upward trend, while a negative slope indicates a downward trend. The trend fluctuation assessment results integrate dynamic indicators from multiple dimensions. The rate of change in trust reflects the direction and speed of user attitude evolution. The amplitude of engagement fluctuation is calculated by calculating the standard deviation of the engagement breadth value series. A larger standard deviation indicates more erratic user engagement. This comprehensive assessment method comprehensively reflects the behavioral characteristics and psychological state changes of fan groups, providing content creators with precise user insights.
[0057] S104. Combine the fluctuations in fan trust and engagement trends to evaluate the impact of changes in content style on fan interaction.
[0058] By using trend fluctuation data on fan trust and engagement, we extract the numerical sequences of trust and engagement intensity for each user at different time periods. The numerical values of the two sequences are combined to form a user behavior feature vector. The feature vectors of all users are then integrated to construct a comprehensive fan behavior profile matrix that includes temporal variation characteristics. Principal component analysis is performed on the fan behavior comprehensive profile matrix to reduce its dimensionality. The principal components whose cumulative contribution rates exceed a preset threshold are retained to obtain a reduced feature matrix. A hierarchical clustering algorithm is then used to group users in the reduced feature matrix to obtain fan group categories with different behavioral patterns and the central eigenvalues of each category. Based on the central eigenvalues of the fan group categories, the content release records corresponding to the active periods of each category are identified. The image-text ratio, color tone, and text length are extracted from the content records as content style parameters. The time interval from the content release moment to the first interaction moment is calculated as the interactive response speed. The number of words in the comments, the number of shares, and the browsing time are weighted and summed according to the preset weight coefficients to obtain the engagement depth index. For adjacent time periods when content style parameters change, calculate the difference in interactive response speed before and after the change. Subtract the participation depth index before the change from the participation depth index after the change and then divide it by the index before the change to obtain the participation depth change rate. Establish an evaluation result table containing three columns of data: style parameter change type, response speed difference, and participation depth change rate, to complete the evaluation of the impact of content style changes on fan interaction.
[0059] Exemplarily, the construction of the comprehensive fan behavior portrait matrix is based on the integration of multi-dimensional time series data.
[0060] Specifically, the trust value series records the changes in users' recognition of content creators over time, with higher values indicating stronger trust. The engagement intensity series reflects users' willingness to actively interact, quantified by the frequency and depth of behaviors such as likes, comments, and shares. By aligning these two series in time and combining them, the resulting feature vector encompasses both attitudinal and behavioral dimensions, comprehensively characterizing user interaction. Principal component analysis plays a key role in dimensionality reduction. The original behavioral feature vector may contain data from dozens of dimensions, and direct processing would result in excessive computational complexity. Principal component analysis uses a linear transformation to project the original features onto the direction of maximum variance, preserving the main patterns of variation in the data. When the cumulative contribution rate reaches 85%, it means that the retained principal components contain 85% of the information in the original data, achieving both data compression and preservation of key features.
[0061] In one embodiment, a hierarchical clustering algorithm calculates user similarities to gradually merge similar user groups. The algorithm first treats each user as a separate category and then calculates the Euclidean distance between any two user feature vectors. Smaller distances indicate more similar behavior patterns. By continuously merging the closest categories, a tree structure is formed, which is then segmented based on a preset distance threshold to create different fan group categories. The central eigenvalue of each category is calculated by calculating the mean of all user features within that category.
[0062] It should be noted that the extraction of content style parameters involves quantification of multiple dimensions. The image-to-text ratio is calculated by calculating the percentage of image area to the total content layout. A high image-to-text ratio indicates a dominant visual element, while a low ratio indicates a predominantly textual content. The color palette is determined by extracting the RGB values of the image's primary hue. Warm tones create a warm atmosphere, while cool tones convey a sense of professionalism. Copy length directly counts the number of words. Short copy facilitates quick scanning, while long copy provides detailed information. The calculation of interactive response speed reflects the immediate appeal of the content. The shorter the interval between content publishing and the first user interaction, the faster the content captures user attention. In the weighted calculation of the engagement depth metric, the number of words in comments reflects the depth of thought and is weighted at 0.5; the number of shares reflects the willingness to spread and is weighted at 0.3; and the viewing time indicates the level of attention and is weighted at 0.2. This weighting balances the depth of interaction and takes into account the diversity of behavior. The construction of the evaluation results table provides a quantitative representation of the influence relationship. When the content style changes from a high image-to-text ratio to a low image-to-text ratio, the corresponding difference in response speed is recorded. A positive value indicates a faster response, while a negative value indicates a slower response. The calculation of the participation depth change rate adopts the relative change method, eliminating the impact of differences in the base numbers of different user groups.
[0063] S105. Based on the evaluation results of the impact of content style changes on fan interaction, dynamically adjust the influencer contribution weights and determine the adjusted weight distribution plan.
[0064] Based on the evaluation results of the impact of content style changes on fan interaction, the corresponding interactive response speed difference and engagement depth change rate data for each influencer are extracted to obtain the influencer's current initial contribution weight. The average time interval between the first fan interaction after the influencer's content is calculated as the feedback speed indicator. The sentiment positivity value is obtained by matching the proportion of positive words in the comments using the sentiment dictionary. For the feedback speed indicator and sentiment positivity value, combined with the interactive response speed difference and engagement depth change rate, if the feedback speed indicator is below the preset time threshold and the sentiment positivity value exceeds the preset ratio threshold, the initial contribution weight is multiplied by the preset amplification factor to obtain the adjusted contribution weight. Otherwise, the initial contribution weight remains unchanged. Using the adjusted contribution weight, the sum of all influencer contribution weights is calculated. Each influencer's contribution weight is divided by the sum to obtain the normalized weight percentage, forming an adjusted weight distribution scheme that includes the influencer ID and the corresponding weight percentage.
[0065] For example, the dynamic adjustment mechanism of the expert contribution weight is based on multi-dimensional behavioral data evaluation.
[0066] Specifically, the initial contribution weight reflects the influencer's historical performance and fundamental influence on the platform. It is typically calculated based on a comprehensive analysis of the influencer's follower count, historical content quality, and conversion rates. This initial value serves as a baseline for adjustments, ensuring the continuity and rationality of weight changes. The calculation of the feedback speed metric involves precise measurement of time. When an influencer publishes new content, the time of the first user interaction is recorded, and the difference from the publishing time is calculated. The influencer's feedback speed metric is calculated by calculating the average first interaction time after multiple publications. A shorter time interval indicates that the content has strong immediate appeal and can quickly spark user interest.
[0067] In one embodiment, the positivity value is derived using sophisticated sentiment lexicon technology. After the review text is segmented, it is matched against a pre-set positive and negative sentiment lexicon. Positive terms include words expressing positive attitudes, such as "like," "wonderful," and "recommended," while negative terms include negative expressions, such as "disappointed," "average," and "not worth it." By calculating the ratio of the number of positive terms to the total number of sentiment terms, a sentiment positivity value between 0 and 1 is obtained.
[0068] It's important to note that the introduction of the interactive response speed difference and engagement depth change rate enhances the comprehensiveness of the evaluation. The interactive response speed difference compares the change in response time before and after content style adjustments, reflecting the impact of style changes on user engagement. The engagement depth change rate measures the actual effectiveness of content optimization from the perspective of user engagement. These two metrics, along with feedback speed and sentiment positivity, form a four-dimensional evaluation system. The conditional judgment for weight adjustment reflects the design philosophy of the incentive mechanism. When feedback speed exceeds the preset threshold and sentiment positivity exceeds the standard, it indicates that the influencer's content is both rapidly engaging users and garnering positive reviews. This high-quality performance is recognized by increasing their weight. The preset increase coefficient is typically set between 1.1 and 1.5, ensuring the significance of the adjustment while avoiding excessive fluctuations. Influencers who do not meet the promotion criteria retain their original weights, maintaining system stability. Normalization ensures the rationality and comparability of weight distribution. By calculating the sum of all influencers' adjusted weights and dividing each influencer's weight by this sum, the resulting weight percentage always sums to 1. This approach allows the relative importance of different influencers to be reflected while preventing the weight values from increasing indefinitely. The resulting weight distribution scheme not only reflects the influencer's real-time performance but also maintains a balanced overall distribution structure.
[0069] S106. Analyze the key mining directions of interactive content themes and sharing channel distribution through the adjusted weight distribution scheme, extract the fan stay time distribution and forwarding propagation path related to purchasing behavior, and obtain the interaction frequency fluctuations under different background environments.
[0070] Using the adjusted weight distribution scheme, we screen high-powered influencers whose weights exceed a preset threshold. We extract the topic tags and keywords for their published content, count the total interactions under each topic, and obtain the distribution of the number of times the content has been shared across platforms. We then identify the topic categories with the highest interactions and the channels with the most shares as key mining targets. Based on the topic categories and channels identified as key mining targets, we screen the initial dataset for content records that meet these topic categories and are distributed through the designated channels. We extract the user dwell time data for those who browse these contents and subsequently make a purchase. This data is sorted by duration and divided into short, medium, and long intervals. We then track the complete content dissemination chain from its original publication to forwarding at all levels, recording the user ID and timestamp of each forwarding level to form a forwarding propagation path. For each node content and dwell time interval data in this forwarding propagation path, we extract the corresponding background environment matching parameters, including color tone values and scene type labels. We then group and count the frequency of interactive behaviors within each dwell time interval according to the background environment matching type. We calculate the difference in interaction frequency between adjacent time periods to obtain interaction frequency fluctuation data, and form a record of the corresponding relationship between background matching type and interaction frequency fluctuation characteristics.
[0071] For example, the screening mechanism for high-powered talents reflects the priority principle of resource allocation.
[0072] Specifically, when an influencer's weight exceeds the preset threshold of 5%, it indicates strong content creation and user influence on the platform. Topic tags are extracted using natural language processing technology. The system identifies noun phrases and trending topics within the content, such as "outfit sharing," "makeup tutorials," and "digital reviews." These tags directly reflect the core attributes of the content. Interaction statistics encompass a comprehensive calculation of various user behaviors.
[0073] In one embodiment, the system counts likes as 1 point, comments as 3 points, and shares as 5 points, using a weighted sum to determine the total number of interactions for each topic. Sharing channel distribution is determined by tracking click logs for share buttons, recording the number of times content is shared on platforms such as WeChat, Weibo, Xiaohongshu, and Douyin. When the total number of interactions for the "outfit sharing" topic reaches 100,000 points, and the number of shares on Xiaohongshu accounts for 60% of the total shares, this topic and channel combination becomes a key area for exploration.
[0074] It should be noted that the dwell time intervals are based on statistical patterns of user behavior. First, the dwell time of all users who made a purchase was collected and sorted from smallest to largest, with the 33rd and 67th percentiles used as interval boundaries. The short interval, typically 0-30 seconds, indicates an impulsive purchase after a quick glance; the medium interval, 30-120 seconds, reflects a rational decision after moderate consideration; and the long interval, exceeding 120 seconds, represents a cautious purchase after in-depth research. This categorization method considers the actual data distribution while also providing clear business implications. Tracking the forwarding propagation path reveals the viral nature of content. When user A posts original content, user B forwards it and adds a comment, and user C then forwards it from user B, forming a chain of transmission from A to B to C. The user ID, forwarding timestamp, and textual content added during the forwarding are recorded for each node. By analyzing the length and number of branches in the propagation path, we can identify content features with high viral value. Extracting background environment parameters involves the application of computer vision technology. Color tone values are calculated by calculating the RGB values of the dominant colors in an image. For example, warm tones may appear as orange-reds with high R values. Scene type labels are determined by image recognition models, classifying backgrounds into typical scenes such as "home," "outdoor," "office," and "café." The combination of these parameters creates a unique visual style identity.
[0075] In one possible implementation, the calculation of interaction frequency fluctuations reflects the temporal changes in user activity. The system counts the number of interactions for each background type on an hourly basis and calculates the difference between the number of interactions between adjacent hours to generate the fluctuation data. For example, when content with warm tones in a home setting sees a surge in interaction frequency between 8 and 9 pm, while content with cool tones in an office setting is more active on weekday mornings, this correspondence provides data support for optimizing content release timing.
[0076] S107. Optimize the platform revenue distribution according to the interaction frequency fluctuation, determine the revenue distribution result, and store the fan interaction preferences, likes peaks, and interaction theme-related data in the database.
[0077] Based on the interaction frequency fluctuation data, the purchase conversion rate under each content theme is calculated, and the conversion ratio is obtained by dividing the number of purchasing users by the total number of interactive users. If the ratio exceeds the preset threshold, it is determined that the interactive content theme has a high promoting effect on the purchase behavior, and the corresponding expert ID and current income ratio are recorded. For experts with a high promoting effect, their current income ratio base is obtained, and the base is multiplied by a fixed increase coefficient to obtain the income increment. The base is added to the increment to obtain the new income ratio. The adjusted income ratio, expert ID and adjustment timestamp are stored, and the data of all experts are summarized to form an income distribution result table. Using the income distribution result table, the ID, new income ratio and adjustment timestamp of each expert are extracted, and the time when the corresponding expert's fans' likes behavior is concentrated is obtained from the historical data as the likes peak, the comment keyword distribution as the interaction preference, and the interaction theme category label. These data are combined with the purchase conversion rate and stored in the database to form a training basic data set. By regularly reading the basic training data set, extracting the peak moments of fan likes and the frequency of interactive preference keywords as feedback indicators, calculating the Pearson correlation coefficient between these indicators and the purchase conversion rate, and comparing the correlation coefficient differences between adjacent periods, we can obtain the changing trend data of the correlation between fan feedback and purchase behavior, which is used for subsequent optimization of dynamic adjustment rules.
[0078] For example, the calculation of purchase conversion rate is a core indicator for evaluating the value of content.
[0079] Specifically, the system calculates the conversion rate by dividing the total number of users who viewed a specific topic within a specific time window by the number of users who actually purchased it. For example, if 1,000 users viewed a beauty tutorial, and 80 of them purchased related products within 72 hours of viewing, the conversion rate would be 8%. If the preset threshold is 5%, the topic is considered highly stimulating. This dynamic adjustment mechanism for the revenue ratio reflects the principle of incentive-based distribution.
[0080] In one embodiment, the influencer's initial revenue ratio is determined based on the number of their fans and historical performance, typically between 1% and 10% of the platform's total revenue. When content created by an influencer is identified as having a high promotional effect, the system uses a fixed increase coefficient of 1.2 for calculation. If the influencer's current revenue ratio is 3%, the incremental revenue is 3% × 0.2 = 0.6%, and the adjusted new revenue ratio is 3.6%. This progressive adjustment not only protects the interests of high-quality creators, but also maintains the balance of overall distribution.
[0081] It should be noted that the identification of likes peaks is based on time series analysis principles. A 24-hour day is divided into 48 half-hour periods, and the number of likes within each period is counted. After smoothing the data using a moving average, the peak is identified when the number of likes is significantly higher than in adjacent periods. When the number of likes reaches its highest point of the day between 8:00 PM and 8:30 PM, this period is marked as the peak likes time for the influencer's content, reflecting the activity patterns of the fan base. Extracting interactive preferences involves the application of text mining technology. All comments under the influencer's content are segmented, and the frequency of each keyword is counted. High-frequency words such as "planting grass," "already purchased," and "link," which indicate purchase intent, as well as product-specific terms such as "texture," "whitening," and "versatile," collectively form a profile of fans' interactive preferences. This keyword distribution not only reflects user focus but also reveals the key factors that drive purchase decisions.
[0082] In one possible implementation, calculating the Pearson correlation coefficient reveals the strength of the linear relationship between variables. The system converts the peak like time into a numeric variable, such as 8 pm, which is recorded as 20. The frequency of preferred interaction keywords is used as another variable and correlation is calculated with the purchase conversion rate. The correlation coefficient ranges from -1 to 1, with values close to 1 indicating a strong positive correlation and values close to 0 indicating no correlation. If the correlation coefficient between peak likes and conversion rate is 0.75 in one period and 0.82 in the next, the difference of 0.07 indicates an increasing correlation. The construction of the training base dataset integrates and stores multi-dimensional features. Each record contains fields such as influencer ID, new revenue ratio, adjustment timestamp, peak like time, preferred interaction keywords and their frequency, interaction topic category, and purchase conversion rate. This structured data organization not only facilitates subsequent query analysis but also provides a standardized input format for training machine learning models. Through regular updates and accumulation, the dataset gradually forms a knowledge base reflecting the evolution of the platform ecosystem.
[0083] Based on the interaction frequency fluctuation data, we calculate the fan participation depth index of the expert content, evaluate the purchase conversion contribution under different interactive content themes, set threshold standards to promote purchasing behavior, increase the revenue weight coefficient for expert content types that exceed the threshold, adjust the platform's commission ratio and expert profit structure, and increase the income of experts in high-conversion interaction models to above the benchmark value, forming a revenue optimization model that encourages the creation of high-quality content.
[0084] Data on the frequency of interaction with influencers' content is collected, including the distribution of likes, comments, and shares over time. The frequency change rate is calculated based on the time interval between interactions. The fan engagement depth index is then calculated by comparing the change rate to the total number of followers. This index reflects the fan base's continued attention to the influencer's content. For each content theme, the fan engagement depth index is correlated with historical purchase records for that theme. Bayesian inference is used to calculate the conditional probability of a purchase at a specific engagement depth. The purchase conversion contribution of each content theme is determined by multiplying the conditional probability value by the engagement depth index. The purchase conversion contribution distribution across all content themes is ranked from high to low. The contribution value in the top 30% percentile is set as the threshold for promoting purchases. If an influencer's contribution to a particular content theme exceeds this threshold, the revenue weight coefficient for that theme is set to the ratio of the contribution to the threshold. This weight coefficient is used to adjust the influencer's share of the original revenue sharing structure. The actual income value of the influencer is calculated based on the adjusted profit-sharing structure and compared with the benchmark income value set by the platform. If the actual income does not reach the benchmark value, the weight coefficient will continue to be adjusted upward according to the proportion of the difference to the benchmark value until the influencer's income exceeds the benchmark value. By periodically executing the above adjustment process, a revenue optimization model that encourages the creation of high-quality content is formed.
[0085] In one possible implementation, obtaining interaction frequency fluctuation data involves collecting user behaviors in different time periods after the influencer publishes content.
[0086] Specifically, when an influencer posts a beauty tutorial video, the system records 500 likes within 0-1 hours, 800 within 1-3 hours, 400 within 3-6 hours, and so on. The frequency of change is calculated by calculating the difference in incremental interactions between adjacent time periods. If an influencer has 100,000 followers and the video accumulates 5,000 interactions within 24 hours, the fan engagement depth index is 0.05, reflecting the activeness of the fan base and the appeal of the content.
[0087] It's important to note that Bayesian inference plays a key role in calculating purchase conversion contribution. First, we count purchase behavior data at a specific engagement depth. For example, if the engagement depth is 0.05 and 50 out of 1,000 users complete a purchase, the conditional probability is 0.05. Multiplying this probability by the engagement depth metric yields a purchase conversion contribution of 0.0025 for this content topic. This calculation method comprehensively considers the correlation between user engagement and actual purchase behavior, avoiding the one-sidedness of judging content value based solely on purchase volume.
[0088] In one embodiment, the threshold standard is set using a dynamic quantile method. The contribution of all content themes is ranked, for example, the contribution of beauty tutorials is 0.0025, outfit sharing is 0.0020, and life vlogs are 0.0015. The value of 0.0020 in the top 30% percentile is set as the threshold. When the contribution of beauty tutorials 0.0025 exceeds the threshold, its income weight coefficient is calculated as 0.0025 / 0.0020=1.25. In the original profit-sharing structure, the influencer received 20% of the sales, which was adjusted to 20%×1.25=25%. This adjustment mechanism can motivate influencers to create more high-conversion content.
[0089] Optimally, the revenue optimization mechanism achieves dynamic balance through periodic adjustments. Assume the platform sets a benchmark monthly revenue of 5,000 yuan. A certain influencer's actual revenue, calculated using the adjusted profit-sharing structure, is 4,500 yuan. The 500 yuan difference represents 10% of the benchmark. Raising the weighting coefficient from 1.25 to 1.375 brings the influencer's revenue to 5,500 yuan. This mechanism is executed monthly, adjusting based on the latest engagement data and purchase conversions. This ensures the profitability of high-quality content creators while promoting the healthy development of the platform's content ecosystem. By directly linking high-converting content with increased revenue, a positive cycle of creative quality and financial returns is formed.
[0090] Obviously, those skilled in the art may make various changes and modifications to the embodiments of the present application without departing from the spirit and scope of the embodiments of the present application. Thus, if these modifications and variations of the embodiments of the present application fall within the scope of the claims of the present application and their equivalents, the present application also intends to include such modifications and variations.
Claims
1. A user behavior data mining method for an e-commerce ecosystem, characterized in that: The method comprises: Collect expert graphic and text content, extract fan interaction frequency fluctuations and comment data, record the proportion of product close-ups and background environment presentation data, and construct an interaction and display feature dataset; combine the interaction and display feature dataset with historical transaction records, extract purchase time points and product category data corresponding to fan interaction frequency, and determine the preliminary correlation mapping between interaction patterns and purchasing behavior; perform text processing on comment data, extract keyword frequency distribution, identify interaction data on user browsing time and sharing channel distribution, analyze fan behavior fluctuations in different time periods, and evaluate trend fluctuations in fan trust and engagement; combine fan trust and engagement trend fluctuations to evaluate the impact of content style changes on fan interaction; based on the evaluation results of the impact of content style changes on fan interaction, dynamically adjust the expert contribution weight and determine the adjusted weight distribution plan; through the adjusted weight distribution plan, analyze the key mining directions of interactive content themes and sharing channel distribution, extract the fan stay time distribution and forwarding propagation path related to purchase behavior, and obtain interaction frequency fluctuations under different background environment combinations; optimize platform revenue distribution based on interaction frequency fluctuations, determine revenue distribution results, and store fan interaction preferences, likes peaks, and interaction theme association data in the database.
2. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that: The collection of influencer graphic content, extraction of fan interaction frequency fluctuations and comment data, recording of product close-up ratios and background environment presentation data, and construction of an interaction and display feature dataset include: Collect the graphic and text content posted by experts, record the number of likes, collections and shares, combine the timestamp marks to obtain the temporal distribution of interactive behavior, calculate the difference between the publishing time and the interaction time as the interaction frequency fluctuation value, and count the number of comments and the distribution of word count; perform image segmentation on the graphic and text content, identify the product area and background area, calculate the ratio of the number of pixels in the product area to the number of pixels in the overall image as the product close-up ratio, and extract the color histogram and texture features of the background area as the environment matching data; establish a corresponding relationship matrix based on the interaction frequency fluctuation value and the product close-up ratio, mark the deep interaction label with the number of words in the comment exceeding the preset threshold, combine the main color value of the color histogram and the texture feature entropy value of the environment matching data, group the data through the clustering algorithm, integrate the grouping results and the corresponding relationship matrix to form an interaction and display feature data set.
3. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that: The method combines the interaction and display feature dataset with historical transaction records, extracts purchase time points and product category data corresponding to fan interaction frequency, and determines a preliminary correlation mapping between interaction patterns and purchase behaviors, including: Extract purchase timestamps and product category identifiers from historical transaction records, match fan accounts in the interaction and display feature dataset, obtain interaction frequency data sequences and purchase behavior sequences, identify time periods when the interaction frequency exceeds a preset threshold as interaction-intensive periods, identify time periods when the number of purchases exceeds a preset threshold as purchase peak periods, calculate the ratio of the intersection duration of the two time periods to the union duration as the overlap, mark strongly associated user groups with an overlap exceeding a preset threshold, and extract a numerical sequence of the product feature ratio of the strongly associated user groups during the interaction-intensive period; construct a category migration matrix based on the product feature ratio numerical sequence, record the frequency of product category conversion, and generate a preliminary association mapping table of interaction patterns and purchase behaviors through an association rule mining algorithm.
4. The user behavior data mining method for an e-commerce ecosystem according to claim 3, characterized in that: The step of extracting a numerical sequence of product feature ratios of the strongly associated user group during the intensive interaction period includes: Product images viewed during the intensive interaction period are extracted from the strongly associated user group, and the ratio of the number of pixels in the product area of each image to the number of pixels in the entire image is calculated to form a numerical sequence of product close-up ratios; for the numerical sequence of product close-up ratios, the product category corresponding to each value in the sequence is counted, and a category distribution vector is constructed to form the numerical sequence of product close-up ratios.
5. The user behavior data mining method for e-commerce ecosystem according to claim 1, characterized in that: The text processing of comment data, extraction of keyword frequency distribution, and identification of interactive data on user browsing time and sharing channel distribution include: The comment data is segmented, the number of occurrences of each word is counted, and a keyword frequency distribution table is constructed; based on the user identifier, the corresponding user's page residence time is extracted as browsing time data, and the number of times the content is shared on each platform is obtained to form sharing channel distribution data; the time period is divided into hours, and the total number of likes in each time period is calculated and divided by the number of independent likers as the like concentration value, and the difference in the number of comments in adjacent time periods is counted as the comment activity change value; through the high-frequency words in the keyword frequency distribution table, the timestamps corresponding to the comments are matched, and the browsing time data and the sharing channel distribution data are combined to generate a temporal interaction feature set containing like concentration and comment activity change values.
6. The user behavior data mining method for e-commerce ecosystem according to claim 1, characterized in that: The above method combines the fluctuations of fan trust and engagement trends to evaluate the impact of content style changes on fan interaction, including: Extract the numerical values of the fan trust and participation trends of each user in different time periods, combine them to form user behavior feature vectors, and construct a fan behavior comprehensive portrait matrix containing time-series change characteristics; perform principal component analysis on the fan behavior comprehensive portrait matrix, retain the principal components whose cumulative contribution rate exceeds a preset threshold, and obtain a reduced-dimensional feature matrix; group the users in the reduced-dimensional feature matrix through a hierarchical clustering algorithm to obtain fan group categories and central eigenvalues; based on the central eigenvalues of the fan group categories, extract the content release records corresponding to the active time period, calculate the time interval from the content release moment to the first interaction moment as the interactive response speed, and complete the evaluation of the impact of content style changes on fan interaction.
7. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that: The evaluation results of the impact of content style changes on fan interaction are used to dynamically adjust the influencer contribution weights and determine the adjusted weight distribution plan, including: Extract data on the impact of each influencer's corresponding content style changes on the response speed and participation depth of fan interactions, and obtain the initial value of the influencer's current contribution weight; calculate the average time interval between the first fan interaction after the influencer releases the content as the feedback speed indicator, and count the proportion of positive words in the comments as the emotional positivity value; based on the feedback speed indicator and the emotional positivity value, if the feedback speed indicator is lower than the preset time threshold and the emotional positivity value exceeds the preset proportion threshold, multiply the initial value of the contribution weight by the preset increase coefficient to obtain the adjusted contribution weight; based on the adjusted contribution weight, calculate the sum of all influencer weights, and normalize them to form an adjusted weight distribution plan.
8. The user behavior data mining method for e-commerce ecosystem according to claim 1, characterized in that: The adjusted weight distribution scheme is used to analyze the key mining directions of interactive content themes and sharing channel distribution, and extract the fan stay time distribution and forwarding propagation paths related to purchasing behavior, including: Through the adjusted weight distribution scheme, high-weight influencers whose weight ratio exceeds the preset threshold are screened, the subject tags and keywords of their published content are extracted, the total amount of interactions under each topic is counted, the distribution of the number of times the content is shared on each platform is obtained, and the subject category with the highest interaction volume and the channel with the most sharing times are determined; content records that meet the subject categories and channels are screened from the interaction and display feature data set, and the user stay time data of users who browse the content that meets the subject categories and channels and then make purchases are extracted, and the time intervals are divided; the propagation chain of the content record from the original release to the forwarding at each level is tracked, and the user ID and timestamp of each forwarding level are recorded to form a forwarding propagation path.
9. The user behavior data mining method for an e-commerce ecosystem according to claim 1, characterized in that: Optimizing the platform revenue distribution based on the interaction frequency fluctuations and determining the revenue distribution results include: Calculate the purchase conversion rate under each content theme, and obtain the conversion ratio by dividing the number of purchasing users by the total number of interactive users; if the conversion ratio exceeds the preset threshold, record the corresponding expert ID and current revenue ratio; for experts whose conversion ratio exceeds the preset threshold, multiply the current revenue ratio base by a fixed increase coefficient to obtain a new revenue ratio; summarize the new revenue ratios and IDs of all experts to form a revenue distribution result table.
10. The user behavior data mining method for e-commerce ecosystem according to claim 9, characterized in that: After the income distribution results table is formed, it includes: Extract the expert ID and the new income ratio from the income distribution result table, obtain the peak moment of the corresponding expert's fans' like behavior as the like peak; extract high-frequency words in the comment keyword distribution, and combine them with the interactive topic category label to form an interactive preference feature set; Calculate the correlation coefficient between the likes peak value and interaction preference feature set and the purchase conversion rate to generate time series correlation data containing the correlation coefficient; store the time series correlation data, expert identification and new revenue ratio in a database, and integrate them to form a training data set containing time series correlation and interaction preference.
Citation Information
Patent Citations
E-commerce precision marketing intelligent system and method fused with AI personalized recommendation
CN120450807A
Method and system for implementing a cloud-based social media marketing method and system
US20140180788A1
Cited By
Information cascade sequence prediction method and device, equipment and medium
CN121526005A
Product upgrading optimization method and system
CN122155781A
Cross-platform e-commerce user behavior joint prediction and recommendation method and system based on artificial intelligence
CN122288841A