E-commerce platform private domain content intelligent generation and distribution method based on data analysis

By building a user multi-interest representation structure using the improved Rocchio algorithm and DeepWalk algorithm and combining it with an exponential smoothing strategy, we can solve the problems of single interest modeling and update lag in content recommendation on e-commerce platforms, and achieve personalized and timely content recommendations, which is suitable for multi-terminal private domain operations.

CN120689117APending Publication Date: 2025-09-23JIANGSU WANGYUE DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510822182.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The content recommendation methods of existing e-commerce platforms are single-minded in modeling user interests, lack the ability to model structural information, are difficult to deal with cold start problems, have lagging update mechanisms and weak deployment adaptability, resulting in insufficient recommendation accuracy and adaptability.

Method used

The improved Rocchio algorithm is combined with user behavior clustering to construct a user multi-interest representation structure. The graph embedding vector is trained through heterogeneous graph modeling and DeepWalk algorithm, and dynamically updated with exponential smoothing strategy to achieve personalized content recommendation.

Benefits of technology

It improves the matching degree between recommended content and user interests, enhances the personalization and timeliness of the recommendation system, has strong adaptability, is suitable for multi-terminal private domain operation scenarios, and improves user experience and operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689117A_ABST
    Figure CN120689117A_ABST
Patent Text Reader

Abstract

The invention discloses an e-commerce platform private domain content intelligent generation and distribution method based on data analysis, and the method comprises the following steps: collecting behavior data, including browsing, clicking, collecting, purchasing and interactive recording, of a user in a private domain scene, and combining the behavior frequency and time information to generate a user behavior representation vector; a user multi-interest representation structure is constructed by using an improved Rocchio algorithm and a clustering mechanism, a heterogeneous graph model is further constructed by combining users, contents and behaviors, random walk is performed based on a graph structure, and node representation is learned by using a DeepWalk algorithm. And carrying out weighted fusion on the user interest vector and the graph embedding result to generate joint user representation, carrying out matching scoring and sorting on candidate contents, and finally realizing accurate pushing and dynamic updating of personalized contents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to a method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis. Background Art

[0002] With the rapid development of e-commerce, private domain operations have become a crucial tool for major e-commerce platforms to enhance user engagement and achieve personalized marketing. Private domain traffic emphasizes the exploration and maintenance of long-term user value, relying on a deep understanding of user behavior and the precise matching of content delivery. Against this backdrop, achieving efficient, personalized, and sustainable content recommendations based on user behavior data has become a core challenge for e-commerce platforms.

[0003] In existing technologies, traditional content recommendation methods mainly rely on collaborative filtering, one-way interest modeling, or explicit feature matching. These methods have the following shortcomings in the complex and ever-changing private domain operation scenarios:

[0004] 1. Single user interest modeling: Most recommendation systems use static or single-center interest modeling methods, which cannot effectively depict users' diverse interest preferences at different times or in different scenarios, resulting in low matching of recommended content.

[0005] 2. Lack of structural information modeling capabilities: Traditional recommendation algorithms mostly rely on explicit behavioral data of users and content, ignoring the implicit graph structure relationship between users, content and their behaviors, which limits the generalization ability and accuracy of recommendations.

[0006] 3. Difficulty in coping with the cold start problem: Existing recommendation technologies perform poorly when user or content interaction data is sparse, especially in private domain scenarios where new users or new content appear frequently, which can easily lead to recommendation blind spots.

[0007] 4. Lagging update mechanism: Many recommendation models lack an efficient mechanism for dynamically updating user interests and are unable to respond to changes in user behavior in a timely manner, resulting in outdated recommendation results and affecting user experience.

[0008] 5. Weak deployment adaptability: Some algorithms have complex structures and high resource consumption, making them difficult to adapt to the actual deployment requirements of private domain operation terminals such as enterprise WeChat and mini-programs, limiting their promotion and application.

[0009] Therefore, how to provide an intelligent generation and distribution method for private domain content on e-commerce platforms based on data analysis is an urgent problem that technicians in this field need to solve. Summary of the Invention

[0010] One purpose of the present invention is to propose a method for intelligently generating and distributing private-domain content on e-commerce platforms based on data analysis. The present invention utilizes user behavior data modeling, graph structure embedding learning, and multi-interest representation construction technology, and describes in detail the entire process of achieving personalized content recommendation through an improved Rocchio algorithm, heterogeneous graph modeling, and DeepWalk algorithm. The method has the advantages of detailed interest characterization, strong recommendation accuracy, and high adaptability to dynamic updates.

[0011] According to an embodiment of the present invention, a method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis includes the following steps:

[0012] S1. Obtain user behavior data in the private domain of the e-commerce platform. The behavior data includes user browsing, clicking, adding to favorites, and interaction records, and includes the frequency and timestamp of each type of behavior. Reconstruct the behavior time series and generate a user behavior embedding representation vector.

[0013] S2. Input the user behavior embedding representation vector into the improved Rocchio algorithm to generate the user interest initial representation vector;

[0014] S3. Perform a clustering operation based on the user behavior embedding representation vector to extract multiple user interest center vectors, wherein the multiple user interest center vectors and the user interest initial representation vector together constitute a user multi-interest representation structure;

[0015] S4. Construct a heterogeneous graph structure including user nodes, content nodes, and behavior nodes. The edge weights in the graph are determined based on the timeliness factor calculated based on the user behavior frequency and the behavior timestamp.

[0016] S5. Perform a random walk operation on the heterogeneous graph structure, use the DeepWalk algorithm to train and obtain the graph embedding representation vectors of user nodes and content nodes, and use the graph embedding representation vectors of content nodes as the representation vectors of candidate content;

[0017] S6. Weighted fusion of the user interest center vector and the graph embedding representation vector to generate a joint user representation vector;

[0018] S7. Sort the candidate contents based on the matching score between the joint user representation vector and the candidate content representation vector, generate a recommendation list and distribute it to the user's private domain operation terminal, and update the user interest center vector based on the exponential smoothing strategy.

[0019] Optionally, the S1 specifically includes:

[0020] S11. Obtaining user behavior data in a private domain scenario on an e-commerce platform, wherein the behavior data includes user identification information, user behavior type information, user behavior object identification information, behavior timestamp, and behavior frequency;

[0021] S12. For each user, sort the user behavior data in ascending order based on the user behavior timestamp information to construct a user behavior original sequence set;

[0022] S13. In the original sequence set of user behavior, based on a preset behavior sequence length threshold, a fixed-length segmentation mechanism is used to segment the user behavior record to obtain a number of behavior fragment subsequences;

[0023] S14. For each user behavior record in each behavior fragment subsequence, a behavior embedding mapping function is constructed to convert the user behavior record containing user behavior type information, user behavior timestamp information, and user behavior frequency information into a user behavior embedding representation vector.

[0024] Optionally, the improved Rocchio algorithm specifically includes:

[0025] Receiving a set of user behavior embedding representation vectors, where the set of user behavior embedding representation vectors is generated by user behavior records generated by the user within a preset time period, as an input set for constructing an initial representation vector of the user's interest;

[0026] According to the user behavior type information contained in each user behavior embedding representation vector, the user behavior embedding representation vector set is divided into a positive interest feedback vector set and a negative interest feedback vector set;

[0027] For each user behavior embedding representation vector in the positive interest feedback vector set, a first time decay weight is calculated according to its corresponding user behavior timestamp information, and for each user behavior embedding representation vector in the negative interest feedback vector set, a second time decay weight is calculated;

[0028] The improved Rocchio algorithm constructs the initial representation vector of user interest based on the user behavior embedding representation vector in the current cycle, and does not include the original centroid vector term in the traditional Rocchio algorithm:

[0029]

[0030] Among them, V u is the initial representation vector of user interest, is the embedding representation vector of the i-th user behavior in the positive interest feedback vector set, is the j-th user behavior embedding representation vector in the negative interest feedback vector set, is the first time attenuation weight corresponding to the positive interest feedback vector, is the second time decay weight corresponding to the negative interest feedback vector, β and γ are the positive interest control coefficient and the negative interest control coefficient respectively, m and n are the number of vectors in the positive interest feedback vector set and the negative interest feedback vector set respectively;

[0031] The user interest initial representation vector is output and provided to the user multiple interest representation structure modeling module for subsequent extraction of multiple user interest center vectors.

[0032] Optionally, the S3 specifically includes:

[0033] S31, receiving the user interest initial representation vector and the user behavior embedded representation vector set;

[0034] S32. Perform a weighted clustering operation on the set of user behavior embedding representation vectors. The weighted clustering operation adopts a clustering method guided by the initial interest vector, uses the initial user interest representation vector as one of the initial clustering guidance centers, and performs sample attribution calculation under a preset number of categories K, combined with the similarity between the user behavior embedding representation vectors and their corresponding behavior time decay weights, and generates K user interest center vectors.

[0035] S33. Construct a user multi-interest representation structure, structurally merge the K user interest center vectors with the user interest initial representation vector to form a set of composite interest vectors, which are used to simultaneously express the user's current dominant interest and distributed interest preferences.

[0036] Optionally, the S4 specifically includes:

[0037] S41. Construct a heterogeneous graph structure, wherein the heterogeneous graph structure includes user nodes, content nodes, and behavior nodes, wherein each user node corresponds to a unique user identification information, each content node corresponds to a content object identification information, and each behavior node represents a record of a user's behavior on specific content at a specific point in time;

[0038] S42. Establishing a node connection relationship in the heterogeneous graph structure, connecting the user node with the behavior node, and connecting the behavior node with the content node, to form a heterogeneous ternary connection relationship among the user node, the behavior node, and the content node;

[0039] S43. Determine an edge weight for each edge between a user node and an action node. The edge weight is calculated based on the user action frequency and the user action timestamp. Specifically, the edge weight is equal to the product of the user action frequency and a time decay weight, where the time decay weight is calculated using an exponential decay function based on the time interval between the time when the user action occurred and the current time.

[0040] S44. Generate a complete heterogeneous graph structure, which includes user nodes, behavior nodes, and content nodes with edge weights, as the input structure of the graph embedding training module for subsequent graph neural network modeling or embedding representation learning.

[0041] Optionally, the S5 specifically includes:

[0042] S61. For each user node and each content node in the constructed heterogeneous graph structure, multiple rounds of biased random walk operations are performed according to a preset random walk step size and number of random walk paths. In each round of the walk, a transition probability distribution is generated based on the edge weight of the edge connecting the current node. Subsequent nodes are sampled based on the transition probability distribution to generate a set of node walk paths.

[0043] S62, converting the node wandering path set into a graph training sample corpus, and constructing a training dataset containing the co-occurrence contextual relationship between user nodes and content nodes in the graph structure;

[0044] S63. Inputting the graph training sample corpus into a graph embedding training module based on the DeepWalk algorithm. The DeepWalk algorithm uses the Skip-gram model to train node representation vectors for user nodes and content nodes, constructs pairs of center nodes and context nodes through a sliding window mechanism, and updates the node representation vectors based on an objective function that maximizes the probability of co-occurring node pairs.

[0045] S64: Use the content node representation vector obtained through training as a candidate content representation vector for subsequent matching calculation and recommendation ranking with the joint user representation vector.

[0046] Optionally, the weighted fusion specifically includes: performing a weighted sum operation on the user interest center vector and the user node graph embedding representation vector according to a preset weight coefficient to generate a joint user representation vector.

[0047] Optionally, the S7 specifically includes:

[0048] S71, respectively calculating the vector similarity between the joint user representation vector and each candidate content representation vector, wherein the vector similarity is determined by calculating the cosine similarity between vectors, and obtaining a matching score for each candidate content representation vector;

[0049] S72. Sort all candidate content representation vectors in descending order according to their matching scores to generate a candidate content ranking list;

[0050] S73: Select a preset number of candidate contents according to the position order in the candidate content ranking list, generate a recommended content list, and distribute the recommended content list to the private domain operation terminal of the target user;

[0051] S74. Collect the target user's interactive behavior data on each recommended content in the recommended content list, and based on the newly generated user behavior embedding representation vector in the current cycle, use an exponential smoothing strategy to update the user interest center vector. The exponential smoothing strategy includes performing a weighted summation of the user interest center vector of the previous cycle and the user behavior embedding representation vector of the current cycle to generate an updated user interest center vector.

[0052] The beneficial effects of the present invention are:

[0053] (1) This paper introduces an improved Rocchio algorithm and user behavior clustering modeling to construct a user multi-interest representation structure, which solves the problem that traditional interest modeling methods are unable to adequately depict the diversity of user interests. This enables the system to accurately express the dynamic preferences of users in private domain scenarios, thereby improving the matching degree between recommended content and users' actual interests.

[0054] (2) The present invention constructs a heterogeneous graph structure containing user nodes, content nodes, and behavior nodes, and combines the DeepWalk algorithm with the skip-word model to train graph embedding vectors. This breaks through the traditional recommendation method's reliance on explicit content similarity and realizes the modeling of structured semantic relationships between users and content. It is particularly suitable for cold-start user or content recommendation scenarios.

[0055] (3) The present invention designs a weighted fusion mechanism of user interest representation vectors and graph embedding vectors, and combines matching degree calculation with content sorting strategy to improve the performance of the recommendation system in multi-dimensional expression and dynamic adaptation, and enhance the accuracy and timeliness of personalized recommendations.

[0056] (4) The present invention introduces an exponential smoothing strategy and a sliding window update mechanism to adaptively and dynamically adjust the user interest center vector, taking into account both long-term preferences and short-term behavior changes, thereby improving the system's ability to model user interest drift.

[0057] (5) The method of the present invention has a clear structure and can be deployed in a modular manner. It is adaptable to multi-terminal private domain operation scenarios and has high real-time performance, strong scalability, and engineering feasibility, filling the practical gap of current recommendation algorithms in the intelligent distribution of private domain content. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0059] Figure 1 This is a flowchart of the method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis proposed by the present invention. DETAILED DESCRIPTION

[0060] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0061] refer to Figure 1 The method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis includes the following steps:

[0062] S1. Obtain user behavior data in the private domain of the e-commerce platform. The behavior data includes user browsing, clicking, adding to favorites, and interaction records, and includes the frequency and timestamp of each type of behavior. Reconstruct the behavior time series and generate a user behavior embedding representation vector.

[0063] S2. Input the user behavior embedding representation vector into the improved Rocchio algorithm to generate the user interest initial representation vector;

[0064] S3. Perform a clustering operation based on the user behavior embedding representation vector to extract multiple user interest center vectors, wherein the multiple user interest center vectors and the user interest initial representation vector together constitute a user multi-interest representation structure;

[0065] S4. Construct a heterogeneous graph structure including user nodes, content nodes, and behavior nodes. The edge weights in the graph are determined based on the timeliness factor calculated based on the user behavior frequency and the behavior timestamp.

[0066] S5. Perform a random walk operation on the heterogeneous graph structure, use the DeepWalk algorithm to train and obtain the graph embedding representation vectors of user nodes and content nodes, and use the graph embedding representation vectors of content nodes as the representation vectors of candidate content;

[0067] S6. Weighted fusion of the user interest center vector and the graph embedding representation vector to generate a joint user representation vector;

[0068] S7. Sort the candidate contents based on the matching score between the joint user representation vector and the candidate content representation vector, generate a recommendation list and distribute it to the user's private domain operation terminal, and update the user interest center vector based on the exponential smoothing strategy.

[0069] In this embodiment, S1 specifically includes:

[0070] S11. Obtaining user behavior data in a private domain scenario on an e-commerce platform, wherein the behavior data includes user identification information, user behavior type information, user behavior object identification information, behavior timestamp, and behavior frequency;

[0071] S12. For each user, sort the user behavior data in ascending order based on the user behavior timestamp information to construct a user behavior original sequence set;

[0072] S13. In the original sequence set of user behavior, based on a preset behavior sequence length threshold, a fixed-length segmentation mechanism is used to segment the user behavior record to obtain a number of behavior fragment subsequences;

[0073] S14. For each user behavior record in each behavior fragment subsequence, a behavior embedding mapping function is constructed to convert the user behavior record containing user behavior type information, user behavior timestamp information, and user behavior frequency information into a user behavior embedding representation vector.

[0074] This implementation method organizes the time sequence of user behavior data in the private domain scenarios of e-commerce platforms and divides them into fixed-length segments, extracts user behavior fragment subsequences, and uses behavior type, behavior time and frequency to construct a behavior embedding mapping function, and represents discrete behavior records in high dimensions as continuous embedding vectors, thereby effectively characterizing the dynamic characteristics of users' interests in different time periods, realizing structured modeling and deep semantic expression of user behavior patterns, enhancing the basic expression capabilities of subsequent interest modeling and recommendation reasoning, and effectively improving the recommendation system's understanding accuracy of user behavior sequences and the efficiency of utilizing behavior characteristics.

[0075] In this embodiment, the improved Rocchio algorithm specifically includes:

[0076] Receiving a set of user behavior embedding representation vectors, where the set of user behavior embedding representation vectors is generated by user behavior records generated by the user within a preset time period, as an input set for constructing an initial representation vector of the user's interest;

[0077] According to the user behavior type information contained in each user behavior embedding representation vector, the user behavior embedding representation vector set is divided into a positive interest feedback vector set and a negative interest feedback vector set;

[0078] For each user behavior embedding representation vector in the positive interest feedback vector set, a first time decay weight is calculated according to its corresponding user behavior timestamp information, and for each user behavior embedding representation vector in the negative interest feedback vector set, a second time decay weight is calculated;

[0079] The improved Rocchio algorithm constructs the initial representation vector of user interest based on the user behavior embedding representation vector in the current cycle, and does not include the original centroid vector term in the traditional Rocchio algorithm:

[0080]

[0081] Among them, V u is the initial representation vector of user interest, is the i-th user behavior embedding representation vector in the positive interest feedback vector set, is the j-th user behavior embedding representation vector in the negative interest feedback vector set, is the first time attenuation weight corresponding to the positive interest feedback vector, is the second time decay weight corresponding to the negative interest feedback vector, β and γ are the positive interest control coefficient and the negative interest control coefficient respectively, m and n are the number of vectors in the positive interest feedback vector set and the negative interest feedback vector set respectively;

[0082] This formula dynamically generates an initial representation vector for a user's interests by weightedly accumulating their positive and negative interests within the current period, and by introducing a time-decay weight and an interest control coefficient. The principle is to emphasize the influence of recent high-frequency, positive user behaviors while suppressing disinterested behaviors. This constructs an interest expression that reflects the user's current preferences without relying on historical interest centroids, effectively integrating time sensitivity and behavioral variability into interest modeling.

[0083] The user interest initial representation vector is output and provided to the user multiple interest representation structure modeling module for subsequent extraction of multiple user interest center vectors.

[0084] In this implementation, a set of user behavior embedding representation vectors generated by users within a preset period is received and divided into a set of positive interest feedback vectors and a set of negative interest feedback vectors based on the user behavior type information contained therein. After calculating the corresponding time-decay weights based on the user behavior timestamp information, an improved Rocchio algorithm is used to model the user's behavior within the current period without introducing the traditional original centroid vector term to generate an initial user interest representation vector. This approach accurately depicts the user's current interest state, avoids redundant interference from historical interests, facilitates the subsequent construction of a multi-interest representation structure, improves the timeliness and dynamic adaptability of user interest modeling, and thus enhances the response accuracy and effectiveness of personalized recommendation systems in private domain content distribution scenarios.

[0085] In this embodiment, S3 specifically includes:

[0086] S31, receiving a user interest initial representation vector and a user behavior embedded representation vector set;

[0087] S32. Perform a weighted clustering operation on the set of user behavior embedding representation vectors. The weighted clustering operation adopts a clustering method guided by the initial interest vector, uses the initial user interest representation vector as one of the initial clustering guidance centers, and performs sample attribution calculation under a preset number of categories K, combined with the similarity between the user behavior embedding representation vectors and their corresponding behavior time decay weights, and generates K user interest center vectors.

[0088] S33. Construct a user multi-interest representation structure, structurally merge the K user interest center vectors with the user interest initial representation vector to form a set of composite interest vectors, which are used to simultaneously express the user's current dominant interest and distributed interest preferences.

[0089] This embodiment receives the initial representation vector of user interests and the set of user behavior embedded representation vectors, adopts a weighted clustering method with the initial representation vector of user interests as the guiding center, and performs sample attribution calculation based on considering user behavior similarity and time decay weight, thereby extracting multiple user interest center vectors, and further constructing a composite interest vector set with the initial interest representation vector, realizing the coordinated expression of the user's dominant interests and potential interests, effectively improving the recommendation system's modeling accuracy of users' diverse preferences and personalized content matching capabilities.

[0090] In this embodiment, the S4 specifically includes:

[0091] S41. Construct a heterogeneous graph structure, wherein the heterogeneous graph structure includes user nodes, content nodes, and behavior nodes, wherein each user node corresponds to a unique user identification information, each content node corresponds to a content object identification information, and each behavior node represents a record of a user's behavior on specific content at a specific point in time;

[0092] S42. Establishing a node connection relationship in the heterogeneous graph structure, connecting the user node with the behavior node, and connecting the behavior node with the content node, to form a heterogeneous ternary connection relationship among the user node, the behavior node, and the content node;

[0093] S43. Determine an edge weight for each edge between a user node and an action node. The edge weight is calculated based on the user action frequency and the user action timestamp. Specifically, the edge weight is equal to the product of the user action frequency and a time decay weight, where the time decay weight is calculated using an exponential decay function based on the time interval between the time when the user action occurred and the current time.

[0094] S44. Generate a complete heterogeneous graph structure, which includes user nodes, behavior nodes, and content nodes with edge weights, as the input structure of the graph embedding training module for subsequent graph neural network modeling or embedding representation learning.

[0095] This implementation constructs a heterogeneous graph structure consisting of user nodes, content nodes, and behavior nodes, establishes ternary connections between users and behaviors, and between behaviors and content, and introduces an edge weighting mechanism based on behavior frequency and time interval calculations. This effectively expresses the dynamic interaction intensity and timeliness between users and content. This heterogeneous graph structure not only enhances the semantic expression of inter-node relationships but also provides a high-quality structural foundation for subsequent graph embedding training. This allows the embedding learning model to capture the deep patterns of user interest evolution, thereby improving the personalized accuracy and timeliness of content recommendations and enhancing the system's ability to model the diversity of user behaviors and changing interests.

[0096] In this embodiment, the S5 specifically includes:

[0097] S61. For each user node and each content node in the constructed heterogeneous graph structure, multiple rounds of biased random walk operations are performed according to a preset random walk step size and number of random walk paths. In each round of the walk, a transition probability distribution is generated based on the edge weight of the edge connecting the current node. Subsequent nodes are sampled based on the transition probability distribution to generate a set of node walk paths.

[0098] S62, converting the node wandering path set into a graph training sample corpus, and constructing a training dataset containing the co-occurrence contextual relationship between user nodes and content nodes in the graph structure;

[0099] S63. Inputting the graph training sample corpus into a graph embedding training module based on the DeepWalk algorithm. The DeepWalk algorithm uses the Skip-gram model to train node representation vectors for user nodes and content nodes, constructs pairs of center nodes and context nodes through a sliding window mechanism, and updates the node representation vectors based on an objective function that maximizes the probability of co-occurring node pairs.

[0100] S64: Use the content node representation vector obtained through training as a candidate content representation vector for subsequent matching calculation and recommendation ranking with the joint user representation vector.

[0101] This implementation performs biased random walks on the constructed heterogeneous graph structure, determines transition probabilities based on edge weights, generates a set of walk paths between user nodes and content nodes, and converts this into a training corpus containing node-context relationships. This corpus is then fed into a graph embedding training module based on the Skip-gram model. A sliding window mechanism is used to establish pairs of central and contextual nodes, maximizing their co-occurrence probabilities and training graph embedding representation vectors for user and content nodes. This approach effectively captures high-order potential relationships between users and content within the graph structure, enhancing the semantic expressiveness of content representation and improving matching accuracy in the subsequent recommendation ranking stage. It is particularly suitable for content association modeling and recommendation optimization in complex behavioral scenarios.

[0102] In this embodiment, the weighted fusion specifically includes: performing a weighted sum operation on the user interest center vector and the user node graph embedding representation vector according to a preset weight coefficient to generate a joint user representation vector.

[0103] In this implementation, a joint user representation vector is generated by weighting the user interest center vector and the user node graph embedding representation vector according to preset weight coefficients. This representation of user preferences simultaneously takes into account the multi-interest characteristics reflected by user behavior clustering and the correlation information captured by graph structure learning. This fusion strategy effectively integrates behavioral semantics and graph structure semantics, improving the completeness and accuracy of user interest modeling. This provides a more accurate representation foundation for subsequent matching calculations and personalized recommendations, significantly enhancing the relevance and adaptability of recommendation results.

[0104] In this embodiment, the S7 specifically includes:

[0105] S71, respectively calculating the vector similarity between the joint user representation vector and each candidate content representation vector, wherein the vector similarity is determined by calculating the cosine similarity between vectors, and obtaining a matching score for each candidate content representation vector;

[0106] S72. Sort all candidate content representation vectors in descending order according to their matching scores to generate a candidate content ranking list;

[0107] S73: Select a preset number of candidate contents according to the position order in the candidate content ranking list, generate a recommended content list, and distribute the recommended content list to the private domain operation terminal of the target user;

[0108] S74. Collect the target user's interactive behavior data on each recommended content in the recommended content list, and based on the newly generated user behavior embedding representation vector in the current cycle, use an exponential smoothing strategy to update the user interest center vector. The exponential smoothing strategy includes performing a weighted summation of the user interest center vector of the previous cycle and the user behavior embedding representation vector of the current cycle to generate an updated user interest center vector.

[0109] This implementation calculates the cosine similarity between the joint user representation vector and the candidate content representation vector to obtain a matching score and sort the candidate content accordingly, generating a recommendation list and distributing it to the user's private domain operation terminal to achieve personalized content distribution; at the same time, it collects the user's actual interactive behavior data on the recommended content, and adopts an exponential smoothing strategy combined with the current period user behavior representation to dynamically update the user interest center vector, so that the system can continuously perceive and adapt to the changing trends of user interests, enhance the relevance and real-time nature of the recommended content, and improve the accuracy of private domain content distribution and user satisfaction.

[0110] Example 1:

[0111] To verify the feasibility of this invention, we applied it to the private marketing system of "Xgou," a large, comprehensive e-commerce platform. This platform boasts over 18 million registered users, approximately 860,000 daily active users, and an average of 730,000 daily content delivery requests. Private marketing primarily utilizes mini-programs and WeChat for Business, pushing content such as promotional information, new product information, articles of interest, and short videos.

[0112] In traditional private domain operations, platforms primarily rely on manual rules and static tagging systems to push content to users. They build basic user profiles based on historical dimensions like consumer preferences and category preferences, and then deliver content in batches within fixed time windows. This model presents significant problems: first, user profile updates lag, resulting in insufficient content matching; second, the diversity of user interests cannot be effectively modeled, and reliance on a single interest center leads to homogenized recommendation results; recommended content is difficult to dynamically adjust, and push strategies fail to respond promptly to changes in user preferences, resulting in a continuous decline in user click-through and conversion rates. Furthermore, operational labor costs are high, and optimization efficiency is low.

[0113] To solve the above problems, the "Xgou" platform decided to introduce the method of the present invention. The system model was fully deployed in the private domain mini-program and enterprise WeChat. The main process includes the following:

[0114] The platform first accesses user behavior data, including private domain behavioral events such as browsing, clicking, adding to favorites, adding to cart, forwarding, and commenting. The system acquires this data in real time and constructs behavior vectors. It also extracts timestamps and behavior frequencies for temporal weighting modeling. The system uses an improved Rocchio algorithm to construct an initial representation vector of user interests based solely on positive and negative interest feedback and time-series weighting, rather than relying on the original centroid vector for the current cycle's behavior vector. A clustering mechanism is then introduced to perform multi-center partitioning of the behavior embedding vector, resulting in multiple user interest centers, thus forming a multi-interest representation structure.

[0115] The system constructs a heterogeneous graph structure, incorporating user nodes, content nodes, and behavior nodes into the graph model. Edge weights are combined with behavior frequency and duration. The DeepWalk algorithm performs random walks on the graph, and the resulting paths are fed into the Skip-gram model to train embedding representations for user and content nodes. Multiple interest vectors are weightedly fused with the graph embedding vector to generate a joint user representation. The representation's similarity score with all candidate content is calculated, and a recommendation list is generated and pushed to the user, sorted in descending order by score.

[0116] The system collects user feedback behavior in real time and updates the user interest center vector through a sliding time window and exponential smoothing strategy to ensure that the system continues to respond to changes in users' short-term interests and retains long-term preference trends.

[0117] To verify the effectiveness of the system, the platform conducted a 30-day online comparative experiment. A total of 40,000 active private domain users were selected and divided into a control group and an experimental group of 20,000 people each. The control group used the existing rule recommendation system, while the experimental group deployed the proposed method. The core indicators are as follows:

[0118] Table 1 Comparison of actual operational effects based on private domain recommendation systems

[0119]

[0120]

[0121] Results show that after deploying this method, the click-through rate of recommended content increased by 73.0%, the conversion rate increased by nearly 80%, the rate of repeated content recommendations decreased by approximately 67%, and the average user dwell time and depth of interaction increased significantly. Recommendation response latency decreased by over 50%, demonstrating the system's enhanced responsiveness and recommendation efficiency. Further analysis revealed that users in the experimental group who used the interest center update strategy consistently improved their content preference prediction accuracy and outperformed the control group in terms of repurchase rate and activity participation rate.

[0122] For example, a high-value user nicknamed "Linlin's Shop" had a click-through rate of only 11.2% for beauty and skin care content pushed in the original system. After deploying the method of the present invention, the system detected a surge in his browsing behavior for "nutritional supplements" and "light meal replacements" in the past week. Through the graph embedding structure, it was found that he had a close relationship with multiple health-related content. The system then adjusted the push direction and generated personalized content including "low-sugar meal replacements" and "nutrition consultation live broadcasts". The click-through rate quickly increased to 23.6% within 3 days, and achieved two high-unit-price conversions. The platform eventually classified the user as a "multi-interest distribution" private domain user to achieve differentiated long-term maintenance.

[0123] In summary, this method achieves a breakthrough in the private domain recommendation scenario of e-commerce platforms, moving from single-interest modeling to dynamic expression of multiple interests. Combining graph structure modeling with embedded learning, it effectively improves recommendation accuracy and response efficiency. By integrating exponential smoothing with behavioral feedback, the system ensures interest updating capabilities, significantly improving user experience and operational conversions. This demonstrates the broad applicability and engineering feasibility of this method in complex, real-world private domain ecosystems.

[0124] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis, characterized in that: The following steps are involved: S1. Obtain user behavior data in the private domain of the e-commerce platform. The behavior data includes user browsing, clicking, adding to favorites, and interaction records, and includes the frequency and timestamp of each type of behavior. Reconstruct the behavior time series and generate a user behavior embedding representation vector. S2. Input the user behavior embedding representation vector into the improved Rocchio algorithm to generate the user interest initial representation vector; S3. Perform a clustering operation based on the user behavior embedding representation vector to extract multiple user interest center vectors, wherein the multiple user interest center vectors and the user interest initial representation vector together constitute a user multi-interest representation structure; S4. Construct a heterogeneous graph structure including user nodes, content nodes, and behavior nodes. The edge weights in the graph are determined based on the timeliness factor calculated based on the user behavior frequency and the behavior timestamp. S5. Perform a random walk operation on the heterogeneous graph structure, use the DeepWalk algorithm to train and obtain the graph embedding representation vectors of user nodes and content nodes, and use the graph embedding representation vectors of content nodes as the representation vectors of candidate content; S6. Weighted fusion of the user interest center vector and the graph embedding representation vector to generate a joint user representation vector; S7. Sort the candidate contents based on the matching score between the joint user representation vector and the candidate content representation vector, generate a recommendation list and distribute it to the user's private domain operation terminal, and update the user interest center vector based on the exponential smoothing strategy.

2. The method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis according to claim 1 is characterized in that: Said S1 specifically includes: S11. Obtaining user behavior data in a private domain scenario on an e-commerce platform, wherein the behavior data includes user identification information, user behavior type information, user behavior object identification information, behavior timestamp, and behavior frequency; S12. For each user, sort the user behavior data in ascending order based on the user behavior timestamp information to construct a user behavior original sequence set; S13. In the original sequence set of user behavior, based on a preset behavior sequence length threshold, a fixed-length segmentation mechanism is used to segment the user behavior record to obtain a number of behavior fragment subsequences; S14. For each user behavior record in each behavior fragment subsequence, a behavior embedding mapping function is constructed to convert the user behavior record containing user behavior type information, user behavior timestamp information, and user behavior frequency information into a user behavior embedding representation vector.

3. The method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis according to claim 1 is characterized in that: The improved Rocchio algorithm specifically includes: Receiving a set of user behavior embedding representation vectors, where the set of user behavior embedding representation vectors is generated by user behavior records generated by the user within a preset time period, as an input set for constructing an initial representation vector of the user's interest; According to the user behavior type information contained in each user behavior embedding representation vector, the user behavior embedding representation vector set is divided into a positive interest feedback vector set and a negative interest feedback vector set; For each user behavior embedding representation vector in the positive interest feedback vector set, a first time decay weight is calculated according to its corresponding user behavior timestamp information, and for each user behavior embedding representation vector in the negative interest feedback vector set, a second time decay weight is calculated; The improved Rocchio algorithm constructs the initial representation vector of user interest based on the user behavior embedding representation vector in the current cycle: Among them, V u is the initial representation vector of user interest, is the i-th user behavior embedding representation vector in the positive interest feedback vector set, is the j-th user behavior embedding representation vector in the negative interest feedback vector set, is the first time attenuation weight corresponding to the positive interest feedback vector, is the second time decay weight corresponding to the negative interest feedback vector, β and γ are the positive interest control coefficient and the negative interest control coefficient respectively, m and n are the number of vectors in the positive interest feedback vector set and the negative interest feedback vector set respectively; Output the initial representation vector of the user interest.

4. The method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis according to claim 1 is characterized in that: The S3 specifically includes: S31, receiving a user interest initial representation vector and a user behavior embedded representation vector set; S32. Perform a weighted clustering operation on the set of user behavior embedding representation vectors. The weighted clustering operation adopts a clustering method guided by the initial interest vector, uses the initial user interest representation vector as one of the initial clustering guidance centers, and performs sample attribution calculation under a preset number of categories K, combined with the similarity between the user behavior embedding representation vectors and their corresponding behavior time decay weights, and generates K user interest center vectors. S33. Construct a user multi-interest representation structure, structurally merge the K user interest center vectors with the user interest initial representation vector to form a set of composite interest vectors, which are used to simultaneously express the user's current dominant interest and distributed interest preferences.

5. The method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis according to claim 1 is characterized in that: The S4 specifically includes: S41. Construct a heterogeneous graph structure, wherein the heterogeneous graph structure includes user nodes, content nodes, and behavior nodes, wherein each user node corresponds to a unique user identification information, each content node corresponds to a content object identification information, and each behavior node represents a record of a user's behavior on specific content at a specific point in time; S42. Establishing a node connection relationship in the heterogeneous graph structure, connecting the user node with the behavior node, and connecting the behavior node with the content node, to form a heterogeneous ternary connection relationship among the user node, the behavior node, and the content node; S43. Determine an edge weight for each edge between a user node and an action node. The edge weight is calculated based on the user action frequency and the user action timestamp. Specifically, the edge weight is equal to the product of the user action frequency and a time decay weight, where the time decay weight is calculated using an exponential decay function based on the time interval between the time when the user action occurred and the current time. S44. Generate a complete heterogeneous graph structure, where the heterogeneous graph structure includes user nodes, behavior nodes, and content nodes with edge weights.

6. The method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis according to claim 1 is characterized in that: The S5 specifically includes: S61. For each user node and each content node in the constructed heterogeneous graph structure, multiple rounds of biased random walk operations are performed according to a preset random walk step size and number of random walk paths. In each round of the walk, a transition probability distribution is generated based on the edge weight of the edge connecting the current node. Subsequent nodes are sampled based on the transition probability distribution to generate a set of node walk paths. S62, converting the node wandering path set into a graph training sample corpus, and constructing a training dataset containing the co-occurrence contextual relationship between user nodes and content nodes in the graph structure; S63. Inputting the graph training sample corpus into a graph embedding training module based on the DeepWalk algorithm. The DeepWalk algorithm uses the Skip-gram model to train node representation vectors for user nodes and content nodes, constructs pairs of center nodes and context nodes through a sliding window mechanism, and updates the node representation vectors based on an objective function that maximizes the probability of co-occurring node pairs. S64: Use the content node representation vector obtained through training as a candidate content representation vector for subsequent matching calculation and recommendation ranking with the joint user representation vector.

7. The method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis according to claim 1 is characterized in that: The weighted fusion specifically includes: performing a weighted sum operation on the user interest center vector and the user node graph embedding representation vector according to a preset weight coefficient to generate a joint user representation vector.

8. The method for intelligently generating and distributing private domain content on an e-commerce platform based on data analysis according to claim 1 is characterized in that: The S7 specifically includes: S71, respectively calculating the vector similarity between the joint user representation vector and each candidate content representation vector, wherein the vector similarity is determined by calculating the cosine similarity between vectors, and obtaining a matching score for each candidate content representation vector; S72. Sort all candidate content representation vectors in descending order according to their matching scores to generate a candidate content ranking list; S73: Select a preset number of candidate contents according to the position order in the candidate content ranking list, generate a recommended content list, and distribute the recommended content list to the private domain operation terminal of the target user; S74. Collect the target user's interactive behavior data on each recommended content in the recommended content list, and based on the newly generated user behavior embedding representation vector in the current cycle, use an exponential smoothing strategy to update the user interest center vector. The exponential smoothing strategy includes performing a weighted summation of the user interest center vector of the previous cycle and the user behavior embedding representation vector of the current cycle to generate an updated user interest center vector.