User portrait-based accurate pushing method and system, and medium

CN122594584APending Publication Date: 2026-08-18BEIJING JIEXINHONGYE NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610749245.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

无基于用户认知负荷的调节机制,无法在某一标签推送过载时抑制该标签及图谱邻近节点的权重,也无法激活知识图谱的远域节点进行合理的内容破圈

Benefits of technology

[0005]This invention extracts historical user behavior features to construct user profiles and establishes a tag migration model encompassing exploratory, stable, and shielded states. It utilizes positive thresholds, negative feedback, and decay cycles to represent the trajectory of user interest changes. Tags in different states are mapped to a knowledge graph, and a bidirectional semantic potential vector is constructed to represent users' positive preferences and negative aversion tendencies. User groups are segmented, and attention mechanisms are used to calculate potential content affinity, exploring users' implicit needs. During the ranking generation stage, a cognitive load index is used to prevent user fatigue with homogeneous content by suppressing the weight of frequently pushed tags and neighboring nodes in the knowledge graph while increasing the weight of distant nodes. This broadens information reach boundaries and optimizes overall push quality and the user's long-term browsing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122594584A_ABST
    Figure CN122594584A_ABST
Patent Text Reader

Abstract

The present application belongs to the field of information pushing, in particular to a precise pushing method and system based on user portrait and a medium, the method specifically comprises: extracting user text interaction and behavior sequence features to construct an initial user portrait, establishing a three-state transition model of exploration state, stable state and shielding state; positive interaction exceeds threshold, label migrates from exploration state to stable state; negative feedback exceeds threshold, label migrates to shielding state; no interaction within the stable state label decay period, then back to exploration state; mapping stable and shielding state labels to knowledge graph, calculating positive activation and negative repulsion components, and constructing a bidirectional semantic potential field vector; clustering the vector to obtain user groups, calculating the potential affinity between content and group centroid through attention mechanism; sorting and pushing by comprehensively matching the initial matching score, potential affinity and stable state label cognitive load index; when the cognitive load exceeds the threshold, suppress the near field and improve the far field weight, and the decay rate is negatively correlated with the user interest breadth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of information push, and in particular relates to a precise push method, system and medium based on user profile. Background Technology

[0002] Building user profiles is fundamental to personalized content delivery. However, user profile tagging systems are static and singular, merely accumulating or extracting features from users' historical interaction texts and behavioral sequences, failing to represent the changing processes of user interests. User interests have a lifecycle and vary in strength. Existing systems cannot distinguish between fleeting curiosity, long-term core preferences, and explicitly disliked negative preferences; existing methods overlook the natural decay of tags over time and lack reasonable tag state transition and rollback mechanisms, causing expired interest tags to continuously influence delivery results, making the profile bloated and distorted. Combining knowledge graphs, group clustering analysis, and attention mechanisms with content recommendation systems leverages external structured knowledge to expand the semantic associations of tags and improves the recommendation effect of potential content by exploring the behavioral characteristics of similar groups. However, most knowledge graph recommendation algorithms are based on positive association mining, extending similarity solely based on tags liked by users, failing to semantically model negative feedback or blocking behaviors, and unable to construct a bidirectional semantic evaluation system encompassing positive activation and negative repulsion. This leads to the easy recommendation of information highly associated with content implicitly disliked by users. Recommendation mechanisms are prone to the "information cocoon" effect and user cognitive fatigue. Once the system detects a user's stable interest, it frequently and intensively pushes highly homogeneous content. Without an adjustment mechanism based on user cognitive load, it cannot suppress the weight of a tag and its neighboring nodes in the knowledge graph when there is an overload of content, nor can it activate distant nodes in the knowledge graph for reasonable content expansion. Furthermore, because it overlooks the differences in the breadth of interests among different users, it fails to achieve personalized, high-quality recommendations. Summary of the Invention

[0003] To optimize overall push notification quality and long-term user browsing experience, one or more embodiments of the present invention provide a precise push notification method based on user profiles, the method comprising: Extract text interaction features and behavior sequence features from user historical behavior data to construct an initial user profile; establish a transition model for profile tags in exploratory, stable, and blocked states. Profile tags migrate from the exploratory state to the stable state based on the frequency of positive interactions exceeding a positive threshold, and profile tags migrate to the blocked state based on the number of negative feedback events of negative interaction events exceeding a negative threshold. Stable state tags regress to the exploratory state if there is no interaction during the decay period. Stable-state labels and masked-state labels are mapped to a knowledge graph, and positive activation components and negative repulsion components are calculated to construct a bidirectional semantic potential vector. The bidirectional semantic potential vectors of users are clustered to obtain groups, and the feature vectors of the content to be pushed are extracted. The similarity between the feature vectors of the content to be pushed and the centroid vectors of the groups is calculated using an attention mechanism to generate potential affinity. A comprehensive factor calculation and ranking process is used to generate a push list. The factors include the initial matching score of the feature vector of the content to be pushed and the bidirectional semantic potential field vector, the potential affinity of the content to be pushed to the user's group, and the stable label cognitive load index. When the push frequency exceeds a threshold, the stable label cognitive load index suppresses the weight of the stable label and the neighboring nodes of the knowledge graph and increases the weight of the distant nodes of the knowledge graph. The decay rate of the stable label cognitive load index is negatively correlated with the breadth of interest represented by the bidirectional semantic potential field vector.

[0004] One or more embodiments of the present invention also provide a precise push system based on user profiles, the system comprising: A module is established to extract text interaction features and behavior sequence features from users' historical behavior data to construct an initial user profile; a transition model is established for profile tags in exploratory, stable, and blocked states. Profile tags migrate from the exploratory state to the stable state based on the frequency of positive interactions exceeding a positive threshold, and profile tags migrate to the blocked state based on the number of negative feedback events of negative interaction events exceeding a negative threshold. Stable state tags regress to the exploratory state if there is no interaction during the decay period. The generation module maps stable-state labels and masked-state labels to the knowledge graph, calculates positive activation components and negative repulsion components to construct bidirectional semantic potential vectors, clusters the bidirectional semantic potential vectors of users to obtain groups, extracts the feature vectors of the content to be pushed, and uses an attention mechanism to calculate the similarity between the feature vectors of the content to be pushed and the centroid vectors of the groups to generate potential affinity. The push module is used to generate a push list by calculating and sorting factors. The factors include the initial matching score of the feature vector of the content to be pushed and the bidirectional semantic potential field vector, the potential affinity of the content to be pushed to the user's group, and the stable label cognitive load index. When the push frequency exceeds a threshold, the stable label cognitive load index suppresses the weight of the stable label and the neighboring nodes of the knowledge graph and increases the weight of the distant nodes of the knowledge graph. The decay rate of the stable label cognitive load index is negatively correlated with the breadth of interest represented by the bidirectional semantic potential field vector.

[0005] This invention extracts historical user behavior features to construct user profiles and establishes a tag migration model encompassing exploratory, stable, and shielded states. It utilizes positive thresholds, negative feedback, and decay cycles to represent the trajectory of user interest changes. Tags in different states are mapped to a knowledge graph, and a bidirectional semantic potential vector is constructed to represent users' positive preferences and negative aversion tendencies. User groups are segmented, and attention mechanisms are used to calculate potential content affinity, exploring users' implicit needs. During the ranking generation stage, a cognitive load index is used to prevent user fatigue with homogeneous content by suppressing the weight of frequently pushed tags and neighboring nodes in the knowledge graph while increasing the weight of distant nodes. This broadens information reach boundaries and optimizes overall push quality and the user's long-term browsing experience. Attached Figure Description

[0006] Figure 1 A flowchart illustrating a user profile-based targeted push notification method; Figure 2 This is a sequence diagram of tag state transitions based on user interaction; Figure 3 This is a comparison chart of the ablation experimental effects of the core modules of the recommendation model. Detailed Implementation

[0007] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0008] In this application, a precise push method based on user profiles is described, such as... Figure 1 As shown, the method includes: S1. Extract text interaction features and behavior sequence features from user historical behavior data to construct an initial user profile.

[0009] We retrieved user browsing, liking, and comment text data from the past 90 days using a relational database MySQL. The NLTK natural language processing toolkit was used to segment and remove stop words from the comment text data. The TF-IDF algorithm was applied to extract the weights of text keywords and construct a text interaction feature vector. The timestamps of browsing and liking records were processed using the Pandas time series processing library, sorting the discrete behaviors chronologically. The temporal dependencies of the behaviors were extracted using a Long Short-Term Memory (LSTM) network, outputting a fixed-dimensional latent feature vector of the behavior sequence. The text interaction feature vector and the latent feature vector of the behavior sequence were concatenated along the feature dimension, and Principal Component Analysis (PCA) was used for dimensionality reduction, outputting a fixed-dimensional comprehensive feature vector as the initial user profile. A tag extraction algorithm based on a large language model was used to extract entities from the comment text data and browsing content, generating a discrete user profile tag set. The comprehensive feature vector and the user profile tag set were used together as the initial user profile.

[0010] In some implementations, the extraction of text interaction features and behavioral sequence features from user historical behavior data to construct an initial user profile includes: Obtain the user's content click records and page interaction timestamp sequence within a preset time window as historical behavior data; The text in the content click record is segmented and vectorized using a natural language processing model to extract the text interaction features. The time difference between adjacent actions is calculated for the page interaction timestamp sequence, and the swipe and click event types are encoded to generate the behavior sequence features that represent the user's reading depth; The text interaction features and the behavior sequence features are concatenated to generate a multi-dimensional initial user vector representation, thereby constructing the initial user profile.

[0011] The preferred range for the preset time window parameter is 7 to 30 days, for example, 14 days. The system retrieves all content click records and page interaction timestamp sequences of the target user within this time window from the backend log database. A natural language processing (NLP) model is then invoked. This model consists of an input layer, multiple self-attention encoding layers, and an output layer. The input is a word sequence vector after word segmentation, and the output is the extracted deep semantic feature vector.

[0012] After segmenting the text using a pre-built Chinese dictionary and removing stop words, the resulting vector is transformed into a 768-dimensional word embedding vector. Average pooling is then used to obtain the overall text interaction feature vector. For behavioral sequence features, page interaction timestamp sequences are extracted, and the time difference between adjacent actions is calculated; typically, the time difference ranges from 1 to 300 seconds. One-hot encoding of swipe events is defined as a sequence of 1s and 0s, and one-hot encoding of click events is defined as a sequence of 0s and 1s. The time difference is discretized and binned; for example, 0s to 10s, 10s to 30s, and greater than 30s are mapped to numerical weights of 0.2, 0.5, and 1.0, respectively. These bins are then concatenated with the one-hot encoding of the event type using a tensor.

[0013] The concatenated feature sequence tensor is input into a Long Short-Term Memory (LSTM) network for temporal modeling. The LTM network structure includes a forget gate, an input gate, and an output gate. The input is the concatenated feature tensor for each time step and the hidden state vector from the previous time step. The output is the hidden layer feature vector for each time step. This extracts a 128-dimensional hidden layer state vector, which serves as the behavioral sequence feature representing the user's reading depth.

[0014] In the feature fusion layer, 768-dimensional text interaction features and 128-dimensional behavioral sequence features are concatenated along the feature channel dimension, and then dimensionality reduction is performed through a multilayer perceptron with a linear rectified activation function. The multilayer perceptron structure includes an input layer, a fully connected hidden layer, and an output layer. The input is the fused feature vector, which performs feature mapping internally using the formula Y=max{0,WX+b}. The output is a multi-dimensional user initial vector representation with a fixed dimension of 512. This multi-dimensional representation is stored as node attribute values ​​in the corresponding user node in the distributed graph database. The Top-N entity words are extracted from the text of the target user's interaction records using the TextRank algorithm as discrete profile labels and added to the label set of the node attributes, thus completing the construction of the initial user profile.

[0015] S2. Establish a transition model for the portrait tag in the exploration state, stable state and shielded state. The portrait tag migrates from the exploration state to the stable state based on the positive interaction frequency exceeding the positive threshold. The portrait tag migrates to the shielded state based on the negative feedback number of negative interaction events exceeding the negative threshold. If the stable state tag has no interaction during the decay period, it regresses to the exploration state.

[0016] In the in-memory database Redis, a state machine model is initialized for each initial user profile's tags, with the initial state of each tag set to the exploratory state. The Apache Flink streaming engine monitors user interaction event data streams in real time. When a positive interaction event (like or favorite) is detected, an accumulator is called to increment the tag's positive interaction count. A conditional statement compares this count with a preset positive threshold of 5. If the count is greater than 5, the tag's state field is updated to a stable state via the state machine. When a negative interaction event (skip or report) is detected, a negative feedback accumulator is called to increment the tag's negative feedback count. A conditional statement compares the negative feedback count with a preset negative threshold of 3. If the count is greater than 3, the tag's state field is updated and locked to a masked state. For all tags in a stable state, a scheduled verification task is registered in the Quartz task scheduling framework and executed every 24 hours. The time difference between the current system time and the last interaction of the tag is calculated. When the time difference exceeds a preset decay period of seven days, the tag's state field is overwritten back to the exploratory state and the positive interaction count is cleared to zero via the state machine's rollback interface.

[0017] In some implementations, the establishment of an exploratory, stable, and shielded state transition model for the image tags includes: Initialize the newly generated profile tags in the initial user profile to the exploration state and add them to the observation queue; Listen for positive user interaction events. If the cumulative number of clicks, favorites, and shares of a certain exploratory state tag during the first observation period exceeds the positive interaction frequency, then the exploratory state tag is migrated to a stable state. Listen for negative user interaction events. If the cumulative number of skips and blacklists of a certain exploration state tag or stable state tag within the same observation period exceeds the negative feedback number, then the exploration state tag or stable state tag will be moved to the blocked state. A periodic scanning mechanism is initiated. If the stable state label does not trigger any relevant interaction events within the decay period, the accumulated count is cleared, and the stable state label is rolled back to the exploration state.

[0018] A distributed state machine is maintained within the caching mechanism. When a new profile tag is parsed and added to the database, it is initialized to the exploratory state, and the tag's unique identifier is added to the observation queue of the asynchronous message queue. The preferred range for the first observation period is set to 48 to 120 hours, for example, 72 hours. Within this period, the frequency of positive user interactions is accumulated using a real-time sliding window through the stream processing engine.

[0019] A positive threshold of 50 points is set for the overall score, and different interaction behaviors are assigned differentiated weights, such as 2 points for reading and clicking, 10 points for adding to favorites, and 15 points for forwarding and sharing. When the accumulated score exceeds the positive threshold of 50 points, the state update interface is triggered to change the status of the tag in the profile database to a stable state and persist it.

[0020] Simultaneously, negative events within the same observation period are monitored, and a negative threshold is set, for example, to 3 times. If, within the same 72-hour observation period, the cumulative number of skip events and active blacklisting events related to the content of the exploration or stable state tag is detected to exceed 3 times, the tag status is set to the blocked state, and the content is written into the Bloom filter for content recall and interception.

[0021] In addition, a periodic scanning mechanism is initiated through a timed scheduling platform. The preferred range for the decay period is 15 to 30 days, for example, 21 days. When the scan detects that the last positive interaction timestamp of a stable-state tag is more than 21 days old, the cumulative score of all its behavioral features is reset to zero, and the tag state is reverted to the exploration state and re-added to the observation queue, thus obtaining a closed loop of tag state transition. Figure 2 As shown, the single-label state transition process is illustrated.

[0022] S3 maps stable-state labels and masked-state labels to the knowledge graph, and calculates positive activation components and negative repulsion components to construct a bidirectional semantic potential vector.

[0023] A pre-built multimodal knowledge graph is stored using the Neo4j graph database. The Node2Vec graph embedding algorithm is used to convert each node in the graph into a low-dimensional, dense semantic vector. A string matching algorithm is used to map stable and masked user profile tags to node entities in Neo4j. For successfully mapped stable-state tag nodes, their corresponding graph embedding semantic vectors are extracted. Positive activation weights are calculated using a logarithmic function based on the positive interaction counts of the stable-state tags. These positive activation weights are multiplied by the graph embedding semantic vectors and then weighted and summed to obtain the positive activation component vector. For successfully mapped masked tag nodes, their corresponding graph embedding semantic vectors are extracted. Negative repulsion weights are calculated using a logarithmic function based on the negative feedback counts of the masked-state tags. These negative repulsion weights are multiplied by the graph embedding semantic vectors and then weighted and summed to obtain the negative repulsion component vector. The positive activation component vector and the negative repulsion component vector are then concatenated along the feature dimension to construct the user's bidirectional semantic potential vector.

[0024] In some implementations, mapping stable-state labels and masked-state labels to a knowledge graph, and calculating positive activation components and negative repulsion components to construct a bidirectional semantic potential vector, includes: In a pre-constructed multi-level semantic knowledge graph of the content domain, locate the entity nodes of the stable state label and the masked state label; Centered on the stable label entity node, the adjacent nodes within a preset number of hops are extracted, and the features of the adjacent nodes are aggregated using a graph convolutional network to generate the positive activation component. Centered on the masked label entity node, the graph association node is extracted, and the penalty weight is calculated through a multilayer perceptron to generate the negative repulsion component; The positive activation component and the negative repulsion component are concatenated and normalized to generate the bidirectional semantic potential vector.

[0025] Based on a graph database, a multi-level semantic knowledge graph of content domains is pre-constructed, comprising four semantic levels: major categories, sub-topics, specific entities, and abstract keywords. When processing profile data, word vector cosine matching is performed using entity linking and the optimal matching algorithm to anchor the current stable-state and masked-state labels of the profile database to the corresponding unique entity nodes in the graph.

[0026] For the stable label node of the mapping, a preset hop count is preferably set to 2 to 3 hops. For example, if set to 2 hops, a breadth-first search algorithm is used to extract all neighboring nodes within 2 hops of the central node to construct a local subgraph. The subgraph topology is input into a two-layer graph convolutional network model. The graph convolutional network structure includes multiple layers of graph convolution operators. The inputs are the subgraph topology adjacency matrix and the initial feature matrix of the nodes. The internal calculation formula is expressed as follows: in Given an adjacency matrix containing self-loops and is the hidden node feature matrix, is the degree matrix, and the output is the node feature representation matrix that aggregates multi-hop graph structure information.

[0027] The first layer of the graph convolutional network aggregates and propagates the initial 128-dimensional embeddings of all neighboring nodes using a linear rectified activation function. The second layer performs semantic aggregation and outputs a 256-dimensional positive activation component through mean pooling. For a masked label node, the one-hop association graph node centered on it is extracted, and the features of multiple association nodes are aggregated using mean pooling to obtain a 128-dimensional attribute vector for the single association node. The negative feedback count value of the corresponding masked label is obtained, and a one-dimensional negative feedback feature is generated after smoothing using a logarithmic function. The 128-dimensional attribute vector of the single-hop association node and the one-dimensional negative feedback feature are concatenated and input into a pre-configured three-layer multilayer perceptron, where the number of fully connected hidden layer neurons is set to 128, 64, and 1 respectively. The last layer of the multilayer perceptron uses a logistic function to output a continuous penalty weight coefficient scalar between 0.01 and 0.99. The 128-dimensional association node feature vector is multiplied by the aforementioned penalty weight coefficient to extract a decayed 128-dimensional negative repulsion component.

[0028] The 256-dimensional positive activation component and the 128-dimensional negative repulsion component are concatenated along the feature dimensions to obtain a 384-dimensional fused vector. A L2 normalization operation is then performed on this concatenated vector, which involves dividing the value of each dimension of the vector by the square root of the sum of the squares of all dimension values, outputting a 384-dimensional bidirectional semantic potential vector with a unit length of 1.

[0029] S4 clusters the bidirectional semantic potential vectors of users to obtain groups, extracts the feature vectors of the content to be pushed, and uses an attention mechanism to calculate the similarity between the feature vectors of the content to be pushed and the centroid vectors of the groups to generate potential affinity.

[0030] Unsupervised clustering of the bidirectional semantic potential vectors of all users is performed using the KMeans clustering algorithm. The silhouette coefficient evaluation method is used to determine the optimal number of clusters, dividing users into multiple interest groups and calculating the bidirectional semantic potential centroid vector of each group. Multimedia objects to be pushed are obtained from the content library. Image frame features are extracted using a deep residual network ResNet, and semantic features of titles and descriptions are extracted using a pre-trained language model BERT. The image frame features and semantic features are fused in a multimodal manner to obtain the feature vector of the content to be pushed. The multi-head attention mechanism module in the TensorFlow deep learning framework is invoked. The feature vector of the content to be pushed is used as the query vector, and the group centroid vector is simultaneously used as the key vector and value vector, input into the attention network. The relevance score between the query vector and the key vector is calculated through inner product operation. The relevance score is normalized into an attention weight matrix using the Softmax function. The normalized values ​​of each dimension of the attention weight matrix are extracted as the potential affinity distribution matrix for each group corresponding to the content to be pushed.

[0031] In some implementations, the clustering of the user's bidirectional semantic potential vectors to obtain groups includes: Initialize a preset number of group centroids, and obtain the bidirectional semantic potential vectors of all users in each scheduling cycle; Calculate the Euclidean distance between each bidirectional semantic potential vector and the centroid of each group, and divide each user into the group with the smallest distance; For each group after division, calculate the arithmetic mean of the bidirectional semantic potential vectors of all users in the group, and generate the updated group centroid vector; When the position offset of the group centroid vector is less than the convergence threshold in multiple consecutive scheduling cycles, the user group partitioning result is output.

[0032] A centroid-based clustering algorithm is used, and a preset number of group centroids is specified during the initialization phase, preferably ranging from 50 to 500, for example, 100 group centroids are set.

[0033] A random sampling algorithm is used to initialize and extract 100 distinct 384-dimensional initial group centroid vectors from the historical feature database of existing active users. A scheduling period parameter is set, for example, to 24 hours. Upon triggering execution, the latest 384-dimensional bidirectional semantic potential vector of all currently active users is obtained. The Euclidean distance between each user's 384-dimensional bidirectional semantic potential vector and the existing 100 group centroid vectors is calculated in parallel, and each user is assigned to the group with the smallest Euclidean distance.

[0034] After all users are allocated, for the 100 newly divided groups, the arithmetic mean of the bidirectional semantic potential vectors of all members within each group is calculated to generate 100 updated 384-dimensional group centroid vectors. A convergence threshold parameter for the displacement offset is set, for example, a L2 norm displacement offset less than 0.001. Convergence is determined when the maximum position offset of the group centroid vectors of the 100 groups is completely below 0.001 in the most recent three consecutive scheduling cycles.

[0035] At this point, the calculation is terminated, and the user group segmentation results and the updated group centroid vectors are output and saved as the basic configuration for recommendation recall.

[0036] In some implementations, the step of extracting the feature vector of the content to be pushed, calculating the similarity between the feature vector of the content to be pushed and the group centroid vector using an attention mechanism, and generating potential affinity includes: The text of the content to be pushed is encoded using a pre-trained language model, and the feature vector of the content to be pushed is extracted. The feature vector of the content to be pushed is used as the query vector, and the centroid vectors of each group are used as the key vector and value vector, respectively, and then input into the attention mechanism model. The attention score is calculated by the dot product of the query vector and each key vector using the attention mechanism model, and then Softmax normalization is performed. The normalized attention score is used as the potential affinity of the content to be pushed to different groups.

[0037] When content to be pushed enters the system, a pre-trained language model is invoked to parse the content in blocks and perform text-level encoding within a length threshold, for example, a fixed limit of 512 tokens. Feature vectors representing the complete text semantics are extracted and reduced from 768 dimensions to 384 dimensions using a feedforward neural network, serving as the feature vector for the content to be pushed.

[0038] An attention mechanism model is constructed, comprising a feature mapping layer, a scaled dot product attention calculation module, and a normalized activation layer. The inputs are a query vector and a set of key-value pair vectors. The 384-dimensional feature vector of the content to be pushed is used as the query vector Q, and 100 pre-constructed 384-dimensional group centroid vectors are used as the key vector K and value vector V, respectively. The dot product of the query vector Q and the key vector K is calculated and divided by a scaling factor to obtain a stable dot product attention score.

[0039] The generated attention scores are compressed to the range of 0 to 1 using Softmax normalization, ensuring that the overall sum is 1. The normalized attention scores are then output and stored in the feature library as the potential affinity of the content to be pushed to different groups.

[0040] S5, a push list is generated by calculating and sorting comprehensive factors. The factors include the initial matching score of the feature vector of the content to be pushed and the bidirectional semantic potential field vector, the potential affinity of the content to be pushed to the user's group, and the stable label cognitive load index. When the push frequency exceeds a threshold, the stable label cognitive load index suppresses the weight of the stable label and the neighboring nodes of the knowledge graph and increases the weight of the distant nodes of the knowledge graph. The decay rate of the stable label cognitive load index is negatively correlated with the breadth of interest represented by the bidirectional semantic potential field vector.

[0041] The feature vector of the content to be pushed is segmented into positive sub-vectors of the same dimension as the positive activation component and negative sub-vectors of the same dimension as the negative repulsion component using a feature alignment mapping model according to the target dimension. The cosine similarity algorithm is used to calculate the positive similarity between each positive sub-vector and the positive activation component of a specific user, and the cosine similarity between each negative sub-vector and the negative repulsion component of a specific user is calculated to obtain the negative similarity. The positive similarity minus the negative similarity is used as the initial matching score. The absolute value of the positive activation component in the user's bidirectional semantic potential vector is taken and transformed into a probability distribution using a Softmax function. The Shannon entropy is then calculated using the information entropy algorithm as the breadth of interest. A distributed caching system is used to count the number of times each stable-state tag is pushed to the client in the past hour as the push frequency. When the push frequency exceeds a preset threshold of ten times, a penalty mechanism is activated to calculate the cognitive load index. The first- and second-order neighboring nodes and third-order and above distant nodes of a stable label in the knowledge graph are obtained through a graph convolutional neural network. An inverse proportional function is used to calculate a negative inhibition multiplier (base e) applied to the stable label and its neighboring nodes to form the initial feature representation of the graph. Simultaneously, a natural exponential function is used to calculate a positive boosting multiplier applied to the activation components of distant nodes excluding the masked nodes, achieving feature weight redistribution. Combined with the user's interest breadth value, a linear mapping formula is used to ensure that the rate of decrease in cognitive load over time is inversely proportional to the interest breadth value. For users with high interest breadth, their profile tags are rich, providing a sufficient pool of cross-domain alternative content. When a cognitive load penalty is triggered for a high-frequency stable label, a slower decay rate is set, thereby forcibly distributing potential interest content from other domains and improving the diversity of recommendations. Conversely, for users with extremely low interest breadth, there is a lack of an effective alternative recommendation pool; if the decay rate is too slow, these users will churn due to a prolonged lack of desired content. Therefore, a faster decay rate is set for vertical groups, allowing them to quickly restore the supply of core preferences after a very brief content interruption, thereby maximizing the overall user stickiness and activity of the platform. The calculated initial matching score, potential affinity, and adjusted steady-state tag cognitive load index are input into the gradient boosting tree LightGBM ranking model for feature cross-referencing and regression scoring. Based on the regression scoring results, all candidate content to be pushed is sorted in descending order, and the top 50 content items with the highest scores are selected to generate a push list sent to the user's terminal.

[0042] In some implementations, the comprehensive factor calculation and sorting to generate the push list includes: Read the current stable-state label cognitive load index of the target user's stable-state label, use the stable-state label cognitive load index to dynamically adjust the attribute representation of the knowledge graph node, suppress the initial feature vector of the node corresponding to the stable-state label and the neighboring node of the knowledge graph, and improve the initial feature vector of the node corresponding to the distant node of the knowledge graph. Re-execute the graph aggregation mechanism to obtain the updated bidirectional semantic potential field vector. After projecting and aligning the feature vector of the content to be pushed to the feature space where the updated bidirectional semantic potential vector is located, the cosine similarity between its positive projection part and the positive activation component in the updated bidirectional semantic potential vector, and the cosine similarity between its negative projection part and the negative repulsion component are calculated respectively. The positive cosine similarity and the negative cosine similarity are subtracted to obtain the initial matching score. Query the attention score of the content to be pushed to the target user's group from the potential affinity; The initial matching score, the attention score of the target user's group, and the steady-state label cognitive load index are input into the regression ranking model to obtain the ranking score of the content to be pushed, and the push list is generated by arranging the content in descending order of the ranking score.

[0043] During recommendation calculations, a stable-state tag cognitive load index, ranging from 0 to 1, is obtained for the target user's stable-state tags. For example, the cognitive load index increases when the content of a certain category tag is displayed more than 5 times in a day. The stable-state tag cognitive load index is used to dynamically assign weights to the graph entity attribute space. The initial graph feature vectors corresponding to stable-state tags and neighboring nodes within 0 to 2 hops of the knowledge graph are multiplied by a suppression parameter, and the initial graph feature vectors corresponding to unmasked nodes in distant nodes at 3 hops or more of the knowledge graph are multiplied by a boosting parameter. The parameter-weighted initial features of each node are then re-inputted into the graph convolutional network and multilayer perceptron to aggregate and obtain an updated 384-dimensional bidirectional semantic potential field vector.

[0044] After aligning the feature vector of the content to be pushed to the feature space of the updated bidirectional semantic potential vector using a multilayer perceptron mapping, it is then divided into a 256-dimensional positive projection part and a 128-dimensional negative projection part according to the feature dimensions. The cosine similarity between the positive projection part of the content to be pushed and the 256-dimensional positive activation components of the updated bidirectional semantic potential vector is calculated, as is the cosine similarity between the negative projection part and the 128-dimensional negative repulsion components. The initial matching score is obtained by subtracting the cosine similarity of the negative part from the cosine similarity of the positive part.

[0045] The system queries the group identifier of the target user and extracts the attention score of the target user's group corresponding to the content to be pushed from the potential affinity. It then extracts the target user's current stable-state label cognitive load index and uses the calculated initial matching score, attention score, and cognitive load index as a three-dimensional combined feature, uniformly inputting them into a pre-trained gradient boosting tree LightGBM ranking model. High-order feature crosses are completed through the non-linear leaf node mapping of the tree model, calculating the regression ranking score of the content to be pushed. All content to be pushed is sorted in descending order of ranking score, and a specific number of top-ranked content are extracted to generate a push list, which is then sent to the user's front-end interface for display.

[0046] The experiment used a dataset of nearly 30 days' complete interaction logs from 800,000 active users randomly sampled from a content distribution platform. The first 23 days' data were used for model training and basic graph construction, while the last 7 days' data were used for offline validation and performance evaluation. Evaluation metrics included content click-through rate (CTR), recommendation result diversity index, and average user page dwell time. The baseline control group used a traditional dual-tower collaborative filtering model, employing only text features and behavioral sequence concatenation for cosine similarity recall and ranking. Ablation group one added forward graph convolution aggregation to the baseline control group but removed the label three-state transfer mechanism and the masked state negative repulsion component. Ablation group two retained the bidirectional semantic potential field vector but removed the cognitive load index suppression mechanism and the group potential affinity weighting module, performing feature space comparison. The experimental groups deployed the complete multi-dimensional collaborative recommendation framework described in this application for testing. The comparison results of content CTR for each experimental group are as follows: Figure 3 As shown, this illustrates the superposition gain of each core module.

[0047] Experimental results showed that the content click-through rate (CTR) of the baseline control group was 5.21%, the diversity index of recommendation results was 12.45, and the average user page dwell time was 135 seconds. In Ablation Group 1, the CTR increased to 6.12%, the diversity index slightly rose to 12.80, and the dwell time reached 158 seconds. In Ablation Group 2, the CTR increased to 6.54%, the diversity index was 13.15, and the dwell time extended to 176 seconds. The experimental group using the complete protocol showed the best performance across all core metrics, with its CTR jumping to 7.35%, the diversity index of recommendation results rising to 16.78, and the average user page dwell time reaching a high of 214 seconds.

[0048] A comparison between the complete experimental group and ablation group one shows that by utilizing the three-state transfer model and the negative repulsion component of the shielded multilayer perceptron, offensive content was blocked, improving user conversion efficiency and reading immersion. Compared to ablation group two, the complete experimental group, by reducing the weight of high-frequency exposure categories and increasing the weight of distant domain graph nodes through the cognitive load index, suppressed the information cocoon effect, resulting in a relative increase of over 27% in the diversity index. Combined with affinity fusion calculation based on the group centroid attention mechanism, personalized recommendations were ensured while smoothly expanding interests, opening up potential reading needs, and increasing the platform's overall user stickiness.

[0049] In this application, a precise push system based on user profiles is disclosed, the system comprising: A module is established to extract text interaction features and behavior sequence features from users' historical behavior data to construct an initial user profile; a transition model is established for profile tags in exploratory, stable, and blocked states. Profile tags migrate from the exploratory state to the stable state based on the frequency of positive interactions exceeding a positive threshold, and profile tags migrate to the blocked state based on the number of negative feedback events of negative interaction events exceeding a negative threshold. Stable state tags regress to the exploratory state if there is no interaction during the decay period. The generation module maps stable-state labels and masked-state labels to the knowledge graph, calculates positive activation components and negative repulsion components to construct bidirectional semantic potential vectors, clusters the bidirectional semantic potential vectors of users to obtain groups, extracts the feature vectors of the content to be pushed, and uses an attention mechanism to calculate the similarity between the feature vectors of the content to be pushed and the centroid vectors of the groups to generate potential affinity. The push module is used to generate a push list by calculating and sorting factors. The factors include the initial matching score of the feature vector of the content to be pushed and the bidirectional semantic potential field vector, the potential affinity of the content to be pushed to the user's group, and the stable label cognitive load index. When the push frequency exceeds a threshold, the stable label cognitive load index suppresses the weight of the stable label and the neighboring nodes of the knowledge graph and increases the weight of the distant nodes of the knowledge graph. The decay rate of the stable label cognitive load index is negatively correlated with the breadth of interest represented by the bidirectional semantic potential field vector.

[0050] In some implementations, the extraction of text interaction features and behavioral sequence features from user historical behavior data to construct an initial user profile includes: Obtain the user's content click records and page interaction timestamp sequence within a preset time window as historical behavior data; The text in the content click record is segmented and vectorized using a natural language processing model to extract the text interaction features. The time difference between adjacent actions is calculated for the page interaction timestamp sequence, and the swipe and click event types are encoded to generate the behavior sequence features that represent the user's reading depth; The text interaction features and the behavior sequence features are concatenated to generate a multi-dimensional initial user vector representation, thereby constructing the initial user profile.

[0051] In some implementations, the establishment of an exploratory, stable, and shielded state transition model for the image tags includes: Initialize the newly generated profile tags in the initial user profile to the exploration state and add them to the observation queue; Listen for positive user interaction events. If the cumulative number of clicks, favorites, and shares of a certain exploratory state tag during the first observation period exceeds the positive interaction frequency, then the exploratory state tag is migrated to a stable state. Listen for negative user interaction events. If the cumulative number of skips and blacklists of a certain exploration state tag or stable state tag within the same observation period exceeds the negative feedback number, then the exploration state tag or stable state tag will be moved to the blocked state. A periodic scanning mechanism is initiated. If the stable state label does not trigger any relevant interaction events within the decay period, the accumulated count is cleared, and the stable state label is rolled back to the exploration state.

[0052] In some implementations, mapping stable-state labels and masked-state labels to a knowledge graph, and calculating positive activation components and negative repulsion components to construct a bidirectional semantic potential vector, includes: In a pre-constructed multi-level semantic knowledge graph of the content domain, locate the entity nodes of the stable state label and the masked state label; Centered on the stable label entity node, the adjacent nodes within a preset number of hops are extracted, and the features of the adjacent nodes are aggregated using a graph convolutional network to generate the positive activation component. Centered on the masked label entity node, the graph association node is extracted, and the penalty weight is calculated through a multilayer perceptron to generate the negative repulsion component; The positive activation component and the negative repulsion component are concatenated and normalized to generate the bidirectional semantic potential vector.

[0053] In some implementations, the clustering of the user's bidirectional semantic potential vectors to obtain groups includes: Initialize a preset number of group centroids, and obtain the bidirectional semantic potential vectors of all users in each scheduling cycle; Calculate the Euclidean distance between each bidirectional semantic potential vector and the centroid of each group, and divide each user into the group with the smallest distance; For each group after division, calculate the arithmetic mean of the bidirectional semantic potential vectors of all users in the group, and generate the updated group centroid vector; When the position offset of the group centroid vector is less than the convergence threshold in multiple consecutive scheduling cycles, the user group partitioning result is output.

[0054] In some implementations, the step of extracting the feature vector of the content to be pushed, calculating the similarity between the feature vector of the content to be pushed and the group centroid vector using an attention mechanism, and generating potential affinity includes: The text of the content to be pushed is encoded using a pre-trained language model, and the feature vector of the content to be pushed is extracted. The feature vector of the content to be pushed is used as the query vector, and the centroid vectors of each group are used as the key vector and value vector, respectively, and then input into the attention mechanism model. The attention score is calculated by the dot product of the query vector and each key vector using the attention mechanism model, and then Softmax normalization is performed. The normalized attention score is used as the potential affinity of the content to be pushed to different groups.

[0055] In some implementations, the comprehensive factor calculation and sorting to generate the push list includes: Read the current stable-state label cognitive load index of the target user's stable-state label, use the stable-state label cognitive load index to dynamically adjust the attribute representation of the knowledge graph node, suppress the initial feature vector of the node corresponding to the stable-state label and the neighboring node of the knowledge graph, and improve the initial feature vector of the node corresponding to the distant node of the knowledge graph. Re-execute the graph aggregation mechanism to obtain the updated bidirectional semantic potential field vector. After projecting and aligning the feature vector of the content to be pushed to the feature space where the updated bidirectional semantic potential vector is located, the cosine similarity between its positive projection part and the positive activation component in the updated bidirectional semantic potential vector, and the cosine similarity between its negative projection part and the negative repulsion component are calculated respectively. The positive cosine similarity and the negative cosine similarity are subtracted to obtain the initial matching score. Query the attention score of the content to be pushed to the target user's group from the potential affinity; The initial matching score, the attention score of the target user's group, and the steady-state label cognitive load index are input into the regression ranking model to obtain the ranking score of the content to be pushed, and the push list is generated by arranging the content in descending order of the ranking score.

[0056] In this specification, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise limited, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. In this document, "a," "an," "the," "the," and "its" may also include plural forms unless the context clearly indicates otherwise. "Multiple" refers to at least two, such as 2, 3, 5, or 8, etc. "And / or" includes any and all combinations of the associated listed items.

[0057] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0058] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A precise push method based on user profiles, characterized in that, include: Extract text interaction features and behavior sequence features from users' historical behavior data to construct an initial user profile; A transition model for exploratory, stable, and shielded states is established for profile tags. Profile tags migrate from the exploratory state to the stable state based on the positive interaction frequency exceeding a positive threshold. Profile tags migrate to the shielded state based on the negative feedback number of negative interaction events exceeding a negative threshold. Stable state tags regress to the exploratory state if there is no interaction during the decay period. Stable-state labels and masked-state labels are mapped to a knowledge graph, and positive activation components and negative repulsion components are calculated to construct a bidirectional semantic potential vector. The bidirectional semantic potential vectors of users are clustered to obtain groups, and the feature vectors of the content to be pushed are extracted. The similarity between the feature vectors of the content to be pushed and the centroid vectors of the groups is calculated using an attention mechanism to generate potential affinity. A comprehensive factor calculation and ranking process is used to generate a push list. The factors include the initial matching score of the feature vector of the content to be pushed and the bidirectional semantic potential field vector, the potential affinity of the content to be pushed to the user's group, and the stable label cognitive load index. When the push frequency exceeds a threshold, the stable label cognitive load index suppresses the weight of the stable label and the neighboring nodes of the knowledge graph and increases the weight of the distant nodes of the knowledge graph. The decay rate of the stable label cognitive load index is negatively correlated with the breadth of interest represented by the bidirectional semantic potential field vector.

2. The method according to claim 1, characterized in that, The initial user profile is constructed by extracting text interaction features and behavior sequence features from historical user behavior data, including: Obtain the user's content click records and page interaction timestamp sequence within a preset time window as historical behavior data; The text in the content click record is segmented and vectorized using a natural language processing model to extract the text interaction features. The time difference between adjacent actions is calculated for the page interaction timestamp sequence, and the swipe and click event types are encoded to generate the behavior sequence features that represent the user's reading depth; The text interaction features and the behavior sequence features are concatenated to generate a multi-dimensional initial user vector representation, thereby constructing the initial user profile.

3. The method according to claim 1, characterized in that, The establishment of an exploration state, stable state, and masked state transition model for the image tags includes: Initialize the newly generated profile tags in the initial user profile to the exploration state and add them to the observation queue; Listen for positive user interaction events. If the cumulative number of clicks, favorites, and shares of a certain exploratory state tag during the first observation period exceeds the positive interaction frequency, then the exploratory state tag is migrated to a stable state. Listen for negative user interaction events. If the cumulative number of skips and blacklists of a certain exploration state tag or stable state tag within the same observation period exceeds the negative feedback number, then the exploration state tag or stable state tag will be moved to the blocked state. A periodic scanning mechanism is initiated. If the stable state label does not trigger any relevant interaction events within the decay period, the accumulated count is cleared, and the stable state label is rolled back to the exploration state.

4. The method according to claim 1, characterized in that, The process of mapping stable-state labels and masked-state labels to a knowledge graph, and calculating positive activation and negative repulsion components to construct a bidirectional semantic potential vector, includes: In a pre-constructed multi-level semantic knowledge graph of the content domain, locate the entity nodes of the stable state label and the masked state label; Centered on the stable label entity node, the adjacent nodes within a preset number of hops are extracted, and the features of the adjacent nodes are aggregated using a graph convolutional network to generate the positive activation component. Centered on the masked label entity node, the graph association node is extracted, and the penalty weight is calculated through a multilayer perceptron to generate the negative repulsion component; The positive activation component and the negative repulsion component are concatenated and normalized to generate the bidirectional semantic potential vector.

5. The method according to claim 1, characterized in that, The process of clustering the bidirectional semantic potential vectors of users to obtain groups includes: Initialize a preset number of group centroids, and obtain the bidirectional semantic potential vectors of all users in each scheduling cycle; Calculate the Euclidean distance between each bidirectional semantic potential vector and the centroid of each group, and divide each user into the group with the smallest distance; For each group after division, calculate the arithmetic mean of the bidirectional semantic potential vectors of all users in the group, and generate the updated group centroid vector; When the position offset of the group centroid vector is less than the convergence threshold in multiple consecutive scheduling cycles, the user group partitioning result is output.

6. The method according to claim 1, characterized in that, The step of extracting the feature vector of the content to be pushed, calculating the similarity between the feature vector of the content to be pushed and the group centroid vector using an attention mechanism, and generating potential affinity includes: The text of the content to be pushed is encoded using a pre-trained language model, and the feature vector of the content to be pushed is extracted. The feature vector of the content to be pushed is used as the query vector, and the centroid vectors of each group are used as the key vector and value vector, respectively, and then input into the attention mechanism model. The attention score is calculated by the dot product of the query vector and each key vector using the attention mechanism model, and then Softmax normalization is performed. The normalized attention score is used as the potential affinity of the content to be pushed to different groups.

7. The method according to claim 1, characterized in that, The comprehensive factor calculation and sorting generates the push list, including: Read the current stable-state label cognitive load index of the target user's stable-state label, use the stable-state label cognitive load index to dynamically adjust the attribute representation of the knowledge graph node, suppress the initial feature vector of the node corresponding to the stable-state label and the neighboring node of the knowledge graph, and improve the initial feature vector of the node corresponding to the distant node of the knowledge graph. Re-execute the graph aggregation mechanism to obtain the updated bidirectional semantic potential field vector. After projecting and aligning the feature vector of the content to be pushed to the feature space where the updated bidirectional semantic potential vector is located, the cosine similarity between its positive projection part and the positive activation component in the updated bidirectional semantic potential vector, and the cosine similarity between its negative projection part and the negative repulsion component are calculated respectively. The positive cosine similarity and the negative cosine similarity are subtracted to obtain the initial matching score. Query the attention score of the content to be pushed to the target user's group from the potential affinity; The initial matching score, the attention score of the target user's group, and the steady-state label cognitive load index are input into the regression ranking model to obtain the ranking score of the content to be pushed, and the push list is generated by arranging the content in descending order of the ranking score.

8. A precise push system based on user profiles, characterized in that, include: A module is established to extract text interaction features and behavioral sequence features from users' historical behavior data to construct an initial user profile. A transition model for exploratory, stable, and shielded states is established for profile tags. Profile tags migrate from the exploratory state to the stable state based on the positive interaction frequency exceeding a positive threshold. Profile tags migrate to the shielded state based on the negative feedback number of negative interaction events exceeding a negative threshold. Stable state tags regress to the exploratory state if there is no interaction during the decay period. The generation module maps stable-state labels and masked-state labels to the knowledge graph, calculates positive activation components and negative repulsion components to construct bidirectional semantic potential vectors, clusters the bidirectional semantic potential vectors of users to obtain groups, extracts the feature vectors of the content to be pushed, and uses an attention mechanism to calculate the similarity between the feature vectors of the content to be pushed and the centroid vectors of the groups to generate potential affinity. The push module is used to generate a push list by calculating and sorting factors. The factors include the initial matching score of the feature vector of the content to be pushed and the bidirectional semantic potential field vector, the potential affinity of the content to be pushed to the user's group, and the stable label cognitive load index. When the push frequency exceeds a threshold, the stable label cognitive load index suppresses the weight of the stable label and the neighboring nodes of the knowledge graph and increases the weight of the distant nodes of the knowledge graph. The decay rate of the stable label cognitive load index is negatively correlated with the breadth of interest represented by the bidirectional semantic potential field vector.

9. The system according to claim 8, characterized in that, The initial user profile is constructed by extracting text interaction features and behavior sequence features from historical user behavior data, including: Obtain the user's content click records and page interaction timestamp sequence within a preset time window as historical behavior data; The text in the content click record is segmented and vectorized using a natural language processing model to extract the text interaction features. The time difference between adjacent actions is calculated for the page interaction timestamp sequence, and the swipe and click event types are encoded to generate the behavior sequence features that represent the user's reading depth; The text interaction features and the behavior sequence features are concatenated to generate a multi-dimensional initial user vector representation, thereby constructing the initial user profile.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program, when executed by a processor, implements the method as described in any one of claims 1-7.