Method and device for pushing content, electronic device, computer medium

CN122845652APending Publication Date: 2026-09-29BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610933702.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-29

Smart Images

  • Figure CN122845652A_ABST
    Figure CN122845652A_ABST
Patent Text Reader

Abstract

This disclosure provides methods, apparatus, electronic devices, and computer media for pushing content, relating to the field of data processing technology, particularly artificial intelligence, content recommendation, and graph neural networks. The specific implementation scheme is as follows: Obtaining a pre-constructed heterogeneous graph; the heterogeneous graph includes multiple types of nodes and multiple types of edges, the nodes including at least user nodes and push content nodes; obtaining a user behavior sequence; in response to the user behavior sequence satisfying a preset trigger condition, performing intent reasoning on the user behavior sequence, and dynamically updating the topology of the heterogeneous graph based on the intent reasoning result; using a heterogeneous graph neural network to update the node embeddings of the updated heterogeneous graph, obtaining updated embedding vectors corresponding to the user nodes and the push content nodes; recalling and ranking the push content based on the updated embedding vectors, and determining the final pushed content based on the ranking result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and more particularly to the fields of artificial intelligence, content recommendation, and graph neural networks. Specifically, this disclosure relates to a method and apparatus for recommending content, an electronic device, and a computer-readable storage medium. Background Technology

[0002] With the rapid development of mobile internet technology, various mobile applications have become the core carriers for users to obtain information and use services. As a key means for apps to reach users and improve activity and retention rates, push notifications directly affect user experience and operational effectiveness.

[0003] The core task of push notifications is to select the push message that the user is most interested in and most likely to click from a massive amount of candidate content. Summary of the Invention

[0004] This disclosure provides a method and apparatus for recommending content, an electronic device, and a computer-readable storage medium.

[0005] According to a first aspect of this disclosure, a method for recommending content is provided, the method comprising: Obtain a pre-built heterogeneous graph; the heterogeneous graph includes multiple types of nodes and multiple types of edges, and the nodes include at least user nodes and push content nodes; Obtain a user behavior sequence; in response to the user behavior sequence satisfying a preset triggering condition, perform intent reasoning on the user behavior sequence, and dynamically update the topology of the heterogeneous graph based on the intent reasoning result; A heterogeneous graph neural network is used to update the node embedding of the updated heterogeneous graph, thereby obtaining the updated embedding vectors corresponding to the user node and the push content node. The updated embedding vector is used to recall and sort the push content, and the final push content is determined based on the sorting results.

[0006] According to a second aspect of this disclosure, an apparatus for recommending content is provided, the apparatus comprising: A heterogeneous graph creation module is used to obtain a pre-built heterogeneous graph; the heterogeneous graph includes multiple types of nodes and multiple types of edges, and the nodes include at least user nodes and push content nodes; An intent reasoning module is used to acquire user behavior sequences; in response to the user behavior sequence satisfying a preset triggering condition, the module performs intent reasoning on the user behavior sequence and dynamically updates the topology of the heterogeneous graph based on the intent reasoning result. The graph update module is used to update the node embedding of the updated heterogeneous graph using a heterogeneous graph neural network to obtain the updated embedding vectors corresponding to the user node and the push content node. The recall module is used to recall and sort the push content based on the updated embedding vector, and determine the final push content to be sent based on the sorting results.

[0007] According to a third aspect of this disclosure, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to at least one of the aforementioned processors; wherein, The memory stores instructions that can be executed by at least one processor, which, when executed by at least one processor, enables the at least one processor to perform the method described above.

[0008] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the methods described above.

[0009] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the methods described above.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a flowchart illustrating a method for pushing content provided in an embodiment of this disclosure; Figure 2 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 3 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 4 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 5 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 6 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 7 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 8 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 9 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 10 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 11 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 12 This is a flowchart illustrating some steps of another method for pushing content provided in this embodiment of the disclosure; Figure 13 This is a schematic diagram of the structure of a device for pushing content provided in an embodiment of this disclosure; Figure 14 This is a block diagram of an electronic device used to implement the method for pushing content according to embodiments of the present disclosure. Detailed Implementation

[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0013] In some related technologies, recommendations can be made by calculating the similarity between users or between pushed content. However, in recommendation scenarios, the user-push interaction matrix is ​​usually extremely sparse, making it difficult for collaborative filtering methods to converge and limiting the recommendation effect.

[0014] Some related technologies utilize user profile features, push content features, and contextual features to rank and score users through models. These methods primarily rely on explicit feature engineering, making it difficult to model high-order relationships between user behaviors and failing to effectively utilize the rich behavioral data generated by users within the application.

[0015] In some related technologies, homogeneous or heterogeneous graphs of users and items can be constructed, and graph convolutional networks, graph attention networks, etc., can be used for message passing and learn node embeddings. However, the graph structure of these solutions is fixed and cannot be dynamically adjusted with real-time user behavior. Furthermore, they can only model static relationships and are difficult to capture instantaneous changes in user interests.

[0016] In some related technologies, LLM (Large Language Model) can be used to perform semantic understanding on push titles and summaries to generate text features to assist in recommendations. However, such solutions only stay at the text semantic level and do not associate with user behavior or graph structure. They cannot uncover users' deep behavioral intentions, nor can they integrate semantic features into global recommendations, resulting in low semantic utilization efficiency.

[0017] The methods, apparatuses, electronic devices, and computer-readable storage media provided in this disclosure are intended to solve at least one of the above-mentioned technical problems of the prior art.

[0018] The method of the recommended content provided in this disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be an in-vehicle device, user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. The method can be implemented by a processor calling computer-readable program instructions stored in memory. Alternatively, the method can be executed by a server.

[0019] Figure 1 A flowchart illustrating the method for providing recommended content according to embodiments of this disclosure is shown. For example... Figure 1 As shown in the figure, the method for providing the recommended content in the embodiments of this disclosure may include steps S110, S120, S130, and S140.

[0020] S110. Obtain a pre-built heterogeneous graph; the heterogeneous graph includes multiple types of nodes and multiple types of edges, and the nodes include at least user nodes and push content nodes.

[0021] Heterogeneous graphs refer to graph structures that contain multiple types of nodes and multiple types of edges.

[0022] In this embodiment of the disclosure, the node types of the heterogeneous graph include at least user nodes and push content nodes. User nodes are used to represent user entities of the application, and they can be associated with user profile features (such as age group, gender, activity level, etc.); push content nodes are used to represent message pushes to be sent, and they can be associated with attributes such as push title, cover image, and content type.

[0023] In some possible implementations, heterogeneous graphs can be pre-built and stored in a graph database or in-memory graph structure during the offline phase. The construction process may include: defining node types and their attribute fields; defining edge types and their weight semantics; and extracting entities and relationships from data sources such as user history behavior logs, app (application) behavior logs, and content tag libraries, and importing them into graph storage.

[0024] In some possible implementations, heterogeneous graphs can also include in-application entity nodes (such as products, videos, and articles), tag nodes, time context nodes, and intent nodes, as well as corresponding edge types (such as click edges, behavior edges, semantic association edges, etc.) for subsequent cross-domain fusion.

[0025] In some specific implementations, obtaining a pre-built heterogeneous graph can involve loading it from persistent storage into memory or a graph computing engine. Node features are stored as tensors, and edge indices and weights are stored as sparse matrices. If the heterogeneous graph is large, distributed graph storage can be used. During subsequent inference, some nodes (such as intent nodes) and edge weights of the heterogeneous graph will be dynamically updated based on real-time user behavior, but the graph's structural skeleton (node ​​type definitions, static nodes such as user and push content nodes) is pre-existing.

[0026] By pre-building heterogeneous graphs, multi-source heterogeneous data (push domain, in-app domain, tag domain) are unified into a knowledge structure, laying the foundation for subsequent cross-domain knowledge transfer and real-time intent reasoning, and avoiding the computational overhead of repeatedly building graph structures online.

[0027] S120. Obtain the user behavior sequence; in response to the user behavior sequence meeting the preset triggering conditions, perform intent reasoning on the user behavior sequence, and dynamically update the topology of the heterogeneous graph based on the intent reasoning result.

[0028] In this context, a user behavior sequence refers to a series of user actions within an application arranged chronologically. Each behavior event includes at least the behavior type (such as click, browse, search, favorite, add to cart, purchase), the behavior object (such as product ID, article ID, search keyword), and the timestamp (or chronological position) of the behavior.

[0029] In some possible implementations, user behavior logs reported by clients can be subscribed to in real time via message queues, or behavioral events can be collected through server-side instrumentation. After cleaning, deduplication, and formatting, the behavioral events are maintained in a sliding window manner to record the N most recent behaviors of each active user (e.g., the most recent 20 behaviors or behaviors within the last 5 minutes), forming a user behavior sequence.

[0030] In some possible implementations, the user behavior sequence can be stored as a list structure, with each element being a triple (type, object, timestamp), where type is the behavior type, object is the behavior object, and timestamp is the timestamp of the behavior occurrence.

[0031] In some specific implementations, when a user launches the app or performs a new action, the client or server gateway receives the action event. The stream processing engine groups the action events by user ID (identifier) ​​and uses a state store to maintain the most recent action sequence for each user (window size is configurable, default 20 entries). Each time a new action is received, it is appended to the end of the sequence, and the oldest action is removed, resulting in an updated user action sequence.

[0032] In some possible implementations, user behavior sequences are acquired in real time to ensure that the latest changes in user interests can be captured, providing raw data for intent inference and avoiding misjudgments based on outdated data.

[0033] In some possible implementations, the triggering condition refers to the rule that initiates the intent reasoning process, which may include event-driven conditions (such as a user performing multiple consecutive actions on the same topic) or behavior escalation conditions (such as escalating from browsing to adding to cart).

[0034] Intent reasoning refers to semantic analysis of user behavior sequences to infer the user's current intent (such as information gathering, comparison, upcoming decision-making, etc.), intent strength, and associated core entities and tags.

[0035] Dynamic graph topology updates refer to operations such as adding new nodes (e.g., intent nodes), establishing new edges, and adjusting the weights of existing edges in a heterogeneous graph based on the intent reasoning results.

[0036] The core of the method for recommending content provided in this disclosure is to perform intent reasoning on user behavior sequences and dynamically update the topology of the heterogeneous graph based on the intent reasoning results. First, it is determined whether the current user behavior sequence meets preset triggering conditions. If it does, intent reasoning is performed (e.g., using a large language model for multi-step reasoning through thought chains or using a pre-trained model) to generate intent analysis results. Then, graph operation instructions are generated based on the analysis results, and these instructions are executed to update the topology of the heterogeneous graph.

[0037] By dynamically updating the graph topology through intent reasoning, the user's real-time intent is explicitly encoded as graph nodes and edges, enabling the graph neural network to perceive and propagate short-term interest signals. This improves the response speed of recommendations to user interest shifts (on a minute-by-minute scale) while avoiding interference from historical behavioral noise.

[0038] S130. Use a heterogeneous graph neural network to update the node embedding of the updated heterogeneous graph, and obtain the updated embedding vectors corresponding to user nodes and push content nodes.

[0039] Among them, the Heterogeneous Graph Neural Network (HeteroGNN) is a graph neural network model capable of handling various types of nodes and edges. Its core feature is that it sets independent parameterized transformations for different types of nodes and edges, and aggregates neighbor messages through a type-aware attention mechanism.

[0040] Node embedding update refers to recalculating the low-dimensional vector representation (embedding) of affected nodes after a change in graph topology, so that the updated embedding vector can reflect the latest graph structure and node attributes.

[0041] In some possible implementations, embodiments of this disclosure employ HGT (Heterogeneous Graph Transformer) as the specific model for the heterogeneous graph neural network. When the graph topology is dynamically updated due to intent reasoning, the entire model is not retrained. Instead, an incremental update strategy is adopted, namely, identifying the affected set of nodes (e.g., newly created intent nodes), extracting the neighborhood subgraph of this set, performing forward propagation of HGT only on this subgraph, updating the embeddings of the affected nodes, while keeping the model parameters unchanged.

[0042] Heterogeneous graph neural networks can fully integrate the semantic information of heterogeneous nodes and edges to generate high-quality node embeddings. An incremental update strategy enables rapid convergence of embeddings after topology changes, with end-to-end latency controlled within seconds, meeting the requirements of real-time recommendations.

[0043] S140. Recall and sort the push content based on the updated embedding vector, and determine the final push content to be sent based on the sorting result.

[0044] Among them, content recall involves quickly filtering out a batch of potentially relevant push notifications from a massive amount of push content as a candidate set; ranking refers to refining the scores of the push content in the candidate set and selecting the best push content as the recommended content to be delivered to the user.

[0045] Among some possible implementations, a cascaded architecture combining multi-way recall and fine sorting can be adopted.

[0046] In the content recommendation method of this disclosure, content is pushed through pre-constructed heterogeneous graphs and real-time intent reasoning. Since heterogeneous graphs can integrate multi-source heterogeneous data (such as the relationships between different types of nodes), even if a user has very few direct interaction records with the pushed content, that user node can still indirectly obtain rich semantic information through other types of nodes and edges in the graph. Therefore, the content recommendation method provided in this disclosure can generate more robust initial representations for users and pushed content even in scenarios with extremely low click-through rates and sparse positive samples, significantly improving the cold-start recommendation effect.

[0047] Meanwhile, compared to existing static graph models that rely solely on offline historical interaction data, this embodiment can proactively adjust the graph structure based on the user's current behavior patterns (e.g., changing the connections between nodes or the weights of edges), enabling the heterogeneous graph to quickly reflect the user's latest interests. Subsequent node embedding updates are based on the updated graph topology, allowing push content retrieval and ranking to promptly capture the user's short-term intent, making the pushed content more aligned with the user's current real needs and reducing invalid pushes caused by interest drift.

[0048] Furthermore, because heterogeneous graph neural networks possess type-aware aggregation capabilities, they can efficiently handle multiple types of nodes and edges. Simultaneously, graph topology updates typically only affect local subgraphs, so node embedding updates can be performed only within the affected region (without requiring a full graph recalculation). This enables the embodiments of this disclosure to significantly reduce the computational latency of online inference while maintaining embedding quality, achieving second-level response times and meeting the stringent real-time requirements of mobile push scenarios.

[0049] Furthermore, since the updated node embedding vectors have fused multiple types of semantic information from the graph through a heterogeneous graph neural network, their representational ability surpasses traditional methods based on collaborative filtering or shallow features. In the recall phase, high-quality embedding vectors can more comprehensively cover push content that users might be interested in; in the ranking phase, combining the updated embedding vectors of user nodes and push content nodes allows for a more accurate prediction of the probability of a user clicking on each push. Ultimately, this significantly improves the click-through rate of push content and user retention.

[0050] The following is a detailed description of the recommended content provided in the embodiments of this disclosure.

[0051] In some possible implementations, the heterogeneous graph in this disclosure includes multiple node types and edge types. These nodes and edges together constitute a multimodal, cross-domain knowledge graph, enabling the content recommendation method provided in this disclosure to fuse multi-source information from different domains and support dynamic encoding of real-time intents.

[0052] In some specific implementations, heterogeneous graphs also include at least one of the following node types: In-application entity nodes are used to represent items or content within the application; Tag nodes are used to represent semantic tags; Time context nodes are used to represent time windows or time periods; Intent nodes are used to represent the user's real-time intent; and at least one of the following edge types: The interaction edge between the user node and the push content node is used to represent the user's click or exposure behavior on the push content; Behavioral edges between user nodes and entity nodes within the application are used to represent user actions within the application. The association edges between push content nodes and tag nodes are used to represent the attribution relationship between push content and semantic tags; The belonging edges between entity nodes and tag nodes within the application are used to represent the association between entities and semantic tags within the application; The active edge between the user node and the time context node is used to represent the user's activity level within a specific time window; The association edges between user nodes and intent nodes are used to represent the association between the user and the real-time intent; Semantic association edges between push content nodes and entity nodes within the application are used to represent cross-modal associations based on semantic similarity.

[0053] In this context, in-application entity nodes represent various items or content within the application, such as goods in e-commerce scenarios, articles in news and information scenarios, videos in short video scenarios, and shops or services in local life scenarios. Each entity node can be associated with its inherent attribute characteristics, such as category, price, sales volume, and descriptive text.

[0054] In some possible implementations, entity IDs and their attributes can be extracted from the application's business database. These attributes can then be converted into initial embedding vectors using feature engineering or pre-trained models, serving as the embedding vectors for the corresponding nodes. Entity nodes can be associated with each other based on business logic (e.g., "same category," "similar products"), or they can be linked to user behavior through subsequent edges.

[0055] In-app entity nodes serve as the core bridge connecting user behavior and push content. Numerous user actions within the app (browsing, searching, saving, purchasing) directly impact these entities, forming a high-density behavioral edge. These entity nodes are further associated with tag nodes, thereby conveying user interests to push content nodes and alleviating the problem of sparse push domain samples.

[0056] Tag nodes are used to represent semantic tags, such as "foldable screen phone", "running gear", "limited-time promotion", etc. Tags can have a hierarchical structure (such as first-level category, second-level category) and can be accompanied by statistical information such as popularity and source (manual annotation or model inference).

[0057] In some possible implementations, tags can be extracted from product descriptions and article titles using knowledge graphs, content classification systems, or natural language processing techniques. Tag nodes can be predefined or dynamically expanded. The embedding vector of a tag node can be obtained by encoding its text description using a language model, or initialized by aggregating features from its associated entity nodes.

[0058] Tag nodes allow heterogeneous entities (push notifications, in-app entities) to be mapped to a unified semantic space, achieving cross-domain alignment. Tag nodes also enable the establishment of a meta-path: "user → in-app entity → tag → push notification content," allowing user in-app behavior signals to propagate along tags to push notification content, significantly improving the efficiency of cross-domain knowledge transfer.

[0059] Time context nodes are used to represent time windows or time periods, such as discrete time periods like "weekday mornings," "weekend evenings," and "holidays." Each time context node can be associated with statistical features such as global user activity and push click-through rates within that time period.

[0060] In some possible implementations, a 24-hour day can be divided into several time periods according to business needs (such as every 2 hours), or nodes can be generated based on weekdays / weekends, holidays, etc. The node embedding vector is encoded from the historical statistical values ​​(such as average click-through rate) of that time period.

[0061] The time context node enables modeling of the relationship between user behavior and time. By using the active edges between user nodes and time context nodes, we can learn the differences in user preferences across different time periods; by using the scheduling edges between push content nodes and time context nodes, we can determine the appropriate time windows for delivering different push content, thereby reaching users at the optimal time and improving click-through rates.

[0062] Intent nodes are used to represent a user's current real-time intent, such as "buy a foldable phone" or "find the typhoon path." Intent nodes are dynamically created online, have a lifecycle, and their attribute vectors are typically obtained by weighted averaging of the features of associated tag nodes, and are updated or decayed with subsequent user actions.

[0063] In some possible implementations, an intent node can be created when intent reasoning determines that the user has a clear intent. The initial embedding vector is the weighted sum of the embedding vectors of the associated tag nodes (the weights being the inferred intent strength). After the intent node is created, edges are established with the user node, tag nodes, and related entity nodes, and initial edge weights are set.

[0064] Intent nodes are the core carrier for achieving real-time interest drift tracking. Traditional static graphs cannot express short-term, dynamic user intents, while intent nodes, as temporary high-dimensional representation units, centrally encode the user's current intent. They can attract attention between user nodes and push content nodes, allowing intent-related push content to achieve higher rankings after embedding updates, thus enabling minute-level interest capture. The lifecycle management of intent nodes (creation, activation, decay, and recycling) ensures that historical intents do not interfere with recommendations in the long term.

[0065] The interaction edge between user nodes and push content nodes represents the user's click or exposure behavior on the push content. This edge can be directed (user → push content) and carries a weight, which is usually calculated based on signals such as the number of clicks, whether it was exposed but not clicked, combined with time decay. It directly connects the user and the push content and is the most direct feedback signal. However, due to the sparse nature of push interactions, it is difficult to fully learn user interests by relying solely on this type of edge. Therefore, the heterogeneous graph in this embodiment of the disclosure also relies on other edge types to supplement information.

[0066] Behavioral edges between user nodes and entity nodes within the application represent user actions within the application, such as browsing, searching, adding to favorites, adding to cart, and purchasing. Different behavior types can be assigned different strength values ​​(e.g., browsing = 1, adding to favorites = 3, purchasing = 8), forming weighted edges. These edges carry dense behavioral data within the app. Since each user generates a large number of actions daily, these edges provide rich supervisory signals for the graph neural network. Through subsequent meta-path propagation, these signals can be indirectly passed to push content nodes, alleviating the sparsity of the push domain.

[0067] The association edges between push content nodes and tag nodes are used to represent the affiliation relationship between push content and semantic tags. For example, a push notification like "Limited-time flash sale on foldable screen phones" can be associated with the tags "foldable screen phones" and "promotion". The association strength can be determined by a content understanding model or manual annotation. Mapping push content to the tag space allows push content to establish indirect connections with entity nodes within the app through tag nodes, achieving cross-domain knowledge alignment.

[0068] The attribution edge between entity nodes and tag nodes within an application represents the association between entities and semantic tags. For example, a product might be associated with tags like "foldable phone" or "high-end phone." Attribution confidence can be provided by a knowledge graph or classifier. Similar to the push content-tag edge, this edge projects entities within the app onto a unified tag space. In this way, user interactions with entities within the app can be passed to push content with the same tags, forming a cross-domain recommendation path.

[0069] The active edges between user nodes and time context nodes represent the user's activity level within a specific time window. The weights can be set as normalized values ​​of the user's historical behavior frequency during that time period, or obtained through statistical calculation. Active edges model temporal patterns of user behavior; for example, some users pay more attention to news feeds on weekday mornings and entertainment feeds on weekend evenings. Active edges enable graph neural networks to learn time-aware user representations, thus taking time factors into account when matching push content.

[0070] The edges connecting user nodes and intent nodes represent the association between the user and the real-time intent. The weights are determined by the intent confidence given during the intent inference process and decay exponentially over time. These edges explicitly record the user's short-term intent in the graph structure. When an intent node is active, the edge has a higher weight, guiding the graph neural network's message passing to focus on the subgraph related to that intent, thereby improving the ranking of relevant pushed content. As the intent fades, the weights decay, and eventually the intent node is recycled, avoiding interference from historical intents.

[0071] Not all of the above-mentioned node and edge types are required to exist, but their collaborative operation constitutes a preferred implementation of the embodiments of this disclosure: Cross-domain data fusion: By using entity nodes and tag nodes within the application as bridges, high-density in-app behavior data (user-entity behavior edge) is transmitted along the path of "user→entity→tag→push content" to sparse push content nodes, thus solving the problem of sample sparsity.

[0072] Real-time intent injection: Intent nodes, as dynamic entities, inject the current intent strength into the graph structure through user-intent association edges, and combined with time context nodes, enable the recommendation system to perceive both short-term intent and long-term time preferences at the same time.

[0073] Multi-type edge weight coordination: All edge weights can decay over time (interaction edges, behavior edges) or be dynamically adjusted (intent-related edges), ensuring that the graph structure always reflects the latest user interests.

[0074] By combining various edge types, the high-density behavioral data within the application can be fully utilized, generating accurate user representations even with minimal interaction in the push domain, effectively alleviating the cold start problem.

[0075] Based on the heterogeneous graph structure described above, Figure 2 This diagram illustrates a flowchart of an implementation method that utilizes a large model to extract and enhance semantic features from push content nodes and entity nodes within the application. This process still falls under the category of constructing heterogeneous graphs, such as... Figure 2 As shown, steps S210 and S220 may be included.

[0076] S210. Utilize a large model to perform semantic analysis on the text information corresponding to push content nodes and / or entity nodes within the application, and generate structured semantic features.

[0077] Large models are pre-trained language models with parameter sizes typically exceeding one billion. Unlike traditional masked language models, LLMs possess powerful text understanding, generation, reasoning, and instruction following capabilities, and can output structured analysis results based on prompts (natural language hints).

[0078] Structured semantic features are semantic information with a fixed format and defined dimensions that LLM outputs from input text. Examples include JSON (JavaScript Object Notation, a lightweight data interchange format) objects containing fields such as core topic entities, user intent types, sentiment tendencies, and associated tags. This feature transforms the unstructured information of the raw text into a structured numerical vector that can be processed by the machine.

[0079] The core of step S210 is to input the text information (such as push title, product description) corresponding to nodes, i.e., push content nodes and / or entity nodes within the application, into the LLM, and guide the LLM to output structured analysis results. There are several ways to achieve this, including: Prompt Engineering: Construct natural language prompts containing specific instructions, requiring the LLM to output them in a specified format. For example, for push content, prompts may require the model to output core topic entities, intent type (promotion / information / interaction), sentiment, and related tags.

[0080] Fine-tuning: Use business-labeled data to efficiently fine-tune the parameters of the LLM, making it more suitable for semantic analysis tasks in specific domains (such as e-commerce and news), and improving output quality and stability.

[0081] Chain thinking: Guide the model to reason step by step in the prompt words. For example, first identify key entities, then judge the intent, and finally generate labels to avoid omissions in one step.

[0082] The output of an LLM is typically natural language text, which needs to be further converted into numerical vectors. Some possible implementations include using a pre-trained language model to encode the structured text output by the LLM, extracting the hidden states as semantic vectors. Alternatively, the semantic vectors can be extracted directly from the LLM's embedding layers (such as the last hidden layer in the GPT series).

[0083] S220. The structured semantic features are fused with the statistical features of the corresponding nodes to generate the initial embedding vectors of the push content nodes and / or the entity nodes in the application.

[0084] Statistical features refer to the inherent numerical attributes of nodes, which do not rely on semantic understanding. Examples include the historical average click-through rate, number of impressions, and content creation time of push content nodes; and sales volume, ratings, and price ranges of entity nodes within the application. These features are typically obtained through statistical calculations or offline data warehouses.

[0085] Semantic features and statistical features describe nodes from different perspectives, therefore, they need to be effectively fused. In some possible implementations, embodiments of this disclosure employ a gating fusion mechanism to fuse structured semantic features with the statistical features of the corresponding nodes.

[0086] Gated fusion is an adaptive weighted fusion technique. It learns a gate vector (with values ​​between 0 and 1) to dynamically determine the proportions of semantic and statistical features in the final initial embedding vector. This mechanism avoids manually setting fixed weights and can automatically adjust based on node type and data quality. Its specific implementation is as follows: in, `h_sem` is the sigmoid function, with an output range of (0,1). `h_sem` is the semantic feature vector, `h_stat` is the statistical feature vector, and `g` is the gating vector. `W_g` is the weight matrix, and `b_g` is the bias vector corresponding to the weight matrix.

[0087] Final initial embedding vector .

[0088] Gating mechanisms can automatically learn the relative importance of semantic and statistical features based on data. For example, for products with very detailed descriptions, h_sem has high credibility, and g tends to be 1; for popular posts with extremely high click-through rates but brief descriptions, h_stat contributes more, and g tends to be 0.

[0089] Besides gating fusion, other fusion methods can also be used, such as: Direct concatenation (Concat), followed by dimensionality reduction through a fully connected layer.

[0090] Weighted summation (fixed weights), with weights set through grid search or experience.

[0091] Based on Attention, the importance weights of each feature are calculated.

[0092] LLM's structured semantic features, combined with statistical features, enable the initial vectors of nodes to possess both deep semantic understanding capabilities and domain statistical experience. Compared to traditional random initialization or Word2Vec (word-to-vector) embedding, it has stronger discriminative and generalization capabilities.

[0093] Figure 3 This diagram illustrates a flowchart of an implementation method for establishing semantic association edges between push content nodes and in-application entity nodes after obtaining the embedding vectors corresponding to the push content nodes and the entity nodes within the application. Figure 3 As shown, steps S310 and S320 may be included.

[0094] S310. Calculate the semantic similarity between push content nodes and entity nodes within the application.

[0095] Semantic similarity refers to the degree of proximity between push content nodes and entity nodes within the application in the semantic space. It uses cosine similarity as a metric, which is the semantic vector of the two nodes, i.e., the semantic feature vector calculated in step S210. Or the initial embedding vector calculated in step S220 The dot product of and is divided by the product of their respective moduli, with a value range of [-1, 1]. The larger the value, the more semantically similar they are.

[0096] In some possible implementations, the semantic feature vector of each node is obtained. Alternatively, the initial embedding vector h_initial can be calculated, and the cosine similarity between the embedding vector pairs of all or some push content nodes and entity nodes within the application can be used as the semantic similarity between the push content nodes and entity nodes within the application.

[0097] S320. In response to the semantic similarity exceeding a preset threshold, establish semantic association edges between push content nodes and application entity nodes in the heterogeneous graph.

[0098] Semantic association edges are a dynamically created edge type used to connect semantically similar push content nodes and in-application entity nodes. Unlike predefined business edges (such as attribution edges and behavior edges), these edges are calculated and added in real-time based on content semantics, primarily for cold-start content distribution and cross-domain knowledge migration.

[0099] In some possible implementations, a similarity threshold can be set. If the semantic similarity between the push content node P_i and the application entity node A_j is greater than the similarity threshold, then an undirected edge from P_i to A_j is added to the heterogeneous graph, and the edge weight is assigned the semantic similarity value.

[0100] The steps S210, S220, S310, and S320 described above can be executed offline periodically (e.g., daily updates) or calculated online in real time (triggered when new content is published). For newly launched push content, it is only necessary to calculate its similarity to all entity nodes within the application, find the most similar entities and establish edges, and it can immediately obtain distribution opportunities (because the entity nodes already have a large number of user behavior signals), thereby achieving a cold start.

[0101] Semantic association edges, built based on semantic similarity, allow new push content to be associated with existing high-popularity entity nodes without any user interaction data. This inherits user preference signals from the entities, enabling rapid distribution within minutes. These semantic association edges directly connect push content and in-app entities, skipping tag nodes, reducing the path length of information transmission, and improving recommendation efficiency. Simultaneously, LLM's multilingual and multi-domain understanding capabilities allow the system to easily handle text of different styles (such as colloquial titles and professional product descriptions).

[0102] The following is a specific implementation example of using a large language model to extract and enhance semantic features of push content nodes and entity nodes within the application, and to establish cross-modal association edges based on semantic similarity. It may include the following steps: Step 1: Construct LLM prompts and extract structured semantic features For the push content node: Assume the push title to be processed is "Limited-time flash sale on foldable screen phones, up to 2000 yuan off," and the summary is "Limited-time offer on branded foldable screen phones." Construct the following prompt keywords: Please analyze the following push notification and extract the key information: Push notification title: {title} Push summary: {summary} Please analyze from the following dimensions and output structured results: 1. Core thematic entities (such as model, category, event) 2. User Intent Types (Promotional Attraction / Information Inquiry / Interactive Invitation / Urgent Reminder) 3. Emotional inclination and urgency (0-1 scalar) 4. Related semantic tags (maximum 5) Calling LLM yields the following output: json { "Core Entity": ["Foldable Phone"], "Intent Type": "Promotional Attraction", "Urgency": 0.9, "Associated Tags": ["Foldable Phone", "Limited-Time Offer", "High-End Phone", "Digital Product"]} The original text is concatenated with the LLM output and input into a pre-trained encoder for encoding. The [CLS] vector is then used as the semantic feature vector h_sem. ;in, It is an encoder, specifically BERT (Semantic Representation Model). For the original text, For LLM output.

[0103] For content nodes within the application: Suppose a product's title is "Foldable Screen Phone," its description is "Foldable Screen Phone Limited-Time Offer," and its category is "Mobile Phone." Construct the following prompt keywords: Please analyze the following product / content information and extract key semantic features: Title: {item_title} Description: {item_description} Category: {category_path} Please output: 1. Key selling points / highlights (maximum 3) 2. Target User Profile Keywords 3. Usage Scenarios Description 4. Related semantic tags (maximum 5) The LLM function outputs the following: json { "Key Selling Points": ["Limited-Time Offer"], "Target Users": ["Interested in the Promotion"], "Usage Scenarios": ["Electronic Product Usage Scenarios"],, "Related Tags": ["Foldable Screen Phones", "Limited-Time Offer", "High-End Phones", "Digital Products"]} The original text is concatenated with the LLM output and input into the pre-trained BERT (Semantic Representation Model) encoding. The [CLS] vector is taken as the semantic feature vector h_sem. Its calculation method can be the same as that of the semantic feature vector of the push content node, which will not be elaborated here.

[0104] Step 2: Obtain and fuse statistical features For each content push node, historical statistical features are extracted from the log system: click-through rate of 0.03 over the past 7 days, impressions of 5000, and content creation time of 2 hours ago. After normalizing these values, they are mapped to a 256-dimensional statistical feature vector h_stat through a single MLP (Multilayer Perceptron).

[0105] The semantic feature vector h_sem has a dimension of 768. It is reduced to 256 dimensions through a node-type-related linear projection layer, resulting in h_sem_proj. Specifically: in, The projection matrix is ​​related to the node type. Different types of nodes use different projection parameters to preserve type specificity. It is the bias vector corresponding to the projection matrix.

[0106] For content nodes within the application, statistical values ​​such as sales volume and page views can also be obtained. After normalizing these values, they are mapped to a 256-dimensional statistical feature vector h_stat through an MLP (Multi-Layer Perceptron). For its semantic feature vector, the node type-related linear projection layer described above is used to reduce the dimension to 256, resulting in h_sem_proj.

[0107] For each push content node or content node within the application, the obtained h_sem_proj and h_stat are gated and fused. The specific implementation is as follows: in, `h_sem` is the sigmoid function, with an output range of (0,1). `h_sem` is the semantic feature vector, `h_stat` is the statistical feature vector, and `g` is the gating vector. `W_g` is the weight matrix, and `b_g` is the bias vector corresponding to the weight matrix.

[0108] Final initial embedding vector .

[0109] The initial embedding vector is stored in the node attributes and used as input for subsequent graph neural networks.

[0110] Step 3: Establish semantic association edges based on semantic similarity Calculate the cosine similarity between the semantic feature vector h_sem (before dimensionality reduction) of the new push content node and the semantic vectors of all in-application entity nodes (such as product nodes). The semantic vector library is pre-stored in the vector index.

[0111] Set the threshold to 0.75. Assuming the cosine similarity between the push content ID_P001 and the entity node ID_A1001 corresponding to a certain product exceeds the threshold (0.82), then perform the following operations in the heterogeneous graph database: Add an edge (Push Node ID_P001, Entity Node ID_A1001), with the edge type being "Semantic Association Edge" and the edge weight being 0.82.

[0112] These semantically related edges are stored in a graph database and serve as the basis for subsequent meta-path retrieval.

[0113] For newly launched push content, when establishing semantic association edges, association edges with tag nodes can also be created simultaneously. For example, based on the associated tags output by LLM, edges can be directly established between push content nodes and tag nodes.

[0114] In some possible implementations, heterogeneous graphs may also include multiple meta-paths.

[0115] Metapaths are specific node sequence patterns defined on heterogeneous graphs, used to capture higher-order semantic relationships between nodes. By traversing the graph along metapaths, push content nodes with specific semantic relationships to user nodes can be found, thus achieving accurate recall without relying on collaborative filtering.

[0116] First-order path: Starting from the user node, it reaches the application entity node via the behavior edge between the user node and the application entity node, then reaches the tag node via the ownership edge between the application entity node and the tag node, and finally reaches the push content node via the association edge between the push content node and the tag node. Secondary path: Starting from the user node, it reaches the entity node in the application through the behavioral edge between the user node and the entity node in the application, and then reaches the push content node through the semantic association edge between the push content node and the entity node in the application. The third path starts from the user node, reaches the time context node via the active edge between the user node and the time context node, and then reaches the push content node via the scheduling edge between the push content node and the time context node.

[0117] Specifically, the first-level path utilizes in-app entity nodes and tag nodes as a bridge to transmit high-frequency user behavior signals within the app to push content nodes. The specific steps are as follows: Starting from the current user node, find all in-app entities the user has interacted with historically (such as browsed products, watched videos, and saved articles) along the "behavior edges between the user node and in-app entity nodes". Behavior edges carry weights, reflecting the intensity of the behavior (such as the intensity values ​​corresponding to different actions like clicking, pausing, saving, and purchasing).

[0118] Starting from each entity node within the application, find the semantic tags associated with that entity along the "attachment edges between the entity node and the tag node" (e.g., mobile phone products are associated with tags "foldable screen phone" and "high-end phone"). The weight of the attachment edge represents the degree of relevance between the tag and the entity.

[0119] Starting from each tag node, find all push content nodes labeled with the same semantic tag along the "association edges between push content nodes and tag nodes". The weight of the association edge can represent the matching strength between the push content and the tag (such as the similarity output based on the content understanding model).

[0120] For example, suppose a user browses a mobile phone product within the app; both products fall under the tag "foldable phone." Searching along the path "User → Product → Tag → Push Content" yields all push content carrying the tag "foldable phone" (e.g., "New foldable phone launch discounts," "Foldable phone reviews and comparisons"). Even if the user has never clicked on this content, it is still associated with the user because it shares the same tag.

[0121] The first-order path uses tags as semantic alignment anchors to migrate rich user behavior data within the app to the push domain, effectively solving the cold start problem caused by sparse push samples. Even when users have very little direct interaction with push content, this path can still recall semantically relevant content, significantly increasing the exposure opportunities for long-tail push content.

[0122] Unlike the first-order path, the second-order path does not pass through tag nodes. Instead, it directly utilizes the semantic association edges between push content nodes and entity nodes within the application. These semantic association edges are dynamically established based on the similarity of the initial embedding vectors of the nodes. When the cosine similarity of the semantic vectors of two nodes exceeds a preset threshold, the edge is automatically added between them.

[0123] For example, a large language model is used to encode the push notification title and product description into semantic vectors, and cosine similarity is calculated. If the similarity exceeds a threshold, a semantic association edge is established between the push notification content node and the entity node within the application. When a user browses the entity node, the associated push notification content is directly retrieved along the "user → entity → push" path.

[0124] The second-order path omits tag nodes, reducing path hops and improving recall efficiency. Meanwhile, semantic association edges are directly built based on content semantics, without relying on a manually labeled tag system, offering strong versatility and scalability. For newly launched push content (cold start content), as long as its semantics are similar to existing entities within the app, associations can be quickly established, achieving minute-level distribution.

[0125] Third-dimensional paths are used to model the matching relationship between user activity patterns across different time periods and the optimal delivery time for push content. Starting from the user node, it finds the user's activity level within a specific time period (e.g., "weekday morning" or "weekend evening") along the "activity edge between the user node and the time context node." The weight of the activity edge reflects the user's historical activity frequency during that time period (e.g., number of app openings, duration of app usage), and is normalized to represent the connection strength.

[0126] Starting from the time context node, the appropriate push content to be delivered during the specified time period is found along the scheduling edge between the push content node and the time context node. The weight of the scheduling edge is set based on the historical click-through rate statistics of the push content during the specified time period; the higher the click-through rate, the greater the edge weight.

[0127] For example, if a user frequently opens a news app between 8:00 and 10:00 AM on a weekday, the active edge weight is 0.9. The corresponding time context node for this period is identified as "weekday morning." During this time, there are multiple push notifications, with "Morning News Express" having a scheduling edge weight of 0.95 and "Late Night Emotional Stories" having a scheduling edge weight of 0.2. The push notifications with higher scheduling edge weights are retrieved along the path "User → Time Period → Push Notifications," with "Morning News Express" being pushed first.

[0128] The third-dimensional path explicitly incorporates the time dimension into the graph structure, enabling personalized optimization of push timing. Even if users' preferences for push content change over time, this path can dynamically adjust the recall results, avoiding disturbing users during inactive periods while improving the click-through rate of pushes.

[0129] In some practical implementations, the three meta-paths mentioned above can be used in parallel during the final recall stage, and the recall results of each path can be combined. Specifically, for each meta-path, a maximum traversal depth (e.g., 3 hops) is set, the score of each path instance is calculated (usually the product of the weights of each edge on the path), and the top-K push content nodes with the highest scores are selected as the recall results for that path. Finally, the union of the recall results from multiple paths is taken and input into the subsequent ranking model.

[0130] Three meta-paths uncover the connection between users and push content from different perspectives: the first path utilizes tag semantics for cross-domain migration, suitable for cold starts and long-tail content; the second path leverages direct semantic similarity, which is computationally lightweight and has broad coverage; and the third path utilizes temporal context to ensure precise push timing. These three approaches complement each other, significantly improving coverage, accuracy, and timeliness in the recall phase.

[0131] The meta-paths in this disclosure are based on a rich set of node and edge types defined in heterogeneous graphs, particularly introducing in-application entity nodes, tag nodes, temporal context nodes, as well as semantic association edges and scheduling edges. This allows the paths to simultaneously capture cross-domain semantics, user interest shifts, and time-sensitive information. Furthermore, these meta-paths collaborate with features such as intent node lifecycle management and incremental training to form a complete technical chain from real-time behavior to precise push notifications.

[0132] In some possible implementations, after the heterogeneous graph is constructed, a full training of the heterogeneous graph neural network is performed on all nodes in the heterogeneous graph to obtain the trained heterogeneous graph.

[0133] Full training involves using complete data from all nodes and edges in the heterogeneous graph to perform multiple rounds of iterative training on the heterogeneous graph neural network model (such as backpropagation based on gradient descent), updating all learnable parameters of the model (such as weight matrices and bias terms), and finally generating the final embedding vector for each node. The computational complexity of full training is typically proportional to the size of the graph, making it suitable for batch execution in offline environments.

[0134] Full training first requires constructing a complete snapshot of the heterogeneous graph at the current time point, which specifically includes: exporting all nodes and their attribute features from the graph database, exporting all edges and their current weights, and serializing the node features and edge weights into the tensor format required by the training framework.

[0135] If the graph is huge (e.g., with billions of nodes), subgraphs can be sampled or training can be performed using distributed graph storage.

[0136] Taking the Heterogeneous Graph Neural Network (HGT) model as an example, the full training process includes: Forward propagation: Initial node features, edge indices, edge types, and edge weights are input into the HGT model. The model aggregates neighbor information through multi-layer type-aware attention and outputs the embedding vector for each node.

[0137] Loss Calculation: The loss is calculated based on the multi-task training objective. Multi-tasks may include cross-domain contrastive learning loss (bringing the embedding distance between the user and the positive sample push, and pushing the negative sample away), push click-through rate prediction loss (binary classification cross-entropy), regularization loss, etc.

[0138] Backpropagation: Calculate the gradient of the loss with respect to the parameters of each layer of the model, and update the parameters using an optimizer (such as Adam).

[0139] Iterative training: Repeat the forward-backward process several times until the loss converges or the preset number of rounds is reached.

[0140] After training, the parameters of the graph neural network model (weight matrix, bias, attention parameters, etc.) are saved to persistent storage (such as a distributed file system or model repository), and incremental updates and inference are performed using these parameters. Simultaneously, the final embedding vectors of all nodes are written to a vector indexing service (such as Faiss or Milvus) for use in the recall phase.

[0141] In some possible implementations, the heterogeneous graph neural network is fully trained on all nodes in the heterogeneous graph according to a preset time period in order to update the embedding vectors of all nodes in the heterogeneous graph.

[0142] After full training, user behavior patterns, content distribution, and tagging systems may change significantly over time. If the model is not updated, it will be unable to adapt to changes in the global data distribution. Therefore, periodic full training is necessary. Retrain the model parameters using accumulated, up-to-date historical data to enable the model to capture long-term statistical trends. Refresh the embeddings for all nodes (especially those that haven't participated in incremental updates for a long time) to prevent outdated embeddings from misleading recommendations.

[0143] In some specific implementations, a full training session is triggered daily (or weekly) to generate globally optimal model parameters and node embeddings.

[0144] After full training, new parameters are loaded, and the local embeddings accumulated during incremental updates are cleared or reset, using the results of full training as a baseline. During the full training interval, whenever the graph topology is dynamically updated (e.g., creating an intent node), forward propagation is performed only on the local subgraph to quickly generate embeddings for the affected nodes, keeping the model parameters unchanged. This dual-track mechanism ensures global model accuracy while achieving real-time response, avoiding the high computational overhead of frequent full training.

[0145] Full training leverages the complete graph structure and accumulated historical data to fully optimize model parameters, ensuring that node embeddings reflect global statistical patterns and long-term user behavior patterns. As time progresses, new users, content, and tags are continuously added, and user interest distributions may shift. Full training at each preset time interval allows the graph neural network model to absorb the latest data distribution changes, preventing the model from becoming outdated. For example, when a new type of product becomes popular, full training can quickly adjust the embeddings of relevant nodes, thus reflecting this in recall and ranking.

[0146] The preset time period (e.g., daily) also allows training tasks to be scheduled during off-peak hours (e.g., early morning), making full use of idle computing resources and avoiding disruption to online services. Meanwhile, the frequency of full-scale training can be flexibly adjusted according to the actual rate of data change (e.g., during promotional seasons, it can be shortened to every 12 hours), providing excellent flexibility.

[0147] In some possible implementations, the weights of at least one type of edge in the heterogeneous graph are dynamically adjusted according to an exponential time decay function, which takes the initial weight of the edge, the interval between the current time and the edge creation time, and the decay rate parameter related to the edge type as independent variables; when the weight of an edge is lower than a preset threshold, the edge is removed from the heterogeneous graph.

[0148] This mechanism enables heterogeneous graphs to reflect the timeliness of user interests and avoids outdated relationships from causing noise interference in recommendation results.

[0149] The exponential time decay function is a mathematical function that causes the weights to decrease exponentially over time. Its general form is: ,in, The initial weight is the initial weight value assigned to an edge when it is created. Different types of edges can be set with different initial weights. For example, the initial weight of an edge that triggers a user click on a push notification is 1.0, while the initial weight of an edge that is exposed but not clicked is negative (-0.5). The initial weight reflects the importance of the behavior type itself. , which is the time interval between the current time and the edge creation time. The longer the interval, the further back in time the action occurred, and the lower its reference value for the current interest. Let be the decay rate parameter, where This indicates the type of edge, with different edge types exhibiting different decay rates: for example, push notification click edges decay more slowly because click behavior reflects longer-term preferences; while in-app browsing edges decay more quickly because browsing behavior is more time-sensitive. The decay rate parameter can be set through statistical experience or learned from historical data. Exponential decay is characterized by a rapid initial decline followed by a gradual decrease, naturally simulating the forgetting curve of user interests.

[0150] The initial weight of an edge is the initial weight value assigned to the edge when it is created. Different types of edges can be set with different initial weights. For example, the initial weight of an interaction edge that triggers a user click on a push notification is 1.0, while the initial weight of an edge that is exposed but not clicked is negative (-0.5). The initial weight reflects the importance of the behavior type itself.

[0151] The preset threshold is a pre-defined lower bound for the weights. When the weight of an edge falls below this threshold after decay, it is considered that the user interest corresponding to that edge has essentially disappeared, and continuing to retain it would waste storage and interfere with the model. Therefore, the edge is marked as inactive and removed from the heterogeneous graph. Specifically, the edge record is deleted from the edge set of the graph database; if the graph structure is stored in memory, the adjacency list and edge weight matrix are updated.

[0152] In some possible implementations, logging removal is used for offline analysis and model diagnostics.

[0153] Edge removal can effectively control the graph size, avoiding storage bloat and computational redundancy caused by the long-term accumulation of a large number of low-weight edges. At the same time, the removal operation does not affect the model accuracy because the weights are already extremely low, and their contribution to message passing is negligible.

[0154] In some possible implementations, the weights of edges in the heterogeneous graph are maintained in real time, automatically decreasing over time, and invalid edges are cleaned up as needed. Edge weight decay can be triggered in two ways: Passive computation: Each time the edge weight needs to be accessed (e.g., during graph neural network message passing or path retrieval calculation), the decayed weight is calculated in real-time based on the current time and the edge's creation time. The advantage is that it eliminates the need for additional scheduled tasks and provides accurate calculations; the disadvantage is that it requires recalculation for each access, increasing overhead.

[0155] Active update: A background scheduled task (e.g., scanning every 10 minutes) updates the weights of all edges in batches. Expired edge weights are multiplied by the decay factor for the current time slice. The advantage is high batch processing efficiency, but the disadvantage is a certain delay.

[0156] This disclosure can employ a hybrid approach: for frequently accessed edges (such as user-push interaction edges), passive computation is used; for low-frequency edges, active batch updates are employed. Regardless of the approach, the decay function form remains consistent.

[0157] The exponential decay function naturally mimics the forgetting curve of human memory, making recent actions contribute significantly more to current recommendations than historical actions. Different edge types are configured with different decay rate parameters, enabling this embodiment to finely control the duration of the impact of various actions based on business semantics.

[0158] By automatically removing edges with excessively low weights using a preset threshold, the size of heterogeneous graphs is effectively controlled, preventing unlimited storage expansion. This also reduces the computational load of message passing in graph neural networks (eliminating the need to traverse edges with near-zero weights). For scenarios with high real-time requirements, this keeps graph operation latency within a manageable range.

[0159] When a heterogeneous graph receives a weight enhancement instruction (such as when the user intent is strong), the edge weight is immediately increased and the decay starting point is reset, so that active intent can continue to influence recommendations. Once the intent fades, the weight decays rapidly and the intent node and its associated edges are eventually deleted, thus realizing the natural lifecycle management of interest.

[0160] After the heterogeneous graph is constructed and trained, in response to the user behavior sequence meeting the preset triggering conditions, the user behavior sequence is subjected to intent reasoning, and the topology of the heterogeneous graph is dynamically updated based on the intent reasoning results.

[0161] In specific implementation methods, the triggering conditions may include: Within a preset time window (e.g., the last 5 minutes), the number of times a user performs the same topic-related behavior reaches a preset threshold (e.g., a threshold of 3). For example, if a user continuously searches, browses, and compares the same type of product (e.g., "foldable screen phone") within 5 minutes, the trigger condition is met, and the intent reasoning process is initiated.

[0162] And / or, When a user's behavior undergoes a significant escalation, intent inference is immediately triggered. Escalation events include, but are not limited to: escalation from "browsing" to "adding to cart," from "adding to cart" to "purchasing," and from "searching" to "browsing." These escalation behaviors typically indicate that the user's intent has shifted from information gathering to decision execution, requiring a real-time response.

[0163] In some possible implementations, throttling control conditions can be set to avoid system overload due to high-frequency behavior (especially when using large language models for inference): If the time interval between two consecutive inference attempts by the same user is not less than a preset throttling threshold, even if the above triggering conditions are met, new inference will not be started temporarily if the time since the last inference does not exceed the throttling threshold, in order to control the consumption of computing resources.

[0164] In some possible implementations, a lightweight pre-judgment can be performed using a rule engine to filter out cases that obviously do not require reasoning, thereby reducing unnecessary computational overhead. Lightweight pre-judgment includes, but is not limited to: Does the user behavior sequence include search behavior? Do the user behaviors belong to the same category? Is the number of user actions too low (e.g., less than 2)?

[0165] Only after the user's behavior sequence passes the pre-judgment is the trigger condition officially determined to be met, and the process enters the intent reasoning stage.

[0166] Figure 4The diagram illustrates a flowchart of one implementation method for inferring intent from user behavior sequences, such as... Figure 2 As shown, steps S410, S420, and S430 may be included.

[0167] S410. Obtain the sequence features of the user behavior sequence, including the behavior type, behavior object, and behavior timing information of each user behavior in the user behavior sequence.

[0168] Here, sequence features are features extracted from the original user behavior sequence to characterize behavioral patterns. In this embodiment, sequence features may specifically include: Behavior type refers to the specific category of action performed by the user. For example, "browse" means viewing a product details page, "search" means entering keywords to search, and "add to cart" means adding the product to the shopping cart. Different behavior types reflect different levels of user engagement and intent.

[0169] A behavior object is the target entity that a user's behavior affects. It can be an item within the application (product, article, video) or an abstract query term, tag, etc. Behavior objects are the direct basis for identifying content of a user's interest.

[0170] Behavioral temporal information refers to the time or sequential position in which a behavior occurs. This includes absolute timestamps or relative order (e.g., the third behavior in a sequence). Temporal information is used to calculate behavioral intervals, identify the degree of behavioral concentration, and analyze trends in behavioral types.

[0171] In some possible implementations, the most recent behavior records of the current user are read from a message queue or real-time storage to obtain a user behavior sequence. For each user behavior in the sequence, the behavior type, behavior object, and behavior timing information are parsed out and stored as a structured tuple.

[0172] In some specific implementations, a sliding window technique can be used to maintain the N most recent actions for each user, with the window length configurable (e.g., N=20 or time window=5 minutes). For computational efficiency, the action type, action object ID, and timestamp can be encoded as integers or embedded vectors, respectively.

[0173] S420. Based on sequence features, determine the intent evolution stage, intent activation intensity, and core entities and tags associated with the user intent corresponding to the user behavior sequence.

[0174] The intent evolution stage refers to the psychological and decision-making process a user goes through before achieving their final goal (such as purchasing a product or finishing reading an article). This stage can include: information gathering (browsing without comparison), active comparison (viewing multiple candidates simultaneously), impending decision (showing strong purchase or consumption signals), and casual browsing (without a clear goal). Different recommendation strategies correspond to different stages.

[0175] Intent activation strength is a numerical value that quantifies the intensity of a user's current intent, typically ranging from 0 to 1. A higher strength indicates that the user is more likely to complete the target action (such as placing an order or making a payment) in the near future. This strength can be calculated by combining signals such as behavior density (the number of actions per unit of time), behavior type escalation trends (e.g., from browsing to adding to cart), and dwell time.

[0176] Core entities are behavioral objects that appear repeatedly or have a high degree of relevance in a user's behavioral sequence. For example, after repeatedly searching for "foldable screen phone," the core entity is "foldable screen phone." Core entities are used to determine the core theme of the recommendation.

[0177] Tags are predefined semantic categories, such as "high-end mobile phones," "limited-time promotions," and "tech news." Tags can supplement core entities to broaden recall or perform cross-domain alignment.

[0178] This step is a high-level abstraction of the original sequence features, which can be achieved through various methods: Rule-based approach: Predefine a series of heuristic rules. For example: if the sequence contains searching and multiple browsing of the same category of products, it is determined to be in the "active comparison" stage; if the behavior type upgrades from browsing to adding to cart or purchasing, it is determined to be in the "imminent decision" stage; calculate the behavior density (number of behaviors / time length) within the time window, the higher the density, the stronger the intent activation; count the frequency of occurrence of behavior objects, and identify objects that appear more than a threshold (e.g., 3 times) as core entities; use a predefined knowledge graph or classifier to map core entities to tags.

[0179] Machine learning-based methods involve training a classifier (such as a gradient boosting tree or a small neural network) that takes sequence features as input and outputs intent stage, intensity value, and entity label. Training data can be weakly supervised by using manually labeled data or final conversion behaviors from historical data (such as placing an order) as positive samples.

[0180] The method based on Large Language Model (LLM) formats the behavior sequence into a natural language description and constructs prompt words to require the LLM to output a structured result. This implementation can leverage the semantic understanding and reasoning capabilities of LLM.

[0181] Regardless of the method used, the final output should be a structured result, containing at least one intent stage category, an intensity score, and one or more core entities and associated labels. This embodiment does not limit the specific implementation technology, as long as it can achieve the function of "determining based on sequence features".

[0182] S430. Based on the intent evolution stage, intent activation intensity, core entities, and labels, generate graph operation instructions. The graph operation instructions include at least one of the following: node type to be created, edge to be established, and edge weight to be set.

[0183] Graph manipulation commands are a set of commands used to modify the topology of heterogeneous graphs, and may include: The type of node to be created: for example, whether an intent node needs to be created, and the type of the node (such as "purchase intent" or "compare intent").

[0184] Edges to be established: Specifies which nodes need to have edges established, such as between a user node and a newly created intent node, or between an intent node and a label node.

[0185] Edge weights to be set: Assign specific weight values ​​to newly added edges or existing edges whose weights need to be adjusted. These weights are usually positively correlated with the intended activation strength.

[0186] This step maps the output of step S420 to specific graph operation commands. The mapping rules can be flexibly configured: Create Node: If the intent evolution stage is "Active Comparison" or "Imminent Decision," a "Create Intent Node" instruction is generated. The node type can be determined based on the intent stage (e.g., "Comparison Intent Node," "Decision Intent Node"). The initial embedding vector of the node can be generated by a weighted average of the embeddings of associated labels, or by using the embeddings of the core entity as initialization.

[0187] Establish Edge: Generate an "Edge Establish" instruction to connect user nodes with newly created intent nodes, as well as intent nodes with nodes or label nodes corresponding to the core entities extracted in step S420. The direction of the edge can be set according to business semantics (e.g., directed or undirected).

[0188] Set edge weights: Assign initial weights to the newly created edges based on the intent activation strength; generally, the higher the strength, the larger the initial weight. Simultaneously, if existing edges exist (e.g., the edge between a user node and a core entity node), a weight enhancement instruction can be generated to increase their weight values ​​(the increment is proportional to the intent activation strength). If the intent strength is low or the intent stage is "casual browsing," no new node is created; only a weight decay instruction is generated to reduce the weight of existing related edges.

[0189] In some possible implementations, the generated graph operation instructions are passed to the graph execution engine in a structured format to perform the "dynamically update the topology of the heterogeneous graph" step. The instructions may contain "at least one" operation (creating nodes only, building edges only, adjusting weights only, or a combination thereof) to retain flexibility.

[0190] By leveraging information such as behavioral type sequences and temporal concentration, it's possible to determine the user's decision-making stage (information gathering / comparison / imminent decision), thereby making more accurate recommendation decisions (e.g., prioritizing evaluation content during the comparison stage and pushing promotional information during the imminent decision stage). By transforming abstract information such as intent evolution stages and intent activation strength into graph operation instructions (creating intent nodes, setting edge weights), downstream heterogeneous graph neural networks can perceive these high-level intents. Compared to using only static features or historical statistics, this allows the graph structure to dynamically reflect the current intent, achieving explicit modeling of intent.

[0191] Figure 5 This diagram illustrates a flowchart of one approach to determining the intent evolution stage, intent activation intensity, and core entities and tags associated with a user intent sequence based on sequence features. Figure 5 As shown, steps S510, S520, and S530 may be included.

[0192] S510. Based on the behavior type and timing information of user behavior in the user behavior sequence, determine the degree of concentration of the timing of the same behavior type in the user behavior sequence, the trend of behavior type change, and the dwell time for the same behavior object.

[0193] The degree of temporal concentration of similar behavior types refers to the degree to which the same type of behavior (such as multiple "searches" or multiple "browsings") clusters along the timeline. Specifically, it can be measured by the reciprocal of the time interval between behavior occurrences, the number of occurrences per unit time (behavior density), or the standard deviation of the time difference between adjacent behaviors. A higher degree of concentration indicates that users repeatedly perform similar operations within a short period, usually signifying strong, focused interest.

[0194] The trend of behavior type change describes the direction of the evolution of a sequence of user behavior types over time. It can be an upgrade from low-engagement behavior (such as browsing) to high-engagement behavior (such as adding to favorites, shopping cart, and purchasing) (positive trend), a downgrade from high-engagement to low-engagement (negative trend), or remain unchanged. The trend can be determined by comparing the differences between preset weights of adjacent behavior types (such as browsing=1, adding to favorites=3, purchasing=8) or by sequence pattern recognition.

[0195] The dwell time for the same behavioral object refers to the total time a user continuously focuses on the same behavioral object (such as a product details page or an article). It can be obtained by summing the differences between adjacent timestamps of the same object appearing in the behavioral sequence (e.g., the difference between the timestamp of the first time the object was viewed and the timestamp of the last time the object was accessed). The longer the dwell time, the deeper the user's interest in the object and the higher the intensity of intent.

[0196] In some specific implementations, the sequence features corresponding to the user behavior sequence are obtained, namely, the list of behavior types, the list of behavior objects, and the list of timestamps.

[0197] For a specific behavior type (such as "browsing"), extract the timestamp of that behavior type and calculate the difference between adjacent timestamps. Behavior temporal concentration = 1 / (average difference) or use "behavior density" = number of behaviors of that type / total time window length to calculate behavior temporal concentration. Alternatively, an overall behavior density can be calculated uniformly for all behavior types as an alternative to temporal concentration.

[0198] Assign numerical weights to each behavior type (e.g., exposure = 0.5, views = 1, searches = 2, favorites = 3, add to cart = 4, purchase = 5). Calculate the difference between the weights of adjacent behaviors in the sequence. If most differences are positive, the behavior type trend is positive (upgrade); if most are negative, the behavior type trend is negative (downgrade); otherwise, the behavior type trend is stable. Alternatively, the behavior type trend can be obtained by calculating the weighted trend score = (last behavior weight - first behavior weight) / sequence length.

[0199] For each behavior object, find the timestamps of its first and last appearance, and calculate the difference as the dwell time of that behavior object. For objects that appear only once in the sequence, the dwell time can be set to the default minimum value (e.g., 0 or 1 second), or estimated based on the typical dwell time for that behavior type. For scenarios requiring precise calculation, the dwell time field reported by the client (e.g., dwell time on a product details page) can be used as the dwell time of that behavior object.

[0200] S520. Calculate the intensity of intent activation based on the degree of concentration of behavior sequence, the trend of behavior type change and / or dwell time.

[0201] The intent activation intensity is a score (e.g., between 0 and 1) that quantifies the urgency of a user to complete a certain goal (such as placing an order or subscription). This embodiment calculates this intensity by integrating the concentration of behavioral time sequence, the trend of behavioral type change, and the duration of dwell time. High concentration, positive change trends, and long dwell times all increase the intent activation intensity.

[0202] In some possible implementations, the degree of concentration of behavioral time sequence, the trend of behavioral type change, and / or dwell time can be fused to obtain the intent activation intensity. The fusion method can be: Weighted summation, i.e., intention activation strength = ×Concentration of Behavioral Timing+ ×Trends in Behavioral Types+ × Duration of stay, of which, These are the weighting coefficients, and all terms have been normalized to the [0,1] interval.

[0203] The product form is: Intent activation intensity = Concentration of behavior time sequence × (1 + Trend of behavior type change) × (1 + Duration of stay).

[0204] Empirical rules state that if the trend is positive and the duration of stay is greater than the threshold, the intensity is directly set to a high value; otherwise, it is linearly mapped according to the degree of concentration.

[0205] This disclosure does not limit the specific fusion function, as long as it can calculate a value reflecting the urgency of the intention based on these three intermediate quantities.

[0206] S530. Based on the behavior objects of user behavior in the user behavior sequence, extract the behavior objects that appear more frequently than a preset threshold in the user behavior sequence and / or the behavior objects in the combination of behavior objects that appear more frequently than a preset threshold in the user behavior sequence as core entities, and obtain the predefined tags associated with the core entities.

[0207] The recurrence frequency refers to the number of times the same action object appears in the entire user behavior sequence. Objects with a recurrence frequency exceeding a preset threshold are considered core entities that best represent the user's current intent. For example, if a user searches for, browses, and favorites the same mobile phone multiple times, then that mobile phone ID is a core entity.

[0208] Co-occurrence frequency (COF) is the number of times two or more different behavioral objects appear simultaneously within the same time window or adjacent behavioral locations. Objects in a combination whose COF exceeds a preset threshold (e.g., appearing simultaneously twice) can also be considered core entities (or core entity candidates). For example, if a user browses two mobile phones simultaneously and alternates between them multiple times, both objects could be core entities, representing the user comparing the two products.

[0209] Predefined tags are a pre-maintained tag library, such as "foldable screen phone," "high-end phone," and "limited-time offer." Each tag can be associated with a set of behavioral objects or keywords. By matching the product category and attributes corresponding to the core entity, a set of associated tags can be obtained for subsequent graph operations (such as establishing edges between intent nodes and tag nodes).

[0210] In some possible implementations, all behavioral objects in the sequence are traversed, and the frequency of each object's occurrence is counted. Objects with an occurrence count greater than or equal to a preset threshold are added to the core entity set. Alternatively, an N×N co-occurrence matrix is ​​constructed (where N is the number of different objects), and the user behavior sequence is scanned. For each time window, each pair of objects that occur simultaneously within the window is recorded, and their co-occurrence count is incremented. Objects in pairs with a co-occurrence count greater than or equal to the preset threshold are added to the core entity set.

[0211] For each core entity, its associated tags are found based on a predefined mapping table (e.g., product ID → category tag). If the tag system is hierarchical, first-level, second-level categories, etc., can be selected. Tags can also be generated for entities using LLM or a classifier. The final output is a set of tags as predefined tags associated with the core entities.

[0212] Unlike related technologies that rely solely on click counts or simple time decay, this embodiment comprehensively utilizes the concentration of behavioral time sequences, trends of change, and dwell time to characterize the intensity of user intent from multiple dimensions.

[0213] By combining repetition frequency and co-occurrence frequency to extract core entities, this approach addresses the issue of missing core entities when a single object appears insufficiently but multiple objects co-occur (e.g., in comparison scenarios). When comparing multiple similar products, co-occurrence frequency can identify the actual objects being compared by the user, while repetition frequency may become ineffective because each product is only viewed once. The two complement each other, improving the recall rate of core entity extraction.

[0214] Figure 6 This diagram illustrates a flowchart of one implementation method for generating graph manipulation instructions based on the intent evolution stage, intent activation strength, core entities, and labels, when the heterogeneous graph also includes intent nodes. Figure 6 As shown, steps S610, S620, and S630 may be included.

[0215] S610. Based on the intent evolution stage, generate a node creation instruction. The node creation instruction is used to create intent nodes in the heterogeneous graph.

[0216] As mentioned above, intent nodes are a dynamic node type in heterogeneous graphs, used to explicitly represent a user's short-term intent at a specific moment (e.g., "buy a foldable phone"). Unlike ordinary content nodes (such as products or articles), intent nodes are not fixed entities, but rather temporary nodes that are dynamically created and recycled based on the user's real-time behavior.

[0217] Each intent node has an initial embedding vector (generated by weighted embeddings of associated tags or core entities), a creation timestamp, and a lifecycle state (created, active, decaying, reclaimed). As an "anchor" for user intent, intent nodes can propagate their signals to other nodes associated with them, thereby influencing downstream node embedding updates and recall ranking.

[0218] The Create Node instruction is a graph operation instruction used to add a new intent node in a heterogeneous graph. This instruction includes the node type (intent node), node identifier (which can be automatically generated), and information for initializing the node's embedding vector (such as associated core entities and labels). Upon execution, a new intent node will be added to the graph database, and an initial embedding vector and timestamp will be assigned.

[0219] In some possible implementations, the intent evolution stage reflects the user's current decision-making progress (such as information gathering, active comparison, imminent decision-making, etc.). An intent node is created only when the intent evolution stage indicates that the user has a clear and actionable intent. For example, if the stage is "casual browsing" or "no intent," no node is created; if the stage is "active comparison" or "imminent decision-making," an intent node is created.

[0220] The specific logic for generating node creation instructions is as follows: Determine the node type: It is fixed as "intent node" (or subdivided into subtypes according to the intent stage, such as "compare intent node" and "purchase intent node").

[0221] Determine the node identifier: typically a unique ID is used.

[0222] Determine the initial embedding vector: Based on the core entities and labels output during the intent inference stage, obtain the node embeddings (or semantic vectors) of these entities / labels in the heterogeneous graph. The initial embedding vector of the intent node is obtained through a weighted average (the weights are positively correlated with the intent activation strength or the frequency of entity occurrence). This embedding vector will serve as the input to the heterogeneous graph neural network.

[0223] Set creation timestamp: Record the time when the node was created for subsequent lifecycle management (such as the start point of the decay phase).

[0224] The instruction format can be a JSON structure, which includes the operation type "CREATE_NODE", node ID, node type, embedding vector, timestamp, etc.

[0225] S620. Based on the core entity and the tag, generate an edge addition instruction. The edge addition instruction is used to establish edges between intent nodes and user nodes, push content nodes, or nodes corresponding to core entities in the heterogeneous graph.

[0226] The "Add Edge" instruction is a graph operation instruction used to establish an edge between two specified nodes. In this embodiment, one end of the edge is typically a newly created intent node, and the other end can be a user node, a push content node, or a node corresponding to a core entity (e.g., a product node or article node). The "Add Edge" instruction requires specifying the edge type (e.g., "intent-related edge"), direction (e.g., directed), and initial weight. After execution, the heterogeneous graph topology changes, and these new edges will be considered during subsequent graph neural network updates.

[0227] The nodes corresponding to the core entities refer to the nodes representing the core entities in the heterogeneous graph, which are usually entity nodes within the application (such as product nodes and video nodes). These nodes already have rich user behavior edges and attribute features. After an intent node establishes an edge with them, it can quickly obtain semantic information and user preferences from these nodes.

[0228] In some possible implementations, core entities and tags are key to connecting user intent with actual content. The "Add Edge" instruction is used to connect newly created intent nodes with user nodes, push content nodes, and nodes corresponding to core entities.

[0229] In some specific implementations, the logic for generating edge instructions can be as follows: Establish an edge between the intent node and the user node: this indicates that "the user currently has this intent". The edge type can be defined as an "intent-associated edge". The weight is determined by the intent activation strength.

[0230] Establish edges between intent nodes and corresponding core entity nodes: this indicates that "the intent is primarily aimed at these entities." For example, if a user compares two products, edges are established from the intent node to each of the two product nodes. The edge type can be defined as "TARGETS." The weights are also based on the intent activation strength.

[0231] Establish edges between intent nodes and label nodes: these edges represent "the intent is associated with these semantic labels". For example, labels "foldable screen phone" and "high-end phone". The edge type can be defined as "RELATES_TO". The weight can be set as the intent activation strength multiplied by the label relevance score.

[0232] Establishing edges between intent nodes and push content nodes: If some push content has already clearly matched the intent (e.g., based on a content understanding model), edges can be established directly, but this is usually not necessary.

[0233] S630. Based on the intent activation strength, generate a weight setting instruction, which is used to set the weight of the edge.

[0234] The weight setting command is also a graph operation command used to assign or modify weights to edges (including newly added edges or existing edges). In this embodiment, the weight setting command is usually used in conjunction with the add edge command to set an initial weight for a new edge; it can also be used independently to adjust the weight of existing edges (such as enhancing or diminishing them). The weight value is generally positively correlated with the intent activation strength; the higher the strength, the greater the edge weight, indicating a closer association between the intent node and the connected node.

[0235] For newly added edges: The weight field is directly included in the edge addition instruction. The weight is regarded as the intent activation strength (or mapped according to the edge type, for example, the weight of the user-intent edge is directly equal to the strength, while the weight of the intent-label edge can be set to the strength multiplied by 0.9). The weight is set when the edge is added.

[0236] For existing edges (such as behavioral edges between user nodes and core entity nodes): if the intent activation strength is high, a weight enhancement instruction can be generated to increase the weight of the existing edge by an increment. This helps to strengthen signals in historical behaviors that are relevant to the current intent.

[0237] The weight setting instruction can be independent of the edge addition instruction, or it can be combined with it (setting the weight at the same time as adding the edge). For clarity, the embodiments of this disclosure treat weight setting as a separate operation, but in actual implementation, it is usually combined with the edge addition instruction.

[0238] By creating dedicated intent nodes, users' short-term intents are encoded into a graph structure, enabling subsequent message passing and embedding learning to directly perceive this intent. Compared to simply adjusting edge weights, node-level representations can carry richer semantic information.

[0239] Based on core entities and tags, connections between intent nodes and related nodes are automatically constructed. This star-shaped structure (with the intent node in the center, connecting users, content, and tags) greatly enhances the propagation efficiency of intent signals: a user's intent is not only transmitted through their own node but also radiates to all related entities through the intent node. When subsequent user behavior changes, only the intent node needs to be updated instead of rebuilding the entire graph.

[0240] By directly mapping intent activation intensity to edge weights, quantitative propagation of intent intensity is achieved. High-intensity intents generate high-weight edges, thus dominating message aggregation in the graph neural network. This results in intent-related push content receiving greater embedding updates and higher rankings. The timestamp attached to the node creation command provides a foundation for subsequent intent node lifecycle management. By recording the creation time, the activity period of intent nodes can be determined, and decay and recycling can be triggered after timeout, preventing the permanent existence of intent nodes from interfering with long-term interests.

[0241] Figure 7The diagram illustrates a specific implementation of generating graph manipulation instructions based on the intent evolution stage, intent activation strength, core entities, and labels. Figure 7 As shown, it may include steps S710, S720, S730, and S740.

[0242] S710. In response to the preset node creation conditions being met during the intent evolution stage, a node creation instruction is generated. The node creation instruction is used to create intent nodes in the heterogeneous graph and initialize the initial embedding vector of the intent node based on the core entity and label.

[0243] The node creation condition is a pre-defined rule that triggers the creation of an intent node. For example, the node creation condition is met when the intent evolution stage is "active comparison" or "imminent decision," while it may not be met in the "casual browsing" or "information gathering" stages. This condition can be a Boolean expression or determined based on an intent activation strength threshold.

[0244] The initial embedding vector is the vector representation assigned when an intent node is created, used for subsequent message passing and embedding updates in the graph neural network. The initial embedding vector is typically generated by a weighted average of the node embeddings of the core entities and labels associated with the intent. For example, for the intent "buy a foldable phone," it can be obtained by fusing the embeddings of the label "foldable phone" and the product name "phone." The weighting coefficients can be the degree of relevance between the entity and the intent (e.g., the confidence level output by an LLM).

[0245] In some possible implementations, it is determined whether the preset node creation conditions are met during the intent evolution stage. If not, the node creation instruction is skipped, and only subsequent edge weight adjustments are performed. When the conditions are met, a node creation instruction is generated, containing all the information required to create the intent node. A unique ID is assigned to the intent node, and the node type is marked as "intent node".

[0246] The initial embedding vector is calculated based on core entities and tags. Specifically, this can be achieved by: obtaining the embedding vector for each core entity's corresponding node (if a core entity corresponds to multiple nodes, such as product nodes and tag nodes), and the embedding vector for each associated tag node; then performing a weighted average on these embedding vectors. The weights can be distributed evenly or assigned differently based on the relevance of the entity / tag to user behavior (e.g., frequency of occurrence, LLM confidence). This ultimately yields the initial embedding vector for the intent node. This vector is stored in the node attributes for subsequent use by the graph neural network.

[0247] S720. Generate an edge addition instruction. The edge addition instruction is used to establish edges between intent nodes and user nodes, push content nodes, or nodes corresponding to core entities in a heterogeneous graph. The edges are determined based on the core entities and tags, and the edges are set with initial weights.

[0248] The node corresponding to the core entity can be a node representing the core entity in a heterogeneous graph, typically an entity node within the application (such as a product node or video node). If the core entity is an abstract concept (such as "foldable phone"), it may correspond to a tag node. This node is the target of the user's behavior and the primary object for establishing connections with intent nodes.

[0249] The initial weight is the weight value assigned to a newly added edge upon creation. In this embodiment, the initial weight is typically proportional to the intent activation strength, but can also be fine-tuned based on the edge type. For example, the user-intent edge weight is directly equal to the intent activation strength; the intent-tag edge weight is equal to the intent activation strength multiplied by the tag relevance.

[0250] The "Add Edge" instruction is used to establish connections between intent nodes and related nodes. This step clarifies that one end of the edge is the intent node, and the other end can be a user node, a push content node, or a node corresponding to a core entity.

[0251] In some possible implementations, an edge between an intent node and a user node indicates that the user has that intent. The type of the edge can be defined as an "intent-associated edge". The initial weight of the edge is set to the intent activation strength.

[0252] In some possible implementations, for each core entity, its corresponding node (which could be a product node or a tag node) is found in the heterogeneous graph. Edges are established from intent nodes to these nodes, and the edge type can be defined as "target entity edge" or "TARGETS". The initial weights are also set to the intent activation strength, or multiplied by the matching degree between the entity and the intent (which can be obtained through LLM).

[0253] In some possible implementations, if there is push content that directly matches the intent (e.g., newly launched content that has been associated with the intent through semantic similarity), an edge can be established between the intent node and the push content node. However, in this embodiment, this edge is not mandatory.

[0254] In some possible implementations, the initial weights of all new edges are linked to the intensity of the intent activation; the higher the intensity, the larger the initial weights, and the stronger the influence of the intent node in the message passing of the graph neural network.

[0255] S730. In response to the intent activation strength being greater than the first threshold, a weight enhancement instruction is generated. The weight enhancement instruction is used to increase the weight of existing edges associated with the core entity in the heterogeneous graph.

[0256] The first threshold is a value used to determine whether the intent activation strength is high enough to trigger a weight enhancement instruction. When the intent activation strength is greater than the first threshold, it indicates that the user's intent is very strong, and the weights of historically relevant behavior edges need to be enhanced to increase the influence of these historical behaviors in subsequent recommendations.

[0257] Weight enhancement commands are also graph operation commands used to increase the weight of existing edges in heterogeneous graphs. The weight increment can be a fixed value or proportional to the intensity of intent activation. The enhanced edge weight does not exceed a preset maximum weight limit. This command typically operates on existing behavioral edges (such as browse or add-to-cart edges) between user nodes and corresponding nodes of core entities, because a strong current intent implies that historical interest in related entities should be assigned a higher weight.

[0258] S740, in response to the intention activation strength being less than the second threshold, generates a weight decay instruction, which is used to accelerate the weight decay rate of existing edges associated with core entities in the heterogeneous graph.

[0259] The second threshold is used to determine whether the intent activation strength is low enough to trigger a weight decay instruction. The second threshold is typically lower than the first threshold. When the intent activation strength is below the second threshold, it indicates that the user's current intent is weak, requiring weight decay of existing related edges to reduce their impact on recommendations and avoid interference from historical noise.

[0260] The weight decay instruction is also a graph operation instruction used to accelerate the decay rate of existing edge weights. In practice, the decay coefficient of an edge can be temporarily amplified by a factor, causing the edge weight to decrease faster over time; or a fixed decay amount can be directly subtracted. This disclosure preferably uses the accelerated decay rate method to respect the original exponential decay mechanism, maintain the smoothness of exponential decay, and avoid abrupt changes. This instruction also primarily operates on existing edges associated with core entities, such as product edges that a user has previously viewed but has not recently viewed. The purpose of weight decay is to reduce interference from historical behaviors related to the current weak intent, allowing the model to focus more on other signals.

[0261] In some possible implementations, when the intention activation strength is between the second threshold and the first threshold, no weight adjustment is made, and the original edge weights remain unchanged to maintain the natural influence of historical behavior.

[0262] Through the above implementation, core entity and label information are fused into the initial embedding vector of the intent node, giving the intent node rich semantic information from its inception. This is more beneficial for subsequent learning of the graph neural network than random initialization or zero vector initialization, enabling faster propagation of intent signals. By setting a first threshold and a second threshold, differentiated processing of user intent is achieved. When the intent is strong, the weights of historically relevant edges are increased, amplifying the user's past interest in core entities, which helps recall the current intent. Figure 1 The content is relevant; when the intent is weak, the weight of historically relevant edges is reduced to avoid outdated interests interfering with recommendations; when the intent is moderate, the weight remains unchanged to maintain natural decay.

[0263] Figure 8 The diagram illustrates a flowchart of one implementation method for dynamically updating the topology of a heterogeneous graph based on the intent reasoning result after generating a graph operation instruction. Figure 8 As shown, steps S810, S820, S830, and S840 may be included.

[0264] S810. Establish the intent node and its associated edges in the heterogeneous graph using the initial embedding vector and initial weights, and initialize the timestamp of the intent node to the creation time.

[0265] The associated edges of an intent node are all edges directly connected to the intent node, including but not limited to: the "intent-related edges" between the intent node and the user node, the "TARGETS" edges between the intent node and the corresponding core entity node (such as the product node), and the "RELATES_TO" edges between the intent node and the tag node. These edges together constitute the local subgraph structure of the intent node.

[0266] A timestamp is an attribute of an intent node that records the time when the node was last "activated". It is initialized to the creation time when the node is created; when subsequent user actions meet the update conditions, the timestamp is updated to the current time. Timestamps are used to determine whether an intent node is active and to account for trigger decay.

[0267] In some possible implementations, this step involves adding a new intent node and storing its initial embedding vector (obtained by fusing the core entity and label). Edges are added between the intent node and the user node, core entity node, and label node, with initial weights set (typically equal to the intent activation strength). A timestamp attribute is set for the intent node, initialized to the current system time.

[0268] At this point, the intent node is in the "creation phase", and the weights of its associated edges are all initial values, with the timestamp being the creation time.

[0269] S820. Within a preset time window starting from the timestamp, in response to the intent node meeting the preset update conditions, update the timestamp of the intent node to the current time and increase the weight of the associated edge.

[0270] The preset time window is a fixed time length, which is used to determine whether the user continues to exhibit behaviors related to the intent within a reasonable time interval.

[0271] The preset update condition is a judgment rule that triggers the refresh of the intent node (timestamp update, edge weight increase). In some possible implementations, it can be that within a preset time window, the number of times the behavior object matching the core entity and label associated with the intent node appears in the user behavior sequence or the repetition pattern of the behavior type reaches a preset threshold.

[0272] Increasing the weight of the associated edges involves updating the timestamps and increasing the weights of all associated edges when the update conditions are met. The increment can be a fixed value or a value related to the current intent activation strength. After increasing the weights, the edge weights should not exceed a preset maximum value. This operation ensures that the influence of the active intent remains high.

[0273] In some possible implementations, user behavior sequences are continuously received. For each active intent node, starting from the current timestamp of that intent node, it is calculated whether the time from the timestamp to the current time exceeds a certain window. If the time window has not been exceeded, whenever a new user behavior is received, it is determined whether the behavior matches the core entity and tag associated with that intent node (e.g., whether the behavior object belongs to the core entity set, or whether the tag associated with the behavior object overlaps with the tag of the intent node). If a match is found, the "last match time" record of that intent node is updated.

[0274] In some specific implementations, the number of times a matching behavior or a recurring pattern of behavior type is counted within a preset time window. Specifically: Count the number of times a matched action object appears. For example, if a user browses a certain phone again, the count is incremented by 1.

[0275] Alternatively, it can identify repetitive patterns of the same behavior type: for example, if a user searches for "foldable phone" twice consecutively within a window, it is considered to meet the repetitive pattern.

[0276] If the number of matches reaches the preset threshold, or the repeating pattern meets the conditions, then the update conditions are considered met.

[0277] When the update conditions are met, the timestamp of the intent node is updated to the current time (i.e., the active period is refreshed). The weights of all associated edges of the intent node are increased. The increase can be a fixed value or a dynamic value related to the strength of the current matching behavior (e.g., the instantaneous intent strength calculated in the matching behavior). After the increase, it must be ensured that the weights do not exceed the upper limit. After the update, the intent node re-enters a complete "active phase", and the next time window is calculated starting from the new timestamp.

[0278] S830. Within a preset time window starting from the timestamp, in response to the intention node not meeting the update conditions, the weight of the associated edge is reduced according to the preset decay function.

[0279] The decay function is a mathematical function used to reduce the weight of the edges associated with the intent node. The preferred embodiment of this disclosure uses an exponential decay function: ,in, This represents the elapsed time since the time point corresponding to the timestamp. Linear or squared decay can also be used. The decay function ensures that the edge weights of the intent nodes gradually decrease over time without any new behavioral reinforcement.

[0280] Reducing the weight of the associated edges involves decreasing the weight of all associated edges according to a decay function when the update conditions are not met within a preset time window. The parameters of the decay function can be configured to ensure that the rate of weight reduction matches business expectations.

[0281] In some possible implementations, if no action meeting the update conditions occurs within a preset time window, decay is triggered, meaning the weights of all associated edges are reduced according to a preset decay function. After reducing the weights, the timestamp corresponding to the intent node can be updated, continuing the calculation from the original timestamp.

[0282] In some possible implementations, if no update condition is met within a preset time window starting from the timestamp, then a decay is performed. For simplicity, the embodiments of this disclosure adopt "decay is performed once per complete window period".

[0283] S840. In response to the weight of all associated edges being lower than a preset threshold, delete the intention node and all its associated edges from the heterogeneous graph.

[0284] The preset threshold is the lower limit of the weight used to determine whether an intent node should be recycled. For example, when the weight of all associated edges is below 0.05, the intent node is considered to no longer represent the user's current interest and should be completely removed from the heterogeneous graph, including the node itself and all its associated edges, in order to free up storage resources and avoid interference.

[0285] In some possible implementations, after each decay operation (or after each weight update), it is checked whether the weights of all associated edges of the intent node are below a preset threshold. If so, the intent node and all its associated edges are deleted from the heterogeneous graph. The deletion operation includes: removing the node and its attributes; removing all edges with that node as the source or destination; and cleaning up the relevant indexes in the graph database.

[0286] Once an intent node is reclaimed, it no longer participates in subsequent graph neural network message passing and recall ranking, thus completely eliminating its impact.

[0287] Users' short-term interests typically exhibit characteristics of "rapid rise, slow decay, and potential reactivation." This disclosure utilizes a timestamp refresh mechanism (resetting the timestamp when update conditions are met) and a decay function (reducing weight when not updated) to dynamically track the intensity of interest. During the active phase, interest has high weight and significant influence; during the decay phase, its weight gradually decreases; and during the recycling phase, it is completely deleted. This aligns with the laws of human cognition.

[0288] Without the aforementioned recycling mechanism, previously created intent nodes would persist indefinitely, causing the recommendation system to consistently favor historical intents. Threshold-based judgment and deletion operations can automatically clean up irrelevant intent nodes, ensuring that recommended content reflects the user's current true needs and improving the timeliness and accuracy of recommendations.

[0289] The update criteria can be either the number of times a behavior object appears or the repetition pattern of a behavior type, covering different forms of user reinforcement behavior (such as searching again, browsing repeatedly, clicking multiple times, etc.). This flexibility allows the system to adapt to diverse user behavior patterns and avoids incorrectly judging active intent as decay due to the absence of a specific behavior.

[0290] By increasing the weights of related edges, the signal of the intent node is amplified in the message passing of the graph neural network, resulting in a larger embedding update amplitude for push content nodes related to the intent, thereby improving their ranking in recall and ranking. Decay weights gradually reduce this influence. This mechanism achieves end-to-end feedback from "user behavior → intent node weight → graph embedding → recommendation ranking" without requiring retraining of model parameters.

[0291] Among some possible implementations, Figures 4 to 7 The process of inferring intent from user behavior sequences can be achieved by leveraging the thought chain reasoning capabilities of large models. Specifically, the thought chain reasoning capabilities of large models can be used to analyze user behavior sequences and generate structured intent analysis results.

[0292] Large language models (LLMs) refer to pre-trained language models based on the Transformer architecture, typically with more than one billion parameters. Unlike traditional masked language models, LLMs possess powerful text generation, context understanding, reasoning, and instruction following capabilities, and can generate complex structured outputs based on natural language prompts.

[0293] The thinking chain reasoning capability refers to the ability of an LLM (Low-Level Model) to solve complex problems by generating intermediate reasoning steps (such as "first...then...therefore..."). Unlike directly outputting the answer, the thinking chain forces a large model to break down the problem step by step, demonstrating the logical derivation process, thereby improving the accuracy and interpretability of the answer. In the embodiments of this disclosure, the thinking chain is used to analyze user behavior sequences, sequentially identifying behavior patterns, determining intent types, evaluating intent strength, extracting entities, and finally providing graph operation suggestions, thus achieving... Figures 4 to 7 A limited process of inferring intent from user behavior sequences.

[0294] Figure 9 This diagram illustrates a process for analyzing user behavior sequences using the reasoning capabilities of a large model to generate structured intent analysis results. Figure 9 As shown, steps S910 and S920 may be included.

[0295] S910. Construct prompt words containing user behavior sequences, including instructions that require the large model to perform multi-step reasoning on the user behavior sequences.

[0296] The prompts are natural language instructions input to the LLM (Limited Language Model) to guide the model in performing specific tasks. In this embodiment, the prompts include descriptions of user behavior sequences, instructions requiring the model to perform multi-step inference, and output format constraints. The design of the prompts directly affects the quality and structure of the model's output.

[0297] In some possible implementations, the raw sequence of user behavior can be converted into a natural language description. Each behavior event should include at least the behavior type (e.g., "search", "browse"), the behavior object (e.g., product name, article title), and time information (e.g., "2 minutes ago"). To improve the model's understanding, the behaviors can be arranged in reverse chronological order (latest first) and listed in a clear format.

[0298] In addition to the natural language description corresponding to the action sequence, the prompt must also contain instructions requiring the model to perform multi-step reasoning. These instructions can explicitly tell the model the reasoning steps to follow, for example: Please analyze the user behavior sequence following these steps: Behavioral pattern recognition: Identifying common themes, repetitive patterns of behavior, and temporal concentration in behavioral sequences.

[0299] Intent type determination: Determine which decision-making stage the user is in (information gathering, active comparison, imminent decision, casual browsing).

[0300] Intent intensity assessment: A score between 0 and 1 is given, based on behavior density, behavior type escalation trend, dwell time, etc.

[0301] Intent entity extraction: Extract core entities and associated tags.

[0302] Graph manipulation suggestions: What kind of intent nodes should be created, which nodes should be connected to them, and what should the edge weights be?

[0303] At the same time, the model is required to output in a structured format (such as JSON) to ensure that the results are parsable.

[0304] S920. Input the prompt words into the large model and obtain the structured intent analysis results output by the large model.

[0305] Structured intent analysis results refer to data objects with fixed fields and formats output by the LLM, such as JSON, containing intent type, intent strength, core entities, associated tags, graph operation suggestions, etc. These results can be directly parsed and used without further natural language processing.

[0306] In some possible implementations, the constructed prompts are sent to the LLM. Based on its internal knowledge and reasoning capabilities, the LLM proceeds step by step according to the instructions, generating a JSON object containing various fields. This JSON is then parsed to extract information such as intent type, intent strength, core entities, tags, and graph operation suggestions.

[0307] If the LLM output contains natural language descriptions, key fields can be extracted through post-processing, but it is best to request plain JSON output in the prompt words.

[0308] In some possible implementations, pre-screening can also be performed, that is, LLM is only called when the user behavior sequence is judged to meet the preset triggering conditions, while inference throttling is set for the same user (such as not calling it repeatedly within 60 seconds).

[0309] The following specific example illustrates the process of analyzing user behavior sequences and generating structured intent analysis results using the reasoning capabilities of a large model.

[0310] Taking a foldable phone comparison scenario as an example, the user behavior sequence (last 5 items) is as follows: [2 minutes ago] Searched for "256GB foldable phone price" [3 minutes ago] Browsing a certain mobile phone model, stayed for 85 seconds. [3.5 minutes ago] Browse other phone models, stay for 120 seconds [4 minutes ago] Searched for "Foldable phone recommendations 2025" [5 minutes ago] Browse "Regular Phone Cases" Based on this user behavior sequence, the following prompt words are generated (example): text You are a user intent analysis expert. Please analyze the following sequence of user behaviors (in reverse chronological order), reason step by step, and output the results in JSON format.

[0311] User behavior sequence: 1. [2 minutes ago] Searched for "256GB foldable phone price" 2. [3 minutes ago] A certain model of mobile phone, paused for 85 seconds. 3. [3.5 minutes ago] Browse other phone models, stay for 120 seconds. 4. [4 minutes ago] Searched for "Foldable phone recommendations 2025" 5. [5 minutes ago] Browse "Regular Phone Cases" Reasoning steps: Step 1: Behavioral pattern recognition: Identify common themes, recurring behaviors, and temporal concentration.

[0312] Step 2: Intent type determination: Information gathering / Active comparison / Imminent decision / Casual browsing.

[0313] Step 3: Intent intensity assessment: Give a score of 0-1, taking into account behavior density, escalation trend, and dwell time.

[0314] Step 4: Intent entity extraction: Extract core entities and associated tags.

[0315] Step 5: Graph operation suggestions: It is recommended to create intent node types, establish edges with which nodes, and define edge weights.

[0316] Output format (output only JSON, no other text): { "intent_type": "", "intent_strength": 0.0, "core_entities": [], "tags": [], "graph_operations": {"create_node":{"node_type": "intent", "initial_weights": []}, "add_edges": [{"src": "user", "dst": "intent", "weight": 0.0},...], "boost_edges": [{"edge_id": "...", "delta": 0.0}]}} Suppose that calling LLM returns the following: json {"intent_type": "Active Comparison", "intent_strength": 0.85, "core_entities": ["Foldable Screen Phone", ], "tags": ["Foldable Screen Phone", "High-End Phone", "256G"], "graph_operations": {"create_node": {"node_type": "intent", "initial_weights":[0.85]}, "add_edges": [{"src": "user", "dst": "intent", "weight": 0.85}, {"src": "intent", "dst": "tag_foldable screen phone", "weight": 0.9}, {"src": "intent", "dst": "tag_high-end phone", "weight": 0.7}], "boost_edges": [{"edge_id": "user_item_huawei", "delta": 0.3}]}} Parsing the JSON, we extract the intent type as "Active Comparison," intent strength as 0.85, core entities and labels as shown above, and graph operation suggestions. This data will be used to create intent nodes, add edges, and enhance edge weights.

[0317] If the user behavior sequence belongs to a scenario without a clear intent (e.g., only browsing a regular product), LLM may output "intent_type": "casual browsing", "intent_strength": 0.2. In this case, the graph operation suggestion may be empty or no node should be created.

[0318] LLM learns massive amounts of human behavioral patterns and common-sense knowledge through pre-training, enabling it to infer deep decision-making stages from finite behavioral sequences—something traditional models cannot do. Through carefully designed prompts and JSON-formatted constraints, LLM directly outputs structured data that can be used directly by downstream programs, eliminating the need for complex natural language parsing and reducing development costs. Furthermore, different users express themselves differently (e.g., diverse search terms, different browsing orders). LLM, with its powerful generalization ability, can handle various variations without requiring separate rules for each pattern, as traditional rule engines do. This improves the system's robustness and scalability. The thought chain reasoning requires LLM to output intermediate reasoning steps (although not explicitly shown in this embodiment, they can be recorded in debug mode). Operations personnel can review the model's reasoning process to understand why a "proactive comparison" was determined, facilitating optimization and auditing. Compared to black-box deep learning models, this interpretability is crucial in real-world business scenarios.

[0319] The graph operation suggestions output by LLM (create nodes, add edges, enhance edges) can be directly mapped to specific instructions, achieving end-to-end automation from natural language understanding to underlying graph structure updates. Compared to traditional methods that require manual design of intent node creation rules, the embodiments disclosed in this disclosure are more flexible and adaptive.

[0320] In some possible implementations, when the topology of the heterogeneous graph is dynamically updated due to intent reasoning (such as adding intent nodes, adding edges, or adjusting edge weights), a heterogeneous graph neural network is used to update the node embeddings of the updated heterogeneous graph. The updated embedding vectors corresponding to user nodes and push content nodes can be obtained by performing forward propagation of the heterogeneous graph neural network only on the local neighborhood subgraphs of the affected nodes, quickly updating the embedding vectors of the relevant nodes, while keeping the model parameters unchanged, so as to achieve a second-level online response while ensuring the embedding quality.

[0321] Figure 10 This diagram illustrates a flowchart illustrating the specific implementation of a heterogeneous graph neural network that performs forward propagation only on the local neighborhood subgraph of the affected nodes to quickly update the embedding vectors of the relevant nodes. Figure 10 As shown, it may include steps S1010, S1020, and S1030.

[0322] S1010: In response to the dynamic update of the topology of the heterogeneous graph, identify the nodes in the heterogeneous graph affected by the update and obtain the set of affected nodes.

[0323] In this context, nodes affected by the update refer to those whose embedding vectors may change during graph topology update operations (adding nodes, adding edges, adjusting edge weights). These typically include newly created nodes (such as intent nodes), nodes directly connected to newly added edges, and neighboring nodes within a certain hop range that can be indirectly affected through message passing in the graph neural network. The purpose of identifying affected nodes is to avoid recalculating the entire graph and to update only the necessary nodes.

[0324] The set of affected nodes is the collection of all nodes affected by the update. This set forms the basis for subsequent extraction of the neighborhood subgraph. The size of this set is typically much smaller than the total number of nodes in the graph (e.g., tens to hundreds).

[0325] After completing the "dynamic update of the heterogeneous graph topology" process through the above steps, it is necessary to determine which node embeddings need to be recalculated.

[0326] Newly created nodes (such as intent nodes) will definitely be affected because they did not exist before.

[0327] Nodes that are directly connected to the newly created node (such as user nodes, core entity nodes, and tag nodes) will also be affected because their new neighbors will change the result of message aggregation.

[0328] Furthermore, the neighbors of these direct neighbors may also be affected, as message propagation in graph neural networks spreads layer by layer. For an L-layer heterogeneous graph neural network, the embedding update of a node depends on the information of all its neighboring nodes within its L hops. Therefore, if a new node or edge affects a node, then all of that node's L-hop neighbors (actually, nodes within the L-hop range of the new node / edge) should be updated.

[0329] In some possible implementations, the set of affected nodes is initialized as {newly created nodes} ∪ {the two endpoints of the new edge}. For each node in the set, all its neighbor nodes are added to the set (breadth-first search), and this is repeated L times (L is the network layer). The final set contains all nodes whose embeddings need to be recalculated.

[0330] Of course, a simpler approach can also be used: directly extract all nodes in the L-jump subgraph of the new node / edge.

[0331] S1020. Extract the k-hop neighborhood subgraph of the affected node set in the heterogeneous graph, where k is the number of layers in the heterogeneous graph neural network.

[0332] For a given set of affected nodes S, its k-hop neighborhood subgraph refers to all nodes reachable by traversing at most k steps along edges (regardless of direction) from each node in S, and the subgraph formed by the edges between these nodes. Here, k can be the same as L, representing the number of layers in the heterogeneous graph neural network. Since the message propagation range of a graph neural network is limited by the number of layers, the final embedding of a node depends only on the information of its k-hop neighbors. Therefore, extracting the k-hop neighborhood subgraph ensures the integrity of the affected node embedding updates.

[0333] In some possible implementations, a local subgraph is formed by extracting all nodes and edges covered by k hops from the original heterogeneous graph, starting from the set of affected nodes.

[0334] Specifically, initialize the subgraph node set to the affected node set.

[0335] For each hop (1 to k): add the direct neighbors of all nodes in the current subgraph node set (via edges of any type) to the subgraph node set.

[0336] Extract all edges (including edge type and edge weight) between nodes in the subgraph node set from the original heterogeneous graph to form a subgraph.

[0337] The size of the subgraph is related to the size of the set of affected nodes and the average degree of the graph. Although intent reasoning is triggered more frequently, the scope of each impact is limited (the subgraph typically contains only a few hundred nodes, a significant reduction compared to the millions of nodes in the full graph), resulting in faster computation speed.

[0338] S1030. Perform forward propagation of the heterogeneous graph neural network on the k-hop neighborhood subgraph to update the embedding vectors of nodes in the affected node set while keeping the model parameters of the heterogeneous graph neural network unchanged.

[0339] Forward propagation refers to the process of inputting node features and edge information into the model and calculating the node embedding vector layer by layer. Forward propagation does not involve gradient calculation or parameter updates, therefore it is computationally fast and suitable for online inference. In this embodiment, forward propagation is performed only on a local subgraph, rather than the entire heterogeneous graph.

[0340] Maintaining unchanged model parameters means keeping the learnable parameters of the heterogeneous graph neural network, such as the weight matrix and bias terms, exactly the same as the values ​​after offline full training. Figure 10 The corresponding incremental updates only perform forward propagation calculations, without backpropagation or parameter updates. This ensures the model's consistency with the global data distribution, while avoiding the instability and high overhead of online training.

[0341] In some possible implementations, the exact same model parameters (weight matrix, bias, etc.) as those used in offline full training are used; the initial embedding vectors of the nodes in the input subgraph are used (for newly created intent nodes, their initial embedding vectors are used; for other nodes, the currently stored embedding vectors are used); and the node embeddings are computed layer by layer according to the standard forward propagation formula for heterogeneous graph neural networks. Due to the small size of the subgraph, the computation can be completed in milliseconds.

[0342] After obtaining the updated embedding vector for each node in the subgraph, for all nodes in the subgraph, replace the old embedding vector with the new one, keeping the model parameters unchanged, and do not perform backpropagation. Then, write the updated embedding vector for each node in the subgraph back to the graph database or vector indexing service for subsequent retrieval and ranking. For nodes not in the subgraph, their embedding vectors remain unchanged, requiring no further action.

[0343] The time complexity of forward propagation across the entire graph is O(|V| + |E|), where |V| can reach millions or even billions of nodes, making it impossible to complete within seconds. However, by calculating only a local subgraph (typically a few hundred nodes), the complexity is reduced to O(|V_sub| + |E_sub|), lowering the online inference latency from minutes to milliseconds / seconds, which meets the requirements of real-time recommendations. Furthermore, since only forward propagation is performed without updating parameters, the issues of parameter fluctuations or forgetting due to online learning are avoided. The model parameters are still managed through daily offline full training, ensuring global optimality.

[0344] By extracting a k-hop neighborhood subgraph (where k equals the network layer number), it is ensured that the embeddings of nodes affected by the update correctly reflect the changes in graph structure, without missing long-distance dependencies. After intent nodes are dynamically created and deleted, the above updates can quickly propagate the intent node's signal to user nodes and content nodes in its neighborhood, enabling downstream recall and ranking to immediately perceive the new intent. Simultaneously, when an intent node decays and is deleted, incremental updates can quickly remove its influence, avoiding the remnants of old intents.

[0345] The above method is applicable to heterogeneous graphs of any size because the cost of each update depends only on the size of the local region, not the overall graph size. It maintains responsiveness even as the number of users and the amount of content increases.

[0346] In some possible implementations, after updating the affected embedding vectors, the push content is recalled and sorted based on the updated embedding vectors, and the final push content is determined based on the sorting results.

[0347] The first stage of a recommendation system is recall, which involves quickly filtering hundreds or thousands of potentially relevant candidates from a massive pool of candidate content (usually in the millions) for further detailed scoring in the subsequent ranking stage. Recall requires high efficiency and a certain level of coverage, while allowing for a certain degree of false positives.

[0348] Figure 11 This diagram illustrates a flowchart of an implementation method for recalling and ranking push content based on updated embedding vectors, and determining the final push content to be delivered based on the ranking results. Figure 11 As shown, at least one of the following recall strategies are executed in parallel, and the recall results are combined. This may include steps S1110, S1120, S1130, and S1140.

[0349] S1110. Calculate the similarity between the updated embedding vector of the user node and the updated embedding vector of each push content node, select the K push content nodes with the highest similarity, and use the push content corresponding to the push content nodes as the recall result.

[0350] This step implements a recall strategy based on the similarity of embedded vectors. Specifically, it calculates the similarity (e.g., cosine similarity) between the updated embedded vector of the user node and the updated embedded vectors of all pushed content nodes, and selects the K pushed content nodes with the highest similarity. Vector recall can capture the global similarity between users and content in the semantic space and is suitable for general interest matching.

[0351] In some possible implementations, since the number of push content nodes can be very large (in the millions), an Approximate Nearest Neighbor (ANN) index can be used to accelerate retrieval. An index is pre-built on the embedding vectors of all push content nodes. When online, the user's embedding vector is queried, and the K nodes with the highest similarity are returned.

[0352] Where K is a positive integer representing the number of push content returned by vector recall, which can be set according to system resources and business needs.

[0353] S1120. Perform path search in the heterogeneous graph according to the preset meta path, obtain the push content node that has a high-order association with the user node, and use the push content corresponding to the push content node as the recall result.

[0354] This step implements a recall strategy based on predefined meta-paths in heterogeneous graphs. Meta-paths define higher-order semantic association patterns between nodes. For example, following the meta-path "user → in-application entity → tag → push content", starting from the user node, push content nodes that match the path pattern are found through graph traversal. Meta-path recall can leverage the rich edge types and cross-domain knowledge in heterogeneous graphs to uncover implicit, indirect higher-order relationships between users and content.

[0355] In some possible implementations, a path search of a finite number of steps is performed on the heterogeneous graph, starting from the user node, based on the preset meta-paths (i.e., the first meta-path, second meta-path, and third meta-path mentioned above).

[0356] In some specific implementations, breadth-first search or random walk can be used to record the push content node at the end of the path, and the path weights (such as the product of the weights of each side) are accumulated as the score. The M nodes with the highest scores are selected (M is also a positive integer, which can be set according to system resources and business needs).

[0357] S1130. Based on the click-through rate statistics of the push content in the global scope, select the N push contents with the highest click-through rate as the recall results.

[0358] This step is implemented based on a recall strategy that leverages the global popularity of the pushed content. Specifically, it can involve calculating the click-through rate (or impressions, conversion rate, etc.) of all pushed content globally and selecting the N pushed contents with the highest click-through rates.

[0359] Hot topic recall ensures that popular, high-quality push content gets exposure, especially suitable for new users or less popular users.

[0360] In some possible implementations, a global push content click rate ranking can be maintained, sorted in descending order by the click rate of the most recent time period (e.g., the past 24 hours), and the top N push content nodes can be selected.

[0361] Where N is a positive integer, representing the number of push content returned by hotspot recall, which can be set according to system resources and business needs.

[0362] In some possible implementations, all three recall strategies mentioned above can be activated simultaneously. Alternatively, at least one strategy can be selected as needed, with each strategy running independently and without interference. Parallel execution can utilize multi-threading, multi-processing, or distributed computing resources to improve recall speed.

[0363] S1140. Sort the results according to the recall results, and determine the final push content based on the sorting results; In some possible implementations, the sets of push content nodes obtained from the three recall strategies are merged, and duplicate node IDs are removed to form a candidate set. The union operation ensures the diversity of the candidate set: vector recall provides semantically similar content, meta-path recall provides content with cross-domain high-order associations, and hotspot recall provides popular content.

[0364] For each push notification in the candidate set, a ranking model is used to calculate a final score, and the content is sorted in descending order of score. The ranking model can be independent of the recall strategy and is typically based on user embedding, push content embedding, and contextual features for prediction. The order of the outputs in the ranking stage determines the final delivery priority.

[0365] Based on the ranking results, select the top-ranked push notifications and deliver them to the user's device through the push channel. Further filtering can be achieved by combining business rules (such as frequency control and deduplication).

[0366] Vector-based recall leverages the global semantic information embedded in nodes, meta-path recall utilizes higher-order relationships and cross-domain knowledge in the graph structure (such as bridging through in-app entities and tags), and trending topic recall utilizes collective intelligence. Combining these three approaches fully leverages the information advantages of heterogeneous graphs. However, vector-based recall may overlook content associated through higher-order paths; meta-path recall may miss purely popular content; and trending topic recall may miss personalized content. By taking the union of these three strategies, they complement each other, enabling a more comprehensive coverage of push content that users might be interested in. Compared to using only one type of recall, the ranking after multi-path recall better reflects the user's overall interests.

[0367] Figure 12 The diagram illustrates a flowchart of one implementation method for determining the final push content based on the sorting results, such as... Figure 12 As shown, steps S1210, S1220, and S1230 may be included.

[0368] S1210. Obtain the updated embedding vector of the user node, the updated embedding vector of the push content node, and the context features. Contextual features include at least one feature related to the user's current context, which includes at least one of the following: current time period, user device type, time interval since the last push, and user activity level.

[0369] Here, the updated embedding vector of the user node is the user node embedding vector updated by the heterogeneous graph neural network, denoted as... This embedded vector integrates the user's historical behavior, cross-domain entity interactions, and the influence of real-time intent nodes, providing a comprehensive representation of the user's current interests.

[0370] The updated embedding vector of the push content node is the push content node embedding vector updated through the heterogeneous graph neural network, denoted as... This embedding vector reflects the semantics of the push content, its association with tags, and its higher-order relationships with other entities.

[0371] Contextual features are features related to the user's current environment and are independent of the static embedding of the user and content. Contextual features include, but are not limited to: current time period (e.g., morning, afternoon, night), user device type (mobile phone / computer), time interval since the last push, user activity level, network environment (wireless network / mobile data), etc. These features describe the context in which the recommendation occurs.

[0372] For each push content node in the candidate set, the updated embedding vector of the user node can be obtained from the vector index. And the updated embedding vector of the push content node At the same time, it collects the contextual characteristics of the current request (time period, device, interval since the last push, etc.).

[0373] S1220. Input the acquired features into the pre-trained ranking model to obtain the ranking score.

[0374] The ranking model can be a machine learning model used to finely score candidate push notifications to predict the probability of a user clicking or interacting with that content. The ranking score is a scalar output of the ranking model, representing the estimated click-through rate. The higher the score, the greater the likelihood that a user will click on the push notification.

[0375] In some possible implementations, the features obtained in step S1210 can be concatenated into a fixed-length embedding vector, which is then input into a pre-trained ranking model to obtain a ranking score.

[0376] In this embodiment, the ranking model input includes the updated embedding vector of the user node, the updated embedding vector of the pushed content node, element-wise multiplication, and contextual features, and outputs a ranking score (typically a value between 0 and 1). The ranking model can be implemented using techniques such as logistic regression, gradient boosting tree, multilayer perceptron (MLP), or factorization machine.

[0377] Assuming the ranking model is an MLP, its input is ,in, For contextual features. Embedding vectors for user nodes and push content nodes can capture the feature interactions between users and content, similar to the Hadamard product in deep cross-networks. It is often used for collaborative filtering and click-through rate prediction. The cross-features of user profiles and pushed content (such as the matching degree between user preference categories and pushed categories) are processed through several fully connected layers and activation functions. The output layer uses Sigmoid activation to obtain the ranking score. ,in, This refers to the function corresponding to the sorting model.

[0378] S1230. Determine the final push content based on the ranking score.

[0379] In some possible implementations, the top-K push notifications can be directly selected and sent in descending order of the ranking score. Alternatively, the ranking score can be multiplied by one or more adjustment factors (timeliness, fatigue level, diversity) to obtain the final score. Weighted adjustment can be performed after the ranking score is output, or the factors can be used as features input into the ranking model.

[0380] In some specific implementations, the ranking score is weighted and adjusted based on at least one of the following factors: Timeliness factor: The timeliness factor decays exponentially with the creation time of the pushed content; Fatigue factor: The fatigue factor is calculated based on the number of push notifications a user has received within a preset time window; Diversity Factor: The diversity factor is calculated based on the semantic difference between the updated embedding vector of the push content node and the updated embedding vector of the push content node recently received by the user. The final push notification content is determined based on the weighted ranking score.

[0381] The timeliness factor is a coefficient that depends on the creation time of the pushed content and is used to penalize outdated content. The timeliness factor is typically designed to decay exponentially with the creation time of the pushed content. ,in, For decay rate, For the current time, The creation time of the push content. The older the push, the lower the timeliness factor, and the lower the ranking score, thus encouraging the recommendation of fresh content.

[0382] The fatigue factor is a coefficient that depends on the number of push notifications a user has received within a preset time window, used to avoid excessively disturbing the user. The fatigue factor can be defined as: ,in, For users in the time window The number of push notifications already received. This represents the maximum number of push notifications allowed. When a user has received the maximum number of push notifications, the fatigue factor is set to 0, preventing further pushes from being sent.

[0383] The diversity factor is a coefficient based on the semantic differences between the pushed content and the user's recently received pushed content, used to avoid recommending homogeneous content. The diversity factor can be defined as: ,in, This is a set of push notification nodes that the user has recently received. For cosine similarity, The semantic feature vector corresponding to the pushed content. This is a semantic feature vector of the push notifications a user has recently received. If a candidate push notification is very similar to a push notification the user has recently received, the diversity factor approaches 0, thus lowering its ranking score and encouraging exploration of different types of content.

[0384] In some possible implementations, for each candidate push, a timeliness factor, a fatigue factor, and a diversity factor are calculated (one or more of these are selected based on business needs). The ranking score is then multiplied by the selected factors to obtain the final score. If a factor is not used, its value defaults to 1.

[0385] Sort all candidate push notifications by their final scores in descending order, and select the top-ranked (or top few) notifications with final scores exceeding a preset delivery threshold. Send these notifications to the user through the push channel. If the final scores of all notifications are below the threshold, do not send any notifications to avoid unnecessary disturbance.

[0386] The ranking model utilizes user embeddings, push embeddings, element-wise multiplication, and contextual features to capture higher-order interactions between users and content (through multiplication) as well as contextual factors (time, device, etc.). Compared to collaborative filtering that only uses user IDs and item IDs, this multi-feature fusion significantly improves the accuracy of click-through rate prediction.

[0387] The timeliness factor ensures that fresh content has a higher probability of being delivered, avoiding user fatigue caused by the system constantly pushing outdated content. The exponential decay function is smooth and natural, conforming to the law that content value decreases over time. The fatigue factor dynamically adjusts based on the number of push notifications a user has received, effectively preventing excessive pushes that could cause user resentment or even lead to the disabling of notification permissions. This mechanism guarantees user experience while adhering to push frequency limits. The diversity factor encourages recommendations of different types of content by penalizing push notifications that are highly similar to content recently received by the user, increasing opportunities for users to explore new interests. This avoids homogenization of the recommendation list and improves long-term user satisfaction. In some possible implementations, which factors can be enabled can be selected based on business strategy. For example, the fatigue factor might be disabled during promotional activities to increase reach, while it can be fully enabled at other times.

[0388] Based on and Figure 1 The method shown follows the same principle. Figure 13 A schematic diagram of the structure of a device for pushing content, as provided in an embodiment of this disclosure, is shown. Figure 13 As shown, the device 13 for pushing content may include: The heterogeneous graph creation module 1310 is used to obtain a pre-built heterogeneous graph; the heterogeneous graph includes multiple types of nodes and multiple types of edges, and the nodes include at least user nodes and push content nodes; The intent reasoning module 1320 is used to acquire user behavior sequences; in response to the user behavior sequence meeting preset trigger conditions, it performs intent reasoning on the user behavior sequence and dynamically updates the topology of the heterogeneous graph based on the intent reasoning results. The graph update module 1330 is used to update the node embedding of the updated heterogeneous graph using a heterogeneous graph neural network to obtain the updated embedding vectors corresponding to user nodes and push content nodes. The recall module 1340 is used to recall and sort the push content based on the updated embedding vector, and determine the final push content to be sent based on the sorting result.

[0389] In the push content device provided in this embodiment, content is pushed through pre-constructed heterogeneous graphs and real-time intent reasoning. Since heterogeneous graphs can integrate multi-source heterogeneous data (such as the relationships between different types of nodes), even if a user has very few direct interaction records with the pushed content, that user node can still indirectly obtain rich semantic information through other types of nodes and edges in the graph. Therefore, the content recommendation device provided in this embodiment can generate more robust initial representations for users and pushed content even in scenarios with extremely low click-through rates and sparse positive samples, significantly improving the cold-start recommendation effect.

[0390] In some possible implementations, the intent reasoning module is specifically used to: obtain the sequence features of the user behavior sequence, including the behavior type, behavior object, and behavior timing information of each user behavior in the user behavior sequence; based on the sequence features, determine the intent evolution stage, intent activation intensity, and the core entity and label associated with the user intent corresponding to the user behavior sequence; and generate graph operation instructions based on the intent evolution stage, intent activation intensity, and core entity and label, including at least one of the node type to be created, the edge to be established, and the edge weight to be set.

[0391] In some possible implementations, the intent reasoning module is specifically used to: determine the degree of concentration of the same behavior type, the trend of behavior type change, and the dwell time for the same behavior object in the user behavior sequence based on the behavior type and behavior sequence information of the user behavior sequence; calculate the intent activation intensity based on the degree of concentration of behavior sequence, the trend of behavior type change, and / or the dwell time; and extract the behavior objects that appear more frequently than a preset threshold in the user behavior sequence and / or the behavior objects that co-occur more frequently than a preset threshold in the user behavior sequence as core entities, and obtain the predefined tags associated with the core entities.

[0392] In some possible implementations, the heterogeneous graph also includes intent nodes; the intent reasoning module is specifically used to: generate node creation instructions based on the intent evolution stage, which are used to create intent nodes in the heterogeneous graph; generate edge addition instructions based on core entities and tags, which are used to establish edges between intent nodes and user nodes, push content nodes, or nodes corresponding to core entities in the heterogeneous graph; and generate weight setting instructions based on intent activation strength, which are used to set the weights of the edges.

[0393] In some possible implementations, the intent reasoning module is specifically used to: generate a node creation instruction in response to the intent evolution stage meeting preset node creation conditions; the node creation instruction is used to create intent nodes in the heterogeneous graph and initialize the initial embedding vector of the intent node based on the core entity and label; generate an edge addition instruction in response to the intent activation intensity being greater than a first threshold; the weight enhancement instruction is used to increase the weight of existing edges associated with the core entity in the heterogeneous graph; and generate a weight decay instruction in response to the intent activation intensity being less than a second threshold; the weight decay instruction is used to accelerate the weight decay rate of existing edges associated with the core entity in the heterogeneous graph.

[0394] In some possible implementations, the intent reasoning module is specifically used to: establish intent nodes and their associated edges in the heterogeneous graph using an initial embedding vector and initial weights, and initialize the timestamp of the intent node to the creation time; within a preset time window starting from the timestamp, in response to the intent node meeting a preset update condition, update the timestamp of the intent node to the current time and increase the weight of the associated edges; within the preset time window starting from the timestamp, in response to the intent node not meeting the update condition, reduce the weight of the associated edges according to a preset decay function; in response to the weight of all associated edges being lower than a preset threshold, delete the intent node and all its associated edges from the heterogeneous graph.

[0395] In some possible implementations, the update conditions include: within a preset time window, the number of times a behavior object matching the core entity and label associated with the intent node appears in the user behavior sequence or the repetition pattern of the behavior type reaches a preset threshold.

[0396] In some possible implementations, the intent reasoning module is specifically used to: analyze user behavior sequences using the thought chain reasoning capabilities of a large model, and generate structured intent analysis results.

[0397] In some possible implementations, the intent reasoning module is specifically used to: construct cue words containing user behavior sequences, the cue words including instructions for the large model to perform multi-step reasoning on the user behavior sequences; input the cue words into the large model, and obtain the structured intent analysis results output by the large model.

[0398] In some possible implementations, the graph update module is specifically used to: identify the nodes in the heterogeneous graph affected by the update in response to the dynamic update of the topology of the heterogeneous graph, and obtain the set of affected nodes; extract the k-hop neighborhood subgraph of the affected node set in the heterogeneous graph, where k is the number of layers of the heterogeneous graph neural network; and perform forward propagation of the heterogeneous graph neural network on the k-hop neighborhood subgraph to update the embedding vectors of the nodes in the affected node set while keeping the model parameters of the heterogeneous graph neural network unchanged.

[0399] In some possible implementations, the recall module is specifically used to: execute at least one of the following recall strategies in parallel and take the union of the recall results: calculate the similarity between the updated embedding vector of the user node and the updated embedding vector of each push content node, select the K push content nodes with the highest similarity, and take the push content corresponding to the push content nodes as the recall results; perform path search in the heterogeneous graph according to the preset meta-path to obtain push content nodes with higher-order associations with the user node, and take the push content corresponding to the push content nodes as the recall results; select the N push content with the highest click-through rate as the recall results based on the global click-through rate statistics of the push content; sort the recall results and determine the final push content to be delivered based on the sorting results; where K and N are positive integers.

[0400] In some possible implementations, the recall module is specifically used to: obtain the updated embedding vector of the user node, the updated embedding vector of the push content node, and context features; the context features include at least one feature related to the user's current context, which includes at least one of the following: current time period, user device type, time interval since the last push, and user activity level; input the obtained features into a pre-trained ranking model to obtain a ranking score; and determine the final push content based on the ranking score.

[0401] In some possible implementations, the recall module is specifically used to: weight the ranking score according to at least one of the following factors: timeliness factor: the timeliness factor decays exponentially with the creation time of the push content; fatigue factor: the fatigue factor is calculated based on the number of pushes received by the user within a preset time window; diversity factor: the diversity factor is calculated based on the semantic difference between the updated embedding vector of the push content node and the updated embedding vector of the push content node recently received by the user; and determine the final push content to be delivered based on the weighted ranking score.

[0402] In some possible implementations, the heterogeneous graph also includes at least one of the following nodes: in-application entity nodes, used to represent items or content within the application; tag nodes, used to represent semantic tags; time context nodes, used to represent time windows or time periods; intent nodes, used to represent the user's real-time intent; and at least one of the following edges: interaction edges between user nodes and push content nodes, used to represent the user's click or exposure behavior on push content; behavior edges between user nodes and in-application entity nodes, used to represent the user's operational behavior within the application; association edges between push content nodes and tag nodes, used to represent the attribution relationship between push content and semantic tags; attribution edges between in-application entity nodes and tag nodes, used to represent the association between in-application entities and semantic tags; activity edges between user nodes and time context nodes, used to represent the user's activity level within a specific time window; association edges between user nodes and intent nodes, used to represent the association between the user and real-time intent; and semantic association edges between push content nodes and in-application entity nodes, used to represent cross-modal associations based on semantic similarity.

[0403] In some possible implementations, the heterogeneous graph includes at least one of the following meta-paths: First meta-path: starting from the user node, reaching the application entity node via the behavior edge between the user node and the application entity node, then reaching the tag node via the affiliation edge between the application entity node and the tag node, and finally reaching the push content node via the association edge between the push content node and the tag node; Second meta-path: starting from the user node, reaching the application entity node via the behavior edge between the user node and the application entity node, and then reaching the push content node via the semantic association edge between the push content node and the application entity node; Third meta-path: starting from the user node, reaching the time context node via the activity edge between the user node and the time context node, and finally reaching the push content node via the scheduling edge between the push content node and the time context node.

[0404] In some possible implementations, the device for pushing content also includes: a training module, used to perform semantic analysis on the text information corresponding to the push content nodes and / or entity nodes within the application using a large model to generate structured semantic features; and to fuse the structured semantic features with the statistical features of the corresponding nodes to generate initial embedding vectors for the push content nodes and / or entity nodes within the application.

[0405] In some possible implementations, the training module is also used to: calculate the semantic similarity between push content nodes and entity nodes within the application; and, in response to the semantic similarity exceeding a preset threshold, establish semantic association edges between push content nodes and entity nodes within the application in the heterogeneous graph.

[0406] In some possible implementations, the weights of at least one type of edge in the heterogeneous graph are dynamically adjusted according to an exponential time decay function, which takes the initial weight of the edge, the interval between the current time and the edge creation time, and the decay rate parameter related to the edge type as independent variables; when the weight of an edge is lower than a preset threshold, the edge is removed from the heterogeneous graph.

[0407] In some possible implementations, the training module is also used to perform full training of the heterogeneous graph neural network on all nodes in the heterogeneous graph according to a preset time period, so as to update the embedding vectors of all nodes in the heterogeneous graph.

[0408] It is understood that the above-described modules of the push content device in the embodiments of this disclosure have the ability to implement... Figure 1 The embodiments shown illustrate the functionality of the corresponding steps in the method for pushing content. This functionality can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-described functionality. These modules can be software and / or hardware, and each module can be implemented individually or multiple modules can be integrated. For a detailed description of the functionality of each module of the above-described content pushing device, please refer to [link to relevant documentation]. Figure 1 The corresponding description of the method for pushing content in the illustrated embodiments will not be repeated here.

[0409] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.

[0410] In the technical solution disclosed herein, the user's authorization or consent is obtained before acquiring or collecting the user's personal information.

[0411] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0412] The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method for pushing content as provided in the embodiments of this disclosure.

[0413] Compared with existing technologies, this electronic device pushes content by pre-constructing a heterogeneous graph and real-time intent reasoning. Because the heterogeneous graph can integrate multi-source heterogeneous data (such as the relationships between different types of nodes), even if a user has very few direct interaction records with the pushed content, that user node can still indirectly obtain rich semantic information through other types of nodes and edges in the graph. Therefore, the content recommendation method provided in this disclosure can generate more robust initial representations for users and pushed content even in scenarios with extremely low click-through rates and sparse positive samples, significantly improving the cold-start recommendation effect.

[0414] The readable storage medium is a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method for pushing content as provided in the embodiments of this disclosure.

[0415] Compared with existing technologies, this readable storage medium uses a pre-constructed heterogeneous graph and real-time intent reasoning for content push. Because the heterogeneous graph can integrate multi-source heterogeneous data (such as the relationships between different types of nodes), even if a user has very few direct interaction records with the pushed content, that user node can still indirectly obtain rich semantic information through other types of nodes and edges in the graph. Therefore, the content recommendation method provided in this disclosure can generate more robust initial representations for users and pushed content even in scenarios with extremely low click-through rates and sparse positive samples, significantly improving the cold-start recommendation effect.

[0416] The computer program product includes a computer program that, when executed by a processor, implements a method for pushing content as provided in embodiments of this disclosure.

[0417] Compared with existing technologies, this computer program product uses pre-built heterogeneous graphs and real-time intent reasoning for content recommendation. Because heterogeneous graphs can integrate multi-source heterogeneous data (such as the relationships between different types of nodes), even if a user has very few direct interaction records with the recommended content, that user node can still indirectly obtain rich semantic information through other types of nodes and edges in the graph. Therefore, the content recommendation method provided in this disclosure can generate more robust initial representations for users and recommended content even in scenarios with extremely low click-through rates and sparse positive samples, significantly improving cold-start recommendation performance.

[0418] Figure 14A schematic block diagram of an example electronic device 1400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0419] like Figure 14 As shown, device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to a computer program stored in read-only memory (ROM) 1402 or a computer program loaded from storage unit 1408 into random access memory (RAM) 1403. The RAM 1403 may also store various programs and data required for the operation of device 1400. The computing unit 1401, ROM 1402, and RAM 1403 are interconnected via bus 1404. Input / output (I / O) interface 1405 is also connected to bus 1404.

[0420] Multiple components in device 1400 are connected to I / O interface 1405, including: input unit 1406, such as a keyboard, mouse, etc.; output unit 1407, such as various types of displays, speakers, etc.; storage unit 1408, such as a disk, optical disk, etc.; and communication unit 1409, such as a network card, modem, wireless transceiver, etc. Communication unit 1409 allows device 1400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0421] The computing unit 1401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1401 performs the various methods and processes described above, such as the method of pushing content. For example, in some embodiments, the method of pushing content may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1408. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1400 via ROM 1402 and / or communication unit 1409. When the computer program is loaded into RAM 1403 and executed by the computing unit 1401, one or more steps of the method of pushing content described above may be performed. Alternatively, in other embodiments, the computing unit 1401 may be configured to perform the method of pushing content by any other suitable means (e.g., by means of firmware).

[0422] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0423] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0424] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0425] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0426] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0427] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0428] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0429] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for pushing content, comprising: Obtain a pre-built heterogeneous graph; The heterogeneous graph includes multiple types of nodes and multiple types of edges, and the nodes include at least user nodes and push content nodes; Obtain a user behavior sequence; in response to the user behavior sequence satisfying a preset triggering condition, perform intent reasoning on the user behavior sequence, and dynamically update the topology of the heterogeneous graph based on the intent reasoning result; A heterogeneous graph neural network is used to update the node embedding of the updated heterogeneous graph, thereby obtaining the updated embedding vectors corresponding to the user node and the push content node. The updated embedding vector is used to recall and sort the push content, and the final push content is determined based on the sorting results.

2. The method according to claim 1, wherein, The intention reasoning of the user behavior sequence includes: Obtain the sequence features of the user behavior sequence, wherein the sequence features include the behavior type, behavior object, and behavior timing information of each user behavior in the user behavior sequence; Based on the sequence features, determine the intent evolution stage, intent activation intensity, and core entities and tags associated with the user intent corresponding to the user behavior sequence; Based on the intent evolution stage, intent activation intensity, and the core entity and label, a graph operation instruction is generated. The graph operation instruction includes at least one of the following: node type to be created, edge to be established, and edge weight to be set.

3. The method according to claim 2, wherein, The step of determining the intent evolution stage, intent activation intensity, and core entities and tags associated with the user intent corresponding to the user behavior sequence based on the sequence features includes: Based on the behavior type and timing information of user behavior in the user behavior sequence, determine the degree of concentration of the timing of the same behavior type in the user behavior sequence, the trend of behavior type change, and the dwell time on the same behavior object; The intensity of intent activation is calculated based on the degree of concentration of the behavior sequence, the trend of the behavior type change, and / or the dwell time. Based on the behavior objects of user behavior in the user behavior sequence, extract the behavior objects that appear more frequently than a preset threshold in the user behavior sequence and / or the behavior objects that co-occur more frequently than a preset threshold in the user behavior sequence as core entities, and obtain the predefined tags associated with the core entities.

4. The method according to claim 2, wherein, The heterogeneous graph also includes intent nodes; the step of generating graph operation instructions based on the intent evolution stage, intent activation strength, and the core entities and labels includes: Based on the intent evolution stage, a node creation instruction is generated, which is used to create an intent node in the heterogeneous graph. Based on the core entity and the tag, an edge addition instruction is generated. The edge addition instruction is used to establish an edge between the intent node and the user node, the push content node, or the node corresponding to the core entity in the heterogeneous graph. Based on the intent activation strength, a weight setting instruction is generated, which is used to set the weight of the edge.

5. The method according to claim 4, wherein, The generation of node creation instructions based on the intent evolution stage includes: In response to the intent evolution stage meeting the preset node creation conditions, a node creation instruction is generated. The node creation instruction is used to create intent nodes in the heterogeneous graph and initialize the initial embedding vector of the intent node based on the core entity and the label. The step of generating the edge-adding instruction based on the core entity and label includes: Generate an edge addition instruction, which is used to establish an edge between the intent node and the user node, the push content node, or the node corresponding to the core entity in the heterogeneous graph; wherein the edge is determined based on the core entity and the tag, and the edge is set with an initial weight; The step of generating a weight setting instruction based on the intent activation strength includes: In response to the intent activation strength being greater than a first threshold, a weight enhancement instruction is generated, which is used to increase the weight of existing edges associated with the core entity in the heterogeneous graph; In response to the intent activation strength being less than a second threshold, a weight decay instruction is generated, which is used to accelerate the weight decay rate of existing edges associated with the core entity in the heterogeneous graph.

6. The method according to claim 5, wherein, The step of dynamically updating the topology of the heterogeneous graph based on the intent reasoning result includes: The intent node and its associated edges are established in the heterogeneous graph using the initial embedding vector and the initial weights, and the timestamp of the intent node is initialized to the creation time. Within a preset time window starting from the timestamp, in response to the intent node meeting the preset update conditions, the timestamp of the intent node is updated to the current time, and the weight of the associated edge is increased. Within a preset time window starting from the timestamp, in response to the intention node not meeting the update condition, the weight of the associated edge is reduced according to a preset decay function; In response to the weight of all the associated edges being lower than a preset threshold, the intention node and all its associated edges are deleted from the heterogeneous graph.

7. The method according to claim 6, wherein, The update conditions include: within the preset time window, the number of times or the repetition pattern of the behavior object that matches the core entity and tag associated with the intent node appears in the user behavior sequence reaches a preset threshold.

8. The method according to claim 1, wherein, The intention reasoning of the user behavior sequence includes: The user behavior sequence is analyzed using the reasoning capabilities of a large model to generate structured intent analysis results.

9. The method according to claim 8, wherein, The analysis of the user behavior sequence using the reasoning capabilities of a large model to generate structured intent analysis results includes: Construct prompt words containing the user behavior sequence, the prompt words including instructions requiring the large model to perform multi-step reasoning on the user behavior sequence; Input the prompt words into the large model and obtain the structured intent analysis results output by the large model.

10. The method according to claim 1, wherein, The step of using a heterogeneous graph neural network to update the node embeddings of the updated heterogeneous graph, obtaining the updated embedding vectors corresponding to the user node and the push content node, includes: In response to the dynamic update of the topology of the heterogeneous graph, the nodes in the heterogeneous graph affected by the update are identified to obtain the set of affected nodes; Extract the k-hop neighborhood subgraph of the affected node set in the heterogeneous graph, where k is the number of layers in the heterogeneous graph neural network; The forward propagation of the heterogeneous graph neural network is performed on the k-hop neighborhood subgraph to update the embedding vectors of the nodes in the affected node set, while keeping the model parameters of the heterogeneous graph neural network unchanged.

11. The method according to claim 1, wherein, The process of recalling and ranking push content based on the updated embedding vector, and determining the final push content based on the ranking result, includes: Execute at least one of the following recall strategies in parallel and take the union of the recall results: Calculate the similarity between the updated embedding vector of the user node and the updated embedding vector of each push content node, select the K push content nodes with the highest similarity, and use the push content corresponding to the push content nodes as the recall result. According to the preset meta-path, a path search is performed in the heterogeneous graph to obtain the push content node that has a high-order association with the user node, and the push content corresponding to the push content node is used as the recall result. Based on the click-through rate statistics of the push content in the global scope, select the N push content with the highest click-through rate as the recall result; The recall results are sorted, and the final push content is determined based on the sorting results; where K and N are positive integers.

12. The method according to claim 1, wherein, The process of recalling and ranking push content based on the updated embedding vector, and determining the final push content based on the ranking result, includes: The updated embedding vector of the user node, the updated embedding vector of the push content node, and context features are obtained; the context features include at least one feature related to the user's current context, and the user's current context includes at least one of the following: current time period, user device type, time interval since the last push, and user activity level; The acquired features are input into a pre-trained ranking model to obtain ranking scores; The final push notification content is determined based on the ranking score.

13. The method according to claim 12, wherein, The step of determining the final push content based on the ranking score includes: The ranking scores are weighted and adjusted according to at least one of the following factors: Timeliness factor: The timeliness factor decreases exponentially with the creation time of the pushed content; Fatigue factor: The fatigue factor is calculated based on the number of push notifications received by the user within a preset time window; Diversity factor: The diversity factor is calculated based on the semantic difference between the updated embedding vector of the push content node and the updated embedding vector of the push content nodes recently received by the user; The final push notification content is determined based on the weighted ranking score.

14. The method according to claim 1, wherein the heterogeneous graph further comprises at least one of the following nodes: In-application entity nodes are used to represent items or content within the application; Tag nodes are used to represent semantic tags; Time context nodes are used to represent time windows or time periods; Intent nodes are used to represent the user's real-time intent; And at least one of the following edges: The interaction edge between the user node and the push content node is used to represent the user's click or exposure behavior on the push content; Behavioral edges between user nodes and entity nodes within the application are used to represent user actions within the application. The association edges between push content nodes and tag nodes are used to represent the attribution relationship between push content and semantic tags; The belonging edges between entity nodes and tag nodes within the application are used to represent the association between entities and semantic tags within the application; The active edge between the user node and the time context node is used to represent the user's activity level within a specific time window; The association edges between user nodes and intent nodes are used to represent the association between the user and the real-time intent; Semantic association edges between push content nodes and entity nodes within the application are used to represent cross-modal associations based on semantic similarity.

15. The method of claim 14, wherein the heterogeneous graph comprises at least one of the following meta-paths: First-order path: Starting from the user node, it reaches the application entity node via the behavior edge between the user node and the application entity node, then reaches the tag node via the ownership edge between the application entity node and the tag node, and finally reaches the push content node via the association edge between the push content node and the tag node. Secondary path: Starting from the user node, it reaches the entity node in the application through the behavioral edge between the user node and the entity node in the application, and then reaches the push content node through the semantic association edge between the push content node and the entity node in the application. The third path starts from the user node, reaches the time context node via the active edge between the user node and the time context node, and then reaches the push content node via the scheduling edge between the push content node and the time context node.

16. The method of claim 14, further comprising: Using a large model, semantic analysis is performed on the text information corresponding to the push content nodes and / or the entity nodes within the application to generate structured semantic features; The structured semantic features are fused with the statistical features of the corresponding nodes to generate the initial embedding vectors of the push content nodes and / or the entity nodes within the application.

17. The method of claim 16, further comprising: Calculate the semantic similarity between the push content node and the entity node in the application; In response to the semantic similarity exceeding a preset threshold, a semantic association edge is established between the push content node and the entity node within the application in the heterogeneous graph.

18. The method according to claim 1, wherein, The weights of at least one type of edge in the heterogeneous graph are dynamically adjusted according to an exponential time decay function, which uses the initial weight of the edge, the interval between the current time and the edge creation time, and the decay rate parameter related to the edge type as independent variables; when the weight of the edge is lower than a preset threshold, the edge is removed from the heterogeneous graph.

19. The method according to claim 1, further comprising: According to a preset time period, the full training of the heterogeneous graph neural network is performed on all nodes in the heterogeneous graph to update the embedding vectors of all nodes in the heterogeneous graph.

20. A device for pushing content, comprising: The heterogeneous graph creation module is used to obtain pre-built heterogeneous graphs; The heterogeneous graph includes multiple types of nodes and multiple types of edges, and the nodes include at least user nodes and push content nodes; An intent reasoning module is used to acquire user behavior sequences; in response to the user behavior sequence satisfying a preset triggering condition, the module performs intent reasoning on the user behavior sequence and dynamically updates the topology of the heterogeneous graph based on the intent reasoning result. The graph update module is used to update the node embedding of the updated heterogeneous graph using a heterogeneous graph neural network to obtain the updated embedding vectors corresponding to the user node and the push content node. The recall module is used to recall and sort the push content based on the updated embedding vector, and determine the final push content to be sent based on the sorting results.

21. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-19.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-19.

23. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-19.