Recommendation Method and System Based on Adaptive Dynamic Knowledge Graph in Heterogeneous Networks

By building a heterogeneous network in the recommendation system and using graph attention network to extract user short-term preferences and dynamically update the knowledge graph, the data sparse, cold start and deviation problems are solved, and high accuracy and personalized recommendation effects are achieved.

CN115329215BActive Publication Date: 2025-06-20BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211001216.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2025-06-20
Estimated Expiration
2042-08-19

AI Technical Summary

Technical Problem

When facing the problems of data sparseness, cold start and deviation, existing recommendation systems are difficult to accurately provide users with personalized recommendation results, and lack timeliness and adaptability.

Method used

By building a heterogeneous network, combining the graph attention network and RippleNet model, users' short-term preference characteristics are extracted, and the knowledge graph is dynamically updated to achieve real-time recommendations.

Benefits of technology

It improves the accuracy and personalization of the recommendation system, solves the problems of data sparseness and cold start, and realizes the timeliness and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115329215B_ABST
    Figure CN115329215B_ABST
Patent Text Reader

Abstract

The present invention relates to a recommendation method and system based on an adaptive dynamic knowledge graph in a heterogeneous network, belonging to the field of recommendation. According to the complex interaction relationships between users and items, a heterogeneous network is constructed to extract implicit features of users. At the same time, multi-head attention in the graph attention network is used to extract the short-term preferences of users, update the knowledge graph, and then cluster the user and item sets to establish a seed cluster set. The RippleNet model is used to calculate probability prediction values to obtain a recommendation result list, realizing timeliness and adaptability, improving the accuracy of the recommendation system, and better solving problems such as data sparsity, cold start, and bias.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of recommendation, and particularly to a recommendation method and system based on an adaptive dynamic knowledge graph in a heterogeneous network. Background Art

[0002] In recent years, with the continuous development of the Internet and big data, the network resources have grown rapidly. While bringing convenience to people, it has also brought many troubles. How to quickly find resources suitable for users among a large amount of data has become a major problem. In order to quickly and accurately provide the most suitable resources for users, recommendation systems have been applied to various fields, such as news recommendation, POI (point of interest) location recommendation, and learning resource recommendation, etc. Due to the improvement of the status of recommendation systems, in order to give users a better experience, different algorithms have been continuously optimized and improved.

[0003] Traditional recommendation methods mainly include collaborative filtering, content-based recommendation methods, and hybrid recommendation methods. However, these algorithms often face the problems of cold start and data sparsity. At the same time, the system will recommend a large number of similar items that have been clicked, causing user disgust. The existence of bias affects the recommendation effect. Therefore, how to properly alleviate and handle the bias problem is very important. Exposure bias means that the attributes of items cannot be fully exposed to users, and no interaction information does not mean negative preference. Selection bias means that explicit feedback data such as user ratings can only interact with some items and is not a representative sample of all ratings. Traditional recommendation systems only use the historical interaction information between users and items as input, and this approach has the following problems:

[0004] In actual scenarios, the interaction information between users and items is often very sparse. For example, a movie APP may contain tens of thousands of movies, yet the average number of movies rated by a user may be only dozens. Using such a small amount of observed data to predict a large amount of unknown information will greatly increase the overfitting risk of the algorithm and there will also be the influence of selection bias. For newly added users or items, since the recommendation system does not have the historical interaction information of this user or item, it is impossible to accurately model and recommend. That is, traditional recommendation systems have the cold start problem, and at the same time, no interaction does not mean rejection, and there is an impact of exposure bias.

[0005] The problem of data sparsity restricts the performance of recommendation systems. A common approach to solving the data sparsity problem in recommendation systems is to introduce auxiliary information, including social networks, user / item attributes, multimedia information such as images / videos / audio / text, and context, etc. In recent years, with the rise of knowledge graphs, more and more researchers have tried to apply knowledge graphs as auxiliary information to recommendation systems to solve the data sparsity problem in recommendation systems. The graph structure can naturally express the rich relationships between entities in the real world. By analyzing, mining, and performing cognitive reasoning on it, implicit connections between things can be captured. In scenarios where the data is relatively sparse, heterogeneous networks can effectively alleviate the data sparsity problem by establishing rich associations for nodes, enabling more implicit information to be included. The nodes in heterogeneous networks are not limited to entity nodes but also include virtual nodes, which can make more full use of auxiliary information and improve the accuracy of recommendation systems.

[0006] Long-term preferences are the long-term interests and hobbies of users, which can be extracted based on the user's relationship network graph. However, due to network popularity trends or sudden public opinion events, users may have short-term preferences that are different from their long-term preferences. If the short-term preferences of users are taken into account, the items and products that users need in a short period can be grasped more accurately, making the recommendation results more accurate and flexible. The attention mechanism has recently become an important part of deep neural networks, enabling deep neural networks to focus on a subset of their inputs (or features), that is, only paying attention to important and meaningful information. Recently, attention mechanisms have been developed to handle different learning tasks, such as reading comprehension, recommendation systems, etc. Some research has applied the attention mechanism in machine translation tasks, significantly improving the accuracy of translation. Some research has developed an attention-based convolutional neural network for Hashtag suggestions in microblogs. The graph attention network is different from some previous graph neural networks based on the spectral domain. It can perform aggregation operations on neighbor nodes through the attention mechanism, consider the relevance of each neighbor node, and achieve adaptive allocation of weights for different neighbor nodes, with advantages such as high efficiency and portability. Since the intimacy between users and their friends varies, different user nodes should have different weights. In addition, different activities that users have interacted with also have different weights, so different user-activity interaction record nodes should also have different weights.

[0007] For the problems of data sparsity and cold start, the current main solution strategies include adding auxiliary information, such as using knowledge graphs to add auxiliary information, and using deep learning for feature extraction.

[0008] Knowledge graphs aim to describe various entities or concepts existing in the real world, as well as the relationships between them. Technologies such as knowledge extraction, knowledge representation, knowledge fusion, and knowledge reasoning are the key technologies for constructing and applying knowledge graphs. To address the above problems, many solutions have been studied to analyze, process, and reduce the sparsity of data for user and product information from different perspectives.

[0009] Common knowledge graph methods mainly include three types: representation-based methods, path-based methods, and fusion methods. Representation-based methods generally first use knowledge graph representation methods to map entities and relationships in the knowledge graph into low-dimensional vectors, and then directly use them to enrich the information of users or items in the recommendation system. The main models include KSR, MKR, KTGAN, KTUP, SED, RCF, BEM, CKE, DKN, entity2rec, ECFKG, SHINE, and DKFM. Path-based methods consider the entity connections of the knowledge graph during the construction of user-item interactions. Such methods are also known as recommendation methods based on heterogeneous information networks (HIN). Usually, the knowledge graph is regarded as a heterogeneous information network, and then some meta-paths are defined to extract the similarity between target nodes. The different weights between different paths reflect different preferences of users in the knowledge graph. Fusion methods integrate representation-based methods and path-based methods, which can be roughly divided into two categories. The first category redefines user representation through user interaction history, and the typical method is RippleNet. The second category of methods redefines item representation by fusing the entities connected to the item in the knowledge graph, and the representative method is KGCN. However, existing research on dynamic knowledge graphs has poor system timeliness and adaptability, and insufficient attention is paid to users' short-term and potential preferences.

[0010] Deep learning-based methods have shown strong capabilities in extracting item features or user social relationships, and thus have been proven to be promising in optimizing recommendation strategies. In the research on applying deep learning to recommendation algorithms, it is mainly divided into rating prediction problems and Top-N recommendations. Researchers use various deep learning models to model and extract feature information from user-item interaction data, including implicit feedback and explicit feedback, as well as auxiliary information such as attribute information and text information, to predict user-item ratings for recommendation. However, deep learning-based recommendation systems often face some dilemmas. The training process of deep learning methods is a black-box operation, with poor interpretability and modifiability; deep learning has high requirements for hardware, and usually requires a long training time, and the model design is relatively complex. Therefore, how to reduce the computational amount to better extract users' preference features remains a hot topic.

[0011] Existing research has achieved certain results in solving data sparsity and cold start problems by adding auxiliary information from knowledge graphs. However, the exploration of users' implicit and potential preferences has always been a research hotspot in the field of recommendation systems. Flexibly extracting user features and improving the accuracy of the system are the pursued goals.

[0012] (1) Construction of heterogeneous knowledge graph network

[0013] In reality, there is a large amount of networked data composed of different but related objects. According to whether the network has multiple node types or edge types, it can be divided into homogeneous networks and heterogeneous networks. Compared with homogeneous networks, heterogeneous networks contain richer information, which can not only naturally integrate different types of objects and their interactions, but also integrate information from heterogeneous data sources. In heterogeneous networks, multiple types of objects and relationships coexist, containing rich structural and semantic information, providing a new and accurate interpretable way to discover hidden patterns. And a knowledge graph is a schema-free heterogeneous network with rich relationship information. In recent years, using the knowledge graph as auxiliary information in recommendation systems has become a hot research topic. The auxiliary information can enrich the description of users and items, can more deeply explore users' potential preferences, and can make appropriate predictions to solve the big problem of data sparsity and reduce the impact of selection bias at the same time.

[0014] Some research uses the CKE framework combined with the TransR heterogeneous network embedding method to obtain the structural representation of items through the heterogeneity of nodes and edges, and applies the embedding techniques of stacked denoising autoencoders and stacked convolutional autoencoders to obtain the text representation and visual representation of items, enabling CKE to obtain the embedding representation of collaborative filtering in the knowledge base. However, this method does not consider users' short-term preferences and the timeliness of the recommendation system.

[0015] Some research in news recommendation uses TransE to learn entity and relationship vectors from a large number of entities and semantic relationships existing in news titles and texts, and then conducts news recommendations. This research only focuses on exploring potential relationships, but does not pay attention to the importance of time for news and does not achieve dynamic extraction, and the recommendation effect needs to be improved.

[0016] Some research constructs a heterogeneous network model for different category objects in tag data, then performs co-space mapping on different category vertices in the heterogeneous network model; finally, based on the network after co-space mapping, a multi-parameter Markov model is introduced for tag scoring and recommendation.

[0017] Existing research has constructed heterogeneous networks to optimize the recommendation system. Establishing a relationship network can not only complete the missing information to a certain extent, but also improve the accuracy of the recommendation system. The fusion of heterogeneous networks and knowledge graphs can perform entity recognition, relationship extraction, knowledge fusion, and prediction, and can solve the problems of cold start and data sparsity. At the same time, it has a good mining effect on the user's implicit information, which can reduce the influence of selection bias in explicit preferences. Selection bias refers to the fact that user ratings and other behaviors only occur in a small number of project samples, not representative samples of all ratings. However, the constructed heterogeneous network cannot be updated in real time. User preferences may change in the short term due to online public opinion and emergencies. At this time, the heterogeneous network information needs to be changed, so that more accurate and flexible recommendations can be made. Therefore, it is important to pay attention to time information, interactive information caused by emergencies, etc. Extracting users' short-term preferences and adding them to heterogeneous networks can make the user experience better.

[0018] (2) Short-term preference

[0019] People's interests can be divided into long-term interests and short-term interests. Long-term interests are caused by individual tendencies, are relatively stable, and are related to factors such as personal growth background, education, outlook on life, and values. Short-term interests are usually caused by certain conditions and stimuli in the current environment. They are relatively unstable and easy to disappear, but they play an important real-time role in affecting users' current preferences. They have become the most concerned part for businesses and a hot topic for research.

[0020] A study proposed a recommendation model (MKASR) that combines knowledge graph information and short-term preferences. The RippleNet algorithm is used to extract relationship pairs between users and knowledge graph entities. A bidirectional GRU network based on the attention mechanism is used to extract users' short-term preferences from the item sequences that users have recently interacted with, and feature representations of users and items are obtained. Comprehensive recommendations are made to users based on these feature representations and users' short-term preferences.

[0021] A study proposed a self-attention metric learning model AttRec, which uses self-attention to learn the relationship between items in users' recent behaviors and their short-term interest tendencies. It also integrates users' long-term preferences through a metric learning framework.

[0022] Some studies have proposed a recommendation algorithm for the long-term and short-term preferences of network users based on knowledge graphs. By building a knowledge graph, the potential semantic information of network users is deeply mined, and timely semantic assistance and supplementation are completed. The historical behavior of network users is matched with the recommendation results, and the project is finally embedded into the long-term and short-term learning of network users to achieve the recommendation of long-term and short-term preferences of network users.

[0023] A study proposed a time interval-aware dynamic knowledge graph representation method TDG2E. This method cuts the dynamic knowledge graph into different static sub-knowledge graphs according to time nodes, and then uses GRU to process each static sub-knowledge graph to capture the temporal dependency, thereby modeling the structural evolution process of the dynamic knowledge graph.

[0024] Many existing studies have focused on time information and users' long-term and short-term preferences, and have achieved certain results in optimizing the recommendation system. Focusing on time information can achieve real-time recommendation results and make the recommendation system more flexible and accurate. With the continuous development of Internet information, online public opinion and emergencies are increasingly affecting the lives of most people, resulting in short-term preference shifts and sudden interest in some projects. Therefore, it is important to pay attention to short-term preferences, but how to flexibly extract users' short-term preferences and achieve system adaptability is very valuable for research.

[0025] (3) Dynamic Adaptation

[0026] In 2018, Petar et al. published an article proposing a graph attention network for graph structured data. The graph attention network is different from some previous spectral-domain-based graph neural networks. It can aggregate neighbor nodes through the attention mechanism, consider the relevance of each neighbor node, and realize the adaptive allocation of weights for different neighbor nodes. It has the advantages of high efficiency and portability. Using attention to extract user features can improve accuracy and better mine user preference information.

[0027] Some studies have used the deep knowledge perception network DKN to predict click-through rate based on project content, integrated the representation relationship between the news semantic level and the knowledge level through a multi-channel project-entity perception network, and added an attention module to dynamically aggregate project information in historical records, thus constructing a deep knowledge news recommendation system.

[0028] A study proposed a model, KG-IGAT, which makes full use of the information of the central node and the adjacent nodes in the embedding propagation process to model, and then aggregates and propagates the information to a higher level. At the same time, the evolution of user interests is integrated into the attention mechanism of the model to more accurately capture the changes in user interests.

[0029] A new method, called Knowledge Graph Attention Network (KGAT), is proposed to explicitly model high-order connectivity in KG in an end-to-end manner. This method recursively propagates embeddings from the neighbors of a node (which can be users, items, or attributes) to optimize the embedding of the node and uses an attention mechanism to distinguish the importance of neighbors.

[0030] Existing research has incorporated attention into recommendation systems for feature extraction and preference mining, which has improved the accuracy of recommendation systems and made the systems adaptive. However, there is relatively little research on combining graph attention networks with time information to process dynamic graphs and achieve deletion and update of heterogeneous graphs. Graph attention networks do not require the entire graph structure and are only related to adjacent nodes, that is, nodes sharing edges, and the model predicts the importance of different adjacent nodes, enabling more flexible extraction of users' important short-term preferences. Summary of the Invention

[0031] The objective of the present invention is to provide a recommendation method and system based on an adaptive dynamic knowledge graph in a heterogeneous network to solve data sparsity, cold start, and bias problems and improve recommendation accuracy.

[0032] To achieve the above objective, the present invention provides the following solution:

[0033] A recommendation method based on an adaptive dynamic knowledge graph in a heterogeneous network, comprising:

[0034] Construct a heterogeneous network according to a dataset of complex interaction relationships between users and items;

[0035] Extract entities and relationships from the heterogeneous network to establish a basic knowledge graph;

[0036] Extract the short-term preference features of users using a graph attention network within a time bin, and calculate multivariate attention coefficients based on the short-term preference features;

[0037] Delete the relationships where the multivariate attention coefficients within the coefficient threshold range are located in the basic knowledge graph to obtain a real-time knowledge graph;

[0038] Cluster users and items in the real-time knowledge graph to obtain multiple user clusters and multiple item clusters;

[0039] Screen item clusters according to the real-time knowledge graph and form a seed set for each user cluster;

[0040] According to the seed set of each user cluster, use the RippleNet model to predict the probability value of each user cluster clicking on each item cluster in the seed set;

[0041] Take the item cluster corresponding to the maximum probability value as the recommendation result for each user cluster to generate a recommendation result list;

[0042] Change the time bin, and return to the step of "extracting the short-term preference features of users using a graph attention network within a time bin and calculating multivariate attention coefficients based on the short-term preference features" to obtain real-time recommendation results.

[0043] Optionally, the construction process of the dataset of the complex interaction relationships between the users and the projects includes:

[0044] Collect the user set and the project set respectively;

[0045] Collect the relationship sets of user-user, user-project, and project-project;

[0046] Use the formula to calculate the weight of each relationship in the relationship set; in the formula, is the weight function of relationship r i , γ is the normalization coefficient, is the establishment time length of relationship r i , is the interaction frequency of relationship r i , is the number of common relationship nodes of two nodes of relationship r i , i ∈ [1, N], and N is the total number of relationships;

[0047] Construct a dataset of the complex interaction relationships between the users and the projects with the user set, the project set, the relationship set, and the weight of each relationship.

[0048] Optionally, the time bin is

[0049] TI a = [ti a , ti a+1

[0050] In the formula, TI a is the time bin, a is a constant, ti a , ti a+1 represent the start time and the end time respectively.

[0051] Optionally, extracting the short-term preference features of the users using the graph attention network within the time bin and calculating the multi-attention coefficients according to the short-term preference features specifically include:

[0052] Use the formula to calculate the latent features of the users; in the formula, represents the latent features of the users, σ represents the non-linear activation function, W represents the neural network weights, AF u-u represents the aggregation function that fuses the explicit friends and implicit friends of the users, represents the interaction of the users with other users under the time bin TI a , Ex u represents the explicit friend feature representation, Im u represents the implicit friend feature representation, and b represents the neural network bias;

[0053] According to the latent features of the user, the formula is used to calculate the attention coefficient of neighboring users; in the formula, represents the attention coefficient of neighboring users, Softmax() represents the normalization function, and W' represents the weight matrix, represents the transpose of the parameters of the attention network, respectively represent the k-th power of the first and second bias terms of the attention network;

[0054] The formula is used to calculate the latent features of user interaction items; in the formula, represents the latent features of user interaction items, and AF u-v represents an aggregation function that fuses the explicit interesting items that the user has interacted with historically and the implicit items that the user has interacted with indirectly through the meta-path, represents the interaction of the user with other items at time TI a , Ex v represents the explicit interesting items that the user has interacted with historically, and Im v represents the implicit items that the user has interacted with indirectly through the meta-path;

[0055] According to the latent features of user interaction items, the formula is used to calculate the attention coefficient of neighboring items; in the formula, represents the attention coefficient of neighboring items;

[0056] The formula is used to calculate the latent features of item-to-item; in the formula, represents the latent features of item-to-item, and AF v-v represents an aggregation function that fuses the directly relevant information and indirectly relevant information with the target item, represents the interaction embedding of the target item with other items in the time bin TI a , Di v represents the items with directly relevant information to the target item, and In v represents the items with indirectly relevant information to the target item;

[0057] According to the latent features of item-to-item, the formula is used to calculate the attention coefficient of interaction items; in the formula, represents the attention coefficient of interaction items;

[0058] The formula is used to calculate the virtual relationship item features independent of user preferences; in the formula, represents the virtual relationship item features independent of user preferences, and Fu…v An aggregation function representing items that are neither directly nor indirectly related to user preferences Represents at time TI a The embedding of random items unrelated to the user, Vi v Represents the item feature representation unrelated to user preferences, which contains 5 randomly selected item features and establishes a virtual relationship

[0059] According to the virtual relationship item features unrelated to user preferences, use the formula Calculate the attention coefficient of the virtual relationship items; in the formula, Represents the attention coefficient of the virtual relationship items

[0060] Optionally, clustering users and items in the real-time knowledge graph to obtain multiple user clusters and multiple item clusters, specifically including:

[0061] Classify users and items according to the real-time knowledge graph into multiple clusters, with the number of nodes in each cluster being 1 - 5, and obtain user clusters as Item clusters as

[0062] Among them, Represents the u r th user cluster, Containing A users; Represents the v r th item cluster, Containing B items, A, B ∈ [1, 5], and r is an arbitrary integer

[0063] Optionally, screening item clusters according to the real-time knowledge graph and forming the seed set of each user cluster, specifically including:

[0064] Determine the interaction matrix between user clusters and item clusters as In the formula, Represents the element of the interaction matrix, The value of is 0, 1, and -1; when Indicates that there is a direct interaction between the user cluster and the item cluster or an indirect interaction along the meta-path of the graph data; when Indicates that there is no interaction information between the user cluster and the item cluster; when Indicates that the relationship between the user cluster and the target user cluster is an aversion relationship or the related item is an item that is not liked; Y represents the interaction matrix, C u Represents the u-th user cluster, C v Represents the v-th item cluster, U represents the set of user clusters, and V represents the set of item clusters;

[0065] Take And The corresponding item clusters form the seed set of each user cluster.

[0066] A recommendation system based on an adaptive dynamic knowledge graph in a heterogeneous network, comprising:

[0067] A heterogeneous network construction module, configured to construct a heterogeneous network according to a dataset of complex interaction relationships between users and items;

[0068] A knowledge graph establishment module, configured to perform entity and relationship extraction on the heterogeneous network to establish a basic knowledge graph;

[0069] An attention coefficient calculation module, configured to use a graph attention network to extract short-term preference features of a user within a time bin, and calculate a multivariate attention coefficient according to the short-term preference features;

[0070] A knowledge graph update module, configured to delete the relationships where the multivariate attention coefficients within the coefficient threshold range are located in the basic knowledge graph to obtain a real-time knowledge graph;

[0071] A clustering module, configured to cluster users and items in the real-time knowledge graph to obtain a plurality of user clusters and a plurality of item clusters;

[0072] A screening module, configured to screen item clusters according to the real-time knowledge graph and form the seed set of each user cluster;

[0073] A prediction module, configured to use the RippleNet model to predict the probability value of each user cluster clicking on each item cluster in the seed set according to the seed set of each user cluster;

[0074] A recommended result generation module, configured to use the item cluster corresponding to the maximum probability value as the recommended result of each user cluster to generate a recommended result list;

[0075] A loop module, configured to change the time bin and call the attention coefficient calculation module to obtain real-time recommended results.

[0076] Optionally, the attention coefficient calculation module specifically includes:

[0077] A first latent feature calculation sub-module, configured to use the formula to calculate the latent feature of the user; in the formula, represents the latent feature of the user, σ represents a non-linear activation function, W represents a neural network weight, AF u-u represents an aggregation function that fuses explicit and implicit friends of the user, represents the interaction of the user with other users at time bin TI a and Ex u represents the explicit friend feature representation, Imu a represents the implicit friend feature representation, and b represents the neural network bias;

[0078] The first attention coefficient calculation sub-module is used to calculate the attention coefficient of neighboring users according to the potential features of the user by using the formula ; in the formula, represents the attention coefficient of neighboring users, Softmax() represents the normalization function, W' represents the weight matrix, represents the transpose of the parameters of the attention network, respectively represent the k-th power of the first and second bias terms of the attention network;

[0079] The second potential feature calculation sub-module is used to calculate the potential features of user interaction items by using the formula ; in the formula, represents the potential features of user interaction items, AF u-v represents the aggregation function that fuses the explicit interested items that the user has interacted with historically and the implicit items that the user has interacted with indirectly through the meta-path, represents the interaction of the user with other items at time TI a , Ex v represents the explicit interested items that the user has interacted with historically, im v represents the implicit items that the user has interacted with indirectly through the meta-path;

[0080] The second attention coefficient calculation sub-module is used to calculate the attention coefficient of neighboring items according to the potential features of user interaction items by using the formula ; in the formula, represents the attention coefficient of neighboring items;

[0081] The third potential feature calculation sub-module is used to calculate the potential features of item-to-item by using the formula ; in the formula, represents the potential features of item-to-item, AF v-v represents the aggregation function that fuses the directly relevant information and indirectly relevant information with the target item, represents the interaction embedding of the target item with other items in the time bin TI a , Di v represents the items with directly relevant information to the target item, In v represents the items with indirectly relevant information to the target item;

[0082] The third attention coefficient calculation sub-module is used to calculate the attention coefficient of item-to-item according to the potential features of item-to-item by using the formula Calculate the attention coefficient of the interaction item; where represents the attention coefficient of the interaction item;

[0083] The fourth latent feature calculation sub-module is used to use the formula to calculate the virtual relationship item features independent of user preferences; where represents the virtual relationship item features independent of user preferences, F u…v represents the aggregation function of items that are not directly or indirectly related to user preferences, represents at time TI a the embedding of random items unrelated to the user, Vi v represents the item feature representation independent of user preferences, and the elements included are 5 random item features to establish a virtual relationship;

[0084] The fourth attention coefficient calculation sub-module is used to calculate the attention coefficient of the virtual relationship item according to the virtual relationship item features independent of user preferences by using the formula ; where represents the attention coefficient of the virtual relationship item.

[0085] Optionally, the clustering module specifically includes:

[0086] The classification sub-module is used to classify users and items according to the real-time knowledge graph into multiple clusters, and the number of nodes in each cluster is 1-5, and the user clusters obtained are The item clusters are

[0087] where represents the u r th user cluster, contains A users; represents the v r th item cluster, contains B items, A, B ∈ [1, 5], and r is an arbitrary integer.

[0088] Optionally, the screening module specifically includes:

[0089] The interaction matrix determination sub-module is used to determine the interaction matrix of the user cluster and the item cluster as where represents the interaction matrix element, takes values of 0, 1, and -1; when it means that there is a direct interaction between the user cluster and the item cluster or an indirect interaction along the meta-path of the graph data; when it means that there is no interaction information between the user cluster and the item cluster; when When it is, it means that the relationship between the user cluster and the target user cluster is an aversion relationship or the related item is a disliked item; Y represents the interaction matrix, and C u represents the u-th user cluster, and C v represents the v-th item cluster, U represents the set of user clusters, and V represents the set of item clusters;

[0090] The seed set composition sub-module is used to and The corresponding item clusters form the seed set of each user cluster.

[0091] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0092] The present invention discloses a recommendation method and system based on an adaptive dynamic knowledge graph in a heterogeneous network. According to the complex interaction relationship between users and items, a heterogeneous network is constructed, implicit user features are extracted, and at the same time, multi-head attention in the graph attention network is used to extract the short-term preferences of users, update the knowledge graph, and then cluster the user and item sets, establish a seed cluster set, calculate the probability prediction value using the RippleNet model, obtain the recommendation result list, realize timeliness and adaptability, improve the accuracy of the recommendation system, and better solve the problems of data sparsity, cold start, and bias. BRIEF DESCRIPTION OF THE DRAWINGS

[0093] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0094] Figure 1 It is a flowchart of the recommendation method based on an adaptive dynamic knowledge graph in a heterogeneous network provided by an embodiment of the present invention;

[0095] Figure 2 It is an overall framework diagram of the recommendation method based on an adaptive dynamic knowledge graph in a heterogeneous network provided by an embodiment of the present invention;

[0096] Figure 3 It is a schematic diagram of the heterogeneous network provided by an embodiment of the present invention;

[0097] Figure 4 It is a schematic diagram of the graph attention network structure provided by an embodiment of the present invention;

[0098] Figure 5 It is a framework diagram of the RippleNet provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0099] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0100] The object of the present invention is to provide a recommendation method and system based on an adaptive dynamic knowledge graph in a heterogeneous network to solve the problems of data sparsity, cold start, and bias, and improve the recommendation accuracy.

[0101] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0102] The embodiment of the present invention provides a recommendation method based on an adaptive dynamic knowledge graph in a heterogeneous network, which is roughly divided into three steps: (1) constructing a heterogeneous network and establishing a basic knowledge graph; (2) using GAT to extract the short-term preferences of users in a time bin, calculating the multi-dimensional attention coefficient, setting a threshold, selecting or rejecting the neighborhood, and obtaining a real-time knowledge graph network; (3) clustering the user and item sets, constructing a RippleNet model, establishing a seed cluster set, calculating the probability prediction value, and obtaining a recommendation result list.

[0103] The overall framework is as Figure 2 shown. (a) Construct a heterogeneous network from the data set, then extract entity nodes, construct a basic knowledge graph, perform short-term feature extraction on the nodes in a time bin, and at the same time randomly establish virtual relationships for irrelevant nodes to calculate weights, and update the relationship network of the basic knowledge graph to obtain a real-time knowledge graph. (b) Cluster the entity nodes, then process them using the RippleNet structure, use the clustered nodes as seeds, do not calculate the probability for items that the user clearly indicates dislike, focus on calculating those with interactions, and at the same time use irrelevant clusters as random seeds to calculate the probability prediction value to obtain the final recommendation result list.

[0104] See Figure 1 , and the implementation process of the recommendation method of the present invention will be described in detail below:

[0105] Step S1: Construct a heterogeneous network according to the data set of the complex interaction relationship between users and items.

[0106] First, construct the data set, and the construction process includes:

[0107] Step 1: Collect the user set U = {u1, u2,..., u n}, where U represents the set of all users, including n individual users, and the collection project set V = {v1, v2, …, v m}, where V represents the set of all projects, including m project quantities, and this project includes some virtual topics, fields, etc.

[0108] Step 2: Collect the complex relationship set R = {r1, r2, …, r N} of user-user, user-project, and project-project, where R represents the set of all types of relationships, including N different relationships.

[0109] Step 3: Calculate the relationship weight function. Use We to represent the implicit relationship between users. Among them, the relationship r i The weight function is represented by to represent, - The establishment time length of relationship r i , - The interaction frequency of relationship r i , - The number of common relationship nodes of two nodes of relationship r i . The specific relationship weight calculation is as follows, and normalization processing is performed.

[0110]

[0111] Among them, γ is the normalization coefficient, and the specific size can be determined according to expert experience to reduce the excessive gap in the weight function and bias the recommendation results.

[0112] Then, construct the heterogeneous network as G HIN = {U, V, R, W}. It includes users, projects, relationships, and weights, making the information utilization more comprehensive and mining users' implicit preferences.

[0113] Due to the diverse node types, the included relationship types are also more diverse and rich. The constructed heterogeneous network is as Figure 3 shown. According to the category of the relationship, the establishment time length of this relationship, the communication frequency between users under this relationship, etc., calculate the relationship weight function of the users. For cold-start and data-sparse users, this kind of interaction relationship can be used to make recommendations according to the nodes with rich information, better solving the long-standing problems.

[0114] Figure 3The arrows in it indicate the initiative of the interaction relationship, including multiple relationship types such as the user actively clicking on an item, the user actively interacting with another user, the item information containing the user, and the item containing another item. The heterogeneous graph has no restrictions and contains richer information. The nodes can also be extended to fictional entities such as environments, domains, and topics, rather than being limited to the entity nodes of users and items. The information in the database is reflected in the topological structure and extended at the same time. According to the meta-path, potential preference relationships can be better mined. Among them, Comment represents comment, Browse represents browse, Friend represents friend, Client represents client, work fanatic represents workaholic, Anxietytendencies represents anxiety tendency, attention represents attention, like represents like, related represents related, interested represents interested, and financial sector represents financial department.

[0115] Step S2: Extract entities and relationships from the heterogeneous network to establish a basic knowledge graph.

[0116] Extract entities and relationships from the heterogeneous network to establish a knowledge graph G, forming a triple relationship group (h, r, t), which are the head, relationship, and tail respectively.

[0117] Step S3: Use the graph attention network to extract the short-term preference features of users in the time bin and calculate the multi-attention coefficients according to the short-term preference features. The multi-attention coefficients refer to multiple types of attention coefficients.

[0118] Under normal circumstances, the time bin is created fixedly to extract short-term features. When there is online public opinion and emergencies, the triggering of the time bin becomes frequent, and the feature extraction also becomes frequent. The time bin is represented as

[0119] TI a =[ti a ,ti a+1

[0120] where TI a represents a time bin, and its time is represented by ti a to ti a+1 , and a is an arbitrary constant.

[0121] For example Figure 4 ​The figure shows a schematic diagram of the structure of a Graph Attention Network (GAT), which is divided into four parts. The attention coefficients are calculated for each pair, and after normalization, they are selected or discarded through a threshold. Paying attention to the interaction information in the above four aspects can not only provide recommendations similar to the user's own preferences for the user, but also extend the preferences of the target user through neighboring items and neighboring users to obtain more personalized and comprehensive recommendation results. Figure 4 In it, threshold represents the threshold.

[0122] Use GAT to extract short-term features of users and calculate the attention weights in four cases.

[0123] 1) User-User GAT

[0124] Calculate the attention coefficient of the user, focusing on the factors that have a greater impact on preferences. For the interaction between users, implicit friends can be mined according to the path relationship of the graph data, and more items can be recommended to the target user. The potential features of the user are represented as:

[0125]

[0126] Among them, σ represents the non-linear activation function, AF u-u is the aggregation function for fusing explicit and implicit friends of the user, b represents the neural network bias, W represents the neural network weight, which can be obtained through iterative training, Ex u represents the explicit friend feature representation, Im u represents the implicit friend feature representation. represents the interaction of the user with other users at time TI a Calculate the attention coefficient of neighboring users according to the potential features and perform normalization calculations.

[0127]

[0128] Softmax() is the normalization function, and W' is the weight matrix, which is obtained through deep learning network training. are the transpose of the parameters of the attention network and the k-th power of the bias term respectively. Finally, the low-dimensional vector feature representation of the user is obtained. To improve the accuracy of its calculation, the multi-head attention mechanism of the GAT model is used for calculation. The specific process is as follows.

[0129]

[0130] Among them, K is the number of heads of multi-head attention, that is, the number of calculations, which can take any positive integer and is set according to application needs. is the low-dimensional vector feature representation of users in the time bin, W represents the neural network weights, which can be iteratively trained, and b represents the neural network bias.

[0131] 2) User-Item GAT

[0132] Focus on factors such as the quality, quality, and price of the item, and calculate the weights of the item. The items that the user has interacted with can be divided into two categories. One is direct interaction by the user, such as evaluating or purchasing the item, and the other is indirect interaction between the user and the item according to the meta-path in the heterogeneous graph.

[0133]

[0134] where σ represents the non-linear activation function, AF u-u is the aggregation function that fuses the explicitly interested items that the user has interacted with in the past and the implicitly interested items that the user has interacted with indirectly through the meta-path. b represents the neural network bias, W represents the neural network weights, which can be iteratively trained, Ex v represents the explicitly interested items that the user has interacted with in the past, Im v represents the implicitly interested items that the user has interacted with indirectly through the meta-path. represents the interaction of the user with other items at time TI a Calculate the attention coefficients of the neighboring items according to the latent features and perform normalization calculations.

[0135]

[0136] Softmax() is the normalization function, W' is the weight matrix, which is obtained through training with a deep learning network, are the transpose of the parameters of the attention network and the k-th power of the bias term respectively. Output the feature representation of the user interaction items:

[0137]

[0138] where K is the number of heads of the multi-head attention, that is, the number of calculations, which can take any positive integer and is set according to application needs, is the low-dimensional vector feature representation of the item in the time bin, W represents the neural network weights, which can be iteratively trained, and b represents the neural network bias.

[0139] 3) Item-Item GAT

[0140] For the interaction information between projects, focus on the degree of association between historical interaction projects and neighborhood projects, give recommendations with higher cost performance of the same type, and calculate the attention coefficient between projects. There is direct relevant information between projects, which belong to similar types of projects. At the same time, there is also a connection between projects through users or other projects, with indirect project information.

[0141]

[0142] where σ represents the non-linear activation function, AF v-v is an aggregation function that fuses direct and indirect relevant information with the target project, b represents the neural network bias, W represents the neural network weight, which can be obtained through iterative training, Di v represents the project with direct relevant information to the target project, In v represents the project with indirect relevant information to the target project. represents the interaction embedding of the target project with other projects at time TI a Calculate the attention coefficient of the project according to the latent features and perform normalization calculations.

[0143]

[0144] Softmax is the normalization function, W' is the weight matrix, which is obtained through deep learning network training, are respectively the transpose of the parameters of the attention network and the k-th power of the bias term. Output the feature representation of relevant projects:

[0145]

[0146] where K is the number of heads of multi-head attention, that is, the number of calculations, which can take any positive integer and is set according to application needs, is the low-dimensional vector feature representation of the project in the time bin, W represents the neural network weight, which can be obtained through iterative training, and b represents the neural network bias.

[0147] 4) Predict GAT

[0148] Since traditional recommendations mostly recommend projects based on users' historical behaviors, it will cause predictions of users' fixed preferences, and the recommended projects are too single. At the same time, since not all items can be exposed to users, and no interaction information does not mean negative preference. Providing more personalized recommendations for users can be eye-catching. Users may discover their potential preferences and at the same time broaden their horizons. Randomly recommend project types that have nothing to do with the user to reduce the impact of exposure bias, calculate the attention coefficient of the project, establish a prediction relationship, and give the project with a larger prediction relationship to the user to improve the flexibility of the recommendation result.

[0149]

[0150] where σ represents a non - linear activation function, AF u…v is an aggregation function for items that have no direct or indirect relation with user preferences, b represents the neural network bias, W represents the neural network weights, which can be obtained through iterative training, and Vi v represents the item feature representation that has nothing to do with user preferences, and the elements it contains are randomly selected 5 item features to establish a virtual relationship. represents at time TI a the embedding of a random item unrelated to the user. Calculate the attention coefficient of the item according to the latent features and perform normalization calculation.

[0151]

[0152] Softmax() is a normalization function, W' is a weight matrix, which is obtained through training with a deep learning network. are respectively the transpose of the parameters of the attention network and the k - th power of the bias term. Output the feature representation of the virtual relationship items:

[0153]

[0154] where K is the number of heads of the multi - head attention, that is, the number of calculation times, which can take any positive integer and is set according to application requirements. is the low - dimensional vector feature representation in the time bin, W represents the neural network weights, which can be obtained through iterative training, and b represents the neural network bias.

[0155] Step S4, in the basic knowledge graph, delete the relationships where the multi - attention coefficients belong to the coefficient threshold range to obtain a real - time knowledge graph.

[0156] Set a threshold [0, q], where q can be determined according to expert experience or the actual system. If then delete the relationships where these coefficients are located, update the knowledge graph G to achieve its real - time performance, and obtain the knowledge graph representation at time TI a

[0157] Step S5, cluster the users and items in the real - time knowledge graph to obtain multiple user clusters and multiple item clusters.

[0158] Classify the users and items according to the real - time knowledge graph into multiple clusters, with the number of nodes in each cluster being 1 - 5, and generate a small - scale recommendation list.

[0159] The user cluster is represented as The item cluster is represented as ​It contains A users and B project clusters, where A, B ∈ [1, 5], and j is an arbitrary integer.

[0160] Step S6: Screen project clusters according to the real-time knowledge graph and form the seed set of each user cluster.

[0161] The interaction matrix between user clusters and project clusters is expressed as:

[0162]

[0163] Among them, when there is a direct interaction between a user cluster and a project cluster or an indirect interaction along the meta-path of the graph data, it is set as When there is no interaction information between the user cluster and the project cluster When the following situations are included. The relationship between the user cluster and the target user cluster is an aversion relationship, or the relevant project is a disliked project. In this case, the project will never be recommended to the user cluster as a seed set. For the project clusters, random extraction will be performed to calculate the relevant probability, reduce the impact of exposure bias on the recommendation effect, and achieve a comprehensive recommendation result. Focus on calculating the prediction probability.

[0164] Step S7: According to the seed set of each user cluster, use the RippleNet model to predict the probability value of each user cluster clicking on each project cluster in the seed set.

[0165] The framework diagram of RippleNet is as Figure 5 shown, which shows the change of the image range of the seed set in the knowledge graph as the number of hops increases, as well as the interaction between the user cluster and the project cluster. Finally, the corresponding prediction probability is calculated to obtain the recommendation result list. Figure 5 In it, ripple set represents the ripple set and Hop represents the number of hops. The seed set is the input point in the RippleNet model.

[0166] The process of using the RippleNet model to predict the probability value is as follows:

[0167] Step 1: Entity set

[0168]

[0169] Its formula represents the relevant entity set for the k-th jump of the user cluster C u , where represents the interaction project and the randomly selected irrelevant project clusters, which are used as the seed set of the user cluster on the knowledge graph, and k is the number of jumps outward from the starting point.

[0170] Step 2: Ripple Set

[0171]

[0172] The potential interest of user clusters in item clusters in the ripple set is like ripples in water, continuously expanding outwards, that is, as the number k increases, the preference also gradually weakens.

[0173] Step 3: By comparing the characteristics C of the item cluster v and the head node h of the triple (h i , r i , t i ), the association probability of each triple in the ripple set i i i i can be obtained, and the formula is as follows.

[0174]

[0175] where R i and h i are the characteristics of the relationship r i and the head node h i respectively. Then calculate the weighted sum of the tail nodes in .

[0176] Step 4: Obtain the vector

[0177]

[0178] where t i is the characteristic of the tail node t i , and the vector represents the first-order response of the user cluster C u to the item cluster in the seed set of the knowledge graph.

[0179] Step 5: Correspondingly expand, calculate the multi-order response, and sum them up to obtain C u as the response of all orders fused.

[0180]

[0181] Step 6: Combine the user cluster and the item cluster, and output the predicted click probability. The calculation formula is as follows.

[0182]

[0183] Step S8, take the item cluster corresponding to the maximum probability value as the recommendation result for each user cluster, and generate a recommendation result list.

[0184] Take the Top-N item set with the highest probability and give a list of recommended results. This list is the recommendation result of a user cluster.

[0185] Step S9, change the time bin, return to step S3, and obtain real-time recommendation results.

[0186] By changing the time bin and repeating steps S3-S8, real-time recommendation results are given, which not only ensures the real-time and accuracy of the recommendation system, but also improves the diversity of recommendation results.

[0187] The present invention proposes a recommender system based on adaptive dynamic knowledge graph in heterogeneous network (ADKHN, Recommender System Based on Adaptive Dynamic Knowledge Graph in Heterogeneous Network). By constructing a heterogeneous network graph for the complex relationships in the recommender system database, the problems of data sparsity and cold start are preliminarily solved through the prediction of the knowledge graph and information completion. In practice, there are different complex relationships between users, users and projects, and projects and projects, and these relationships can reflect the potential preferences of users and extract the implicit preferences of users. For example, when browsing a page, if the user only reads the beginning, it may be a missed click. At this time, if a large number of related projects are recommended, it will cause user disgust. In traditional recommendation systems, for cold start users, most of the processing methods are to recommend some projects with the highest click rate to users, without personalized recommendations, which are too popular. Paying attention to the relationship between users can well solve the problems of data sparsity and cold start, and at the same time can provide users with more personalized recommendation results and improve the user experience.

[0188] Referring to the basic idea of ​​collaborative filtering algorithm, people are divided into groups and things are gathered into categories, and the system establishes a heterogeneous graph of relationship networks. When data is sparse and cold-start users join, similar users are found according to the relationship paths of the heterogeneous network, and then the project sets that meet the preferences are found to mine the implicit preferences of users. The construction of complex heterogeneous network-knowledge graph can not only complete the relationship, but also make appropriate predictions. Construct a domain knowledge graph, extract knowledge semantic information from the data set, and pre-process the knowledge semantic information. There is a lot of noise information such as wrong information and blank information in the information in the database. Process this information in an appropriate way. For the extraction of relationships between entities, entities can be extracted on the basis of predefined relationships to find matching relationships. For example, there is a writing relationship between movies and screenwriters, a performance relationship between movies and actors, and a relationship between movies and styles. The entities and relationships are stored in the relational database as the table structure of the relational database.

[0189] The key point of this method is to consider the influence of time and unexpected events on user preferences, use multi-head attention in the graph attention network to extract the short-term preferences of users, and embed them into the heterogeneous network according to their weights to achieve the timeliness and adaptability of the recommendation system. The main key points are as follows:

[0190] (1) The heterogeneous network incorporates the differences of different types of nodes into node representations, reducing information loss and enabling more accurate mining of users' implicit preference information. Considering the complex relationships among products, buyers, and sellers can improve the accuracy of recommending suitable products for target users and enhance the recommendation quality of the recommendation system. The nodes in the heterogeneous network have a richer variety of types, not limited to entities, and also have fewer restrictions such as quantity restrictions and existence restrictions, enabling better mining of the implicit information of relationships and reducing the influence of selection bias and exposure bias.

[0191] (2) By choosing to focus on users' short-term preferences, taking into account time information and the impact of online public opinion on users' short-term behavior, the understanding of users' preferences can be more accurate. Short-term preferences are characterized by suddenness, deviation, and timeliness. Users usually exhibit short-term preference behavior characteristics different from their long-term preferences due to online public opinion. For example, the influence of online popularity on users' subconscious minds can lead to a herd mentality. However, the duration of short-term preferences may be short, and they will return to their original preferences after the online public opinion storm has passed. Therefore, accurately grasping users' short-term preferences can provide recommendation results for users at different times, greatly enhancing the adaptability of the recommendation system and improving users' experience.

[0192] (3) Different from some previous spectral-domain graph neural networks, the graph attention network can perform an aggregation operation on neighbor nodes through an attention mechanism, consider the relevance of each neighbor node, and achieve an adaptive allocation of weights for different neighbor nodes, with advantages such as high efficiency and portability. The graph attention network is a component of deep learning. Applying the component to the system can not only reduce the computational cost but also better optimize the algorithm. Multi-head attention can more deeply explore the potential of node data by calculating attention multiple times, enabling the model to better understand the characteristic meaning contained in the nodes and extract the short-term preference characteristics of users more precisely.

[0193] (4) When using RippleNet to process graph data, clustering is performed, and at the same time, unrelated clusters are randomly selected as seeds to reduce exposure bias and increase the personalization of the recommendation system. Items that are clearly indicated as disliked are suppressed and not recommended, improving the user experience.

[0194] The above four points have a certain effect on the performance change of the recommendation system. Pay attention to the complex interaction relationships of different types of nodes, construct a heterogeneous network graph, mine the implicit preferences of users, and increase the interpretability of the system; add time information to the recommendation system and pay attention to the short-term preference behavior characteristics of users; use multi-head attention in the graph attention network to extract the short-term preferences of users and improve the accuracy of the recommendation system.

[0195] An embodiment of the present invention further provides a recommendation system based on an adaptive dynamic knowledge graph in a heterogeneous network, including:

[0196] A heterogeneous network construction module, configured to construct a heterogeneous network according to a data set of complex interaction relationships between users and items;

[0197] A knowledge graph establishment module, configured to perform entity and relationship extraction on the heterogeneous network to establish a basic knowledge graph;

[0198] An attention coefficient calculation module, configured to extract the short-term preference features of users using a graph attention network within a time bin, and calculate a multivariate attention coefficient according to the short-term preference features;

[0199] A knowledge graph update module, configured to delete the relationships where the multivariate attention coefficients within the coefficient threshold range are located in the basic knowledge graph to obtain a real-time knowledge graph;

[0200] A clustering module, configured to cluster users and items in the real-time knowledge graph to obtain multiple user clusters and multiple item clusters;

[0201] A screening module, configured to screen item clusters according to the real-time knowledge graph and form a seed set for each user cluster;

[0202] A prediction module, configured to, according to the seed set of each user cluster, use the RippleNet model to predict the probability value of each user cluster clicking on each item cluster in the seed set;

[0203] A recommended result generation module, configured to use the item cluster corresponding to the maximum probability value as the recommended result for each user cluster to generate a recommended result list;

[0204] A loop module, configured to change the time bin, call the attention coefficient calculation module, and obtain real-time recommended results.

[0205] The attention coefficient calculation module specifically includes:

[0206] A first latent feature calculation sub-module, configured to use the formula to calculate the latent feature of the user; in the formula, represents the latent feature of the user, σ represents a non-linear activation function, W represents a neural network weight, AFu-u An aggregation function that combines explicit and implicit friends of a user Indicates the interaction of a user with other users in time bin TI a Ex, representing the interaction of a user with other users u Im, representing the explicit friend feature representation u Im, representing the implicit friend feature representation; b represents the neural network bias

[0207] The first attention coefficient calculation sub-module, which is used to calculate the attention coefficient of neighboring users according to the latent features of the user using the formula ; where represents the attention coefficient of neighboring users, Softmax() represents the normalization function, and W' represents the weight matrix represents the transpose of the parameters of the attention network respectively represent the k-th power of the first and second bias terms of the attention network

[0208] The second latent feature calculation sub-module, which is used to calculate the latent features of user interaction items using the formula ; where represents the latent features of user interaction items, AF u-v An aggregation function that combines the explicitly interesting items that the user has interacted with historically and the implicit items that the user has interacted with indirectly through the meta-path Indicates the interaction of a user with other items at time TI a Ex, representing the interaction of a user with other items v represents the explicitly interesting items that the user has interacted with historically v Im, representing the implicit items that the user has interacted with indirectly through the meta-path

[0209] The second attention coefficient calculation sub-module, which is used to calculate the attention coefficient of neighboring items according to the latent features of user interaction items using the formula ; where represents the attention coefficient of neighboring items

[0210] The third latent feature calculation sub-module, which is used to calculate the latent features of item-to-item using the formula ; where represents the latent features of item-to-item, AF v-v An aggregation function that combines directly relevant information and indirectly relevant information with the target item Indicates the interaction embedding of the target item with other items in time bin TI a Di, representing the interaction embedding of the target item with other items v In, representing the items with directly relevant information to the target itemv Items that represent information indirectly related to the target item;

[0211] The third attention coefficient calculation sub-module, which is used to calculate the attention coefficient of the interaction item according to the potential features of the item and the item, using the formula ; in the formula, represents the attention coefficient of the interaction item;

[0212] The fourth potential feature calculation sub-module, which is used to calculate the virtual relationship item feature independent of the user preference using the formula ; in the formula, represents the virtual relationship item feature independent of the user preference, F u…v represents the aggregation function of items that are not directly or indirectly related to the user preference, represents at time TI a the embedding of random items unrelated to the user, Vi v represents the item feature representation independent of the user preference, and the elements included are 5 random item features, establishing a virtual relationship;

[0213] The fourth attention coefficient calculation sub-module, which is used to calculate the attention coefficient of the virtual relationship item according to the virtual relationship item feature independent of the user preference, using the formula ; in the formula, represents the attention coefficient of the virtual relationship item.

[0214] The clustering module specifically includes:

[0215] The classification sub-module, which is used to classify users and items according to the real-time knowledge graph into multiple clusters, with the number of nodes in each cluster being 1-5, and obtaining the user cluster as the item cluster as

[0216] where represents the u r th user cluster, containing A users; represents the v r th item cluster, containing B items, A, B ∈ [1, 5], and r is an arbitrary integer.

[0217] The screening module specifically includes:

[0218] The interaction matrix determination sub-module, which is used to determine the interaction matrix of the user cluster and the item cluster as in the formula, represents the element of the interaction matrix, takes values of 0, 1, and -1; when When , it means that the user cluster and the item cluster have direct interaction or indirect interaction along the meta-path of the graph data; when When , it means that there is no interactive information between the user cluster and the item cluster; when When , it means that the relationship between the user cluster and the target user cluster is a dislike relationship or the related items are disliked items; Y represents the interaction matrix, C u represents the u-th user cluster, C v represents the vth item cluster, U represents the user cluster set, and V represents the item cluster set;

[0219] The seed set is composed of submodules, which are used to and The corresponding item clusters constitute the seed set of each user cluster.

[0220] This paper proposes a research on a recommendation system based on an adaptive dynamic knowledge graph in a heterogeneous network. According to the complex interactive relationship between users and items, a heterogeneous network is constructed to extract implicit features of users. At the same time, attention is paid to users' short-term preferences, and the knowledge graph is updated to achieve the timeliness and adaptability of the system, improve the accuracy of the recommendation system, and better solve the problems of data sparsity, cold start and deviation. The advantages of this solution are mainly:

[0221] (1) Focus on the complex interactive relationships between users, between users and projects, and between projects, build heterogeneous networks, and reduce dependence on information such as user historical ratings.

[0222] (2) Using the GAT’s multi-head attention framework, we set up the calculation of four types of attention coefficients to form a dual task, extract the user’s short-term preference characteristics, assign different weights to the neighborhood nodes, and update the knowledge graph.

[0223] (3) User clusters are constructed according to the weights, and the RippleNet model framework is used for probability calculation. At the same time, disgusting items are suppressed, and unrelated clusters are randomly used as seeds to reduce the exposure bias problem. A list of user cluster recommendation results is given to reduce the amount of algorithm calculation.

[0224] (4) The knowledge graph is dynamic, taking into account the impact of time and online public opinion on user preferences, better combining users' long-term preferences with short-term preferences, and exploring users' potential preferences.

[0225] By utilizing the above points, the dependence of the adaptive dynamic knowledge graph recommendation system on data labels can be solved, which improves the accuracy of the recommendation system. At the same time, it makes the system timely, making people's lives more efficient and faster.

[0226] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method section.

[0227] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A recommendation method based on an adaptive dynamic knowledge graph in a heterogeneous network, characterized in that, Including: Construct a heterogeneous network based on a dataset of complex interaction relationships between users and projects; Extract entities and relationships from the heterogeneous network to establish a basic knowledge graph; Use a graph attention network in a time bin to extract the short-term preference features of users, and calculate multivariate attention coefficients based on the short-term preference features; Delete the relationships where the multivariate attention coefficients within the coefficient threshold range are located in the basic knowledge graph to obtain a real-time knowledge graph; Cluster the users and projects in the real-time knowledge graph to obtain multiple user clusters and multiple project clusters; Screen the project clusters according to the real-time knowledge graph and form the seed set of each user cluster; According to the seed set of each user cluster, use the RippleNet model to predict the probability value of each user cluster clicking on each project cluster in the seed set; Take the project cluster corresponding to the maximum probability value as the recommendation result of each user cluster, and generate a recommendation result list; Change the time bin, and return to the step "Use a graph attention network in a time bin to extract the short-term preference features of users, and calculate multivariate attention coefficients based on the short-term preference features" to obtain real-time recommendation results; The short-term preference features of the user include the user's latent features, the latent features of the user's interaction items, the latent features between projects, and the virtual relationship item features unrelated to the user's preferences; Calculating the multivariate attention coefficients according to the short-term preference features specifically includes: According to the latent features of users, the formula is used to calculate the attention coefficient of neighboring users; in the formula, represents the attention coefficient of neighboring users, Softmax() represents the normalization function, W' represents the weight matrix, represents the transpose of the parameters of the attention network, respectively represent the k-th power of the first and second bias terms of the attention network; σ represents the non-linear activation function, represents the interaction of the user with other users in the time bin TI a ; represents the latent features of the user. According to the potential features of user interaction items, the formula is used to calculate the attention coefficient of neighborhood items; in the formula, represents the attention coefficient of neighborhood items; represents the interaction of the user with other items at time TI a , represents the potential features of user interaction items. According to the potential features of the project and the project, the formula is used to calculate the attention coefficient of the interaction project; in the formula, represents the attention coefficient of the interaction project; represents the interaction embedding of the target project with other projects in the time bin TI a ; represents the potential features of the project and the project; Based on the virtual relationship item features independent of user preferences, use the formula to calculate the attention coefficient of the virtual relationship item; in the formula, represents the attention coefficient of the virtual relationship item; represents the embedding of the random item independent of the user at time TI a , and represents the virtual relationship item features independent of user preferences.

2. The recommendation method according to claim 1, characterized in that, The construction process of the dataset of complex interaction relationships between the user and the project includes: Collect the user set and the project set respectively; Collect the relationship sets of user-user, user-project, and project-project; Use the formula to calculate the weight of each relationship in the relationship set; in the formula, is the weight function of relationship r i , γ is the normalization coefficient, is the establishment time length of relationship r i , is the interaction frequency of relationship r i , is the number of common relationship nodes of two nodes of relationship r i , i ∈ [1, N], and N is the total number of relationships; Construct a dataset of complex interaction relationships between users and projects by combining the user set, the project set, the relationship set, and the weight of each relationship.

3. The recommendation method according to claim 1, characterized in that, The time bin is TI a = [ti a , ti a+1 ​ Wherein, TI a is the time bin, a is a constant, and ti a , ti a+1 represent the start time and the end time respectively.

4. The recommendation method according to claim 1, characterized in that, Using a graph attention network in a time bin to extract the short-term preference features of users specifically includes: Use the formula to calculate the potential features of the user; where W represents the neural network weights, AF u-u represents the aggregation function that fuses the user's explicit friends and implicit friends, Ex u represents the explicit friend feature representation, Im u represents the implicit friend feature representation, and b represents the neural network bias; Use the formula to calculate the potential features of user interaction items; where AF u-v represents an aggregation function that fuses the explicitly interesting items that the user has interacted with in the past and the implicit items that the user has interacted with indirectly through the meta-path, Ex v represents the explicitly interesting items that the user has interacted with in the past, Im v represents the implicit items that the user has interacted with indirectly through the meta-path; Using the formula to calculate the potential features between projects; in the formula, AF v-v represents an aggregation function that fuses information directly related and indirectly related to the target project, Di v represents a project with information directly related to the target project, In v represents a project with information indirectly related to the target project; Using the formula calculate the virtual relationship item characteristics that have nothing to do with the user's preferences; in the formula, F u…v represents the aggregation function of items that are not directly or indirectly related to the user's preferences, Vi v represents the item characteristics representation that has nothing to do with the user's preferences, and the elements included are randomly selected 5 item characteristics to establish a virtual relationship.

5. The recommendation method according to claim 1, characterized in that, Clustering the users and projects in the real-time knowledge graph to obtain multiple user clusters and multiple project clusters specifically includes: Classify users and projects according to the real-time knowledge graph into multiple clusters, with the number of nodes in each cluster being 1 - 5, and obtain the user clusters as The project clusters are Among them, represents the u r th user cluster, which contains A users; represents the v r th project cluster, which contains B projects, where A, B ∈ [1, 5] and r is an arbitrary integer.

6. The recommendation method according to claim 1, characterized in that, Screening the project clusters according to the real-time knowledge graph and forming the seed set of each user cluster specifically includes: The interaction matrix between user clusters and item clusters is determined as In the formula, represents the element of the interaction matrix, takes values of 0, 1, and -1; when it means that there is a direct interaction between the user cluster and the item cluster or an indirect interaction along the meta-path of the graph data; when it means that there is no interaction information between the user cluster and the item cluster; when it means that the relationship between the user cluster and the target user cluster is an aversion relationship or the related item is a disliked item; Y represents the interaction matrix, C u represents the \(u\)-th user cluster, C v represents the \(v\)-th item cluster, U represents the set of user clusters, and V represents the set of item clusters; Combine and the corresponding item clusters to form the seed set of each user cluster.

7. A recommendation system based on an adaptive dynamic knowledge graph in a heterogeneous network, characterized in that, Including: A heterogeneous network construction module for constructing a heterogeneous network based on a dataset of complex interaction relationships between users and projects; A knowledge graph establishment module for extracting entities and relationships from the heterogeneous network to establish a basic knowledge graph; An attention coefficient calculation module for using a graph attention network in a time bin to extract the short-term preference features of users and calculating multivariate attention coefficients based on the short-term preference features; A knowledge graph update module for deleting the relationships where the multivariate attention coefficients within the coefficient threshold range are located in the basic knowledge graph to obtain a real-time knowledge graph; A clustering module for clustering the users and projects in the real-time knowledge graph to obtain multiple user clusters and multiple project clusters; A screening module for screening project clusters according to the real-time knowledge graph and forming the seed set of each user cluster; A prediction module for using the RippleNet model to predict the probability value of each user cluster clicking on each project cluster in the seed set according to the seed set of each user cluster; A recommended result generation module, which is used to take the item cluster corresponding to the maximum probability value as the recommended result of each user cluster and generate a recommended result list; A loop module, which is used to change the time bin, call the attention coefficient calculation module, and obtain real-time recommended results; The short-term preference features of the user include the latent features of the user, the latent features of the user interaction items, the latent features of the items and items, and the virtual relationship item features irrelevant to the user preferences; The attention coefficient calculation module includes: The first attention coefficient calculation sub-module is used to calculate the attention coefficients of neighboring users according to the potential features of the user, using the formula ; where, represents the attention coefficient of the neighboring user, Softmax() represents the normalization function, W' represents the weight matrix, represents the transpose of the parameters of the attention network, respectively represent the k-th power of the first and second bias terms of the attention network; σ represents the non-linear activation function, represents the interaction of the user with other users in the time bin TI a ; represents the potential feature of the user; The second attention coefficient calculation sub-module is used to calculate the attention coefficient of neighborhood items according to the potential features of user interaction items by using the formula ; in the formula, represents the attention coefficient of the neighborhood item; represents the interaction of the user with other items at time TI a ; represents the potential feature of the user interaction item; The third attention coefficient calculation sub-module is used to calculate the attention coefficient of the interaction item according to the potential features of the item and the item, using the formula ; in the formula, represents the attention coefficient of the interaction item; represents the interaction embedding of the target item with other items in the time bin TI a ; represents the potential features of the item and the item; The fourth attention coefficient calculation sub-module is used to calculate the attention coefficient of the virtual relationship item according to the virtual relationship item characteristics independent of user preferences, using the formula ; in the formula, represents the attention coefficient of the virtual relationship item; represents the embedding of the random item independent of the user at time TI a , and represents the virtual relationship item characteristics independent of user preferences.

8. The recommendation system according to claim 7, characterized in that, The attention coefficient calculation module further includes: The first potential feature calculation sub-module is used to calculate the potential features of the user by using the formula ; where W represents the neural network weights, AF u-u represents the aggregation function that fuses the user's explicit friends and implicit friends, Ex u represents the explicit friend feature representation, Im u represents the implicit friend feature representation, and b represents the neural network bias; The second potential feature calculation sub-module is used to utilize the formula to calculate the potential features of user interaction items; in the formula, AF u-v represents an aggregation function that fuses the explicitly interested items that the user has interacted with in the past and the implicit items that the user has indirectly interacted with through the meta-path, Ex v represents the explicitly interested items that the user has interacted with in the past, Im v represents the implicit items that the user has indirectly interacted with through the meta-path; The third potential feature calculation sub-module is used to calculate the potential features between items by using the formula ; where AF v-v represents an aggregation function that fuses direct and indirect relevant information with the target item, Di v represents an item with direct relevant information to the target item, In v represents an item with indirect relevant information to the target item; The fourth potential feature calculation sub-module is used to utilize the formula to calculate the virtual relationship item features that have nothing to do with user preferences; in the formula, F u…v represents the aggregation function of items that have no direct or indirect relationship with user preferences, and Vi v represents the item feature representation that has nothing to do with user preferences, and the elements it contains are randomly selected 5 item features to establish a virtual relationship.

9. The recommendation system according to claim 7, characterized in that, The clustering module specifically includes: A classification sub-module, which is used to classify users and projects according to a real-time knowledge graph into multiple clusters, with the number of nodes in each cluster being 1-5, and obtain the user cluster as The project cluster is Among them, represents the $u$-th r user cluster, which contains $A$ users; represents the $v$-th r item cluster, which contains $B$ items, where $A, B \in [1, 5]$ and $r$ is an arbitrary integer.

10. The recommendation system according to claim 7, characterized in that, The screening module specifically includes: An interaction matrix determination sub-module, which is used to determine that the interaction matrix between user clusters and item clusters is In the formula, represents an element of the interaction matrix, The value of is 0, 1, and -1; when , it means that there is a direct interaction between the user cluster and the item cluster or an indirect interaction along the meta-path of the graph data; when , it means that there is no interaction information between the user cluster and the item cluster; when , it means that the relationship between the user cluster and the target user cluster is an aversion relationship or the related item is a disliked item; Y represents the interaction matrix, C u represents the u-th user cluster, C v represents the v-th item cluster, U represents the set of user clusters, and V represents the set of item clusters; The seed set composition sub-module is used to and The corresponding item clusters form the seed set of each user cluster.

Citation Information

Patent Citations

  • Knowledge graph-based user modeling method and sequence recommendation method

    CN110516160A

  • Sequence recommendation method and device and computer readable storage medium

    CN111522962A