Knowledge graph recommendation algorithm oriented to large language model
By constructing a hierarchical knowledge graph and combining user historical data, and using weighted calculation formulas to calculate recommendation scores, the problem of low accuracy of recommendation results in the existing technology is solved, more accurate and personalized recommendation results are achieved, and the response speed of the recommendation system is improved.
Patent Information
- Application Number
- CN202510203828.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing knowledge graph recommendation algorithm cannot effectively reflect users' common habits and domain problems, resulting in low accuracy of recommendation results and cannot meet users' personalized needs.
Terms and concepts in subject text are extracted through natural language processing technology, entities are marked using deep learning models and hierarchical knowledge graphs are constructed, and recommendation scores are calculated using weighted calculation formulas, and weight coefficients are dynamically adjusted to improve the accuracy and personalization of recommendations.
It realizes more accurate and personalized knowledge graph recommendation results, which can capture changes in user preferences and interests in real time, improve user stickiness and satisfaction, and at the same time, improve the response speed and query efficiency of the recommendation system through hierarchical management of memory cache.
Smart Images

Figure CN119938941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph recommendation technology, and specifically to a knowledge graph recommendation algorithm for a large language model. Background Art
[0002] The knowledge graph recommendation algorithm is a recommendation algorithm based on graph structure. It combines domain knowledge and user needs by constructing a graph containing entities and their relationships to provide personalized recommendation services. As a graphical data storage method, the knowledge graph can intuitively express the multi-level relationships between different entities and use the nodes and edges in the graph to obtain rich contextual information.
[0003] The existing knowledge graph recommendation algorithm only considers limited recommendation factors and relies solely on large language models to generate sentences about user characteristics based on context. It cannot reflect users' common habits and domain issues, resulting in low accuracy of recommendation results and failure to meet users' personalized needs.
[0004] In view of the above technical defects, a solution is now proposed. Summary of the invention
[0005] In view of the deficiencies in the prior art, the present invention provides a knowledge graph recommendation algorithm for large language models.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a knowledge graph recommendation algorithm for a large language model, comprising the following steps:
[0007] Step 1: Extract terms and concepts from subject texts through natural language processing (NLP) technology, use deep learning models to annotate entities and set them as nodes, use GCN methods to extract relationships between entities and annotate them in the knowledge graph, and build a hierarchical graph. The first layer is the basic knowledge layer, and the second layer is the advanced knowledge layer.
[0008] Step 2: Use Neo4j graph database to store data, use Cypher query language to store and retrieve graph structure data, regularly scan academic literature, news reports and other data sources, update knowledge graph, analyze and delete duplicate and redundant nodes and edges;
[0009] Step 3: When the large language model queries the knowledge in a certain subject area, it uses the depth-first search algorithm to recursively explore the adjacent nodes from the starting node K_i, and finally outputs the path from K_i to K_j, and transmits the node information in the local knowledge graph associated with the path to the large language model;
[0010] Step 4: When the user sets preference information, analyze the user's historical data and extract the preference factor K. Combine similarity, timeliness and relevance to calculate the recommendation score S through a weighted calculation formula. Multiple queries are performed to transmit the node with the highest recommendation score S and the node information in the local knowledge graph associated with the path to the large language model.
[0011] Step 5: Through hierarchical management of memory cache, high-access data is stored in the first-level cache, where data is frequently updated; the second-level cache stores less frequently accessed but important data; the third-level cache stores the data with the lowest access frequency, which is updated the slowest. The LRU replacement strategy is used to eliminate infrequently accessed data and optimize cache space.
[0012] Extract certain subject texts from existing databases and academic papers, use natural language processing technology NIP to identify terms, concepts and technical terms in the text, annotate each entity and set it as a node, use the GCN method to extract the relationship between entities, annotate the relationship between entities in the subject field in the graph, define edges and their types, and build different hierarchical knowledge graphs for different subject fields. The first layer is the basic knowledge layer, and the second layer is the advanced knowledge layer. Each node and edge includes its identifier, name, description and other metadata.
[0013] Use the graph database Neo4j for data storage, use the graph database query language Cypher to store and retrieve graph structured data, regularly scan data sources such as academic literature and news reports, regularly update and analyze the knowledge graph, and identify and delete duplicate and redundant nodes and edges.
[0014] When the large language model needs to query the knowledge data in a certain subject area, the graph search depth-first search algorithm is used to find the path between the required nodes. The graph search depth-first search algorithm is: input the starting node K_i, the target node K_j and the graph G, recursively explore the adjacent nodes starting from node K_i, record the path until the target node K_j is reached or there are no further nodes, and finally output the path from K_i to K_j, and then transmit the node information in the local knowledge graph associated with the path to the large language model.
[0015] When users set preference information, we analyze user historical data and extract user preference factors K. We introduce multiple dimensions of scoring factors such as similarity, timeliness, and relevance, and calculate the final recommendation score S through a weighted calculation formula. We perform multiple queries to transmit the node with the highest recommendation score S and the node information in the local knowledge graph associated with the path to the large language model, and calculate the recommendation score S according to the formula. The specific formula is as follows:
[0016] S=f S × 1 +f R ×2 +f T × 3
[0017] Among them, f S is the similarity value between query Q and user preference factor K; f R is the correlation value between the queried node knowledge and the query Q, obtained based on the path length; f T It is the timeliness value of the user preference factor K, W1 represents the weight coefficient of the similarity value between the query Q and the user preference factor K, W2 represents the weight coefficient of the correlation value between the queried node knowledge and the query Q, and W3 represents the weight coefficient of the timeliness value of the user preference factor K. The weight coefficient is obtained through historical data training, and the weight coefficient in the formula is not fixed.
[0018] Store highly accessed data in the memory cache to reduce access to the graph database during each query. Manage the cache in a hierarchical manner. The first-tier cache stores the most frequently queried nodes and relationships, which are updated frequently. The second-tier cache stores nodes and relationships that are less frequently accessed but still important. The data update speed of this cache is lower than that of the first-tier cache. The third-tier cache stores the nodes and relationships with the lowest access frequency. These data need to be accessed occasionally. The data update speed of this cache is lower than that of the second-tier cache. Use the LRU cache replacement strategy to manage the cache space and intermittently eliminate infrequently accessed data in each cache layer.
[0019] Design experimental and control groups for testing and verification, collect experimental data and analyze indicators such as accuracy, recall, and coverage of recommendation results, analyze the impact of different coefficients on these indicators, identify key factors, and adjust the weight coefficients of optimization similarity, timeliness, and relevance in the recommendation formula based on comparative experimental results. Analyze the differences in algorithm performance in different scenarios and adjust algorithm parameters.
[0020] The present invention provides a knowledge graph recommendation algorithm for large language models. Compared with the prior art, it has the following beneficial effects:
[0021] The present invention takes into account the user's historical behavior, preference information and subject areas, combines multi-dimensional scoring factors such as similarity, timeliness and relevance, and dynamically adjusts the weight coefficient to ensure that each recommendation result is more accurate and personalized. In particular, in terms of changes in user preferences and interests, the algorithm can capture and adjust the recommended content in real time to ensure that the recommended content is highly matched with user needs, thereby increasing user stickiness and satisfaction.
[0022] The present invention can reduce the access to the graph database during each query through hierarchical management of the memory cache, thereby improving the response speed and query efficiency of the recommendation system. The hierarchical management of the cache reduces the database load by giving priority to the storage of frequently accessed data, while reducing the read and write pressure on the hard disk, effectively improving the real-time performance and accuracy of the large language model when processing complex queries. In addition, the LRU cache replacement strategy is adopted to further ensure the efficient use of the cache space and avoid redundant and invalid storage of data in the cache. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic diagram of the principle framework of the present invention. DETAILED DESCRIPTION
[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0025] See also Figure 1 , this application provides a knowledge graph recommendation algorithm for a large language model, including the following steps:
[0026] Step 1: Extract terms and concepts from subject texts through natural language processing (NLP) technology, use deep learning models to annotate entities and set them as nodes, use GCN methods to extract relationships between entities and annotate them in the knowledge graph, and build a hierarchical graph. The first layer is the basic knowledge layer, and the second layer is the advanced knowledge layer.
[0027] Step 2: Use Neo4j graph database to store data, use Cypher query language to store and retrieve graph structure data, regularly scan academic literature, news reports and other data sources, update knowledge graph, analyze and delete duplicate and redundant nodes and edges;
[0028] Step 3: When the large language model queries the knowledge in a certain subject area, it uses the depth-first search algorithm to recursively explore the adjacent nodes from the starting node K_i, and finally outputs the path from K_i to K_j, and transmits the node information in the local knowledge graph associated with the path to the large language model;
[0029] Step 4: When the user sets preference information, analyze the user's historical data and extract the preference factor K. Combine similarity, timeliness and relevance to calculate the recommendation score S through a weighted calculation formula. Multiple queries are performed to transmit the node with the highest recommendation score S and the node information in the local knowledge graph associated with the path to the large language model.
[0030] Step 5: Through hierarchical management of memory cache, high-access data is stored in the first-level cache, where data is frequently updated; the second-level cache stores less frequently accessed but important data; the third-level cache stores the data with the lowest access frequency, which is updated the slowest. The LRU replacement strategy is used to eliminate infrequently accessed data and optimize cache space.
[0031] Extract certain subject texts from existing databases and academic papers, use natural language processing technology NIP to identify terms, concepts and technical terms in the text, annotate each entity and set it as a node, use the GCN method to extract the relationship between entities, annotate the relationship between entities in the subject field in the graph, define edges and their types, and build different hierarchical knowledge graphs for different subject fields. The first layer is the basic knowledge layer, and the second layer is the advanced knowledge layer. Each node and edge includes its identifier, name, description and other metadata.
[0032] During the construction process, the knowledge graphs of different subject areas will be designed differently according to their complexity and hierarchy. The basic knowledge layer covers the core concepts, basic terms and common entities in the subject area, and is mainly used to provide the basic framework of the subject. The advanced knowledge layer is built on the basic knowledge layer and contains more complex and in-depth academic theories, cutting-edge research results and expert-level knowledge nodes. In this process, each node and edge includes the entity's identifier, name, description information, and also covers information such as the entity's creation time, related research fields and sources.
[0033] Use the graph database Neo4j for data storage, use the graph database query language Cypher to store and retrieve graph structured data, regularly scan data sources such as academic literature and news reports, regularly update and analyze the knowledge graph, and identify and delete duplicate and redundant nodes and edges.
[0034] By regularly analyzing graph data, the system can identify and eliminate redundant nodes and edges. For example, when multiple nodes represent the same entity, they are automatically merged to reduce data duplication and edges with overly lengthy or invalid relationships are deleted to ensure that the graph structure is concise and efficient. In addition, by introducing the graph algorithm PageRank system, the system can deeply analyze the graph structure, identify important nodes and potential hidden associations, thereby enhancing the insight and reasoning ability of the graph.
[0035] When the large language model needs to query the knowledge data in a certain subject area, the graph search depth-first search algorithm is used to find the path between the required nodes. The graph search depth-first search algorithm is: input the starting node K_i, the target node K_j and the graph G, recursively explore the adjacent nodes starting from node K_i, record the path until the target node K_j is reached or there are no further nodes, and finally output the path from K_i to K_j, and then transmit the node information in the local knowledge graph associated with the path to the large language model.
[0036] When users set preference information, we analyze user historical data and extract user preference factors K. We introduce multiple dimensions of scoring factors such as similarity, timeliness, and relevance, and calculate the final recommendation score S through a weighted calculation formula. We perform multiple queries to transmit the node with the highest recommendation score S and the node information in the local knowledge graph associated with the path to the large language model, and calculate the recommendation score S according to the formula. The specific formula is as follows:
[0037] S=f S × 1 +f R × 2 +f T × 3
[0038] Among them, f S is the similarity value between query Q and user preference factor K; f R is the correlation value between the queried node knowledge and the query Q, obtained based on the path length; f T It is the timeliness value of the user preference factor K, W1 represents the weight coefficient of the similarity value between the query Q and the user preference factor K, W2 represents the weight coefficient of the correlation value between the queried node knowledge and the query Q, and W3 represents the weight coefficient of the timeliness value of the user preference factor K. The weight coefficient is obtained through historical data training, and the weight coefficient in the formula is not fixed.
[0039] Store highly accessed data in the memory cache to reduce access to the graph database during each query. Manage the cache in a hierarchical manner. The first-tier cache stores the most frequently queried nodes and relationships, which are updated frequently. The second-tier cache stores nodes and relationships that are less frequently accessed but still important. The data update speed of this cache is lower than that of the first-tier cache. The third-tier cache stores the nodes and relationships with the lowest access frequency. These data need to be accessed occasionally. The data update speed of this cache is lower than that of the second-tier cache. Use the LRU cache replacement strategy to manage the cache space and intermittently eliminate infrequently accessed data in each cache layer.
[0040] Design experimental and control groups for testing and verification, collect experimental data and analyze indicators such as accuracy, recall, and coverage of recommendation results, analyze the impact of different coefficients on these indicators, identify key factors, and adjust the weight coefficients of optimization similarity, timeliness, and relevance in the recommendation formula based on comparative experimental results. Analyze the differences in algorithm performance in different scenarios and adjust algorithm parameters.
[0041] When the weight of similarity is increased, the system will recommend more content that is consistent with the user's historical behavior and interests, but it will sacrifice a certain degree of recommendation diversity. Increasing the weight of timeliness will give priority to recommending the latest content, so that the recommendation system can stay updated. Adjustment of relevance weight can affect the matching degree between the recommendation results and the content described by the user. Appropriately increasing the relevance weight can usually improve accuracy, but will reduce the recall rate. Users can set the normal mode to adjust the weight coefficient through the multi-objective optimization algorithm to achieve the best balance among various indicators.
[0042] Specific workflow:
[0043] Through NLP technology, we extract terms and concepts from subject texts and annotate entities, and use the GCN method to extract the relationship between entities, build knowledge graphs in different subject areas, define nodes and edges and their attributes, use the Neo4j graph database to store data, regularly scan academic literature, news reports and other data sources to update the knowledge graph, and analyze, identify and delete duplicate or redundant nodes and edges. When the large language model needs to query knowledge in a certain subject area, we use the depth-first search algorithm to recursively explore adjacent nodes, find the path from the start node to the target node, and transmit the local knowledge graph node information associated with the path to the large language model. We further analyze the user's historical data, extract user preference factors, and calculate the recommendation score based on scoring factors such as similarity, timeliness and relevance. According to the query, the node information with the highest score is transmitted to the large language model multiple times. We design experimental and control groups, collect experimental data through A / B testing, analyze the accuracy, recall and coverage of the recommendation results, and adjust the weight coefficient in the recommendation formula to optimize the recommendation effect.
[0044] Furthermore, the present invention dynamically adjusts the weight coefficient by comprehensively considering the user's historical behavior, preference information, and subject areas, combining multi-dimensional scoring factors such as similarity, timeliness, and relevance, to ensure that each recommendation result is more accurate and personalized. In particular, in terms of changes in user preferences and interests, the algorithm can capture and adjust the recommended content in real time to ensure that the recommended content is highly matched with user needs, thereby increasing user stickiness and satisfaction.
[0045] Furthermore, the present invention can reduce the access to the graph database during each query through hierarchical management of the memory cache, thereby improving the response speed and query efficiency of the recommendation system. The hierarchical management of the cache reduces the database load by giving priority to the storage of frequently accessed data, while reducing the read and write pressure on the hard disk, effectively improving the real-time and accuracy of the large language model when processing complex queries. In addition, the LRU cache replacement strategy is adopted to further ensure the efficient use of the cache space and avoid redundant and invalid storage of data in the cache.
[0046] Some of the data in the above formulas are dimensionless and numerically calculated. Meanwhile, the contents not described in detail in this specification belong to the prior art known to those skilled in the art.
[0047] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. Knowledge graph recommendation algorithm for large language models, characterized by: The following steps are involved: Step 1: Extract terms and concepts from subject texts through natural language processing (NLP) technology, use deep learning models to annotate entities and set them as nodes, use GCN methods to extract relationships between entities and annotate them in the knowledge graph, and build a hierarchical graph. The first layer is the basic knowledge layer, and the second layer is the advanced knowledge layer. Step 2: Use Neo4j graph database to store data, use Cypher query language to store and retrieve graph structure data, regularly scan academic literature, news reports and other data sources, update knowledge graph, analyze and delete duplicate and redundant nodes and edges; Step 3: When the large language model queries the knowledge in a certain subject area, it uses the depth-first search algorithm to recursively explore the adjacent nodes from the starting node K_i, and finally outputs the path from K_i to K_j, and transmits the node information in the local knowledge graph associated with the path to the large language model; Step 4: When the user sets preference information, analyze the user's historical data and extract the preference factor K. Combine similarity, timeliness and relevance to calculate the recommendation score S through a weighted calculation formula. Multiple queries are performed to transmit the node with the highest recommendation score S and the node information in the local knowledge graph associated with the path to the large language model. Step 5: By managing memory cache in a hierarchical manner, store high-access data in the first-tier cache, where data is frequently updated; The second layer stores less frequently accessed but important data; the third layer stores the data with the lowest access frequency and is updated the slowest. The LRU replacement strategy is used to eliminate infrequently accessed data and optimize cache space.
2. The knowledge graph recommendation algorithm for large language models according to claim 1, characterized in that: Extract certain subject texts from existing databases and academic papers, use natural language processing technology NIP to identify terms, concepts and technical terms in the text, annotate each entity and set it as a node, use the GCN method to extract the relationship between entities, annotate the relationship between entities in the subject field in the graph, define edges and their types, and build different hierarchical knowledge graphs for different subject fields. The first layer is the basic knowledge layer, and the second layer is the advanced knowledge layer. Each node and edge includes its identifier, name, description and other metadata.
3. The knowledge graph recommendation algorithm for large language models according to claim 1, characterized in that: Use the graph database Neo4j for data storage, use the graph database query language Cypher to store and retrieve graph structured data, regularly scan data sources such as academic literature and news reports, regularly update and analyze the knowledge graph, and identify and delete duplicate and redundant nodes and edges.
4. The knowledge graph recommendation algorithm for large language models according to claim 1, characterized in that: When the large language model needs to query the knowledge data in a certain subject area, the graph search depth-first search algorithm is used to find the path between the required nodes. The graph search depth-first search algorithm is: input the starting node K_i, the target node K_j and the graph G, recursively explore the adjacent nodes starting from node K_i, record the path until the target node K_j is reached or there are no further nodes, and finally output the path from K_i to K_j, and then transmit the node information in the local knowledge graph associated with the path to the large language model.
5. The knowledge graph recommendation algorithm for large language models according to claim 1, characterized in that: When users set preference information, we analyze user historical data and extract user preference factors K. We introduce multiple dimensions of scoring factors such as similarity, timeliness, and relevance, and calculate the final recommendation score S through a weighted calculation formula. We perform multiple queries to transmit the node with the highest recommendation score S and the node information in the local knowledge graph associated with the path to the large language model, and calculate the recommendation score S according to the formula. The specific formula is as follows: S=f S ×w1+f R ×w2+f T ×w3 Among them, f S is the similarity value between query Q and user preference factor K; f R is the correlation value between the queried node knowledge and the query Q, obtained based on the path length; f T It is the timeliness value of the user preference factor K, W1 represents the weight coefficient of the similarity value between the query Q and the user preference factor K, W2 represents the weight coefficient of the correlation value between the queried node knowledge and the query Q, and W3 represents the weight coefficient of the timeliness value of the user preference factor K. The weight coefficient is obtained through historical data training, and the weight coefficient in the formula is not fixed.
6. The knowledge graph recommendation algorithm for large language models according to claim 1, characterized in that: Store highly accessed data in the memory cache to reduce access to the graph database during each query. Manage the cache in a hierarchical manner. The first-tier cache stores the most frequently queried nodes and relationships, which are updated frequently. The second-tier cache stores nodes and relationships that are less frequently accessed but still important. The data update speed of this cache is lower than that of the first-tier cache. The third-tier cache stores the nodes and relationships with the lowest access frequency. These data need to be accessed occasionally. The data update speed of this cache is lower than that of the second-tier cache. Use the LRU cache replacement strategy to manage the cache space and intermittently eliminate infrequently accessed data in each cache layer.
7. The knowledge graph recommendation algorithm for large language models according to claim 1, characterized in that: Design experimental and control groups for testing and verification, collect experimental data and analyze indicators such as accuracy, recall, and coverage of recommendation results, analyze the impact of different coefficients on these indicators, identify key factors, and adjust the weight coefficients of optimization similarity, timeliness, and relevance in the recommendation formula based on comparative experimental results. Analyze the differences in algorithm performance in different scenarios and adjust algorithm parameters.
Citation Information
Cited By
Property low-frequency transaction-oriented recommendation method and system based on knowledge graph
CN121210769A
Ultra-high performance concrete knowledge graph construction and graph retrieval enhanced generation-based mechanism interpretation method
CN121390245A
Construction of ultra-high performance concrete knowledge graph and machine interpretation method based on graph retrieval enhanced generation
CN121390245B