Model-knowledge dual-driven tourism project recommendation system and method
Through a dual-driven tourism project recommendation system with a model-knowledge drive, combined with a graph convolution network and attention gating mechanism, a tourism knowledge graph is built, which solves the shortcomings of the existing system in terms of the relevance and personalization of recommended content, and realizes efficient and personalized tourism project recommendations.
Patent Information
- Application Number
- CN202411954065.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-30
AI Technical Summary
The existing tourism recommendation system has limitations in the relevance and richness of recommended content, and lacks in the timeliness and accuracy of data updates, which cannot meet the personalized needs of tourists.
A tourism project recommendation system driven by a model-knowledge dual drive is adopted, combining graph convolution networks, attention gating mechanisms, large language models and network crawling technology to build a tourism knowledge graph, capture high-order semantic information of users and items through heterogeneous information aggregation and dissemination modules, and use attention weighting mechanisms to improve the personalization of recommendations.
It significantly improves the recall rate and normalized loss cumulative gain of tourism project recommendations, improves the relevance and personalization of recommended content, and ensures the accuracy and timeliness of recommendation results and actual situations.
Smart Images

Figure CN120067433A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a model-knowledge dual-driven tourism project recommendation system and method, belonging to the field of intersection of scenic spot recommendation and artificial intelligence technology. Background Technique
[0002] With the rapid development of social economy and the continuous improvement of people's living standards, the tourism industry has gradually become an important pillar industry of the national economy. People's demand for tourism is increasing day by day, and the scale of the tourism market is constantly expanding. The demand of tourists for tourism products and services gradually shows a trend of diversification and personalization. Traditional tourism recommendation methods are often limited to popular scenic spots and cannot meet the growing personalized needs of tourists. Tourists hope to discover more novel destinations and experiences in food, accommodation during the journey according to their own interests and preferences.
[0003] From e-commerce, social platforms to news websites, recommendation systems have become an indispensable part of many applications providing information services. Traditional recommendation algorithms are mainly based on collaborative filtering, which assumes that users with similar behaviors may have similar preferences for items, and completes recommendations based on this assumption. However, the recommendation algorithm based on collaborative filtering has problems of sparsity and cold start in the absence of side information. Knowledge graphs have been widely studied and applied in the field of recommendation systems because they can effectively solve the sparsity and cold start problems of collaborative filtering.
[0004] In recent years, the algorithms of recommendation systems mainly include the following: collaborative filtering algorithm, content-based recommendation algorithm, and knowledge graph-based recommendation algorithm. The collaborative filtering algorithm mainly relies on users' behavioral data, such as browsing records, ratings, collections, etc., to find other users with similar interests to the target user, and then makes recommendations based on the behaviors of these similar users. The paper "LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation" by X. He, K. Deng, X. Wang, etc. proposed a simplified graph convolution network model LightGCN, which only retains the neighbor aggregation operation to improve the performance of the recommendation system and simplify the model structure. By removing the feature transformation and non-linear activation operations in GCN, LightGCN only uses linear propagation to learn the embedding vectors of users and items on the user-item interaction graph, achieving significant improvements compared to the state-of-the-art GCN-based recommendation models. The content-based recommendation algorithm is based on the content features of items and makes recommendations by analyzing the matching relationship between item features and user preferences. Compared with the collaborative filtering algorithm, content recommendation has better ability to solve the cold start problem of new users and new items. With the development of knowledge graph technology, the knowledge graph-based recommendation algorithm has gradually become a research hotspot. This method constructs a knowledge graph containing entities and their relationships, captures the complex relationships between users and items from the semantic level, and thus realizes more interpretable recommendations. The paper "Knowledge-Enhanced Hierarchical Graph Transformer Network for Multi-Behavior Recommendation" by L. Xia, C. Huang, Yong. Xu, etc. proposed a knowledge-enhanced hierarchical graph transformation network (KHGT) to solve the multi-behavior recommendation problem, which can effectively capture multiple types of interaction patterns between users and items, and integrate knowledge-aware item relationships, improving the accuracy and generalization ability of the recommendation system. Specifically, KHGT adopts a graph-based neural architecture, captures the semantic of specific types of behaviors, and clearly distinguishes the importance of different types of user-item interactions in predicting the target behavior. In addition, by combining the multi-modal graph attention layer with the time encoding strategy, the learned embeddings can reflect the dedicated multi-way user-item and item-item collaborative relationships, as well as potential interaction dynamics.
[0005] In summary, recommendation algorithms based on knowledge graphs are becoming increasingly mature, but their application exploration in the tourism field is insufficient. On the one hand, there are obvious limitations in the relevance and richness of the recommended content in the existing systems. Usually, only single tourism items such as scenic spots, hotels, or restaurants are recommended, ignoring the internal relevance and mutual influence among them. On the other hand, the existing recommendation systems based on knowledge graphs also lack in the timeliness and accuracy of data updates. Due to the frequent dynamic changes in the tourism market, information such as the opening hours and ticket prices of scenic spots, as well as the business conditions of hotels and restaurants, may change at any time, resulting in the inconsistency between the recommended results and the actual situation. Summary of the Invention
[0006] The purpose of the present invention is to address the defects or deficiencies in the existing technologies in the field of tourism recommendation, and propose a model-knowledge dual-driven tourism item recommendation system and method. This system combines the advantages of technologies such as graph convolutional networks, attention gating mechanisms, large language models, and web crawlers to fully mine and utilize the complex correlations among entities such as tourism scenic spots, hotels, and restaurants, and to a certain extent improve the recommendation performance such as recall rate (Recall) and normalized discounted cumulative gain (NDCG), effectively alleviating the problems of single recommended content, insufficient relevance, and lack of personalization of tourism interest points.
[0007] The technical solution adopted by the present invention to solve its technical problems is: a model-knowledge dual-driven tourism item recommendation system, which includes a data collection module, a knowledge graph construction module, a heterogeneous information aggregation / propagation module, an attention weighting module, and a prediction module.
[0008] The data collection module introduces large language models (LLMs) into web scraping technology, and uses key information prompting words to obtain the triple information required for constructing a tourism knowledge graph, including information such as scenic spot names, scenic area tickets, and tourist ratings, which is the basis for subsequent modules.
[0009] The knowledge graph construction module preprocesses the information collected by the data collection module, and constructs a tourism-related knowledge graph based on the preprocessed data file, facilitating subsequent recommendation system models to perform semantic enhancement in combination with the knowledge graph.
[0010] The heterogeneous information aggregation / propagation module aggregates user-item interaction information into entity information in the knowledge graph based on the preprocessed user-item interaction matrix and knowledge graph, and then propagates it along the connection relationships in the knowledge graph to capture the high-order semantic information of users and items based on the knowledge graph, further enriching the representations of users and items.
[0011] The attention weighting module operates on the triples in the knowledge graph, and uses the attention mechanism to assign different weights to the tail entities according to the differences in the head entities and relationships.
[0012] The prediction module is the output module of the system, which aggregates and performs dot product operations on the weighted user and item representation vectors to achieve tourism project recommendations.
[0013] The present invention also provides a method for implementing a model-knowledge dual-driven tourism project recommendation system, which includes the following steps:
[0014] Step 1: Acquisition and storage of tourism-related data sets. Use the scrapy framework and large language models in the Python language to accurately extract effective tourism-related information from Qunar.com web pages and convert it into structured data for storage in a text file.
[0015] Step 2: Data preprocessing and knowledge graph construction. Preprocess the extracted data set to obtain a rating data file, a knowledge graph file, and a ripple set file containing a multi-hop relationship set.
[0016] Step 3: Heterogeneous information aggregation and propagation based on the tourism knowledge graph. Use the diffusion mechanism of the knowledge graph to capture high-order semantic information and enhance the embedding representations of users and items.
[0017] Step 4: Triple relevance weighting based on the attention mechanism. Use the attention mechanism to assign different weights to the tail entity according to the different head entities and relationships, so as to weight the impacts on different entities (users, items) and relationships, and generate the embedding representation of each hop.
[0018] Step 5: Update of user and item embeddings. Aggregate the multi-hop embeddings of users or items to generate the final embedding representation.
[0019] Step 6: Tourism interest point recommendation and effect testing. According to the updated embedding representations of users and items, use dot product calculation to achieve tourism project recommendations. Then evaluate the prediction effect of the model on user click behavior through the Area Under Curve (AUC) metric, and evaluate the accuracy and ranking quality of the recommendation results through the Recall@ and NDCG@K metrics.
[0020] Further, step 1 of the present invention includes: First, deploy the Scrapy framework to build the basic environment of the web crawler. In the initialization stage of the Scrapy project, construct the corresponding Spider (crawler class). In the Spider, define the initial URL (Uniform Resource Locator) of the Qunar website to be crawled, and the Scrapy framework automatically initiates an HTTP request to obtain the HTML content of the corresponding web page. To obtain structured data convenient for subsequent construction of the knowledge graph and prevent problems caused by non-standard or dynamically modified web source HTML designs, the present invention pre-designs a prompt word to clearly indicate which information the large language model (LLMs) should extract. For example, in the present invention, it is necessary to crawl the tourism data of the Qunar website, including information such as title (scenic spot name), address (scenic spot address), price (scenic spot ticket), hot_num (heat index), review (traveler rating), etc. Then input the HTML content into the LLMs for reading, understanding its semantics, and operating according to the prompt word and structured instructions, outputting the structured data and storing it in CSV format.
[0021] Further, step 2 of the present invention includes: Preprocess the csv-format data obtained in step 1 to generate a rating data file, a knowledge graph file, and a ripple set file containing a multi-hop relationship set. Specifically, the rating data file is stored in the format of (user_id, item_id, label), where label represents the label of whether the user has positive feedback or interaction with the item. The present invention believes that when the user does not have positive interaction with the item or is not interested in the item, label = 0; when the user has clear positive feedback (rating higher than 3) on the item, that is, is interested, label = 1. Obtain the rating data file from the csv-format file collected in step 1 according to the above settings for storing the interaction records between users and items. The knowledge graph file is stored in the format of (head_id, tail_id, relation_id), representing the knowledge graph triple, that is, the head entity, the tail entity, and the relationship. The present invention first defines different relationship types, and then extracts the corresponding entities and relationships from the csv-format file. This knowledge graph file forms a visual tourism knowledge graph through Neo4j software. Finally, use the above rating data file in combination with the knowledge graph for multi-hop expansion to construct a semantically enhanced user interest set, that is, a ripple set file containing a multi-hop relationship set.
[0022] Further, step 3 of the present invention includes: Assume U = {u 1 , u 2 , … u M} and I = {i 1 , i 2 , … i N} represent the user and item sets respectively. The present invention obtains the user-item interaction matrix Y from the scoring data file obtained in step 2, where y ui = 1 indicates that there is an interaction between user u and item i, otherwise y ui = 0; The knowledge graph G is organized in the form of triples (h, r, t), where h is the head entity, t is the tail entity, and r is the relationship between them; In addition, there is an alignment set A for explaining the alignment relationship between items and entities in the knowledge graph, that is, A = {(i, e)|i ∈ I, e ∈ G}, where (i, e) indicates that item i can be aligned with entity e in the knowledge graph G.
[0023] First, according to the user-item interaction matrix Y, the present invention converts the items interacted with user u into an initial entity set through the alignment set A, and this process can be expressed as:
[0024]
[0025] Among them, u represents the user, i represents the item, e represents the entity in the knowledge graph, and A represents the alignment set. Then, for item i, find other items that have common interacting users with it to form a collaborative neighbor set Then obtain the initial entity set of item i through the alignment set A This process can be expressed as:
[0026]
[0027] where i u represents the item that has an interaction with user u. By converting the user-item interaction matrix into a set of entities in the knowledge graph and considering the collaborative neighbor set of items, the present invention enables the representation of items to incorporate collaborative information from other related items, further enhancing the embedded representation of items.
[0028] Then, starting from the obtained initial entity set, the present invention propagates information along the connection relationships in the knowledge graph. For user u and item i, use the entity set recursion formula for propagation, which is expressed as follows:
[0029]
[0030] where represents after the l-th step of propagation, represents after the (l - 1)-th step of propagation, o represents the entity set corresponding to (user or item), l represents the number of steps of propagation, where L is the maximum number of propagation steps on the knowledge graph G, and in each step of propagation, the entity set is expanded according to the relationships in the knowledge graph, and at the same time, a corresponding triple set is generated As l increases, the entity set is gradually expanded in the knowledge graph to capture the high-order interaction information of users and items, further enriching the embedded representations of users and items.
[0031] Further, step 4 of the present invention includes: taking the triple set obtained in step 3 as input, and for each triple (h, r, t), calculating the attention embedding a k of the tail entity. First, concatenate the head entity embedding with the relation embedding r k through a vector concatenation operation and then pass it through a fully connected layer (the learnable weight matrix is W 0 , and the bias is b 0 ) and the ReLU activation function to obtain the intermediate result z 0 . The calculation formula is as follows:
[0032]
[0033] z 0 Then pass it through a second fully connected layer (the learnable weight matrix is W 1 , and the bias is b 1 ) and the ReLU activation function to obtain z 1 . The calculation formula is as follows:
[0034] z 1 = ReLU(W 1 z 0 + b 1 )
[0035] Finally, pass it through a third fully connected layer (the learnable weight matrix is W 2 , and the bias is b 2 ) and use the Sigmoid as the activation function to obtain the attention weight. This process can be expressed as:
[0036]
[0037] where is the attention weight calculated by the neural network, and σ(·) represents the Sigmoid activation function.
[0038] The present invention uses the Softmax non-linear activation function to normalize the attention weight, and finally obtains that for each triple in the triple set of the l-th layer , its normalized attention weight is According to the normalized attention weight, perform a weighted sum of all the tail entity embeddings in the triple set of the l-th layer to obtain the representation of the triple set of this layer and the representation of the initial entity set (o represents the entity set corresponding to the user or item) and the initial entity representation for item i Incorporate, and finally form the representation set of item i and the representation set of user u
[0039] Furthermore, step 5 of the present invention includes: taking the user and item representation sets obtained in step 4 as input, using three aggregators (sum, pooling, and concatenation) to perform aggregation operations on the user and item representation sets, integrating multiple representation vectors into a single vector, and obtaining the updated final embedding representation vectors e of the user and item u and e i .
[0040] Further, step 6 of the present invention includes: according to the updated embedding representations of the user and the item e u and e i , using dot product calculation, we can recommend tourist projects. Specifically, by calculating their inner product The predicted user preference score for the item is obtained, which is used to determine the potential interaction possibility between the user and the non-interacted item.
[0041] To test the model effect, the present invention randomly selects 60% of the user-item data from the open source Yelp dataset as the training set, the remaining 20% of the data as the test set, and 20% as the validation set to adjust the model hyperparameters. The present invention mainly tests the performance of the model in terms of indicators such as recall rate (Recall@K) and normalized discounted cumulative gain (NDCG@K).
[0042] Beneficial effects:
[0043] 1. The present invention introduces a large language model (LLM) in the process of crawling tourism-related network data. With its powerful semantic understanding and natural language processing capabilities, the model can obtain valuable information content from static and dynamic web page data and convert it into structured data based on set prompt words, which is easy to store and update.
[0044] 2. The present invention constructs a tourism knowledge graph based on the information content obtained from web page data and has a visualization effect. The recommendation system uses the knowledge graph diffusion mechanism to assist in capturing high-order interactive information of tourism-related users and items, thereby improving the interpretability of the system and the recommendation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 The present invention is a flow chart of the method.
[0046] Figure 2 This is a visualization effect diagram of the tourism knowledge graph constructed by the present invention.
[0047] Figure 3 This is a comparison chart of the performance of the present invention based on the Yelp dataset on the training set, validation set, and test set.
[0048] Figure 4 This is the performance effect chart of the present invention on the Yelp dataset under different top-K recommendation rankings.
[0049] Identification instructions: Figure 4 (a), Figure 4 (b) are respectively the performance effect charts of the recall rate and the normalized discounted cumulative gain of the present invention on the Yelp dataset.
[0050] Figure 5 This is a comparison chart of the normalized discounted cumulative gain of the present invention and the method without introducing a knowledge graph on the Yelp dataset. Detailed implementation manners
[0051] The following further describes the present invention in detail with reference to the accompanying drawings of the specification.
[0052] It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.
[0053] As Figure 1 shown, the present invention provides a method for implementing a model-knowledge dual-driven tourism project recommendation system, and the method includes the following steps:
[0054] Step 1, obtaining and storing tourism-related datasets.
[0055] The present invention first uses web crawling technology and large language models to collect tourism-related datasets and store them. Specifically, first deploy the Scrapy framework to build the basic environment of the web crawler. In the initialization stage of the Scrapy project, construct the corresponding Spider (crawler class). In the Spider, define the initial URL (Uniform Resource Locator) of the Qunar website to be crawled. The Scrapy framework automatically initiates an HTTP request to obtain the HTML content of the corresponding web page. To obtain structured data that is convenient for subsequent construction of the knowledge graph and prevent problems caused by non-standard or dynamically modified web source HTML, the present invention pre-designs a prompt word to clearly indicate which information the large language model (LLMs) should extract. For example, in the present invention, it is necessary to crawl tourism data from the Qunar website, including information such as title (scenic spot name), address (scenic spot address), price (scenic spot ticket), hot_num (heat index), review (passenger rating), etc. Then input the HTML content into the LLMs to read, understand its semantics, and perform operations according to the prompt word and structured instructions, output the structured data and store it in CSV format.
[0056] Step 2: Data preprocessing and knowledge graph construction.
[0057] The present invention preprocesses the csv-formatted data obtained in the above step 1 to generate a rating data file, a knowledge graph file, and a ripple set file containing a multi-hop relationship set. Specifically, the rating data file is stored in the format of (user_id, item_id, label), where label represents the label of whether the user has positive feedback or interaction with the item. The present invention believes that when the user has no positive interaction with the item or is not interested in the item, label = 0; when the user has clear positive feedback (rating higher than 3) on the item, that is, is interested, label = 1. Obtain the rating data file from the csv-formatted file collected in step 1 according to the above settings for storing the interaction records between users and items. The knowledge graph file is stored in the format of (head_id, tail_id, relation_id), representing the knowledge graph triple, that is, the head entity, the tail entity, and the relationship. The present invention first defines different relationship types, and then extracts the corresponding entities and relationships from the csv-formatted file. This knowledge graph file forms a visual tourism knowledge graph through Neo4j software. Finally, use the above rating data file in combination with the knowledge graph for multi-hop expansion to construct a semantically enhanced user interest set, that is, a ripple set file containing a multi-hop relationship set.
[0058] Step 3: Heterogeneous information aggregation and dissemination based on the tourism knowledge graph.
[0059] Suppose U = {u1 , u 2 , … u M}, and I = {i 1 , i 2 , … i N}, respectively represent the user and item sets. The present invention obtains the user-item interaction matrix Y from the scoring data file obtained in step 2, where y ui = 1 indicates that there is an interaction between user u and item i, otherwise y ui = 0; the knowledge graph G is organized in the form of triples (h, r, t), where h is the head entity, t is the tail entity, and r is the relationship between them; in addition, there is an alignment set A for explaining the alignment relationship between items and entities in the knowledge graph, that is, A = {(i, e)|i ∈ I, e ∈ G}, where (i, e) means that item i can be aligned with entity e in the knowledge graph G.
[0060] First, according to the user-item interaction matrix Y, the present invention converts the items interacted with user u into an initial entity set through the alignment set A, and this process can be expressed as:
[0061]
[0062] where u represents the user, i represents the item, e represents the entity in the knowledge graph, and A represents the alignment set. Then, for item i, find other items that have common interacting users with it to form a collaborative neighbor set and then obtain the initial entity set of item i through the alignment set A This process can be expressed as:
[0063]
[0064] where i u represents the item that has an interaction with user u. The present invention converts the user-item interaction matrix into an entity set in the knowledge graph and considers the collaborative neighbor set of items, so that the representation of items integrates the collaborative information from other related items, further enhancing the embedded representation of items.
[0065] Then, starting from the obtained initial entity set, the present invention propagates information along the connection relationships in the knowledge graph. For user u and item i, use the entity set recursion formula for propagation, which is expressed as follows:
[0066]
[0067] where represents after the l-th step of propagation, It represents that after the propagation in the (l - 1)-th step, o represents the entity set corresponding to (users or items), l represents the number of propagation steps, where L is the maximum number of propagation steps on the knowledge graph G. In each step of propagation, the entity set is expanded according to the relationships in the knowledge graph, and at the same time, the corresponding triple set is generated. As l increases, the entity set is gradually expanded in the knowledge graph to capture the high-order interaction information of users and items, further enriching the embedding representations of users and items.
[0068] Step 4: Triple Relevance Weighting Based on the Attention Mechanism
[0069] The present invention takes the triple set obtained in step 3 as the input, and for each triple (h, r, t), calculates the attention embedding a of the tail entity k . First, concatenate the head entity embedding with the relation embedding r k through a vector concatenation operation , and then pass it through a fully connected layer (the learnable weight matrix is W 0 , and the bias is b 0 ) and the ReLU activation function to obtain the intermediate result z 0 . The calculation formula is as follows:
[0070]
[0071] z 0 Then pass it through the second fully connected layer (the learnable weight matrix is W 1 , and the bias is b 1 ) and the ReLU activation function to obtain z 1 . The calculation formula is as follows:
[0072] z 1 = ReLU(W 1 z 0 + b 1 )
[0073] Finally, pass it through the third fully connected layer (the learnable weight matrix is W 2 , and the bias is b 2 ) and use the Sigmoid as the activation function to obtain the attention weight. This process can be expressed as:
[0074]
[0075] where is the attention weight calculated by the neural network, and σ(·) represents the Sigmoid activation function.
[0076] The present invention uses the Softmax non - linear activation function to normalize the attention weights, and finally obtains the normalized attention weights for each triple in the l - th layer triple set in According to the normalized attention weights, perform a weighted sum of all the tail - entity embeddings in the l - th layer triple set to obtain the representation of the triple set at this layer And incorporate the representation of the initial entity set (where o represents the entity set corresponding to the user or item) and the initial entity representation for item i to finally form the representation set for item i and the representation set for user u
[0077] Step 5: Update the user and item embeddings.
[0078] The present invention takes the user and item representation sets obtained in Step 4 as inputs, and uses three aggregators (sum, pooling, concatenation) to perform aggregation operations on the user and item representation sets, integrating multiple representation vectors into a single vector to obtain the updated final embedding representation vectors e u and e i .
[0079] Step 6: Recommendation of tourist attractions and effect testing. According to the updated embedding representations e u and e i of the user and item, use the dot - product calculation to achieve tourist project recommendation. Specifically, calculate their inner product to obtain the predicted preference score of the user for the item, and this score is used to judge the potential interaction possibility between the user and the un - interacted item.
[0080] To test the model effect, the present invention randomly selects 60% of the user - item data from the open - source Yelp dataset as the training set, the remaining 20% as the test set, and 20% as the validation set to adjust the model hyperparameters. The present invention mainly tests the performance of the model on indicators such as recall rate (Recall@K), normalized discounted cumulative gain (NDCG@K), etc. The calculation formulas for these indicators are as follows:
[0081]
[0082] Among them, U test represents the users in the test set, |U test | represents the number of users in the test set, I u represents the items interacted by user u in the test set, |I u$n_{u}$ represents the number of items interacted with by user $u$ in the test set, and $\omega(k)$ is the $k$-th item in the recommendation list. The effects of the present invention will be further described in detail below in combination with simulation experiments.
[0083] 1. Simulation Conditions and Parameter Settings:
[0084] The simulation experiment of the present invention is carried out on a simulation platform of Python 3.8 and Pytorch 1.10. The computer CPU model is AMD EPYC 7302, and the GPU model is NVIDIA Geforce RTX 3090. The dataset used in the present invention is Yelp, the learning rate is set to 0.01, and the batch size is set to 1024.
[0085] 2. Simulation Content:
[0086] Figure 2 Shown is the performance of the AUC metric of the model used in the technical solution of the present invention on the training set, test set, and validation set based on the Yelp dataset. The Yelp dataset is a publicly available real-world dataset widely used in the research of recommendation systems, which contains merchants, user-merchant interaction records, and user reviews, mainly involving local business information such as restaurants, hotels, scenic spots, and services, and is suitable for the training and testing of tourism recommendation systems. The rating range of the dataset is 0 - 5 points. In the present invention, it is set that if the user's rating for a product is greater than 3 points, it is considered that the user is interested in the product, otherwise not interested. Figure 2 The abscissa is the number of training epochs, and the ordinate is the area under the curve (AUC), which is used to evaluate the ranking ability of the recommendation system of the present invention. The higher the AUC value, the stronger the ability of the model to distinguish positive and negative samples, that is, the stronger the ranking ability. From the changing trends of the three curves, it can be seen that as the number of training epochs increases, the AUC values of the training set, validation set, and test set show a gradually increasing and stabilizing trend, indicating that the tourism recommendation system proposed by the present invention has good convergence and generalization performance.
[0087] Figure 3 Shown is the visualization effect diagram of the tourism knowledge graph constructed by the present invention. The knowledge graph is constructed based on the dataset obtained by the web crawling technology combined with LLMs, and presents rich tourism-related information through an intuitive visualization effect. The figure contains multiple nodes and edges connecting these nodes. Each node represents a tourism-related entity, and the edge represents the relationship between entities. The knowledge graph can provide strong support for the tourism recommendation system and enhance the interpretability of the system.
[0088] Figure 4 Shown is the performance effect diagram of the present invention on the Yelp dataset under different top-K recommendation rankings. Figure 4(a) shows the change curve of the Normalized Discounted Cumulative Gain (NDCG) based on the Yelp dataset with respect to the top K recommended rankings. Figure 4 (b) shows the change curve of the Recall based on the Yelp dataset with respect to the top K recommended rankings. From observing the two figures, it can be seen that as the number of recommended items K increases, the model performs well in terms of both the Normalized Discounted Cumulative Gain and Recall. However, when the number of recommended items is small, there is still room for further optimization.
[0089] Figure 5 Comparison graph of the Normalized Discounted Cumulative Gain between the method of the present invention and the method without introducing a knowledge graph on the Yelp dataset. The results show that using a travel knowledge graph to assist in recommendations can significantly improve the Normalized Discounted Cumulative Gain of the recommendation results, that is, it can better rank the items that the user may be interested in at the front of the recommendation list, thereby improving the recommendation quality, proving the effectiveness of the solution of introducing a knowledge graph in the present invention.
[0090] It should be understood that the above specific embodiments of the present invention are only used for exemplary illustration or explanation of the principle of the present invention, and do not constitute a limitation to the present invention. Therefore, any modifications, equivalent replacements, improvements, etc. made without departing from the spirit and scope of the present invention shall be included within the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modification examples falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.
Claims
1. A model-knowledge dual-driven tourism project recommendation system, characterized in that: The system includes a data collection module, a knowledge graph construction module, a heterogeneous information aggregation / propagation module, an attention weighting module and a prediction module; The data collection module introduces large language models (LLMs) into web crawling technology and uses key information prompts to obtain the triple information required to build a tourism knowledge graph, including information such as scenic spot names, scenic spot tickets, and passenger ratings. It is the basis for subsequent modules; The knowledge graph construction module preprocesses the information collected by the data collection module and constructs a tourism-related knowledge graph based on the preprocessed data files, so that the subsequent recommendation system model can be combined with the knowledge graph for semantic enhancement; The heterogeneous information aggregation / propagation module aggregates the user-item interaction information into entity information in the knowledge graph based on the user-item interaction matrix and knowledge graph obtained by preprocessing, and then propagates it along the connection relationship in the knowledge graph to capture the high-order semantic information of users and items based on the knowledge graph, further enriching the representation of users and items; The attention weighting module operates on triples in the knowledge graph and uses the attention mechanism to assign different weights to the tail entities according to the differences in the head entities and relations. The prediction module is the output module of the system, which aggregates and performs dot product operations on the weighted user and item representation vectors to achieve travel project recommendations.
2. A method for implementing a model-knowledge dual-driven tourism project recommendation system, characterized in that: The method comprises the following steps: Step 1: Acquisition and storage of tourism-related datasets; Use the scrapy framework and large language model in Python to accurately extract effective tourism-related information from Qunar.com web pages and convert it into structured data and store it in text files; Step 2: Data preprocessing and knowledge graph construction; Preprocess the extracted dataset to obtain the scoring data file, knowledge graph file, and ripple collection file containing multi-hop relationship sets; Step 3: Aggregation and dissemination of heterogeneous information based on tourism knowledge graph; Utilize the diffusion mechanism of knowledge graph to capture high-level semantic information and enhance the embedding representation of users and items; Step 4: Triple relevance weighting based on attention mechanism; The attention mechanism is used to assign different weights to the tail entities according to the differences in the head entities and relations, so as to weight the influence of different entities (users, items) and relations and generate an embedded representation for each hop. Step 5: User and item embedding update, aggregating the multi-hop embedding of users or items to generate the final embedding representation; Step 6: Recommend tourist points of interest and test the effect.
3. The method for implementing a model-knowledge dual-driven tourism project recommendation system according to claim 2 is characterized in that: Step 1 of the method includes: first deploying the Scrapy framework to build a basic environment for web crawlers, building a corresponding Spider (crawler class) in the initialization phase of the Scrapy project, defining the initial URL (Uniform Resource Locator) of the Qunar.com web page to be crawled in the Spider, and automatically initiating an HTTP request by the Scrapy framework to obtain the HTML format content of the corresponding web page. In order to obtain structured data that is convenient for the subsequent construction of a knowledge graph and to prevent problems caused by irregular design or dynamic modification of the source HTML of the web page, the present invention pre-designs a prompt word to clearly indicate which information the large language model (LLMs) should extract, and needs to crawl the travel data of the Qunar.com website, including title (attraction name), address (attraction address), price (attraction ticket), hot_num (heat index), and review (passenger rating) information, and then input the HTML content into the LLMs for reading, understanding its semantics, and operating according to the prompt word and structured instructions, outputting the structured data and storing it in CSV format.
4. The method for implementing a model-knowledge dual-driven tourism project recommendation system according to claim 2 is characterized in that: Step 2 of the method includes: preprocessing the data in csv format obtained in the above step 1 to generate a scoring data file, a knowledge graph file, and a ripple collection file containing a multi-hop relationship set. The scoring data file is stored in the format of (user_id, item_id, label), where label represents a label indicating whether the user has positive feedback or interaction with the item. The present invention believes that when the user does not have a positive interaction with the item or is not interested in the item, label=0; when the user has clear positive feedback on the item (the score is higher than 3), that is, is interested, label=1. According to the above settings, the scoring data file is obtained from the csv format file collected in step 1 to store the interaction records between users and items. The knowledge graph file is stored in the format of (head_id, tail_id, relation_id), which represents the knowledge graph triple, namely the head entity, the tail entity and the relationship. First, different relationship types are defined, and then the corresponding entities and relationships are extracted from the csv format file. The knowledge graph file is used to form a visual tourism knowledge graph through Neo4j software. Finally, the above-mentioned rating data file is combined with the knowledge graph for multi-hop expansion to construct a semantically enhanced user interest set, that is, a ripple collection file containing a multi-hop relationship set.
5. The method for implementing a model-knowledge dual-driven tourism project recommendation system according to claim 2 is characterized in that: Step 3 of the method includes: assuming that U = {u1, u2, ... M } and I={i1,i2,…i N } represent the user and item set respectively. The present invention obtains the user-item interaction matrix Y from the rating data file obtained in step 2, where y ui =1 means that user u interacts with item i, otherwise y ui = 0; the knowledge graph G is organized in the form of triples (h, r, t), where h is the head entity, t is the tail entity, and r is the relationship between them; in addition, there is an alignment set A used to illustrate the alignment relationship between items and entities in the knowledge graph, that is, A = {(i, e) | i∈I, e∈G}, where (i, e) means that item i can be aligned with entity e in the knowledge graph G; First, according to the user-item interaction matrix Y, the items that have interacted with user u are converted into the initial entity set through the alignment set A. The process can be expressed as: Among them, u represents the user, i represents the item, e represents the entity in the knowledge graph, and A represents the alignment set. Then, for item i, find other items that have common interacting users with it to form a collaborative neighbor set Then, the initial entity set of item i is obtained by aligning set A The process can be expressed as: where i u Representing items that interact with user u, by converting the user-item interaction matrix into an entity set in the knowledge graph and considering the collaborative neighbor set of the item, the representation of the item is integrated with the collaborative information from other related items, further enhancing the embedded representation of the item; Then, the present invention starts from the obtained initial entity set and propagates information along the connection relationship in the knowledge graph. For user u and item i, the entity set recursive formula is used for propagation, which is expressed as follows: in represents after the lth step of propagation, represents after the l-1th step of propagation, o represents the entity set corresponding to (user or item), l represents the number of propagation steps, where L is the maximum number of propagation steps on the knowledge graph G. Each step of propagation will expand the entity set according to the relationship in the knowledge graph and generate the corresponding triple set. As l increases, the entity set is gradually expanded in the knowledge graph to capture the high-order interaction information of users and items, further enriching the embedding representation of users and items.
6. The method for implementing a model-knowledge dual-driven tourism project recommendation system according to claim 2 is characterized in that: Step 4 of the method comprises: As input, for each triple (h, r, t), the attention embedding a of the tail entity is calculated k , first embed the head entity With relational embedding k Perform vector concatenation operations Then, after a fully connected layer (the learnable weight matrix is W0 and the bias is b0) and the ReLU activation function, the intermediate result z0 is obtained. The calculation formula is as follows: z0 then passes through the second fully connected layer (the learnable weight matrix is W1 and the bias is b1) and the ReLU activation function to obtain z1. The calculation formula is as follows: z1=ReLU(W1z0+b1) Finally, after passing through the third fully connected layer (the learnable weight matrix is W2, the bias is b2) and using Sigmoid as the activation function, the attention weight is obtained. The process can be expressed as: in is the attention weight calculated by the neural network, σ(·) represents the Sigmoid activation function; The attention weights are normalized using the Softmax nonlinear activation function, and finally the l-th layer triple set is obtained For each triple in , the normalized attention weight is According to the normalized attention weights, all tail entity embeddings in the l-th layer triple set are weighted summed to obtain the representation of the triple set of this layer And the representation of the initial entity set (o represents the entity set corresponding to the user or item) and the initial entity representation for item i Incorporate, and finally form the representation set of item i and the representation set of user u 7. The method for implementing a model-knowledge dual-driven tourism project recommendation system according to claim 2 is characterized in that: Step 5 of the method includes: taking the user and item representation sets obtained in step 4 as input, using three aggregators (sum, pooling, and concatenation) to perform aggregation operations on the user and item representation sets, integrating multiple representation vectors into a single vector, and obtaining the updated final embedding representation vectors e of the user and item u and e i .
8. The method for implementing a model-knowledge dual-driven tourism project recommendation system according to claim 2 is characterized in that: Step 6 of the method comprises: according to the updated embedding representations of the user and the item u and e i , using dot product calculation, to achieve tourism project recommendation, by calculating their inner product Get the predicted user preference score for the item, which is used to determine the potential interaction possibility between the user and the uninteracted item; To test the model effect, 60% of the user-item data were randomly selected from the open source Yelp dataset as the training set, the remaining 20% of the data were used as the test set, and 20% were used as the validation set to adjust the model hyperparameters and test the model's performance in indicators such as recall (Recall@K) and normalized discounted cumulative gain (NDCG@K).
Citation Information
Cited By
Article recommendation method based on non-negative sampling attention gating graph convolutional network
CN120765350A