A method for recommending points of interest based on time-series knowledge maps
By constructing dynamic and static knowledge maps and using a heterogeneous mutual attention mechanism to fuse multimodal information, the method addresses the limitations of existing POI recommendation techniques, achieving improved personalized recommendations.
Patent Information
- Application Number
- JP2024184448
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2024-06-24
- Filing Date
- 2024-10-18
- Publication Date
- 2025-05-12
- Estimated Expiration
- 2044-10-18
AI Technical Summary
Existing methods for recommending points of interest (POIs) fail to fully utilize multimodal information and effectively integrate behavioral pattern information from user trajectories, leading to suboptimal personalized recommendations.
The proposed method constructs dynamic time series knowledge maps and static group knowledge maps based on user historical behavior trajectories, using a heterogeneous mutual attention mechanism to fuse multimodal information, including user comments, to predict next interest points.
This approach effectively learns user dynamic behavior preferences and static features, providing personalized interest point recommendations by integrating multiple heterogeneous information sources and user emotion tendencies.
Smart Images

Figure 0007674640000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a personalized recommendation technology field, and in particular to a method for recommending points of interest based on a time-series knowledge map. [Background technology]
[0002] The breakthrough development of mobile Internet has made users all over the world more and more connected, and people can share their daily activities and convey the joy of life on social platforms based on geographic location. The large amount of user interaction information also prompted the emergence of geotagged datasets such as Foursquare, Gowalla, and Yelp, providing new vitality and impetus to point-of-interest (POI) recommendation. POI recommendation can utilize the user's past check-in information to predict where the user may visit next, and at the same time, it can utilize multi-modal information such as time, geographic location, POI category, and social relationship to achieve better prediction ability, providing convenience for users' daily outings.
[0003] Among the existing technologies, many methods based on map neural networks have achieved good results by learning global user and POI features. However, most of the existing methods only use one of the elements such as trajectory information, geographic location, social network, and user comments, and do not fully utilize the advantages of multi-modal information in real situations. In addition, the behavior pattern information contained in the user's past trajectory is not effectively separated, which is disadvantageous to learning the user's personalized preferences.
[0004] Therefore, how to build a proven model to establish multimodal information relationships and fully fuse them is currently a technical problem that needs to be urgently solved. Summary of the Invention
[0005] In order to overcome the shortcomings of the prior art, the present invention provides an interest point recommendation method based on time-series knowledge map, which can effectively solve the above problems.
[0006] The technical solutions specifically adopted in the present invention are as follows: In a first aspect, the present invention provides a method for recommending points of interest based on a time-series knowledge map, which includes the following steps: S1, construct a dynamic time series knowledge map and a static group knowledge map based on the complete historical behavior trajectory of all users. The dynamic time series knowledge map is a map set consisting of dynamic relationship knowledge maps of different historical time slices, each dynamic relationship knowledge map records dynamic relationships between all users and points of interest in a historical time slice, the dynamic relationships include access relationships for recording users' access behaviors of points of interest, and follow relationships for recording users' neighboring access behaviors of different points of interest. The static group knowledge map records static relationships between all users and points of interest in all historical time slices. The static relationships include social relationships for recording friendship relationships between users, location relationships for recording spatial regions in which points of interest are located, adjacent relationships for recording whether different points of interest belong to neighboring points, category relationships for recording interest point categories to which points of interest belong, and group relationships for recording that users are grouped according to the points of interest accessed and the spatial regions accessed. S2, obtain a substring of the historical behavior trajectory of the target user before the predicted waiting time, and sequentially extract the user comment text of each interest point accessed by the user from it, and use an aspect-based sentiment analysis module constructed based on the pre-training model to perform word embedding on the user comment text, and stitch together the sentiment embeddings of all the user comment texts to obtain a user comment sentiment embedding sequence. S3, the historical behavior trajectory substring, the dynamic time series knowledge map, the static group knowledge map and the user comment emotion embedding sequence are input into an interest point recommendation model, and the embedding module first performs a word embedding operation on the input data, and then the multimodal knowledge fusion module performs a fusion operation on the dynamic time series knowledge map and the static group knowledge map based on a heterogeneous mutual attention mechanism, and fuses the interest points, the user and other multimodal information to obtain an interest point fusion feature representation and a user fusion feature representation, and finally the decoding module combines the interest point fusion feature representation and the user fusion feature representation, and then inputs them into a cascaded recurrent neural network and a multi-layer sensor to predict the interest points that the target user can access at the next time.
[0007] As a preferred embodiment of the first aspect, in the dynamic time series knowledge map, the access relationship is recorded by a four-element tuple consisting of a user, an access relationship identifier, an access interest point, and an access time, and the following relationship is recorded by a four-element tuple consisting of a preceding access location, a following relationship identifier, a subsequent access location, and an access time.
[0008] In the static group knowledge map, social relations are recorded by a three-element tuple consisting of a user, a social relation identifier, and a user, location relations are recorded by a three-element tuple consisting of an interest point, a location relation identifier, and a geohash-5 spatial region where the interest point is located, adjacent relations are recorded by a three-element tuple consisting of an interest point, an adjacent relation identifier, and an interest point, category relations are recorded by a three-element tuple consisting of an interest point, a category relation identifier, and an interest point category to which the interest point belongs, and group relations are recorded by a three-element tuple consisting of a user, a group relation identifier, and a user group to which the interest point belongs. Here, the user group is divided into two categories, one is an interest point level group obtained by performing cluster division based on the interest points accessed by the user, and the other is an area level group obtained by performing cluster division based on the geohash-5 spatial region to which the user belongs.
[0009] As a preferred embodiment of the first aspect, the aspect-based sentiment analysis module is obtained by cascading pre-trained DistilBERT models into one multi-class classifier and further fine-tuning the overall configuration. In the aspect-based sentiment analysis module, an embedding representation is first generated for the user comment text by the DistilBERT model, and then the embedding representation is input into a multi-class classifier to obtain the comment dimension corresponding to the user comment text and the positive and negative scores for each comment dimension, and the positive and negative scores for all comment dimensions are combined to output as the sentiment embedding corresponding to the user comment text.
[0010] Furthermore, the comment dimension includes three dimensions: product, price, and service.
[0011] As a preferred embodiment of the first aspect, the process flow in the multimodal knowledge fusion module is as follows. S31, the dynamic time series knowledge map and the static group knowledge map are input into the heterogeneous map attention network respectively to perform information fusion, and a fused dynamic time series knowledge map and a static group knowledge map are obtained. S32, arrange all interest points in the historical behavior trajectory substring in the order of user access, sequentially extract hidden layer vectors corresponding to each interest point from the fused dynamic time series knowledge map, construct a user behavior trajectory embedding with global time slice information, sequentially extract hidden layer vectors corresponding to each interest point from the fused static group knowledge map, construct a user behavior trajectory embedding with global static information, and extract interest point level group features and area level group features from the fused static group knowledge map. S33, the user comment emotion embedding sequence is used as a query, and the user behavior trajectory embedding with global time slice information and the original user behavior trajectory embedding are fused through an attention mechanism to obtain a fused user behavior trajectory embedding. The fused user behavior trajectory embedding is used as a value, the user behavior trajectory embedding with global static information is used as a query, and the interest point level group feature is used as a key to input into the Encoder module of the Transformer model for fusion encoding to obtain an interest point fusion feature representation. S34: The interest point level group features and the area level group features are connected and fused, and the obtained fused group features are used as keys, the user embedded feature representations of all users in the fused static group knowledge map are used as queries, and the user embedded feature representations of all users in the original static group knowledge map are used as values, which are input into the Encoder module of the Transformer model to perform fusion encoding and obtain a user fused feature representation.
[0012] As an advantage of the first aspect, the interest point recommendation model needs to be pre-optimized by a total loss function obtained by weighting an interest point prediction loss and a static map loss.
[0013] JPEG0007674640000002.jpg62170
[0014] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides a method for recommending points of interest based on a time series knowledge map for the problem of recommending next points of interest in a multimodal scene. The method respectively constructs a dynamic time series knowledge map and a static group knowledge map based on the user's historical action trajectory, and learns the user's dynamic action preferences and static features. The present invention uses a heterogeneous mutual attention mechanism to aggregate information on the dynamic time series knowledge map and the static group knowledge map, learns the multi-dimensional heterogeneous information in the knowledge map, and introduces a multimodal knowledge fusion module to cross-learn user features and points of interest features under the assistance of high-quality semantic information reflecting the user's emotional tendency extracted from the user's comment text by the aspect-based sentiment analysis module, effectively addresses the problem of information fusion under real scenes, and provides guidance for recommending next points of interest. The present invention has the characteristics of high accuracy and strong scalability, and can timely grasp the direction of user actions, and provides technical support for realizing personalized user action trajectory prediction. [Brief description of the drawings]
[0015] [Figure 1] FIG. 2 is a schematic diagram illustrating the steps of a method for recommending points of interest based on a time-series knowledge map in an embodiment of the present invention; [Diagram 2] FIG. 1 is a schematic diagram of a network architecture of an interest point recommendation model in an embodiment of the present invention. [Diagram 3] FIG. 2 is a schematic diagram of a dynamic time-series knowledge map according to an embodiment of the present invention; [Figure 4] FIG. 2 is a schematic diagram of a dynamic time-series knowledge map of user B as an example and a user behavior trajectory embedding extracted therefrom in an embodiment of the present invention. [Diagram 5] 1 is a schematic diagram of a configuration of a computer electronics device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0017] In a preferred embodiment of the present invention, as shown in FIG. 1, the method for recommending points of interest based on the above-mentioned time-series knowledge map is specifically realized by the following steps S1 to S3. The method for recommending points of interest of the present invention uses a deep learning method to build a network frame based on the user's past check-in data and multi-modal information, extracts the user's favorite features as shown in FIG. 2, and accurately predicts the place that the user is most likely to visit at present. The present invention can use the dynamic time-series knowledge map and the static group knowledge map to learn and organically integrate the trajectory sequence and the multi-modal relationship, respectively. At the same time, the dynamic time-series knowledge map effectively divides the behavior mode information contained in the user's history trajectory, and objectively learns the inherent mode and evolutionary relationship of each time slice, thereby realizing the next interest point prediction that matches the user's personalized preferences.
[0018] A specific method for implementing each of steps S1 to S3 will be described in detail below. S1, construct a dynamic time series knowledge map and a static group knowledge map based on the complete historical behavior trajectory of all users. The dynamic time series knowledge map is a map set composed of dynamic relationship knowledge maps of different historical time slices, and each dynamic relationship knowledge map records the dynamic relationship between all users and points of interest in the historical time slice. The dynamic relationship includes an access relationship for recording the access behavior of the user to the points of interest, and a follow relationship for recording the neighboring access behavior of the user to different points of interest. The static group knowledge map records the static relationship between all users and points of interest in all historical time slices, and the static relationship includes a social relationship for recording the friendship relationship between users, a location relationship for recording the spatial region where the points of interest are located, an adjacent relationship for recording whether different points of interest belong to a neighboring point, a category relationship for recording the interest point category to which the points of interest belong, and a group relationship for recording that users are grouped according to the points of interest accessed and the spatial region accessed.
[0019] The construction of the two kinds of maps requires the complete historical behavior trajectories of all users, and the complete historical behavior trajectories of each user can be obtained from social platforms that record the user's access behavior to POIs, such as Foursquare, Gowalla, Yelp, etc. Note that the construction of the two kinds of maps requires the trajectories of all users on the social platform, and the time span of one complete historical behavior trajectory is a specified history period. The specific length of the history period can be reasonably selected based on the actual data situation, for example, the most recent six months or the most recent year can be selected.
[0020] The specific method of constructing a knowledge map belongs to the prior art, and the relationships between entities can be recorded by rationally designed tuples.
[0021] In an embodiment of the present invention, for two types of dynamic relationships existing in the dynamic time-series knowledge map, the access relationship can be recorded by a four-element tuple consisting of a user, an access relationship identifier, an access interest point, and an access time, and the following relationship can be recorded by a four-element tuple consisting of a preceding access location, a following relationship identifier, a subsequent access location, and an access time.
[0022] JPEG0007674640000003.jpg138170
[0023] In the embodiment of the present invention, constructing a static group knowledge map simply requires constructing all data in a complete time span as a single map, without the need to consider different historical time slices as in the case of constructing a dynamic time series knowledge map. For the five types of static relationships present in the static group knowledge map here, the social relationship is recorded by a three-element tuple of user, social relationship identifier, and user, the location relationship is recorded by a three-element tuple of interest point, location relationship identifier, and the located Geohash-5 spatial area, the adjacent relationship is recorded by a three-element tuple of interest point, adjacent relationship identifier, and interest point, the category relationship is recorded by a three-element tuple of interest point, category relationship identifier, and the belonging interest point category, and the group relationship is recorded by a three-element tuple of user, group relationship identifier, and the belonging user group. Here, the user group is divided into two categories, one is an interest point level group obtained by performing cluster division based on the interest points accessed by the user, and the other is an area level group obtained by performing cluster division based on the Geohash-5 spatial area accessed by the user. The cluster division here essentially involves dividing users into groups based on the interest points and Geohash-5 spatial regions they access; that is, users who access the same interest point are placed in an interest point level group for that interest point, and users who access the same Geohash-5 spatial region are placed in a region level group for that Geohash-5 spatial region.
[0024] JPEG0007674640000004.jpg119170
[0025] The Geohash-5 spatial region adopted by the present invention can be obtained by encoding the entire geographic space by the Geohash algorithm, and the algorithm is an address encoding method, which can encode two-dimensional spatial latitude and longitude data into a string, which belongs to the prior art and can be calculated directly by adopting the conventional correlation function or program. The basic steps of the algorithm are as follows: first, convert the longitude and latitude into binary according to different accuracy requirements, then combine the longitude and latitude, where the longitude and latitude occupy even digits and the latitude occupy odd digits, and finally encode the binary string by Base32. The longer the encoding, the smaller the display range and the more accurate the position, and the specific value of the accuracy can be optimized according to the actual situation, and in the embodiment of the present invention, the accuracy is 5, that is, the encoding length is 5.
[0026] The dynamic time series knowledge map is constructed based on user entities, interest point entities and the dynamic relationships between them, and can be used to learn the user's behavior patterns in each time slice and the behavior preferences that change over time. The static group knowledge map is constructed based on the static relationships between entities, and can be used to learn multi-dimensional heterogeneous information and stable feature dependencies that do not change over time. Both can provide the user's preference information for future interest point selection from different dimensions.
[0027] S2, obtain a substring of the historical behavior trajectory of the target user before the predicted waiting time, and sequentially extract the user comment text of each interest point accessed by the user from it, and use an aspect-based sentiment analysis module constructed based on the pre-training model to perform word embedding on the user comment text, and stitch together the sentiment embeddings of all the user comment texts to obtain a user comment sentiment embedding sequence.
[0028] It should be noted that the historical behavior trajectory substring before the predicted waiting time of the target user refers to a behavior trajectory consisting of a series of interest points recently accessed by the target user before the predicted waiting time for which interest point recommendations need to be made, and the number of interest points constituting the historical behavior trajectory substring can be adjusted according to actual needs, and in the embodiment, the most recent 20 interest points can be adopted, that is, the length of the historical behavior trajectory substring is 20.
[0029] Theoretically, the above aspect-based sentiment analysis module can be trained and fine-tuned based on any pre-trained language model. In an embodiment of the present invention, considering the requirements for model scale and execution speed in practical execution scenarios, the aspect-based sentiment analysis module preferably adopts a DistilBERT model pre-trained on a large corpus to construct the aspect-based sentiment analysis module, and after the pre-trained DistilBERT model, one more multi-classifier needs to be cascaded, and then the cascaded models are fine-tuned together on the sentiment analysis dataset to obtain the aspect-based sentiment analysis module. The processing flow of the aspect-based sentiment analysis module is as follows: First, an embedding expression is generated for the input user comment text by the DistilBERT model, and then the embedding expression is input into the multi-classifier, and the comment dimension corresponding to the user comment text and the positive and negative scores (which are two-dimensional vectors that record positive and negative scores) in each comment dimension are obtained, and the positive and negative scores in all comment dimensions are combined to output as the sentiment embedding corresponding to the user comment text.
[0030] The sentiment analysis dataset adopted in the fine-tuning process is obtained by manually labeling the task of the present invention. The training samples in the sentiment analysis dataset include user comment texts for points of interest, and truth labels for the comment dimensions and comment positive / negative scores of the user comment texts. After cascading the DistilBERT model and the multi-classifier, supervised learning is performed on the sentiment analysis dataset, and fine-tuning is completed after convergence. The specific comment dimensions can be designed according to the actual situation of the user comment text data actually collected. For example, in a general review site, the dimensions related to the user comment texts for points of interest cover three dimensions of product, price, and service, so these three dimensions can be considered as the comment dimensions output by the multi-classifier.
[0031] The product in the comment dimension above refers to the service product provided by the merchant corresponding to the point of interest, for example, for a restaurant, the product is food, and for an amusement park, the product is an attraction.
[0032] In an embodiment of the present invention, the entire process of pre-training and fine-tuning may be specifically implemented by the following steps. S21,We pre-train the DistilBERT pre-trained model on large-scale corpora,,i.e., the BookCorpus dataset and the English Wikipedia data.
[0033] JPEG0007674640000005.jpg88170
[0034] JPEG0007674640000006.jpg29170
[0035] JPEG0007674640000007.jpg61170
[0036] S3, the target user's historical behavior trajectory substring before the predicted waiting time, the dynamic time series knowledge map, the static group knowledge map and the user comment emotion embedding sequence are input into the interest point recommendation model, the embedding module first performs a word embedding operation on the input data, and then the multimodal knowledge fusion module performs a fusion operation on the dynamic time series knowledge map and the static group knowledge map based on the heterogeneous mutual attention mechanism, and fuses the interest points, user and other multimodal information to obtain an interest point fusion feature representation and a user fusion feature representation, and finally the decoding module combines the interest point fusion feature representation and the user fusion feature representation, and then inputs them into the cascaded recurrent neural network and multi-layer sensor to predict the interest points accessible by the target user at the next time.
[0037] JPEG0007674640000008.jpg45170
[0038] In an embodiment of the present invention, the multi-modal knowledge fusion module performs fusion operations on the dynamic time-series knowledge map and the static group knowledge map respectively, and can fuse POIs, users and other multi-modal information to obtain richer context information. The processing flow in the multi-modal knowledge fusion module is as follows:
[0039] JPEG0007674640000009.jpg70170
[0040] JPEG0007674640000010.jpg171170
[0041] S33, the user comment emotion embedding sequence is used as a query, and the user behavior trajectory embedding with global time slice information and the original user behavior trajectory embedding are fused through an attention mechanism to obtain a fused user behavior trajectory embedding. The fused user behavior trajectory embedding is used as a value, the user behavior trajectory embedding with global static information is used as a query, and the interest point level group feature is used as a key to input into the Encoder module of the Transformer model for fusion encoding to obtain an interest point fusion feature representation.
[0042] The query, key, and value in the present invention are Q (Query), K (Key), and V (Value) in the attention mechanism, respectively.
[0043] JPEG0007674640000011.jpg94170
[0044] JPEG0007674640000012.jpg57170
[0045] S34: The interest point level group features and the area level group features are connected and fused, and the obtained fused group features are used as keys, the user embedded feature representations of all users in the fused static group knowledge map are used as queries, and the user embedded feature representations of all users in the original static group knowledge map are used as values, which are input into the Encoder module of the Transformer model to perform fusion encoding and obtain a user fused feature representation.
[0046] JPEG0007674640000013.jpg68170
[0047] JPEG0007674640000014.jpg86170
[0048] JPEG0007674640000015.jpg67170
[0049] JPEG0007674640000016.jpg85170
[0050] JPEG0007674640000017.jpg27170
[0051] The specific model training process belongs to the prior art. By combining the total loss function and the optimizer, the learnable parameters can be continuously optimized, and the interest point recommendation model after optimization can be used for actual inference.
[0052] It should be noted that the method steps shown in S1 to S3 above can be essentially realized in the form of a computer program.
[0053] In order to facilitate understanding of the essence of the present invention, the detailed implementation process and technical effects of the interest point recommendation method based on the time-series knowledge map shown in S1 to S3 above on a specific dataset will be illustrated below through further specific examples.
[0054] [Example] The steps of this embodiment are the same as the method of recommending points of interest based on the time-series knowledge map shown in the above-mentioned steps S1 to S3, and will not be repeated here, but will mainly show the specific data set, some specific parameter settings, and implementation results of this embodiment.And, for the sake of convenience, the method abbreviated as steps S1 to S3 below will be the method of the present invention, and the interest point recommendation model used will be described as MINet.
[0055] The original data used in this embodiment are three widely used real scene datasets: Foursquare, Gowalla, and Yelp. The Gowalla dataset has 3,300,986 check-in actions occurred in a total of 121,851 locations by 52,979 users. The Foursquare dataset has 9,447,873 check-in actions occurred in a total of 69,005 locations by 46,065 users. The Yelp dataset has 632,476 check-in actions occurred in a total of 8,696 locations by 9,627 users. Since the Foursquare and Gowalla datasets do not contain user comment information, this embodiment also builds a default version model based on MINet without inputting user comment sentiment embedding sequences, which is denoted as MINet-Revised, and is used as one of the control methods in subsequent experiments.
[0056] JPEG0007674640000018.jpg36170
[0057] Furthermore, the experiment in this embodiment compares the method of the present invention with several conventional prediction methods. The conventional prediction methods are as follows: (1) FPMC: a conventional Markov chain model based on user personalized behavior; (2) RNN: a recurrent neural network for prediction based on historical trajectory sequences; (3) DeepMove: a recurrent neural network based on the fusion of attention information and trajectory sequences; (4) STAN: a two-layer attention network based on space-time attention; (5) TiSASRec: an attention network based on sequence position and time interval; (6) Flashback: a recurrent neural network based on context information of past hidden layer and current hidden layer; (7) GETNext: a Transformer model that predicts user behavior based on global transition probability; (8) a model that learns POI transition relations based on space-time knowledge map; (9) MARAN: a model based on aggregation of local central trajectories and user behavior patterns. In this embodiment, the mean precision (Acc@K) and mean reciprocal rank (MRR) are used as evaluation indices for prediction models. Acc@K calculates the ratio of true positive samples in the top K predicted samples. In the experiment, K = {5, 10}. MRR can reflect the overall performance of the recommendation and pays more attention to the prediction ranking.
[0058] The final experimental results are shown in Table 1, which shows that the MINet in the method of the present invention has better results than the control model in the Yelp dataset. Specifically, in the Yelp dataset, MINet has 9.86%, 6.48%, and 7.21% improvement over the best-represented control model MARAN in Acc@5, Acc@10, and MRR indicators, respectively. In addition, the default version of MINet-Revised has better results than the control model in Gowalla, with improvements of 2.56%, 0.90%, and 3.45%, respectively. And in the Foursquare dataset, it is slightly inferior to the best-represented control model MARAN, with an average difference of 1.40%. In particular, in the Yelp dataset, MINet has 47.17%, 45.24%, and 33.53% improvement over the default version in each indicator, respectively, which shows the effectiveness of the method of the present invention. [Table 1] Comparison of experimental results between the method of the present invention and the control method JPEG0007674640000019.jpg89165
[0059] The above embodiment is one of the preferred solutions of the present invention, and does not limit the present invention. Those skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by adopting the method of equivalent replacement or equivalent conversion is included in the protection scope of the present invention.
Claims
1. A method for recommending points of interest based on a time-series knowledge map, which is realized in the form of a computer program and executed by a computer electronics device, comprising: S1: constructing a dynamic time series knowledge map and a static group knowledge map based on the complete historical behavior trajectories of all users, the dynamic time series knowledge map being a map set consisting of dynamic relationship knowledge maps of different historical time slices, each dynamic relationship knowledge map recording dynamic relationships between all users and points of interest in the historical time slice, the dynamic relationships including access relationships for recording the access behavior of the users to the points of interest and following relationships for recording the neighboring access behavior of the users to different points of interest, the static group knowledge map recording static relationships between all users and points of interest in all historical time slices, the static relationships including social relationships for recording friend relationships between users, location relationships for recording the spatial regions in which the points of interest are located, adjacent relationships for recording whether different points of interest belong to neighboring points, category relationships for recording the interest point categories to which the points of interest belong, and group relationships for recording that users are grouped according to the points of interest accessed and the spatial regions accessed; S2. Obtain a substring of the target user's historical behavior trajectory before the predicted waiting time, and sequentially extract the user comment text of each interest point accessed by the user from the substring. Use an aspect-based sentiment analysis module constructed based on the pre-training model to perform word embedding on the user comment text, and stitch together the sentiment embeddings of all the user comment texts to obtain a user comment sentiment embedding sequence; S3, the historical behavior trajectory substring, the dynamic time series knowledge map, the static group knowledge map and the user comment emotion embedding sequence are input into an interest point recommendation model, and the embedding module first performs a word embedding operation on the input data, and then the multimodal knowledge fusion module performs a fusion operation on the dynamic time series knowledge map and the static group knowledge map based on the heterogeneous mutual attention mechanism, and fuses the interest point, the user and the multimodal information to obtain an interest point fusion feature representation and a user fusion feature representation, and finally the decoding module combines the interest point fusion feature representation and the user fusion feature representation, and then inputs them into the cascaded recurrent neural network and the multi-layer sensor to predict the interest points that the target user can access at the next time; The multimodal information includes dynamic relations including the access relations and the following relations, and static relations including the social relations, the location relations, the adjacent relations, the category relations, and the group relations. The method for recommending points of interest based on a time-series knowledge map is characterized by the following.
2. In the dynamic time-series knowledge map, an access relationship is recorded by a four-element tuple consisting of a user, an access relationship identifier, an access interest point, and an access time, and a follow-up relationship is recorded by a four-element tuple consisting of a preceding access location, a follow-up relationship identifier, a subsequent access location, and an access time; The method for recommending points of interest based on a time-series knowledge map according to claim 1, characterized in that in the static group knowledge map, social relationships are recorded by a three-element tuple consisting of a user, a social relationship identifier, and a user; location relationships are recorded by a three-element tuple consisting of a point of interest, a location relationship identifier, and a Geohash-5 spatial region in which the point of interest is located; adjacent relationships are recorded by a three-element tuple consisting of a point of interest, an adjacent relationship identifier, and a point of interest; category relationships are recorded by a three-element tuple consisting of a point of interest, a category relationship identifier, and an interest point category to which the point of interest belongs; and group relationships are recorded by a three-element tuple consisting of a user, a group relationship identifier, and a user group to which the point of interest belongs, among which user groups are divided into two categories, one of which is an interest point level group obtained by performing cluster division based on the points of interest accessed by the user, and the other is an area level group obtained by performing cluster division based on the Geohash-5 spatial region to which the user belongs.
3. The aspect-based sentiment analysis module is obtained by cascading a pre-trained DistilBERT model into a multi-class classifier, and further fine-tuning the overall configuration of the cascaded models; 2. The method for recommending points of interest based on a time-series knowledge map according to claim 1, characterized in that, in the aspect-based sentiment analysis module, first, an embedding representation is generated for the user comment text by the DistilBERT model, then the embedding representation is input to the multi-class classifier to obtain comment dimensions corresponding to the user comment text and positive and negative scores in each comment dimension, the positive and negative scores in all comment dimensions are stitched together to output as a sentiment embedding corresponding to the user comment text, and then, by the fine-tuning process, the model is optimized by performing supervised learning on an emotion dataset including the user comment text, the comment dimensions corresponding to the user comment text, and truth-value labels in the positive and negative scores in each comment dimension, and the method is completed after convergence.
4. The method for recommending points of interest based on a time-series knowledge map according to claim 3, wherein the comment dimensions include three dimensions: product, price and service.
5. The process flow of the multimodal knowledge fusion module is as follows: S31, the dynamic time series knowledge map and the static group knowledge map are input into the heterogeneous map attention network respectively to perform information fusion, and obtain the fused dynamic time series knowledge map and the static group knowledge map; S32, arrange all interest points in the historical behavior trajectory substring in the order of user access, sequentially extract hidden layer vectors corresponding to each interest point from the fused dynamic time series knowledge map, construct a user behavior trajectory embedding with global time slice information, sequentially extract hidden layer vectors corresponding to each interest point from the fused static group knowledge map, construct a user behavior trajectory embedding with global static information, and extract interest point level group features and area level group features from the fused static group knowledge map; S33, the user comment emotion embedding sequence is used as a query, and the user behavior trajectory embedding with global time slice information and the original user behavior trajectory embedding are fused through an attention mechanism to obtain a fused user behavior trajectory embedding, the fused user behavior trajectory embedding is used as a value, and the user behavior trajectory embedding with global static information is used as a query, and the interest point level group feature is used as a key to input into the Encoder module of the Transformer model for fusion encoding, and obtain an interest point fusion feature representation; S34: The method for recommending points of interest based on a time-series knowledge map according to claim 1, characterized in that the interest point level group features and the area level group features are connected and fused, the obtained fused group features are used as keys, the user embedded feature representations of all users in the static group knowledge map after the fusion are used as queries, and the user embedded feature representations of all users in the original static group knowledge map are used as values, which are inputted into the Encoder module of the Transformer model to perform fusion encoding, and a user fused feature representation is obtained.
6.
7. The method for recommending points of interest based on time-series knowledge map according to claim 1, wherein the recurrent neural network in the decoding module adopts a long short-term memory network.
Citation Information
Patent Citations
Sequence recommendation method fusing dynamic knowledge graph
CN113590900A
Knowledge graph and time sequence feature fused interpretable interest point recommendation method
CN113656709A
POI recommendation method, device and equipment based on multivariate relation space-time network
CN116503588A
Multi-modal knowledge graph unified representation learning framework
CN117035078A
Interest point recommendation method and device based on node relation, equipment and medium
CN117390285A