Next interest point recommendation method based on global trajectory flow graph and graph contrast learning

Through the method based on global trajectory flow graph and graph comparison learning, the problem of data sparseness and short trajectory space-time context information in the POI recommendation system is solved, and a more accurate and personalized user next POI selection prediction is achieved.

CN119940489APending Publication Date: 2025-05-06CHONGQING UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510001759.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Due to the sparseness of user check-in data, it is difficult to accurately predict the user's next POI selection, and the short-track space-time context information is scarce, and the model performance is degraded, and there are problems such as "cold start" and the problem of insufficient modeling of the time mode.

Method used

Using a method based on global trajectory flow graph and graph comparison learning, the initial k-nearest-neighbor-aware point of interest graph is constructed, and information-rich POI embedding is generated through the spectrum convolution network and attention mechanism module, and the recommended results are optimized through the multi-dimensional data fusion module and the Transformer encoder.

Benefits of technology

It effectively alleviates the challenges brought by data sparseness, improves the performance of short trajectory prediction, enhances the recommendation accuracy for active users who have few check-in records, and significantly improves the overall performance of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940489A_ABST
    Figure CN119940489A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of interest recommendation, in particular to a next interest point recommendation method based on global trajectory flow graph and graph contrast learning. The method comprises the following steps: S1, constructing a global trajectory flow chart; s2, based on the constructed global trajectory flow chart, constructing an initial k neighbor perceptual interest point diagram; and S3, training the constructed perceptual interest point diagram through a spectrogram convolutional network (GCN). According to the next point-of-interest recommendation method based on global trajectory flow graph and graph contrast learning, the global trajectory flow graph irrelevant to a user is introduced, trajectory data of all the users are integrated, and a universal moving mode between POIs is extracted, so that personalized features are faded; then, a k-nearest neighbor sparsification technology is adopted to accurately screen out key relations in the trajectory flow graph; finally, a graph-based comparative learning framework is designed, and wide cooperative signals can be utilized more comprehensively, so that accurate prediction of the next POI is realized, and the cold start problem is effectively relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of interest recommendation, and in particular to a next point of interest recommendation method based on global trajectory flow graph and graph comparison learning. Background Art

[0002] With the widespread popularity of GPS-enabled mobile devices, location-based services (LBS) have made significant progress in recent years. Service platforms such as Foursquare, Gowalla, and Yelp have accumulated huge spatiotemporal data sets by encouraging users to share their experiences, suggestions, and moments at various points of interest (POIs). This data resource provides strong support for the development of a more advanced next POI recommendation system and promotes the in-depth development of related research.

[0003] The core of the POI recommendation system is to accurately capture user preferences, which mainly relies on the user's current and past location records (i.e., check-in data). In recent years, the research focus has shifted to predicting the next location that a user may be interested in. Such recommendation systems not only greatly facilitate individuals' exploration of new areas, but also provide businesses with accurate target marketing strategies. Generally speaking, these recommendation models are mainly trained based on the interaction data between users and POIs, i.e., the user's check-in records. However, since users usually only visit a few POIs that meet their preferences, these locations only account for a very small proportion in the huge POI database, which leads to the problem of data sparsity.

[0004] Given the sparsity of user check-in data, it is often difficult to achieve ideal results by relying solely on such data for POI recommendation, which seriously hinders the accurate prediction of the user's next POI choice. In order to overcome this problem, the current research trend is to integrate a variety of auxiliary information, including time features, geographic coordinates, classification labels, and social network relationships, to effectively alleviate the challenges brought by data sparsity. In addition, some studies have also explored the use of hypergraphs and knowledge graphs to deeply explore more complex associations between users and POIs, or use specific sampling techniques to further optimize the data processing process to cope with the problem of data sparsity.

[0005] Given that short-term temporal patterns in user trajectories play a key role in predicting their future actions, current mainstream methods tend to treat the next POI recommendation as a sequence prediction task. Therefore, many cutting-edge models generally adopt various recurrent neural network (RNN) variants to effectively encode spatiotemporal information. However, such methods still face several limitations. First, the model performance drops significantly when dealing with short trajectories, which is attributed to the relatively scarce spatiotemporal context information that short trajectories can provide. Second, for users who are active but have check-in records at only a few POIs, the recommendation accuracy is usually low, which is particularly significant in real-world recommendation systems, namely the so-called "cold start" problem. Third, POI categories often show significant temporal correlations, for example, check-in activities at bars are more frequent at night than during the day. These temporal patterns are crucial for achieving accurate recommendations, but may not be fully modeled and utilized in some scenarios.

[0006] In order to break through the above limitations to a certain extent, it is particularly important to effectively integrate the behavior data of other users. Considering that there may be partial overlap of activity trajectories between different users, and that the same user may repeatedly show a specific movement pattern, a more complete and coherent movement path can be constructed by aggregating the trajectory information of multiple users. These aggregated trajectory streams can deeply reveal the general movement patterns among user groups, which in turn helps to improve the prediction performance in the case of short trajectories and enhance the recommendation accuracy for those active users with sparse check-in records. However, how to efficiently and accurately use this trajectory stream information in the next POI recommendation is still a challenging task.

[0007] To this end, a next point of interest recommendation method based on global trajectory flow graph and graph comparative learning is designed to provide another technical solution to the above technical problems. Summary of the invention

[0008] Based on this, it is necessary to provide a next point of interest recommendation method based on global trajectory flow graph and graph comparative learning to solve the technical problems raised in the above background technology.

[0009] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0010] Here are the steps:

[0011] S1: construct global trajectory flow chart;

[0012] S2: Based on the constructed global trajectory flow chart, an initial k-nearest neighbor perception interest point map is constructed;

[0013] S3: Train the constructed perceptual POI graph through the spectral graph convolutional network to obtain informative POI embedding;

[0014] S4: Attention mechanism module based on step S3;

[0015] S5: Based on step S4, add a multi-dimensional data fusion module;

[0016] S6: Perform residual connection adjustment on the predicted POI and the POI result obtained through the graph-based contrastive learning framework to optimize the recommendation result.

[0017] As a preferred implementation of the next point of interest recommendation method based on global trajectory flow graph and graph comparison learning provided by the present invention, in the step S1, a global trajectory flow graph is constructed, and the steps are as follows:

[0018] Define check-in as a tuple q=<u,p,t> , where u is the user, p is the point of interest, and t is the timestamp, indicating that user u visited point of interest p at time t;

[0019] All check-in activities of each user u constitute a check-in sequence in is the i-th check-in record;

[0020] The set of check-in sequences of all users is denoted as Q U = {Q u1 ,Q u2 ,…,Q uM};

[0021] In the data preprocessing stage, the check-in sequence Q of each user u is u Split into a series of continuous trajectories, represented as in represents a connection operation; and each track contains a list of check-ins within a shorter time interval (e.g., 24 hours).

[0022] As a preferred implementation of the next point of interest recommendation method based on global trajectory flow graph and graph comparison learning provided by the present invention, in the step S2, based on the constructed global trajectory flow graph, an initial k-nearest neighbor perception point of interest graph is constructed, and the steps are as follows:

[0023] Given a set of historical trajectories and a specific user u i The current trajectory S′=(q1,q2,…,q M );

[0024] Predict user u i Points of interest that may be visited in the next short period of time m+1 ,q m+2 ,…,q m+k, where k is a small integer greater than or equal to 1 (usually k = 1).

[0025] As a preferred implementation of the next point of interest recommendation method based on global trajectory flow graph and graph comparison learning provided by the present invention, the initial Knn perception graph S is constructed by the global trajectory flow graph. m ;

[0026] The similarity score of the interest point pair (i, j) is calculated by the cosine similarity function The specific calculation is as follows:

[0027]

[0028] The fully connected graph is subjected to k-nearest neighbor sparse processing. For each poi node p i , only retain the k edges with the highest similarity scores, the expression is as follows:

[0029]

[0030] in, is the element in row i and column j, is the sparse graph adjacency matrix.

[0031] As a preferred implementation of the next point of interest recommendation method based on global trajectory flow graph and graph contrast learning provided by the present invention, in the step S3, the constructed perceptual point of interest graph is trained by a spectral graph convolutional network to obtain an information-rich POI embedding, and the steps are as follows:

[0032] Through the spectral graph convolutional network, mining graph S m Topological structure information;

[0033] When A∈R N×N Figure S m The adjacency matrix of , calculates the corresponding normalized Laplace matrix, the expression is as follows:

[0034]

[0035] in, is the normalized matrix of the adjacency matrix of the perceptual graph, D is the degree matrix, and I is S m The identity matrix of

[0036] Let H (0) =X∈R N×C , the propagation rule between the layers of the spectral graph convolutional network is:

[0037]

[0038] Among them, H (l-1)represents the input signal of the l-1th layer, W (l) ∈R C×Ω represents the model weight matrix of layer l, and the corresponding bias term is b (l) ∈R C×Ω , σ is a leaky ReLU activation function for nonlinearity;

[0039] In each iteration, the spectral convolutional network layer updates the embedding of the node by aggregating the neighborhood information of the node and the embedding of the node itself. The output of the GCN module is expressed as:

[0040]

[0041] Among them, e Ρ is the embedding vector of all POIs after the perceptual map passes through GCN, ~L is the normalized Laplacian matrix of the perceptual map, W (l*+1) is the R^(Ω×Ω) matrix, which represents the weight matrix of the l*+1th layer GCN; b (l*+1) It is the R^(Ω×1) matrix, which represents the bias vector of the l*+1th layer GCN;

[0042] Points of Interest i The embedding vector is the embedding matrix e Ρ The i-th row of the matrix has a dimension of N×Ω. The embedding vector of interest point p reflects the position of p in the user's historical trajectory and captures the general movement trend at p.

[0043] As a preferred implementation of the next point of interest recommendation method based on global trajectory flow graph and graph comparison learning provided by the present invention, in the step S4, based on the step S3, the attention mechanism module has the following steps:

[0044] Based on the input node features and graph G, the attention graph Φ is calculated as follows:

[0045] Φ1=(X×W1)×a1∈R N×1

[0046] Φ2=(X×W2)×a2∈R N×1

[0047]

[0048] Where W1 and W2∈R C×h are two trainable feature transformation matrices; a1 and a2 are two learnable vectors used to construct an N×N attention matrix through broadcast addition operations; 1 is a shape of R N×1 All 1 vectors of J Nis an all-1 matrix, and ⊙ represents element-by-element multiplication. Φ1 and Φ2 are two N×1 vectors, representing the attention weights from the current POI to other POIs respectively; X is an N×C matrix, representing the node feature matrix, where N is the number of POIs and C is the dimension of the node feature; W1 and W2 are two R^(C×h) matrices, representing two learnable feature transformation matrices respectively; a1 and a2 are two R^h vectors, representing two learnable vectors, used to construct the attention matrix; ~L: the normalized Laplace matrix of G; JN: an N×N matrix with all elements set to 1.

[0049] As a preferred implementation of the next point of interest recommendation method based on global trajectory flow graph and graph comparison learning provided by the present invention, in the step S5, on the basis of step S4, a multidimensional data fusion module is added, and the steps are as follows:

[0050] Train an embedding layer f(·) to map each user into a low-dimensional vector space;

[0051] The embedding vector corresponding to each user is learned based on its historical check-in data sequence, and the expression is as follows:

[0052] e u =f(u)∈R Ω

[0053] Among them, e u is the embedding vector representing user u, f(u) is the function that maps user u to a low-dimensional vector space;

[0054] The embedding vector of the point of interest is concatenated with the embedding vector of the user, and the concatenated vector is fed into a fully connected layer. The expression is as follows:

[0055] e p,u =σ(w p,u [e p ;e u ]+b p,u )∈R Ω×2

[0056] Among them, e p 、e u They represent the embedding of POI and user respectively, and e p,u is the embedded representation after fusion of POI and user embedding, w p,u and b p,u are the learnable weight vector and bias respectively, [·;·] represents concatenation;

[0057] For access time, a 24-hour day is divided into 48 time slots, each of which takes 30 minutes, and the time embedding is obtained as follows:

[0058]

[0059] Among them, e t [i] is the embedding vector of a node i at time t, ω and It is a learnable parameter. The sin activation function is used to capture periodic patterns.

[0060] Another embedding layer is used to process the poi category, and the expression is as follows:

[0061] e c =f(c)∈R Ψ

[0062] Among them, e c Represents the embedding vector of category c, f(c) is the function that maps category c to a low-dimensional vector space;

[0063] Use a dense layer to embed the time t and category embedding c Concatenate them to form a new time-category, the expression is as follows:

[0064] e t,c =σ(w t,c [e t ;e c ]+b t,c )∈R Ψ×2

[0065] Among them, w t,c and b t,c are the learnable weight vector and bias, e t,c is the embedding representation after the fusion of time and category embedding, e t and e c denote the temporal and category embeddings respectively.

[0066] As a preferred implementation of the next point of interest recommendation method based on global trajectory flow graph and graph comparison learning provided by the present invention, on the basis of step S5, a Transformer encoder is added to capture the global user movement pattern to provide users with accurate POI recommendations, and the steps are as follows:

[0067] Given a user check-in sequence, which contains several check-in events, these check-in embeddings are stacked in order to form an input tensor As the first layer input of the Transformer encoder;

[0068] Multiple multi-layer perceptrons are used to predict the next point of interest, visit time and category of the point of interest.

[0069] As a preferred implementation of the next point of interest recommendation method based on global trajectory flow graph and graph contrast learning provided by the present invention, in the step S6, the predicted POI and the POI result obtained by the graph-based contrast learning framework are adjusted by residual connection to optimize the recommendation result, and the steps are as follows:

[0070] After being processed by the Transformer encoder, the embedding of POI is obtained;

[0071] A contrastive learning framework based on graph convolutional networks is used for feature learning and representation of graph data.

[0072] It can be seen without a doubt that the above-mentioned technical solution of the present application can definitely solve the technical problem to be solved by the present application.

[0073] At the same time, through the above technical solutions, the present invention has at least the following beneficial effects:

[0074] The next point of interest recommendation method based on global trajectory flow graph and graph contrastive learning provided by the present invention first introduces a user-independent global trajectory flow graph, which aims to model the general access order and conversion information, and uses a graph convolutional network to effectively encode it as POI embedding; then, the k-nearest neighbor sparsification technology is used to accurately screen out the key relationships in the trajectory flow graph; finally, a graph-based contrastive learning framework is designed, which is used for feature learning and representation of graph data, and improves the performance of graph matching or other related tasks through contrastive learning; the framework can more comprehensively utilize a wide range of collaborative signals, thereby achieving accurate prediction of the next POI and effectively alleviating the cold start problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0076] Figure 1 It is a schematic diagram of the overall framework of the model of the present invention;

[0077] Figure 2 This is a schematic diagram of the NYC dataset ablation experiment of the present invention;

[0078] Figure 3 It is a schematic diagram of the performance ratio of different τ of the present invention;

[0079] Figure 4 This is a schematic diagram of the effect of different k values ​​on the performance of the present invention. DETAILED DESCRIPTION

[0080] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0081] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings.

[0082] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features and technical solutions in the embodiments may be combined with each other.

[0083] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0084] Reference Figure 1-Figure 4 ,Next POI recommendation method based on global trajectory flow graph and graph contrast learning.

[0085] 1. Conception

[0086] A novel user-independent global trajectory flow graph method is introduced, which abstractly represents the mobility patterns and characteristics between POIs in the form of a graph structure. Specifically, the nodes in the graph represent each POI, and the node attributes cover key information such as geographic location, category, and check-in frequency. If two POIs are visited sequentially in the same check-in sequence, a directed edge is constructed between the two nodes, and the weight of the edge reflects the frequency of co-occurrence of this pair of POIs.

[0087] Different from the existing graphs in POI and product recommendation models that are used to characterize the relationship between users and items, the trajectory flow graph proposed in this paper focuses on capturing the transformation relationship and its impact between POIs. In order to make full use of the collective information in the trajectory flow graph, a graph convolutional network (GCN) is introduced to map POIs into a latent space that can contain global transformation information.

[0088] Specifically, GCN updates the embedding representation of each node by aggregating the embedding information of its incoming neighbor nodes. Therefore, the embedding representation of each POI in the trajectory flow graph is affected by its previously visited POIs, effectively preserving the global transition pattern.

[0089] This method allows the most likely next destination from the current POI to be recommended even when the user's historical check-in information is missing. The basic assumption of the global trajectory flow graph is that, although recommending the most popular next POI may not always be the best choice, it still has certain advantages over random guessing. In particular, using the global trajectory flow graph is expected to significantly improve the accuracy of top-k recommendations, especially in the case of larger k values.

[0090] To achieve this goal, we are committed to building a specific similarity structure graph to simulate the potential semantic associations between POIs. The k-nearest neighbor sparsification technique can screen out the most critical relationships, which is helpful to build the perceptual structure graph, thereby further optimizing the connectivity of the graph structure. At the same time, in order to alleviate the "cold start" problem caused by "users who are active but have only checked in at a few POIs", a graph-based contrastive learning framework is designed. Specifically, the framework trains the spectral graph convolutional network through contrastive learning, especially for node representation learning of graph data. By extracting node features and mapping the features to a low-dimensional space using a projection head, the contrast loss between different views is calculated. Furthermore, the loss function is innovatively improved under this framework. Based on the traditional InfoNCE method, the POIs generated by the multi-head Transformer model are embedded into the corresponding graph structure and random noise is added to it. Through the above processing, positive and negative samples are effectively utilized, thereby enhancing the effectiveness of the feature representation learned by the model. In this way, not only can the data sparsity problem be better solved, but the overall performance of the model is also improved.

[0091] At the same time, user preferences and spatiotemporal context information play a vital role in achieving personalized recommendations. Long-term preferences can reveal users' general preferences, such as preference for specific restaurants or cinemas; while short-term preferences are reflected through users' recent check-in records, providing more specific and immediate spatiotemporal information for the recommendation system. To this end, the embedding layer is used to effectively capture the user's overall preferences, the POI category is embedded, and the time2vec model is introduced to finely characterize the temporal features, thereby comprehensively improving the personalization level of recommendations.

[0092] 2. Solution

[0093] 2.1 Next point of interest recommendation

[0094] In the field of traditional point of interest (POI) recommendation, the recommendation method for the next point of interest (Next POI) pays special attention to the time series influence of the user's recent trajectory, in order to accurately predict the user's next action. Early studies widely borrowed from mature methods in other sequence recommendation tasks, such as Markov chains. For example, Cheng et al. pioneered the matrix factorization method (FPMC) fused with personalized Markov chain for the recommendation of the next POI. Zhang et al. introduced additive Markov chains to characterize the transfer influence between sequences. At the same time, there are also studies exploring the possibility of applying matrix decomposition or metric embedding technology to the next POI recommendation.

[0095] With the booming development of deep learning and advanced embedding techniques, researchers have gradually turned their attention to these emerging methods. The spatial-temporal recurrent neural network (ST-RNN) proposed by Liu et al. is one of the best. This model cleverly integrates spatiotemporal context information into the RNN layer. Specifically, the spatial context is represented by the geographic distance transfer matrix, while the temporal context is encoded by the time transfer matrix. In addition, LSTM is also widely used to simulate users' long-term and short-term preferences. In LSPL and PLSPL, researchers trained a standard LSTM model to mine short-term trajectory information and used a universal embedding layer to capture user preferences. Zhao et al. designed a novel LSTM unit, the spatiotemporal gated network (STGN), which simulates the time interval and distance interval in short-term and long-term sequences respectively through two time gates and two distance gates.

[0096] However, these studies often ignore the potential value of leveraging users' common mobility patterns. Recent works have adopted a combination of recurrent neural networks (RNNs) and self-attention mechanisms. For example, LSTPM uses the contextual information of POIs to model users' long-term preferences and combines a geographically expanded RNN to capture short-term preferences. STAN uses a dual-attention architecture to improve recommendation effects by learning explicit spatiotemporal relationships within user trajectories. Although these methods have achieved remarkable results in capturing the spatiotemporal relationships between POIs in a single check-in sequence, they have failed to fully exploit the relationships between multiple check-in sequences.

[0097] To make up for this shortcoming, the latest research has begun to try to combine graph representation learning technology. GETNext constructs a directed trajectory graph to characterize the correlation between multiple check-in sequences and applies a graph convolutional network (GCN) to learn the representation of POIs. AGRAN combines geographic dependencies and spatiotemporal information learned from adaptive graphs to simultaneously capture users' dynamic preferences. STHGCN introduces a hypergraph structure that aims to learn more refined trajectory granularity information from users' historical trajectories and other users' collaborative trajectories.

[0098] 2.2 Graph-based location recommendation methods

[0099] In this paper, we are committed to optimizing the recommendation effect of the next point of interest (Next POI) through graph-based methods. Graph-based methods, especially those combined with location-based social network (LBSN) data, have opened up a powerful new paradigm for traditional (non-serialized) recommendation tasks. For example, some studies have accurately captured users' check-in behaviors and preferences by constructing a geo-temporal influence-aware graph (GTAG), a three-part graph containing POI, session, and user nodes. However, the introduction of session nodes may lead to a significant increase in the size of the graph, which in turn leads to an increase in computational complexity. To solve this problem, some studies have proposed a solution to construct four bipartite graphs, namely the POI-POI graph, the POI-region graph, the POI-time graph, and the POI-word interaction graph, which are used to capture sequence effects, geographic influences, temporal dynamics, and semantic features, respectively. The builders of these graphs further extended the LINE network embedding model to bipartite graphs, and used conditional probability and other statistical tools to train graph embeddings to improve recommendation performance.

[0100] Although existing studies have achieved certain results, they are still insufficient in fully utilizing the advantages of graph structures. Although existing studies have applied graph convolutional networks (GCNs) to learn POI representations, they have not fully explored the potential of graph feature representations.

[0101] 3. Implementation Methods

[0102] The structure of the framework is as follows Figure 1 As shown, its core consists of several key components.

[0103] The first step is to construct a comprehensive global trajectory flow graph, which serves as the cornerstone to build an initial k-nearest neighbor (KNN)-aware point-of-interest (POI) graph based on the original feature set. This carefully constructed POI graph is then deeply trained using a spectral graph convolutional network (GCN) to generate informative POI embeddings. These embeddings not only accurately capture the common movement patterns of users between POIs, but also fully incorporate multiple key dimensions such as the category characteristics, geographical distribution, and check-in frequency of POIs. In addition, the training process effectively filters out noisy information and deeply explores the important structural correlations between the latent features of POIs.

[0104] On this basis, an efficient attention mechanism module is innovatively introduced. This module takes the adjacency matrix and node features of the global trajectory flow graph as input, and outputs an accurate transition attention map through sophisticated calculations. This mapping can efficiently model the transition probability between POIs, and then generate a detailed POI transition probability map.

[0105] In order to further enrich the capture and utilization of contextual information, a multi-dimensional data fusion module is introduced. This module can integrate user embedding, POI category embedding, and time encoding converted through a time-to-vector model. In order to achieve more accurate personalized recommendations, user embedding is deeply fused with POI embedding in the corresponding trajectory. At the same time, in order to capture and reflect the user's preferences for different POI categories over time, POI category embedding and time encoding are also cleverly combined.

[0106] Finally, a cutting-edge adaptive Transformer encoder was used to model the user's mobility pattern and provide accurate POI recommendations to users. In order to ensure the accuracy and reliability of the recommendation results, the predicted POIs were adjusted with the POI results obtained through the graph-based contrastive learning framework, thereby further optimizing and improving the recommendation results. This method not only improves the accuracy of the recommendation, but also provides users with a more personalized POI recommendation experience.

[0107] 3.1 Global Trajectory Flow Graph Learning

[0108] Define a user set U = {u1,u2,…,u M}, which contains M users and a set of points of interest (POIs) P = {p1, p2, …, p N}, which contains N points of interest, such as a specific shopping mall or restaurant. In addition, let T = {t1, t2, …, t K} is a timestamp set, K is a positive integer. Each point of interest p∈P can be represented by a tuple containing latitude, longitude, category and check-in frequency.<lat,lon,cat,freq> The category cat is selected from a predefined list of POI categories, such as "subway station" or "bar".

[0109] In the present invention, sign-in is defined as a tuple q=<u,p,t> , where u is the user, p is the point of interest, and t is the timestamp, indicating that user u visited point of interest p at time t. All check-in activities of each user u constitute a check-in sequence where q iu is the i-th check-in record. The check-in sequence set of all users is represented by Q U = {Q u1 ,Q u2 ,…,Q uM}.

[0110] In the data preprocessing stage, the check-in sequence Q of each user u is u Split into a series of continuous trajectories, represented as in Represents a join operation. The length of each trajectory may vary, and each trajectory contains a list of check-ins within a short time interval (e.g., 24 hours).

[0111] The purpose of next POI recommendation is to predict the points of interest that the user may visit next based on the current trajectory and the user's historical check-in records. More specifically, given a set of historical trajectories and a specific user u i The current trajectory S′=(q1,q2,…,q M ), the goal is to predict user u i Points of interest that may be visited in the next short period of time m+1 ,q m+2 ,…,q m+k , where k is a small integer greater than or equal to 1 (usually k = 1).

[0112] (1) POI embedding representation of global trajectory flow graph

[0113] Considering that different users may have similar itineraries and the same user may repeat the same itinerary, a user-independent global trajectory flow graph is introduced. This graph provides a global perspective to reveal the common flow patterns of users between points of interest (POIs). By deeply analyzing this graph, it is possible to effectively extract the widespread user movement patterns from historical check-in data.

[0114] Using a collection of historical traces The global trajectory flow graph is a weighted directed graph G = (V, E, l, ω), which is specifically expressed as follows: the vertex set V corresponds to the interest point set P. For each interest point p in the set P, its attributes L(p) include (lat, lon, category, freq), where (lat, lon) represents the geographic coordinates of p, category represents the category of p, and freq represents the total number of times p has been visited in the historical trajectory set S. If any trajectory in the set S In the graph, interest points p1 and p2 are visited consecutively, then there is an edge from p1 to p2 in the graph G. The weight ω(p1,p2) of any edge (p1,p2) in the graph G represents the number of times this pair of interest points appear consecutively in all trajectories in the set S.

[0115] First, we construct an initial Knn perception graph S by using the global trajectory flow graph m Based on the assumption that similar interest points are more likely to influence each other than dissimilar interest points, the semantic relationship between two interest points is quantified by their similarity. The cosine similarity function is chosen to calculate the similarity score of interest point pair (i, j). The specific calculation is as follows:

[0116]

[0117] Normally, the adjacency matrix of a graph should be non-negative, but S ij The value of may be in the range [-1,1]. Therefore, S ij Negative values ​​in are set to zero. In addition, compared with common graph structures, fully connected graphs are usually more computationally intensive and may contain noise and unimportant edges. To address this issue, k-nearest neighbor (knn) sparsification is performed on the fully connected graph. For each poi node p i , only retain the k edges with the highest similarity scores, the expression is as follows:

[0118]

[0119] in, is the element in row i and column j, is the sparse graph adjacency matrix;

[0120] in is the resulting sparse, directed graph adjacency matrix. In order to alleviate the problem of gradient explosion or vanishing, the adjacency matrix is ​​normalized:

[0121]

[0122] in, is the sparse graph adjacency matrix after normalization, D m ∈R N×N yes The diagonal matrix of

[0123] Process the given global trajectory flow graph G into a perception graph S m ,The next goal is to build a vectorized interest point representation that can capture common interest point transfer patterns and interest point attributes. To achieve this, a spectral graph convolutional network (GCN) is used. m Specifically, let A∈R N×N Figure S m The adjacency matrix of , first calculate its corresponding normalized Laplace matrix, the expression is as follows:

[0124]

[0125] in, is the normalized matrix of the adjacency matrix of the perceptual graph, D is the degree matrix, and I is S m Next, let H(0) =X∈R N×C , the propagation rule between GCN layers is:

[0126]

[0127] For any l>0, H (l-1) represents the input signal of the l-1th layer, W (l) ∈R C×Ω represents the model weight matrix of layer l, and the corresponding bias term is b (l) ∈R C×Ω , and σ is a leaky ReLU activation function (with a leakiness rate of 0.2) for nonlinearity.

[0128] From a spatial perspective, in each iteration, the GCN layer updates the embedding of the node by aggregating the neighborhood information of the node and the embedding of the node itself. In order to improve the expressiveness of the model, l*GCN layers are stacked. Before the last layer, the dropout technique is applied. The output of the GCN module can be expressed as:

[0129]

[0130] Among them, e Ρ is the embedding vector of all POIs after the perceptual map passes through GCN, ~L is the normalized Laplacian matrix of the perceptual map, It is the R^(Ω×Ω) matrix, which represents the weight matrix of the l*+1th layer GCN; It is an R^(Ω×1) matrix, which represents the bias vector of the l*+1th layer GCN.

[0131] Finally, the point of interest p i The embedding vector is the embedding matrix e Ρ The i-th row of the matrix has a dimension of N×Ω. In general, the embedding vector of the interest point p reflects the position of p in the user's historical trajectory and captures the general movement trend at p. These embedding vectors will be sent to the downstream Transformer model to simulate the user's access pattern. It should be noted that even if the current trajectory is short, the embedding vector of the interest point can provide sufficient information for the prediction model.

[0132] (2) POI transfer learning with attention mechanism

[0133] The POI embeddings obtained from the graph G implicitly capture common mobility patterns. To enhance the role of group signals, a new transition attention mechanism is introduced to explicitly model the transition probability from one POI to another. As mentioned before, these transition probabilities will be used to optimize the final prediction results. Based on the input node features and the graph G, the attention map Φ is calculated as follows:

[0134] Φ1=(X×W1)×a1∈R N×1

[0135] Φ2=(X×W2)×a2∈R N×1

[0136]

[0137] Where W1 and W2∈R C×h are two trainable feature transformation matrices; a1 and a2 are two learnable vectors that are used to construct an N×N attention matrix through broadcast addition operations; 1 is a shape of R N×1 All 1 vectors of J N is an all-1 matrix, and ⊙ represents element-by-element multiplication. Φ1 and Φ2 are two N×1 vectors, representing the attention weights from the current POI to other POIs respectively. X is an N×C matrix, representing the node feature matrix, where N is the number of POIs and C is the dimension of the node feature.

[0138] W1 and W2 are two R^(C×h) matrices, representing two learnable feature transformation matrices respectively.

[0139] a1 and a2 are two R^h vectors, representing two learnable vectors, used to construct the attention matrix;

[0140] ~L: normalized Laplace matrix of G;

[0141] JN: An N×N matrix with all elements set to 1.

[0142] The normalized Laplacian matrix The range of is adjusted from [0, 1] to [1, 2] to avoid zero values. The i-th row of the transition attention map Φ represents the transition from interest point p i The (unnormalized) probability of moving to each interest point. Given the last interest point in the current trajectory, look up the transition probabilities stored in the corresponding row in Φ and use these probabilities to adjust the recommendations produced by the subsequent transformation module.

[0143] 3.2 Multi-dimensional Data Fusion

[0144] Many POI recommendation methods have demonstrated that spatial-temporal context and user preferences are key factors in personalizing the next POI recommendation. Next, we introduce a multi-dimensional data fusion module, which is used to fuse context information, including user embedding, POI embedding, POI category embedding, and temporal encoding.

[0145] (1) POI-User Embedding Fusion

[0146] POI embedding is learned from the perception graph obtained by processing the trajectory flow graph, and user-specific patterns are omitted. In order to grasp the general behavior pattern of a specific user u′s, an embedding layer f(·) is trained, which is responsible for mapping each user to a low-dimensional vector space. The embedding vector corresponding to each user is learned based on its historical check-in data sequence. Specifically, the embedding vector of user u is obtained in the following way:

[0147] e u =f(u)∈R Ω (8)

[0148] Among them, e u is the embedding vector representing user u, and f(u) is the function that maps user u to a low-dimensional vector space.

[0149] To form a comprehensive representation of each check-in event, a simple approach is to concatenate the embedding vector of the point of interest (POI) with the embedding vector of the user. The concatenated vector is then fed into a fully connected layer to adjust the fused embedding and enhance its expressiveness. The output of this process can be expressed as:

[0150] e p,u =σ(w p,u [e p ;e u ]+b p,u )∈R Ω×2 (9)

[0151] Among them, e p 、e u They represent the embedding of POI and user respectively, and e p,u is the embedded representation after fusion of POI and user embedding, w p,u and b p,u are the learnable weight vector and bias, respectively, and [·;·] represents concatenation. After the above processing, the dimension of the output embedding is twice that of the POI embedding or user embedding. In other words, the size of the fused embedding vector remains unchanged.

[0152] (2) Time-category embedding fusion

[0153] Users' visit behaviors to points of interest (POIs) have obvious time period characteristics. For example, users usually visit bars and similar places at night, while subway stations are busier during peak commuting hours. Therefore, in order to model the relationship between user movement patterns and time and category, the visit time and POI category are encoded separately.

[0154] Similar to the POI-user embedding fusion module, the POI category and time are first encoded. For the visit time, a 24-hour day is divided into 48 time slots, each of which takes 30 minutes. Then the time2vector method is applied to embed these time slots to obtain the time embedding representation:

[0155]

[0156] Among them, e t [i] is the embedding vector of a node i at time t, ω and is a learnable parameter, and the sin activation function is used to capture periodic patterns.

[0157] Use another embedding layer to process the poi category:

[0158] e c =f(c)∈R Ψ (11)

[0159] Among them, e c represents the embedding vector of category c, and f(c) is the function that maps category c to a low-dimensional vector space.

[0160] Next, a dense layer is used to embed the time t and category embedding c Concatenate them to form a new time-category representation:

[0161] e t,c =σ(w t,c [e t ;e c ]+b t,c )∈R Ψ×2 (12)

[0162] Among them, w t,c and b t,c are the learnable weight vector and bias, e t,c is the embedding representation after the fusion of time and category embedding, e t and e c Representing time and category embeddings respectively;

[0163] Finally, for a check-in tuple q =<p,u,t> , where the interest point p belongs to category c and its embedding is done by embedding the POI into ep 、User Embed u and time code t The splicing result is e q =[e p ,e u ;e t ,e c ]. Therefore, each input trajectory (q1,…,q s ) is embedded in a series of check-ins These encoded sequences will be input into the Transformer encoder.

[0164] (3) Transformer Encoder

[0165] To avoid the problem of over-smoothing, graph convolutional networks (GCNs) usually keep a low number of layers, which limits them to capturing only local features in the check-in sequence. Therefore, a Transformer encoder is introduced to capture the global user mobility pattern. Given a user check-in sequence, which contains several check-in events, the embeddings of these check-ins are stacked in sequence to form an input tensor As the first layer input of the Transformer encoder. Considering that accurately predicting the category of the next point of interest can enhance the model's understanding of user preferences and mobility trends. In addition to the main task of predicting the next POI, several auxiliary tasks are designed to assist in training the proposed next point of interest recommendation framework based on global trajectory flow graph and graph contrast learning. In the decoding stage, in order to predict the user's next behavior, multiple multi-layer perceptron (MLP) decoders are used to replace the traditional Transformer decoder. Specifically, three MLP heads are deployed, which are used to predict the next point of interest (POI), visit time, and category of the point of interest, respectively. Let the output of the encoder be represented as These MLP heads can be represented as:

[0166]

[0167] in, represents the output matrix of the Transformer encoder, Represents the output matrix of POI prediction; represents the output matrix of time prediction; Represents the output matrix of category prediction; W poi and b poi are the weight matrix and bias vector of POI prediction, W time and b time are the weight matrix and bias vector for time prediction, W cat and b catiare the weight matrix and bias vector for category prediction, W poi ∈R d×N , W time ∈R d×1 and W cat ∈R d×Γ are the weights in the multi-layer perceptron (MLP), and Γ represents the number of interest point categories. For the output matrix For example, we only focus on the last row because this row represents the recommendation of points of interest in future actions. In addition, we combine this point of interest recommendation with the transition attention map. The final recommendation result is:

[0168] in, express The kth row in is the pth k OK, p k It means the sign-in record Corresponding points of interest.

[0169] At the same time, in order to reduce the impact of time fluctuations on prediction and better predict the category of the next POI, in addition to the POI header, a time header and a category header are also added.

[0170] 3.3 Image Contrastive Learning

[0171] After being processed by the Transformer encoder, the POI embedding is obtained. In order to obtain more effective information from the embedded representation, learn a more robust and meaningful node representation, and seek to maximize the consistency of representations between different views. A contrastive learning framework based on a graph convolutional network (GCN) is designed for feature learning and representation of graph data, as well as improving the performance of related tasks through contrastive learning.

[0172] First, we learn the SimGCL approach to create different perspectives by slightly rotating the refined node embeddings in space. This approach not only preserves the original information, but also enhances the robustness of the model by introducing the InfoNCE loss as an additional self-supervised learning task. The specific method is as follows:

[0173]

[0174] in and represents the noise vector added to the lth layer, which needs to satisfy ||Δ||2=ε and You can use ε to control and Relative to The rotation angle of and The noise is always in the same super-octant, so the introduced noise does not cause significant deviation. Noise is added in each convolutional layer and the output of all layers is averaged to obtain the final embedding of the node. To simplify the representation, h i represents the final node embedding after L layers, and and represents the two views generated by this method. Next, we use InfoNce and the designed GRACE as additional self-supervised tasks to improve the robustness of the transition attention map.

[0175] For InfoNce, after obtaining the generated view of each interest point node, it is used to maximize the consistency of positive pairs and minimize the consistency of negative pairs:

[0176]

[0177] Among them, h′ i and h" i represents different views of the same node, h′ j and h" j represents different views of different nodes, s(·) represents the cosine similarity function, which calculates the similarity between two vectors; τ is a hyperparameter.

[0178] Then, the GRACE module is used to calculate the contrast loss between the two views again. In each iteration, noise is added to the view, as shown in formula (8), and then two views are generated, represented as G1 = t(h' i ) and G2 = (h" i ), and the interest point nodes in the two generated views are represented as and in and They are the feature matrix and adjacency matrix of the view respectively.

[0179] First, we use the designed spectral graph convolutional network (GCN) encoder model f(X,A)∈R N×F' , which is used to learn the representation of nodes in the graph. Specifically, its function is to process the input graph data and generate a low-dimensional node embedding representation, that is, F<<F', and represent H=f(X, A) as the learned representation of the node. After that, contrast loss is used as a discriminator, that is, a loss function used to distinguish the embedding representation of the same node in two different views from the embedding representation of other nodes. For any node i, its embedding representation a in one view i is regarded as an anchor point, and the embedding b generated in another view iis considered as a positive sample, while the other node embeddings in the two views are naturally considered as negative samples. In the multi-view contrastive learning setting, we borrow the idea of ​​InfoNCE objective and i ,b i ) defines a pair of optimization objectives.

[0180]

[0181] in, It is sample a i and b i (positive sample pair) contrast loss; θ(a i ,b i ) is a metric function used to calculate sample a i and b i The similarity between them can be dot product, cosine similarity or any other function that can quantify similarity; τ is a hyperparameter, temperature parameter, which is used to control the smoothness of the distribution. A smaller τ value will make the model pay more attention to positive sample pairs, while a larger τ value will make the model more sensitive to negative sample pairs; is the exponent of the similarity of the positive sample pair scaled by the temperature parameter; is a i With all other positive samples b k The exponential sum of similarities (k≠1); is a i With all other positive samples a k (k≠1) similarity index and

[0182] Define θ(a,b)=s(g(a),g(b)), where s(·) represents the cosine similarity function, which calculates the similarity between two vectors, and g(·) represents a nonlinear mapping function, the purpose of which is to enhance the expressive power of the function θ(·). A two-layer perceptron is used to construct g(·).

[0183] For a given positive sample pair, all nodes in the other two views are naturally defined as negative samples. Therefore, negative samples come from two sources, namely, cross-view and intra-view nodes, corresponding to the second and third terms in the denominator in formula (10), respectively. Since the two views are symmetric, the loss of the other view is similarly defined as l(b i ,a i ). The overall objective to be maximized is formally defined as the average of all positive sample pairs, namely:

[0184]

[0185] in, Represent the contrast loss of ai and bi respectively.

[0186] Next, the loss functions calculated by formulas (15) and (17) are jointly optimized to obtain the final task objective function:

[0187] L CL =αL InfoNce +βL Grace (18)

[0188] Among them, L CL The final contrastive loss, α, β are hyperparameters.

[0189] 3.4 Loss

[0190] During the training phase of the model, the losses of multiple decoder predictions are combined. For the next point of interest, the cross entropy is used as the loss function. At the same time, KL divergence is used to measure the performance of category prediction. For the performance of temporal prediction, the mean squared error (MSE) is used. The loss learned from the graph representation is then calculated as the weighted sum of the losses of all MLP heads.

[0191] In order to balance the importance of different losses, many POI recommendation models usually use weighted linear combinations to integrate the losses of each independent task, which usually requires manual weight adjustment. However, the performance of the model is often affected by these parameter settings, and manually adjusting these parameters is time-consuming and challenging in practice. Therefore, a multi-task learning strategy based on task-related uncertainty is introduced, which can automatically adjust the weights between multiple tasks to efficiently train the model. Finally, the overall loss function of the proposed model can be expressed as:

[0192]

[0193] Among them, L, L poi , L time , L cat , L CL They represent the final model loss, POI loss, time loss, category loss, and contrastive learning loss, respectively. σ1, σ2, and σ3 represent the learnable uncertainties in the three prediction tasks. It can be inferred that the larger σ is, the higher the uncertainty of the task, and therefore the smaller the weight assigned to that specific task. InfoNce ,L Grace They represent the losses calculated by the InfoNce and Grace modules in graph representation learning, and α and β are hyperparameters.

[0194] IV. Experiment

[0195] 4.1 Experimental Setup

[0196] 4.1.1 Dataset Description

[0197] The experiments are based on three public datasets, all of which are derived from location-based service systems: FourSquare-NYC, FourSquare-TKY, and Gowalla-CA. The FourSquare-NYC dataset covers user check-in data in New York City from April 2012 to February 2013, while the FourSquare-TKY dataset covers user check-in data in Tokyo during the same period. The Gowalla-CA dataset contains user check-in information on the Gowalla platform in California and Nevada from February 2009 to October 2010. Each data record includes user information, point of interest (POI), POI category, GPS location coordinates, and check-in timestamp. When processing these three datasets, 10-core filtering is first performed, that is, POIs and users with less than 10 check-ins are deleted. Then, the user's check-in sequence is divided into different trajectories according to 24-hour time intervals, and trajectories containing only a single check-in record are excluded. Finally, the dataset is divided into training set, validation set and test set according to the time order, where the training set contains the first 80% of the check-in data, which is used to generate the trajectory flow graph G; the validation set contains the middle 10% of the data; and the test set contains the remaining 10% of the data. When evaluating the prediction performance, if the user or POI does not appear in the training set but appears in the test set, the data of these users or POIs will be ignored. The specific statistical information of these datasets is shown in Table 1.

[0198] Table 1: Dataset statistics

[0199]

[0200] 4.1.2 Evaluation criteria.

[0201] In order to measure the performance of the next point of interest (POI) recommendation system, two commonly accepted ranking evaluation metrics are selected: Hit Ratio (HR) and Normalized Discounted Cumulative Gain (NDCG). These metrics are used to evaluate the accuracy of the recommendation list. HR@K is a metric frequently used in previous POI recommendation systems. It is used to check whether the next check-in location of the target user is included in the top K recommended results. NDCG@K is a widely recognized evaluation metric used to measure the performance of the recommendation system. It not only considers the accuracy of the recommendation, but also the position order of the recommended items in the list.

[0202] 4.1.3 Baseline Comparison Model

[0203] In this study, the following baseline models were selected for performance comparison:

[0204] MF is a traditional method in recommender systems, which uses matrix factorization techniques to mine implicit features of users and points of interest (POIs).

[0205] FPMC combines matrix factorization and Markov chain to capture users’ long-term interests and continuous behavior patterns simultaneously.

[0206] LSTM is a variant of recurrent neural network (RNN) that is specifically designed to process sequence data and can handle both short-term and long-term sequence patterns.

[0207] PRME proposes a personalized ranking metric embedding method to capture users’ preferences for POIs and the sequential transitions between POIs.

[0208] ST-RNN combines the time and distance transition matrices to simulate the local spatiotemporal context and uses the RNN model to recognize the user's sequential behavior.

[0209] STGN introduces spatial gates and temporal gates based on standard LSTM to better capture user preferences in the spatial and temporal dimensions.

[0210] STGCN is an improved version of STGN, which adopts a coupled input and forget gate mechanism.

[0211] PLSPL combines the attention mechanism to learn users’ long-term preferences and LSTM to learn short-term preferences, and integrates these two preferences through a personalized linear layer.

[0212] STAN exploits the spatiotemporal information in the check-in trajectory and adopts a self-attention mechanism to identify the interactions between non-consecutive check-in points.

[0213] GETNext constructs a directed trajectory graph to represent the relationships between different check-in sequences, and uses a graph convolutional network (GCN) to learn the feature representation of POIs.

[0214] 4.1.4 Experimental Configuration

[0215] The model was built using the PyTorch framework and tested on a hardware platform equipped with a Core(TM) i7-13700KF processor and an NVIDIA GeForce RTX 4090 graphics card. The following are the key hyperparameter configurations of the model: the POI and user embedding dimensions are set to Ω = 128, and the time and POI category embedding dimensions are set to Ψ = 32. The GCN model contains three hidden layers, configured with 32, 64, and 128 channels respectively. The transformation attention module is responsible for mapping the input node features to a 128-dimensional vector space. In the graph representation learning module, two graph convolutional layers are used, with the number of input and output feature channels being 128, and feature fusion is performed through skip connections. The dimensions of the hidden layer and projection layer are both 1024, and the temperature parameter is set to 0.2 by default. For the Transformer architecture, two layers of encoders are stacked. In these encoder layers, the dimension of the feedforward network is 1024, and two attention heads are used in the multi-head attention module. The Adam optimizer is also used, with a learning rate of 1e-3 and a weight decay rate of 5e-4. In the GCN model and Transformer encoder, Dropout was enabled with a ratio of 0.3. The weight α of the temporal loss was set to 10 to balance the impact of POI loss and category loss. These settings were applied on all three datasets and each model was run for 200 epochs with a batch size of 20 samples.

[0216] 4.2 Performance Comparison

[0217] Table 2 Performance comparison of Acc@k and MRR on three datasets

[0218]

[0219] Table 2 compares the performance of the proposed model with the baseline model on three different datasets. The top-1, top-5, top-10, top-20 accuracy and MRR (mean reciprocal rank) of the models are evaluated. Overall, all models outperform the CA dataset on the NYC and TKY datasets. This may be because the points of interest (POIs) in the NYC and TKY datasets are more concentrated and located in smaller areas in New York City and Tokyo, respectively.

[0220] In the comparison between the model of the present invention and the baseline model, the performance of the model of the present invention is better on all data sets. For example, on the NYC data set, the model of the present invention achieved a top-1 accuracy of 25.2%, exceeding the 24.35% of the best baseline model GETNext. In terms of top-5 accuracy, the improvement reached 10%, and in terms of top-20 accuracy, the improvement was 14.2%. The TKY data set also showed a similar trend. In addition, recommendation systems using attention mechanisms or LSTM (such as STAN, PLSPL, STGCN) are significantly better than traditional models based on Markov chains or matrix decomposition (such as FPMC, PRME). Taking STAN as an example, the top-1 accuracy on the NYC data set is 22.31%, while FPMC is only 10.03%. At the same time, POI recommendation models using graph convolution (such as GETNext) outperform models based on attention mechanisms or LSTM (such as STAN, PLSPL, STGCN) in performance. For example, the top-1 accuracy of GETNext on the NYC data set is 24.35%, while that of STAN is 22.31%.

[0221] The CA dataset has 9.9k POIs and 250k check-in records, which are distributed over a vast area of ​​more than 400,000 square kilometers, while the TKY dataset contains 7.8k POIs and 361k check-in records, which are concentrated in an area of ​​approximately 2,000 square kilometers. Since the CA dataset faces more serious data sparsity problems in the number of check-ins and the spatial distribution of POIs, the performance of the model in these aspects is generally not as good as on the NYC and TKY datasets. For example, STAN has a top-1 accuracy of 22.31% on the NYC dataset, but it drops to 11.04% on the CA dataset. The top-1 accuracy of the model of the present invention on the CA dataset is 14.04%, which exceeds all baseline models, but is still far lower than the performance on the NYC dataset.

[0222] The performance improvement is due to the effective integration of POI content information and check-in sequences, the use of a designed graph representation learning method, and the application of an adaptive multi-task Transformer.

[0223] 4.3 Ablation Experiment

[0224] An ablation experiment was conducted to evaluate the specific impact of each proposed component on the performance of the model on the NYC dataset.

[0225] Seven different configurations were designed for the experiment: 1) complete model; 2) removing the KNN perception graph, the graph convolution layer directly processes the trajectory flow graph; 3) no graph representation learning, no maximization consistency processing is performed on the graph data; 4) removing the trajectory flow graph, POI embedding is learned through the embedding layer; 5) no transformer encoder, the time series data is completely processed by LSTM; 6) omitting time and category information; 7) no graph convolutional network (GCN), node embedding is learned from the transition graph through random walks; the experimental results are detailed in Table 3.

[0226] The complete model shows the best performance. The experimental results also reveal that graph representation learning, trajectory flow graph, and GCN play a more important role in the overall performance than other components. For example, in the absence of graph representation learning or using random walks to learn POI embedding, the top-1 accuracy drops from 25.20% to 20.57% and 21.34%, respectively. Other components also have an impact on the final recommendation effect. For example, using k-nearest neighbors to construct a perceptual graph can capture the important structural relationships between POI potential features and improve performance by removing noise from historical data.

[0227] Table 3 Ablation experiment on NYC dataset

[0228]

[0229] 4.4 Parameter Sensitivity

[0230] In order to achieve the best performance of the model, important hyperparameters need to be adjusted. The present invention conducts experiments on the temperature parameter τ and knn-k parameters on the NYC dataset.

[0231] Sensitivity analysis of temperature parameter τ: Different temperature parameters τ will directly affect the effect of the model. The temperature coefficient can be used to control the distribution shape of logits. For a given logits distribution shape, when the τ value increases, 1 / τ decreases, s(h′ i ,h″ i ) / τ will make the values ​​in the original logits distribution smaller, and after the exponential operation, it will become even smaller, causing the original logits distribution to become smoother. On the contrary, if τ is small, 1 / τ will become larger, and the values ​​in the original logits distribution will become larger accordingly. After the exponential operation, it will become even larger, making the distribution more concentrated and more peaky. In order to select the best τ, in the experiment, the τ of the present invention increases from 0.15 to 0.50 every 0.05 units. Taking the NYC data set as an example, the experimental results are as follows: Figure 3As shown, it can be observed that the performance of the model designed by the present invention changes with the change of τ. When τ is adjusted to a suitable value, the performance of the model reaches the best. Excessive adjustment of τ may weaken the auxiliary ability of contrastive learning in improving recommendation performance.

[0232] Knn-k. As shown in Table 4, a parameter sensitivity experiment of k value was conducted on the NYC dataset to process the trajectory flow graph and construct the perceptual graph. The influence of k was studied by changing [1, 5, 25, 50, 100]. The results show that the model of the present invention performs best when k=5. And as can be seen in Table 4, the method of the present invention has the best performance on all datasets. It is found that choosing a suitable k value in constructing the perceptual graph can filter out noise and capture important structural relationships between the potential features of POIs. Therefore, choosing an appropriate k value can maximize the overall benefit.

[0233] Table 4 The impact of Knn-k on recommendation performance, measured by Acc@20

[0234]

[0235]

[0236] V. Conclusion

[0237] In the present invention, a novel recommendation framework is creatively proposed, which integrates global trajectory flow graph and graph comparison learning technology to optimize the point of interest (POI) recommendation system. This framework cleverly integrates the content information of POI into the recommendation algorithm of the next POI through the strategy of global graph structure learning, effectively coping with the challenges brought by data sparsity. Specifically, the graph representation learning model is deeply used to mine the rich information implicit in the graph data and to maximize the consistency of representations between different views, thereby significantly enhancing the understanding and analysis capabilities of the model. In order to further improve the performance of the model, the knn sparsification technology is introduced, which can efficiently filter out noise information, thereby constructing a perceptual map with a clear structure and easy to understand. At the same time, in order to accurately capture the general mobility pattern of users, the concept of trajectory flow graph is innovatively introduced. This design not only cleverly solves the problems brought by inactive users and short trajectories, but also significantly improves the generalization ability of the model, so that it can better adapt to the needs of different user groups. In order to comprehensively and accurately model the mobility pattern of users, an adaptive Transformer structure is carefully designed. This structure can make full use of the user's mobile history data, capture the temporal dependency and spatial correlation therein, and thus provide users with more personalized recommendation services. Through a series of experiments conducted on three sets of real-world datasets, the significant advantages of the proposed model over the current advanced models are fully verified. At the same time, through meticulous ablation experiments, the effectiveness of each component of the model is verified one by one, fully demonstrating the superiority and advancement of this method from multiple dimensions.

[0238] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the present invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. The next point of interest recommendation method based on global trajectory flow graph and graph contrast learning is characterized by: Here are the steps: S1: construct global trajectory flow chart; S2: Based on the constructed global trajectory flow chart, an initial k-nearest neighbor perception interest point map is constructed; S3: Train the constructed perceptual POI graph through the spectral graph convolutional network to obtain informative POI embedding; S4: Attention mechanism module based on step S3; S5: Based on step S4, add a multi-dimensional data fusion module; S6: Perform residual connection adjustment on the predicted POI and the POI result obtained through the graph-based contrastive learning framework to optimize the recommendation result.

2. The next point of interest recommendation method based on global trajectory flow graph and graph comparative learning according to claim 1 is characterized in that: In the step S1, a global trajectory flow chart is constructed, and the steps are as follows: Define check-in as a tuple q=<u,p,t> , where u is the user, p is the point of interest, and t is the timestamp, indicating that user u visited point of interest p at time t; All check-in activities of each user u constitute a check-in sequence in is the i-th check-in record; The set of check-in sequences of all users is denoted as Q U = {Q u1 ,Q u2 ,…,Q uM }; In the data preprocessing stage, the check-in sequence Q of each user u is u Split into a series of continuous trajectories, represented as in represents a connection operation; and each track contains a list of check-ins within a shorter time interval (e.g., 24 hours).

3. The next point of interest recommendation method based on global trajectory flow graph and graph comparative learning according to claim 2 is characterized in that: In the step S2, based on the constructed global trajectory flow chart, an initial k-nearest neighbor perception interest point map is constructed, and the steps are as follows: Given a set of historical trajectories and a specific user u i The current trajectory S′=(q1,q2,…,q M ); Predict user u i Points of interest that may be visited in the next short period of time m+1 ,q m+2 ,…,q m+k , where k is a small integer greater than or equal to 1 (usually k = 1).

4. The next point of interest recommendation method based on global trajectory flow graph and graph comparative learning according to claim 3 is characterized in that: Construct the initial Knn perception graph S through the global trajectory flow graph m ; The similarity score of the interest point pair (i, j) is calculated by the cosine similarity function The specific calculation is as follows: The fully connected graph is subjected to k-nearest neighbor sparse processing. For each poi node p i , only retain the k edges with the highest similarity scores, the expression is as follows: in, is the element in row i and column j, is the sparse graph adjacency matrix.

5. The next point of interest recommendation method based on global trajectory flow graph and graph comparative learning according to claim 3 is characterized in that: In the step S3, the constructed perceptual interest point graph is trained through a spectral graph convolutional network to obtain an informative POI embedding, and the steps are as follows: Through the spectral graph convolutional network, mining graph S m Topological structure information; When A∈R N×N Figure S m The adjacency matrix of , calculates the corresponding normalized Laplace matrix, the expression is as follows: in, is the normalized matrix of the adjacency matrix of the perceptual graph, D is the degree matrix, and I is S m The identity matrix of Let H (0) =X∈R N×C , the propagation rule between the layers of the spectral graph convolutional network is: Among them, H (l-1) represents the input signal of the l-1th layer, W (l) ∈R C×Ω represents the model weight matrix of layer l, and the corresponding bias term is b (l) ∈R C×Ω , σ is a leaky ReLU activation function for nonlinearity; In each iteration, the spectral convolutional network layer updates the embedding of the node by aggregating the neighborhood information of the node and the embedding of the node itself. The output of the GCN module is expressed as: Among them, e Ρ is the embedding vector of all POIs after the perceptual map passes through GCN, ~L is the normalized Laplacian matrix of the perceptual map, It is the R^(Ω×Ω) matrix, which represents the weight matrix of the l*+1th layer GCN; It is the R^(Ω×1) matrix, which represents the bias vector of the l*+1th layer GCN; Points of Interest i The embedding vector is the embedding matrix e Ρ The i-th row of the matrix has a dimension of N×Ω. The embedding vector of interest point p reflects the position of p in the user's historical trajectory and captures the general movement trend at p.

6. The next point of interest recommendation method based on global trajectory flow graph and graph comparative learning according to claim 1, characterized in that: In the step S4, based on the step S3, the attention mechanism module has the following steps: Based on the input node features and graph G, the attention graph Φ is calculated as follows: Φ1=(X×W1)×a1∈R N×1 Φ2=(X×W2)×a2∈R N×1 Where W1 and W2∈R C×h are two trainable feature transformation matrices; a1 and a2 are two learnable vectors used to construct an N×N attention matrix through broadcast addition operations; 1 is a shape of R N×1 All 1 vectors of J N is an all-1 matrix, and ⊙ represents element-by-element multiplication. Φ1 and Φ2 are two N×1 vectors, representing the attention weights from the current POI to other POIs respectively. X is an N×C matrix, representing the node feature matrix, where N is the number of POIs and C is the dimension of the node feature. W1 and W2 are two R^(C×h) matrices, representing two learnable feature transformation matrices respectively. a1 and a2 are two R^h vectors, representing two learnable vectors, used to construct the attention matrix. ~L: the normalized Laplace matrix of G. JN: an N×N matrix with all elements set to 1.

7. The next point of interest recommendation method based on global trajectory flow graph and graph comparative learning according to claim 1, characterized in that: In step S5, based on step S4, a multi-dimensional data fusion module is added, and the steps are as follows: Train an embedding layer f(·) to map each user into a low-dimensional vector space; The embedding vector corresponding to each user is learned based on its historical check-in data sequence, and the expression is as follows: have been u =f(u)∈R Ω Among them, e u is the embedding vector representing user u, f(u) is the function that maps user u to a low-dimensional vector space; The embedding vector of the point of interest is concatenated with the embedding vector of the user, and the concatenated vector is fed into a fully connected layer. The expression is as follows: e p,u σ(w p,u [e p 10. The u ]+b p,u )∈R Ω×2 Among them, e p 、e u They represent the embedding of POI and user respectively, and e p,u is the embedded representation after fusion of POI and user embedding, w p,u and b p,u are the learnable weight vector and bias respectively, [·;·] represents concatenation; For access time, a 24-hour day is divided into 48 time slots, each of which takes 30 minutes, and the time embedding is obtained as follows: Among them, e t [i] is the embedding vector of a node i at time t, ω and It is a learnable parameter. The sin activation function is used to capture periodic patterns. Another embedding layer is used to process the poi category, and the expression is as follows: e c =f(c)∈R Ψ Among them, e c Represents the embedding vector of category c, f(c) is the function that maps category c to a low-dimensional vector space; Use a dense layer to embed the time t and category embedding c Concatenate them to form a new time-category, the expression is as follows: e t,c σ(w t,c [e t 10. The c ]+b t,c )∈R Ψ×2 Among them, w t,c and b t,c are the learnable weight vector and bias, e t,c is the embedding representation after the fusion of time and category embedding, e t and e c denote the temporal and category embeddings respectively.

8. The next point of interest recommendation method based on global trajectory flow graph and graph comparative learning according to claim 7, characterized in that: Based on step S5, add a Transformer encoder to capture the global user movement pattern and provide users with accurate POI recommendations. The steps are as follows: Given a user check-in sequence, which contains several check-in events, these check-in embeddings are stacked in order to form an input tensor As the first layer input of the Transformer encoder; Multiple multi-layer perceptrons are used to predict the next point of interest, visit time and category of the point of interest.

9. The next point of interest recommendation method based on global trajectory flow graph and graph comparative learning according to claim 1, characterized in that: In step S6, residual connection adjustment is performed between the predicted POI and the POI result obtained by the graph-based contrastive learning framework to optimize the recommendation result. The steps are as follows: After being processed by the Transformer encoder, the POI embedding is obtained; A contrastive learning framework based on graph convolutional networks is used for feature learning and representation of graph data.

Citation Information

Cited By

  • Interest point recommendation method and system based on multistage denoising

    CN122489852A

  • A point of interest recommendation method and system based on multi-stage denoising

    CN122489852B