Next POI Recommendation System Based on Spatiotemporal Information Representation
By designing a next point of interest recommendation system based on spatiotemporal information representation, using embedded representations of multi-grained periodic information, time intervals and geographical distances, combined with causal convolution and Transformer model, the problem of insufficient utilization of spatiotemporal information in the prior art is solved, and the recommendation performance is significantly improved.
Patent Information
- Application Number
- CN202211552543.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-12-05
AI Technical Summary
The current next point of interest recommendation technology is difficult to effectively utilize space-time information, resulting in insufficient recommendation performance, especially in capturing time or geographical distance effects.
A next point of interest recommendation system based on spatiotemporal information representation is designed. Through the time personalization module, the time encoding module, the causal convolution enhancement module and the output module, combined with the attention mechanism and the Transformer model, the embedded representation of multi-grained periodic information, time intervals and geographical distances is used to enhance the local information and spatiotemporal perception of the user's check-in sequence.
It significantly improves the performance of recommendations for the next point of interest, can more accurately capture the user's time and space movement patterns, and improves the accuracy and user experience of recommendations.
Smart Images

Figure CN115827974B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of recommendation technologies, and particularly to a next point of interest recommendation system based on spatio-temporal information representation. Background Art
[0002] The popularization of smart devices and the mobile Internet has promoted the growing prosperity of location-based social networks (LBSNs) such as Foursquare, Gowalla, Yelp, Twitter, WeChat, and Weibo, attracting a large number of users. People check in at locations through the check-in function provided by the LBSNs platform, share their dynamics and real-time locations with friends, and interact with friends by posting information such as opinions, photos, and comments related to the location. This social way has increasingly penetrated into the daily life of the public and gradually evolved into an important communication method in people's lives.
[0003] As one of the core functions of LBSNs, the next point of interest recommendation technology has a relatively wide range of application scenarios and plays a crucial role in people's lives. For the government, by predicting the next point of interest that people will visit, the government can design more reasonable traffic planning and dispatching strategies to alleviate traffic jams and crowd gatherings; for platforms such as carpooling and food delivery, the next point of interest prediction technology can accurately help drivers or food delivery riders effectively avoid congested roads and plan trips in advance; for merchants, store information and coupons can be accurately distributed to potential target users who may visit, thereby avoiding blind large-scale advertising and achieving targeted advertising delivery, saving advertising costs; for users, the next point of interest recommendation technology can assist users in making decisions and improve the user experience. As an independent sub-field of the recommendation system, the next point of interest recommendation has wide applications, can provide better user experiences and third-party services for users, and has thus received extensive attention from the academic and industrial communities in recent years.
[0004] In the next point-of-interest recommendation task, a user's interests can be divided into long-term preferences and short-term preferences. Long-term preferences refer to the comprehensive interests of the user mined from the user's historical trajectory, which rely on all historical check-in records; while short-term preferences refer to the user's interest preferences within a short period, which are affected by the recently visited points of interest and are more inclined to the most recent check-in records. The time of the user's check-in contains two aspects of information. On the one hand, the timestamp reflects the absolute time when the user visits the point of interest, including different granularity cycles such as year, month, week, day, hour, etc.; on the other hand, the time interval between check-ins reflects the degree of association between two check-ins. Therefore, reasonably using time information can better mine the user's movement patterns. To model the periodicity of human movement using time information, some models propose an embedded representation of time. For example, TMCA divides time into hourly granularity and distinguishes between weekdays and weekends, and thus represents time as a one-hot vector with a dimension of 48 as the input of the model. However, this model only considers cycles of hours and weeks, ignores information with cycles of months and days, and assumes that adjacent time periods are independent of each other without considering their proximity. In addition, in the recommendation task, methods such as TMCA, TiSASRec, and STAN in model design utilize spatio-temporal information by learning the representations of time intervals and geographical distances. However, the current models have limited ability to learn the representations of time intervals and geographical distances and cannot reveal the impact of real time or geographical distance effects on human movement.
[0005] Different from digital product recommendations such as commodity recommendations, news recommendations, and music recommendations, which are pure online interaction-based, user check-in behaviors do not have implicit feedback data such as browsing and clicking. Human activities are affected by real-world factors and exhibit complex transfer patterns, and the dataset itself is sparse. Therefore, the next point-of-interest recommendation task is highly challenging. Summary of the Invention
[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is to provide a next point-of-interest recommendation system based on spatio-temporal information representation to improve the performance of the next point-of-interest recommendation.
[0007] The present invention provides a next point-of-interest recommendation system based on spatio-temporal information representation, including:
[0008] A time personalization module that personalizes the cycle information of the four granularities of month, week, day, and hour of the check-in timestamp. For the cycle representation of any one of the four granularities, an attention mechanism is used to adaptively combine the representation of the time period where the check-in moment t is located with the representation of the adjacent time period to obtain the personalized representation of the time period where the granularity is located for the user u. Using the The time personalization representation T of the moment t for the user u is calculated using an attention mechanism. u(t);
[0009] A time encoding module that uses a time encoding function to map the check-in timestamp to a vector space, obtaining a time encoding representation Φ(t);
[0010] A user check-in sequence module that takes the T u (t), the Φ(t), and the embedded representation of the point of interest where user u checks in as inputs, and calculates the embedded representation S of the user check-in sequence u ;
[0011] A causal convolution enhancement module that takes the embedded representation of the point of interest where user u checks in as an input, combines causal convolution with Transformer to enhance the local information of the user check-in sequence, obtaining an enhanced embedded representation S' after causal convolution u ;
[0012] A new user check-in sequence module that takes the S u , the S u ', and the embedded representation E of the spatial relationship Δ , and calculates the representation Z of the new user check-in sequence u ;
[0013] An output module that uses the representation of the new user check-in sequence of user u to calculate the preference of user u for the point of interest at the moment t, and predicts and recommends the next point of interest.
[0014] Technical effects:
[0015] In view of the problem of recommending the next point of interest based on spatio-temporal information representation, the present invention designs four personalized granularity period representations, combines the encoded representation of time and the embedded representation of geographical distance, and considers the representation of time interval and geographical distance when calculating the attention between check-ins, so as to utilize spatio-temporal information when modeling the long-term preferences of users, and uses causal convolution for local information enhancement, improving the performance of recommending the next point of interest.
[0016] The following will further illustrate the concept, specific structure and technical effects generated by the present invention with reference to the accompanying drawings, so as to fully understand the purpose, features and effects of the present invention. Description of the drawings
[0017] Figure 1 is the structural diagram of the next point of interest recommendation system in the embodiment of the present invention;
[0018] Figure 2 is the schematic diagram of the continuous check-in activities of a single user in the embodiment of the present invention
[0019] Figure 3 is the schematic diagram of the STIRSAN model in the embodiment of the present invention;
[0020] Figure 4.a It is a heat map of the check-in interest point categories at the monthly granularity in the NYC dataset of the embodiments of the present invention;
[0021] Figure 4.b It is a heat map of the check-in interest point categories at the daily granularity in the NYC dataset of the embodiments of the present invention;
[0022] Figure 4.c It is a heat map of the check-in interest point categories at the weekly granularity in the NYC dataset of the embodiments of the present invention;
[0023] Figure 4.d It is a heat map of the check-in interest point categories at the hourly granularity in the NYC dataset of the embodiments of the present invention;
[0024] Figure 5 It is a schematic diagram of causal convolution in the embodiments of the present invention;
[0025] Figure 6.a It is a trend chart of Recall@10 varying with the dimension in the NYC dataset of the embodiments of the present invention;
[0026] Figure 6.b It is a trend chart of Recall@20 varying with the dimension in the NYC dataset of the embodiments of the present invention;
[0027] Figure 7 It is a schematic diagram of the influence of dimensions on STIRSAN in the TKY dataset of the embodiments of the present invention;
[0028] Figure 8 It is a schematic diagram of the influence of the sequence length on STIRSAN in the TKY dataset of the embodiments of the present invention;
[0029] Figure 9.a It is a heat map of similarity represented by the Month cycle in the embodiments of the present invention;
[0030] Figure 9.b It is a heat map of similarity represented by the Date cycle in the embodiments of the present invention;
[0031] Figure 9.c It is a heat map of similarity represented by the DayofWeek cycle in the embodiments of the present invention;
[0032] Figure 9.d It is a heat map of similarity represented by the Hour cycle in the embodiments of the present invention;
[0033] Figure 10.a It is a heat map of Example 1 of the geographical distance and spatial relationship in the embodiments of the present invention;
[0034] Figure 10.b It is a heat map of Example 2 of the geographical distance and spatial relationship in the embodiments of the present invention;
[0035] Figure 10.c It is the heat map of Example 3 of the geographical distance and spatial relationship in the embodiment of the present invention;
[0036] Figure 10.d It is the heat map of Example 4 of the geographical distance and spatial relationship in the embodiment of the present invention. Specific Embodiments
[0037] The following introduces multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification, making its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0038] In the embodiments of the present invention, the variables and their mathematical expressions are defined as follows: The set of users in the check-in data is defined by U = {u1, u2,... u |U|}, where u represents one of the elements, and the total number of users is |U|; the set of points of interest in the check-in data is defined by L = {l1, l2,... l |L|}, and the total number of check-in points of interest is |L|; p k = (lon k , lat k ) represents the geographical coordinates of the point of interest l k , that is, longitude and latitude. The k-th check-in of user u is denoted as , indicating that user u visits location k at time t . The trajectory of each user is denoted as . The user trajectory is clipped to a fixed length , where N is the maximum length of the specified trajectory. If N < m, only consider the last N check-ins; if N > m, pad 0s on the left side of the sequence until the length of the sequence is N. In addition, the time interval between the i-th and j-th check-ins is denoted as ΔT ij = |t i - t j |, and the geographical distance is denoted as ΔD ij = Haversine(p i , p j ), where the Haversine formula is as follows:
[0039]
[0040] where R represents the radius of the earth, which is 6371 km. Therefore, the geographical interval ΔD between any two check-ins in the check-in sequence is expressed as:
[0041]
[0042] Next point - of - interest recommendation refers to predicting the point of interest that a user u will visit at the next moment based on the user's check - in record tra(u).
[0043] As Figure 1 shown, an embodiment of the present invention provides a next point - of - interest recommendation system based on spatio - temporal information representation, including:
[0044] A time personalization module that personalizes the periodic information of four granularities of the check - in timestamp, namely month, week, day, and hour. For the periodic representation of any one of the four granularities, an attention mechanism is used to adaptively combine the representation of the time period where the check - in moment t is located with the representation of the adjacent time period to obtain the personalized representation of the time period where the granularity is located for the user u Using the The time personalization representation T u (t) of the moment t for the user u is calculated by using an attention mechanism;
[0045] A time encoding module that maps the check - in timestamp to a vector space using a time encoding function to obtain a time encoding representation Φ(t);
[0046] A user check - in sequence module that takes the T u (t), the Φ(t), and the embedded representation of the points of interest where the user u checks in as inputs, and calculates the embedded representation S u ;
[0047] A causal convolution enhancement module that takes the embedded representation of the points of interest where the user u checks in as an input, combines causal convolution with Transformer to enhance the local information of the user check - in sequence, and obtains an enhanced embedded representation S′ u ;
[0048] A new user check - in sequence module that takes the S u , the S u ′ and the embedded representation E Δ of the spatial relationship as inputs, and calculates the new check - in sequence representation Z u ;
[0049] An output module that uses the new check - in sequence representation of the user u to calculate the preference of the user u for the points of interest at the moment t, and predicts and recommends the next point of interest.
[0050] In the embodiments of the present invention, it is utilized that people's travel shows multi-granularity periodicity. Considering the personalized influence of multi-granularity period information on users, four kinds of personalized period representations of different granularities are designed. Combining the encoded representation of time and the embedded representation of geographical distance, the time interval and geographical distance between check-ins are respectively encoded and embedded based on the Bochner theory and the AutoDis method, overcoming the deficiencies of existing work. When calculating self-attention, the representations of the time interval and geographical distance between check-ins are considered to capture spatio-temporal effects. In addition, this method uses causal convolution to enhance the Transformer's perception of local sequences, and uses causal convolution for local information enhancement, improving the performance of the next point-of-interest recommendation.
[0051] The following introduces this embodiment in combination with the accompanying drawings. Figure 2 It is a schematic diagram of the check-in activities of a single user in the embodiment dataset. Among them, p1, p2, p3,..., p i is the sequence of points of interest; Δt1, Δt2,…, Δt i-1 is the time interval between two adjacent check-in records; Δd1, Δd2,…, Δd i-1 is the geographical distance between two adjacent points of interest. The embodiments of the present invention only consider using the user's most recent N check-in data to model and learn the spatio-temporal information representation, and predict the next point of interest.
[0052] Figure 3 It is a schematic diagram of the model structure of a preferred embodiment of the present invention. The markings in the figure are defined as: DistanceInterval Embedding: Embedded representation of geographical distance; Inputs: Input; Input Embedding: Embedded representation of the check-in point of interest; Causal Convolution Layer: Causal convolution layer; User embedding: User embedded representation; Time Interval Encoding: Encoded representation of time; PMGP: Encoding of the personalized representation of user u at time t; PositionalEncoding: Positional encoding representation; Self-Attention Layer: Self-attention layer; Feed-ForwardNetwork: Feed-forward network layer; Dropout: Model generalization; LayerNorm Layer: Layer normalization. Among them, Dropout means that during forward propagation, the activation value of a certain neuron stops working with a certain probability p, which can make the model more general and alleviate the occurrence of overfitting.
[0053] In the embodiment of the present invention, the original data is the user's check-in record, which includes information such as user ID, interest location ID, interest location type, and interest location longitude and latitude. The user's check-in record can be regarded as a sequence of interest points arranged in time, and the next interest point recommendation task can essentially be regarded as a sequence prediction problem. The original data set is grouped according to the user ID and sorted in the order of check-in time, generating the check-in sequence data of each user. For each user, the [2,N]th check-in is used as the test set, and the [2,N - 1]th check-in is used as the input sequence to predict the Nth check-in interest point.
[0054] For user u, extract the sub-attributes of the check-in time, namely Month, DayOfWeek, Date, and Hour. Obtain the embedded representation e of user u u , and randomly initialize the embedded representations of the four granularities of month, week, day, and hour to obtain the personalized four-granularity periodic information representation of user u.
[0055] Refer to Figure 4.a to Figure 4.d It can be seen that the categories of the check-in interest points shown in adjacent time periods are similar. Therefore, adjacent time periods should have relatively similar representations. So in the embodiment of the present invention, for granularity s, the attention mechanism is used to adaptively combine the representation of the time period where time t is located with the representations of adjacent time periods to obtain the personalized representation of the time period where the granularity is located for user u.
[0056] To obtain the final representation of time t for user u, that is, to obtain the personalized multi-granularity periodic representation, for the periodic representations of the four granularities, the attention mechanism is used to calculate the time-personalized representation of time t for user u.
[0057] Since the time interval plays an important role in expressing the time effect and revealing the sequence pattern, the embodiment of the present invention considers the representation of the time interval, calculates the time interval with the time encoding representation Φ(t), and uses it to measure the connection between two timestamps.
[0058] Since Transformer is naturally good at capturing long-term and global dependencies in sequences, but it is not good at extracting fine-grained local information, while convolutional operations can well explore local information. Therefore, to overcome the drawbacks of Transformer, STIRSAN combines causal convolution with Transformer to enhance local information in the sequence. Different from traditional convolutional operations, causal convolution aims to ensure that there is no information leakage from future moments during the calculation process. To implement causal convolution, it is necessary to pad (kernel size - 1) zeros at the beginning of the sequence and truncate the redundant output at the end of the sequence to ensure that the sequence lengths before and after convolution are the same. In this way, the input of the convolution at time t only contains information of itself and previous moments, and does not contain information of time t + 1 and later moments. Figure 5 Schematic diagram of causal convolution with a kernel size of 3, and the result after causal convolution is Input of causal convolution Is the embedded representation of the POIs of each check-in of user u.
[0059] The embedded representation of geographical distance is another important factor in the problem of next POI recommendation. For the embedded representation of continuous numerical features, a simple solution is to treat numerical features as categorical features and assign an independent embedding vector to each numerical feature. However, this method has serious defects, such as a large number of parameters and insufficient training of low-frequency features. To reduce the model parameters, the domain embedding method shares one to three embedded representations for all feature values of the same type of feature, and obtains the representations of each feature value through transformations such as multiplying with the feature value and linear interpolation method. However, the capacity of this type of method model is relatively low, resulting in a decline in performance. Therefore, the present invention designs a meta-embedding H ∈ R K×d Shared by all geographical distances. Each meta-embedding h v ∈ R d Can be regarded as a subspace in the latent space to improve the expressive ability and capacity of the model. To capture the complex relationship between geographical distance and meta-embedding, a differentiable automatic discretization module is designed. By means of weighted average, the geographical distance ΔD ij Between the i-th check-in and the POI of the j-th check-in is obtained. The way of weighted average will make the relevant meta-embeddings more conducive to providing rich information, while the irrelevant meta-embeddings will be largely ignored. Thus, the embedded representation E ΔD Of the geographical distance ΔD between each check-in of user u is obtained, and further the embedded representation E Δ Of the spatial relationship is obtained.
[0060] STIRSAN aggregates the representations of historical check-ins by calculating self-attention to assign different weights to each check-in in the trajectory. The input of the spatio-temporal aware self-attention is the embedded representation of the user check-in sequence as S u and the embedded representation of the user check-in sequence after enhancing the local information causal convolution as S′ u and the embedded representation E of the spatial relationship Δ . During the process of calculating the attention, in addition to calculating the similarity between the check-in points of interest, the influence of the time interval and geographical distance between each check-in is also considered, and the new check-in sequence representation Z of the user u is obtained u .
[0061] Using the new check-in sequence representation of the user u, the latent factor model is used to calculate the preference of the user u for the point of interest at the moment t; the model is trained to learn the model parameters and predict the next point of interest
[0062] Let S = {Month, DayOfWeek, Date, Hour} represent the set of four granularity period information, s ∈ S, s represents the element in the set, that is, one of the four granularities. In a preferred embodiment of the present invention, for the granularity s, the embedded representation of the user u at the moment t is where, when s represents the month, it is expressed as
[0063]
[0064] where, E Month (t) ∈ R d represents the embedded representation of the month at the moment t, R represents the real number field, d is the dimension of the embedded representation, ° represents the Hadamard product, and W Month ∈ R d×d is the weight coefficient of the linear transformation, which is randomly initialized and continuously optimized during iteration. W Month can perform corresponding transformation on the embedded representation of the user to be used for personalizing the month representation. In the same way, the personalized embedded representations of the personalized week, day, and hour granularity period information can be obtained
[0065] For the granularity s, the personalized representation of the adjacent m time periods of the moment t for the user u is Then there is the following relational expression:
[0066]
[0067] where,
[0068]
[0069] Among them, is the embedded representation of the adjacent m time periods of the user u at the time t with the granularity s. When m < 0, it represents the -m-th time period before the time period where the time t is located; when m > 0, it represents the m-th time period after the time period where the time t is located; when s represents month, t s represents the month where the time t is located, and f is 12; when s represents week, t s represents the week where the time t is located, and f is 7; when s represents day, t s represents the day where the time t is located, and f is 31; when s represents hour, t s represents the hour where the time t is located, and f is 24. However, different from the calculation methods of month, week, and day, for hours, since a day starts from 0 o'clock, so when s represents hour, Δ s (t, m) = (t s + m) mod f, where mod is the modulo operator.
[0070] For the granularity s, the attention mechanism is used to adaptively combine the representation of the time period where the time t is located with the representation of the adjacent m time periods. The personalized representation of the time period where the granularity s is located for the user u is:
[0071]
[0072] Among them, corresponding to the granularity s, is the personalized representation of the time period where the time t is located for the user u. Corresponding to the four granularities of month, week, day, and hour, the personalized representations for the user u are respectively obtained as: is the personalized representation of the adjacent m time periods of the time t for the user u: the attention coefficient α t,m represents and similarity.
[0073] Preferably, the attention coefficient α t,m calculation formula is:
[0074]
[0075] Among them, r s represents the window size of the adjacent time periods of the granularity s. r s is set as: when s represents month, r s = 2; when s represents week, r s = 1; when s represents day, r s = 6, when s represents hour, r s = 5, exp represents the exponential function with the natural constant e as the base, and n represents the value within the window of the adjacent time periods of the granularity s.
[0076] In another preferred embodiment of the present invention, an attention mechanism is used to calculate the time personalized representation T of the moment t for the user u u (t) ∈ R d×1 as follows:
[0077]
[0078] wherein, represents the periodic representation attention coefficient of the user u to the granularity s, and W u ∈ R d×d represents the transformation of the user representation, and e u ∈ R d is the embedded representation of the user u;
[0079] T u (t) ∈ R d×1 As the final time personalized representation of the moment t for the user u, it contains the representation of personalized multi-granularity periodic information.
[0080] The time interval between user check-ins reflects the degree of association between two check-ins and plays an important role in expressing time effects and revealing sequence patterns. Reasonably using time information can better mine the movement patterns of users. The timestamp reflects the absolute time when the user visits the point of interest. It is necessary to process the timestamp information to express the time interval. In another preferred embodiment of the present invention, a time encoding function is used to map the timestamp to a vector, i.e., Φ: t → R d , then the time interval can be expressed as the dot product of the corresponding time encoding representations:
[0081] ψ(t1 - t2) = K(t1, t2): = <Φ(t1), Φ(t2)> (1)
[0082] wherein, K is the time kernel function, and <,> represents the dot product operation. ψ(t1 - t2) measures the connection between two timestamps. Since this time encoding function directly encodes the timestamp, it can be generalized to any timestamp, and then the representation of any time interval can be obtained:
[0083]
[0084] Based on Bochner's theory and Monte Carlo integration, Φ(t) in Equation (1) can be realized by Φ d (t) in Equation (2). Wherein, ω = [ω1,..., ω d T are the parameters to be learned by the model. Therefore, the encoded representation of time t is Φ d (t), and the time interval is obtained by the inner product of the vectors shown in Equation (1).
[0085] In another preferred embodiment of the present invention, by using the time personalization representation of the user u, the time-coded representation, and the embedded representation of each check-in point of interest of the user u, the embedded representation of the check-in sequence of the user u is calculated as:
[0086]
[0087] where N is the number of check-ins of the user u, is the representation of the k-th check-in of the user u:
[0088]
[0089] where, is the point of interest checked in by the user u at time t k , T u (t k ) is the time personalization representation at time t k , which integrates the periodic information of the four granularities, Φ(t k ) is the check-in time coding representation at time t k , e pk is the coding representation of the check-in location at time t k , e u ∈R d is the embedded representation of the user u.
[0090] In another preferred embodiment of the present invention, in order to capture the complex relationship between geographical distance and meta-embedding, a differentiable automatic discretization module is designed:
[0091] γ ij = ReLU(W1ΔD ij )
[0092]
[0093] where V is the number of meta-embeddings, ΔD ij is the geographical distance between the i-th check-in and the j-th check-in,
[0094] W1∈R V×1 , W2∈R V×V are the parameters to be learned by the model. α controls the proportion of the residual connection. γ ij represents the result obtained through one-layer neural network transformation, represents the correlation between the geographical distance ΔD ij and the meta-embedding H, where, represents the geographical distance ΔD ijThe correlation with the v-th meta-embedding h v .
[0095] Therefore, the geographical distance ΔD between the i-th check-in and the point of interest of the j-th check-in ij The embedded representation of is:
[0096]
[0097] By means of weighted average, the embedded representation of the geographical distance ΔD between the i-th check-in and the point of interest of the j-th check-in ij is obtained Then the embedded representation of the geographical distance ΔD between each check-in of user u is E ΔD ∈R N×N×d , then the embedded representation E of the spatial relationship Δ ∈R N×N is:
[0098] E Δ = E ΔD W D
[0099] where the mapping matrix W D ∈R d×1 is used to transform the dimension of E ΔD .
[0100] The weighted average method will make the relevant meta-embeddings more conducive to providing rich information, while the irrelevant meta-embeddings will be largely ignored.
[0101] In another preferred embodiment of the present invention, combining the embedded representation of the user check-in sequence as S u , the embedded representation of the local information causal convolution enhancement of the user check-in sequence as S' u , and the embedded representation E of the spatial relationship Δ , calculate the new check-in sequence representation Z of user u u :
[0102]
[0103] where The mapping matrices W Q , W K , W V ∈R d×d , ° represents the Hadamard product. In order to ensure causality when calculating self-attention, that is, the information at time t+1 and later cannot be used when predicting the point of interest at time t, a lower triangular mask matrix M with elements of 1 is introduced, and T represents the transpose operation.
[0104] In another preferred embodiment of the present invention, a Feed-Forward Network (FFN) is used to represent the new check-in sequence Z of the user u u Introduce non-linearity:
[0105]
[0106] wherein, The weight parameters of the neural network, is the threshold parameter of the neural network, and f1 and f2 represent the first layer network and the second layer network of the FFN.
[0107] In order to accelerate the training process and prevent phenomena such as overfitting and gradient disappearance during the training process, techniques such as LayerNorm, Dropout, and residual connections are introduced, and the optimized representation of the check-in sequence of the user u is:
[0108]
[0109] where N is the number of check-ins of the user u, is the time t k The optimized check-in representation of the user u, k is the k-th check-in, t k is the time of the k-th check-in.
[0110] Preferably, a latent factor model is used to calculate the preference of the user u for the point of interest k at time t preference
[0111]
[0112] where, is the representation of the point of interest representation. is the time t k-1 The optimized check-in representation of the user u, t k-1 is the time of the (k - 1)-th check-in.
[0113] After model training and predicting the next point of interest, and obtaining the preferences of the user u for each point of interest at time t, the points of interest are output in descending order of preference values, which can be used to recommend the next point of interest that the user is going to visit or predict each point of interest.
[0114] In order to learn the parameters of the model, the BPR loss function is adopted:
[0115]
[0116] where the training set sample D = (u, li , l j , t)|(u, l i , t) ∈ I u ,l j ∈ I \ I u ,where the positive sample (u, l i , t) represents the point of interest l that user u checks in at time t i ,while the negative sample (u, l j , t) is randomly sampled from the points of interest that user u does not check in at contains all the learnable parameters of the model, and the penalty term is used to prevent the model from overfitting. σ is the sigmoid function
[0117] Next, for the STIRSAN model established in the embodiments of the present invention, two evaluation metrics, the recall rate Recall@K of the top K and the mean reciprocal rank (MRR), are used to evaluate the performance of the model
[0118] (1) Recall rate Recall@K
[0119] The recall rate refers to the ratio of the positive samples predicted by the model to all positive samples. In the next point of interest recommendation, it refers to the ratio of the labels in the top K after sorting the predicted values of the model for each point of interest in descending order. The calculation method of Recall@K is as follows
[0120]
[0121] where K ∈ {1, 5, 10, 20} and S label respectively represent the top K points of interest recommended by the model to the user and the points of interest actually visited by the user, that is, the labels. Obviously, in the next point of interest recommendation task, there is only one point of interest that the user visits at the next time step, that is, |S label | = 1
[0122] (2) MRR
[0123] MRR is the average of the reciprocals of the ranks of positive samples among all samples, reflecting the overall ranking ability of the model. The calculation method of MRR is as follows
[0124]
[0125] where rank u represents the rank of the label of user u in the model recommendation list. The higher the rank of the label in the recommendation list, the higher the MRR value and the better the model performance
[0126] The embodiments of the present invention evaluate the STIRSAN model in two real-world datasets, Foursquare-NYC and Foursquare-TKY. The datasets are preprocessed by deleting inactive users with less than 5 check-ins and points of interest with less than 5 check-ins. The dataset statistics are as follows:
[0127] Table 1 Dataset Statistics
[0128] dataset #User #POI #Check-in Foursquare-NYC 1083 9989 179468 Foursquare-TKY 2293 15177 494807
[0129] In the experiment, only the most recent N visits of each user are used. The dataset is divided into a training set and a test set. The [1, N-1] check-ins of each user are used as the training set, and the [2, N] check-ins are used as the test set. The [2, N-1] check-ins are used as the input sequence to predict the Nth check-in point of interest.
[0130] Experimental Results
[0131] The present invention mainly uses spatio-temporal information representation for the recommendation of the next point of interest. There are mainly two experiments in this experimental part: (1) A comparative experiment between the model of the embodiments of the present invention and 5 baseline models. (2) Design an ablation experiment to verify the effectiveness of each module of the model of the embodiments of the present invention. (3) Design an experiment to verify the robustness and interpretability of the model.
[0132] In the comparative experiment of the performance of the next point of interest recommendation, we compare the STIRSAN model established by the embodiments of the present invention with the following models:
[0133] (1) TMCA model: Adopts an encoder-decoder structure based on LSTM, and proposes two attention mechanisms to adaptively select relevant historical check-ins and context factors.
[0134] (2) DeepMove model: Proposes to use the attention mechanism to obtain relevant information from historical check-ins to model long-term user preferences, and uses RNN to model short-term preferences.
[0135] (3) LSTPM model: Utilizes the temporal and spatial correlations between the current trajectory and historical trajectories to model long-term preferences, and uses geo-dilated RNN to capture the geographical connections between non-consecutive check-ins when modeling short-term preferences.
[0136] (4) STAN model: A next point of interest recommendation model based on Transformer, which considers the time interval and geographical distance between non-consecutive check-ins, and uses linear interpolation to learn the representations of different time intervals and geographical distances.
[0137] (5) TiSASRec model: A sequence recommendation model based on Transformer that personalizes the time intervals for different users to obtain relative time intervals and takes into account the influence of different time intervals when calculating self-attention.
[0138] The hyperparameters of the STIRSAN model are set as: r M = 2, r W = 1, r D = 6, r H = 5, α = 0.1, the number K of meta-embedding representations is set to 30, and the convolutional kernel size of causal convolution is set to 5. The input sequence length is 100, the dimension of the embedded representation is 100, the hidden layer dimension is also set to 100, the number of layers of CAC-Transformer is 2, and λ is 5e-5. The parameter quantities of each model are shown in Table 2.
[0139] Table 2 Comparison of the parameter quantities of each model
[0140] Models TMCA DeepMove LSTPM STAN TiSASRec STIRSAN Number of parameters 2,420,753 4,207,563 3,257,863 1,795,000 1,172,400 1,268,492
[0141] In the two datasets of NYC and TKY, the experimental results of STIRSAN and the baseline model in the next point-of-interest recommendation task are shown in Tables 3 and 4. The data in bold in the tables are the highest experimental results.
[0142] Table 3 Comparative experiment on recommendation performance on the NYC dataset
[0143]
[0144] Table 4 Comparative experiment on recommendation performance on the TKY dataset
[0145]
[0146] As can be seen from Tables 3 and 4, the performance of the STIRSAN model in the embodiments of the present invention is significantly higher than that of the baseline model. In the NYC dataset, STIRSAN is higher than the baseline model by 1.7%, 4.07%, 3.32%, 0.56% and 3.14% respectively in the five evaluation metrics of Recall@1, Recall@5, Recall@10, Recall@20 and MRR; in the TKY dataset, STIRSAN is higher than the baseline model by 2.48%, 2.23% and 1.11% respectively in the Recall@5, Recall@10 and MRR metrics, and is similar to the optimal baseline model in Recall@1 and Recall@20, with differences of 0.37% and 0.92% respectively.
[0147] As shown in Table 2, the number of parameters of the STIRSAN model is significantly lower than that of TMCA, DeepMove, LSTPM, and STAN, and is comparable to that of the TiSASRec model. STIRSAN achieves significantly superior performance with fewer parameters, indicating that the performance improvement of the STIRSAN model is achieved by mining the internal patterns in the check-in data rather than fitting a large number of parameters.
[0148] To verify the effectiveness of each module, the embodiments of the present invention set the following variants of the STIRSAN model for ablation experiments:
[0149] (1) STIRSAN w / o CAC: A variant of the STIRSAN model, which removes the causal convolution for enhancing local sequence information;
[0150] (2) STIRSAN w / o PMGP: A variant of the STIRSAN model, which removes the personalized multi-granularity periodic representation;
[0151] (3) STIRSAN w / o TIE: A variant of the STIRSAN model, which removes the influence of the time interval between check-ins;
[0152] (4) STIRSAN w / o DIE: A variant of the STIRSAN model, which removes the embedded representation of the geographical distance between check-ins.
[0153] The results of the ablation experiments are shown in Tables 5 and 6.
[0154] Table 5 Results of ablation experiments on the NYC dataset
[0155] Evaluation metrics R@1 R@5 R@10 R@20 MRR STIRSAN 0.1837 0.4414 0.5272 0.5753 0.2954 STIRSAN without PMGP 0.1801 0.4201 0.5115 0.5734 0.2899 STIRSAN without TIE 0.1810 0.4146 0.4940 0.5457 0.2881 STIRSAN without DIE 0.1837 0.4183 0.5032 0.5688 0.2906 STIRSAN without CAC 0.1625 0.4090 0.4903 0.5531 0.2722
[0156] Table 6 Results of ablation experiments on the TKY dataset
[0157]
[0158]
[0159] Tables 5 and 6 list the experimental results of each variant model of STIRSAN. It can be seen that a relatively consistent trend is shown on both the NYC and TKY datasets, that is, each module such as causal convolution, multi-granularity periodic information representation learning, time interval, and geographical distance representation contributes to the performance improvement of the STIRSAN model.
[0160] Using causal convolution to perform local sequence enhancement on Transformer has the greatest improvement on the model performance, with Recall@1 and MRR on the NYC dataset increased by 2.12% and 2.22% respectively; on the TKY dataset, they are increased by 2.18% and 2.1% respectively. Combining Tables 3 and 4, it can be seen that for STIRSAN w / oCAC without using causal convolution in the two datasets, Recall@10 and Recall@20 are higher than those of DeepMove and LSTPM based on LSTM, but the Recall@1, Recall@5 and MRR metrics are lower than those of the LSTM-based comparison models. In addition to STIRSAN proposed in this paper, the TiSASRec model based on Transformer also shows the same phenomenon, that is, Recall@10 and Recall@20 are significantly higher than those of comparison models such as DeepMove and LSTPM based on LSTM, but for the metrics Recall@1 and Recall@5 that can better illustrate the recommendation accuracy, the Transformer-based model is inferior to the LSTM-based model. This shows that Transformer is good at capturing long-term dependencies and is insensitive to local information. For the next POI recommendation task, the user's recent access, that is, short-term preference, is crucial for predicting the next POI. Therefore, in terms of accurate recommendation, the Transformer-based model is inferior to the LSTM-based model. Therefore, STIRSAN combines causal convolution with Transformer to strengthen the local information in the check-in sequence, so as to further explore the user's short-term preference, and the experimental results also prove that the Transformer enhanced by causal convolution greatly improves the recommendation accuracy.
[0161] Next, to verify the robustness of the model, the embodiments of the present invention compare SIRSAN with the baseline model under different dimensions and sequence lengths respectively. And the periodic representation and spatial relationship learned by the STIRSAN model are visualized to prove the interpretability of the model.
[0162] (1) Influence of dimension on experimental results
[0163] Figure 6.a to Figure 6.b Intuitively shows the trends of Recall@10 and Recall@20 of each model on the NYC dataset with the change of dimension; Figure 7 Shows the performance of the STIRSAN model in the TKY dataset under different vector dimensions. Generally speaking, the results of STIRSAN in the two datasets change relatively smoothly with the dimension change, and are higher than the baseline model in most cases, indicating that the STIRSAN model is relatively robust to the change of vector dimension.
[0164] (2) Influence of sequence length on experimental results
[0165] Figure 8 Intuitively shows the influence of the change in sequence length in the TKY dataset on the STIRSAN model. As the sequence length decreases, the performance of STIRSAN drops, but the change is relatively gentle. So the STIRSAN model is relatively robust to sequence length.
[0166] (3) Visualization of the relationships between different granularity periods
[0167] Figure 9.a to Figure 9.d Respectively show the heatmaps of the cosine similarities of the representations of different granularity periods learned by the STIRSAN model. It can be observed from the figures that the representations of adjacent time periods in each period are relatively similar. Taking Figure 9.d as an example, it can be clearly seen that 24 hours is roughly divided into 4 - 14 o'clock and 15 - 3 o'clock, and it can be further subdivided into 8 - 11 o'clock, 11 - 14 o'clock, 14 - 18 o'clock, 19 - 23 o'clock, etc. This is also relatively consistent with the division of time in daily life.
[0168] (4) Visualization of the spatial relationship E Δ of
[0169] Figure 10.a to Figure 10.d Show the heatmap of the geographical distance between the check - ins of four randomly selected users and the heatmap of the spatial relationship E Δ learned by the STIRSAN model. It can be clearly seen from the figures that the texture of the heatmap of the geographical distance is close to that of the heatmap of the spatial relationship E Δ learned by the model, while the numerical values show an opposite relationship. That is, Figure 10.a and Figure 10.b the greater the geographical distance, the smaller the corresponding value of E Δ ; the smaller the geographical distance, the greater the corresponding value of E Δ . This indicates that the model learns that the association between two check - ins with a smaller geographical distance is greater, and vice versa.
[0170] In summary, the next interest point recommendation model based on spatio - temporal information representation proposed in the embodiments of the present invention is superior to other comparative experiments in terms of recommendation performance, thus proving the effectiveness of the embodiments of the present invention and enabling it to be applied to the recommendation task of the next interest point; in addition, through ablation experiments, the effectiveness of the multi - granularity period representation, the representation of the time interval between check - ins, the geographical distance representation, and the causal convolution for enhancing local sequence information proposed in the embodiments of the present invention is verified.
[0171] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A next interest point recommendation system based on spatio-temporal information representation, characterized in that, Including: The time personalization module personalizes the periodic information of the monthly, weekly, daily, and hourly granularities of the personalized check-in timestamp. For the periodic representation of any one of the four granularities, the attention mechanism is used to adaptively combine the representation of the time period where the check-in moment t is located with the representation of the adjacent time period, so as to obtain the personalized representation of the time period of the granularity for the user u Using the The time personalization representation T of the moment t for the user u is calculated by using the attention mechanism u (t); A time encoding module that uses a time encoding function to map the check-in timestamp to a vector space to obtain a time encoding representation Φ(t); User check-in sequence module, with the T u (t), the Φ(t), and the embedded representation of the user u's check-in interest points as inputs, calculates the embedded representation S of the user check-in sequence u ; Causal Convolution Enhancement Module, which takes the embedded representation of the user u's check-in points of interest as input, combines causal convolution with Transformer to enhance the local information of the user check-in sequence, and obtains the enhanced embedded representation S' after causal convolution u ; The user's new check-in sequence module, using the S u , the S' u and the embedded representation E of the spatial relationship Δ , calculates the user's new check-in sequence representation Z u ; An output module that uses the new check-in sequence representation of the user u to calculate the preference of the user u for the point of interest at the moment t, and predicts and recommends the next point of interest.
2. The next point of interest recommendation system based on spatio-temporal information representation according to claim 1, wherein The personalized representation of the user u for the time period in which the granularity lies is as follows: where s ∈ S, S = {Month, DayOfWeek, Date, Hour} represents a set of four granularity period information, s represents an element in the set, that is, one of the four granularities, R represents the real number field, d is the dimension of the embedded representation, and for the granularity s, is the personalized representation for user u in the adjacent m time periods of the granularity: when m < 0, it represents the -m-th granularity period before the time period where the moment t is located; when m > 0, it represents the m-th granularity period after the time period where the moment t is located; the attention coefficient α t,m represents the similarity with , and r s represents the window size of adjacent time periods of the granularity s.
3. The next point of interest recommendation system based on spatio-temporal information representation according to claim 2, characterized in that, The personalized representation of user u for the adjacent m time periods of the granularity is as follows: Wherein, Among them, is the embedded representation of user u at granularity s at the moment t. is the embedded representation of user u in m adjacent time periods at granularity s at the moment t, Δ s (t, m) represents the m-th granularity time period adjacent to the moment t; when s represents month, t s represents the month in which the moment t is located, and f is 12; when s represents week, t s represents the week in which the moment t is located, and f is 7; when s represents day, t s represents the day in which the moment t is located, and f is 31; when s represents hour, t s represents the hour in which the moment t is located, and f is 24. However, different from the calculation methods of month, week, and day, for hour, since a day starts from 0 o'clock, so when s represents hour, Δ s (t, m) = (t s + m) mod f.
4. The next point of interest recommendation system based on spatio-temporal information representation according to claim 3, wherein The attention coefficient α t,m The calculation formula is as follows: Among them, r s represents the window size of adjacent time periods of the particle size s.
5. The next point of interest recommendation system based on spatio-temporal information representation according to claim 1, characterized in that The time personalization representation T of the moment t for the user u u (t) ∈ R d×1 is: Among them, s ∈ S, where S = {Month, DayOfWeek, Date, Hour} represents a set of four types of granularity period information, and s represents an element in the set, that is, one of the four granularities described above. represents the periodic representation attention coefficient of user u for granularity s. represents the personalized representation of the time period where the granularity s is located for user u, W u ∈ R d×d represents the transformation of the user representation, e u ∈ R d is the embedded representation of user u.
6. The next point of interest recommendation system based on spatio-temporal information representation according to claim 1, wherein The embedded representation S of the check-in sequence of the user u u is as follows: Where N is the number of check-ins of user u, For user u in t k Time sign-in means: in, For the user u at t k Points of interest to check in at any time The embedded representation of T u (t k ) is the t k The personalized time representation of the moment combines the periodic information of the four granularities, Φ(t k ) is the t k The sign-in time code indicates that pk For the t k The encoding representation of the sign-in location at the moment, e u ∈R d is the embedded representation of user u.
7. The next point of interest recommendation system based on spatio-temporal information representation according to claim 1, wherein The embedded representation E of the spatial relationship Δ ∈R N×N is as follows: E Δ = e ΔD W D Among them, E ΔD ∈R N×N×d is the embedded representation of the geographical distance between each check-in of user u, and the mapping matrix W D ∈R d×1 For E ΔD Perform dimensional transformation, where N is the number of check-ins of user u.
8. The next point of interest recommendation system based on spatio-temporal information representation according to claim 1, characterized in that The new check-in sequence representation Z of the user u u : Among them, mapping matrix W Q 、W K 、W V ∈R d×d , denotes the Hadamard product, M is a lower triangular masking matrix with elements all being 1, and T represents the transpose operation.
9. The next point of interest recommendation system based on spatio-temporal information representation according to any one of claims 1 or 8, characterized in that, A feed-forward neural network (FFN) is used to introduce non-linearity into the model: Among them, the weight parameters of the neural network are the threshold parameters of the neural network, and f1 and f2 represent the first layer network and the second layer network of the FFN; The optimized representation of the check-in sequence of the user u is as follows: where N is the number of check-ins of user u, denotes at time t k the optimized check-in representation of user u.
Citation Information
Patent Citations
Recurrent neural network interest site recommendation method based on space-time periodic attention mechanism
CN110399565A
Magellan: a context-aware itinerary recommendation system built only using card-transaction data
US20200387988A1