Next Point of Interest Recommendation Method Based on Spatiotemporal Information Representation
By adopting personalized spatiotemporal information representation and combination of causal convolution and Transformer in the recommendation technology of the next point of interest, the problem of insufficient utilization of spatiotemporal information in the prior art is solved, and the accuracy and performance of recommendations are significantly improved.
Patent Information
- Application Number
- CN202211552518.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-12-05
AI Technical Summary
The current next point of interest recommendation technology is difficult to effectively utilize space-time information, resulting in insufficient recommendation performance, especially in capturing time or geographical distance effects.
Using a recommendation method based on spatiotemporal information representation, through personalized four granular periodic representations (month, week, day, and hour) and time-encoding representations, combining causal convolution and Transformer model, local information of user check-in sequences is enhanced, and the geographical distance embedded representation of check-in points of interest is calculated to improve recommendation performance.
By effectively utilizing space-time information, the performance of recommendations for the next point of interest can be improved, and the user's movement patterns and geographical distance effects can be more accurately captured, improving the accuracy and user experience of recommendations.
Smart Images

Figure CN115982477B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of recommendation, and particularly to a method for recommending the next point of interest based on spatio-temporal information representation. Background Art
[0002] The popularization of intelligent devices and mobile Internet has promoted the growing prosperity of location-based social networks (LBSNs) such as Foursquare, Gowalla, Yelp, Twitter, WeChat, and Weibo, attracting a large number of users. People check in at locations through the check-in function provided by the LBSNs platform, share their dynamics and real-time locations with friends, and interact with friends by posting information such as opinions, photos, and comments related to the location. This social way has increasingly penetrated into the daily life of the public and gradually evolved into an important communication method in people's lives.
[0003] As one of the core functions of LBSNs, the next point of interest recommendation technology has a relatively wide range of application scenarios and plays a crucial role in people's lives. For the government, by predicting the next point of interest that people will visit, the government can design more reasonable traffic planning and scheduling strategies to alleviate traffic congestion and crowd gathering; for platforms such as ride-sharing and food delivery, the next point of interest prediction technology can accurately help drivers or food delivery riders effectively avoid congested roads and plan their trips in advance; for merchants, store information and coupons can be accurately distributed to target users who may visit, thereby avoiding blind large-scale advertising and achieving targeted advertising, saving advertising costs; for users, the next point of interest recommendation technology can assist users in making decisions and improve the user experience. As an independent sub-field of the recommendation system, the next point of interest recommendation has wide applications and can provide better user experience and third-party services for users. Therefore, it has received extensive attention from the academic and industrial communities in recent years.
[0004] In the next point of interest (POI) recommendation task, a user's interests can be divided into long-term preferences and short-term preferences. Long-term preferences refer to the comprehensive interests of the user mined from the user's historical trajectory, which rely on all historical check-in records; while short-term preferences refer to the user's interest preferences within a short period, which are affected by the recently visited POIs and are more inclined to the most recent check-in records. The time of user check-in contains two aspects of information. On the one hand, the timestamp reflects the absolute time when the user visits the POI, including cycles of different granularities such as year, month, week, day, hour, etc.; on the other hand, the time interval between check-ins reflects the degree of association between two check-ins. Therefore, reasonably using time information can better mine the user's movement pattern. To model the periodicity of human movement using time information, some models propose to represent time in an embedded way. For example, TMCA divides time into hourly granularity and differentiates between weekdays and weekends, and thus represents time as a one-hot vector with a dimension of 48 as the input of the model. However, this model only considers the cycles of hour and week, ignores the information with cycles of month and day, and assumes that adjacent time periods are independent of each other without considering their proximity. In addition, in the recommendation task, methods such as TMCA, TiSASRec, and STAN in model design utilize spatio-temporal information by learning the representations of time intervals and geographical distances. However, the current models have limited ability to learn the representations of time intervals and geographical distances and cannot reveal the impact of real time or geographical distance effects on human movement.
[0005] Different from digital product recommendations such as commodity recommendations, news recommendations, and music recommendations which are pure online interactions, user check-in behavior does not have implicit feedback data such as browsing and clicking. Human activities are affected by real-world factors and show complex transfer patterns, and the dataset itself is sparse. Therefore, the next POI recommendation task is highly challenging. Summary of the Invention
[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is to provide a method for recommending the next POI based on spatio-temporal information representation to improve the performance of the next POI recommendation.
[0007] On the one hand, the present invention provides a method for recommending the next POI based on spatio-temporal information representation, including the following steps:
[0008] Personalize the cycle information of the four granularities of month, week, day, and hour of the check-in timestamp. For the cycle representation of any one of the four granularities, use the attention mechanism to adaptively combine the representation of the time period where the check-in moment t is located with the representation of the adjacent time period to obtain the personalized representation of the time period of the granularity for user u;
[0009] Using the personalized representation of user u during the period when the granularity is located, the time personalized representation of the moment t for the user u is calculated by using the attention mechanism;
[0010] Use the time encoding function to map the check-in timestamp to the vector space, and calculate the time interval with the time encoding representation Φ(t) to measure the connection between two timestamps;
[0011] According to the time personalized representation of the user u, the time encoding representation, and the embedded representation of the user u's check-in point of interest, calculate the embedded representation S of the user's check-in sequence u ;
[0012] Combine causal convolution and Transformer to enhance the local information of the user's check-in sequence. Using the embedded representation of the user u's check-in point of interest as the input, obtain the embedding enhanced by causal convolution
[0013] ′
[0014] The formula represents S u ;
[0015] Calculate the embedded representation of the geographical distance of the check-in point of interest to obtain the embedded representation E of the spatial relationship Δ ;
[0016] ′
[0017] Combine the above S u 、the above S u 、the above E Δ to calculate the new check-in sequence representation Z of user u u ;
[0018] Use the new check-in sequence representation of the user u to calculate the preference of the user u for the point of interest at the moment t.
[0019] On the other hand, the present invention also provides an electronic device, including:
[0020] At least one processor; and,
[0021] A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the foregoing method.
[0022] Technical effects:
[0023] Aiming at the next point of interest recommendation problem based on spatiotemporal information representation, the present invention designs personalized four-granularity period representations, combines the encoded representation of time and the embedded representation of geographic distance, considers the representation of time interval and geographic distance when calculating the attention between check-ins, utilizes spatiotemporal information when modeling users' long-term preferences, and uses causal convolution for local information enhancement, thereby improving the performance of next point of interest recommendation.
[0024] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a schematic diagram of a single user's continuous sign-in activity in an embodiment of the present invention;
[0026] Figure 2 is a schematic diagram of a STIRSAN model according to an embodiment of the present invention;
[0027] Figure 3.a is a heat map of check-in POI categories with monthly granularity in the NYC dataset of an embodiment of the present invention;
[0028] Figure 3.b is a heat map of check-in POI categories with daily granularity in the NYC dataset of an embodiment of the present invention;
[0029] Figure 3.c is a heat map of check-in POI categories with weekly granularity in the NYC dataset of an embodiment of the present invention;
[0030] Figure 3.d is a heat map of check-in POI categories with hourly granularity in the NYC dataset of an embodiment of the present invention;
[0031] Figure 4 is a schematic diagram of causal convolution according to an embodiment of the present invention;
[0032] Figure 5.a is a trend chart of Recall@10 changing with dimension in the NYC dataset according to an embodiment of the present invention;
[0033] Figure 5.b is a trend chart of Recall@20 changing with dimension in the NYC dataset according to an embodiment of the present invention;
[0034] Figure 6 Schematic diagram of the impact of the dimension of the TKY dataset on STIRSAN in an embodiment of the present invention;
[0035] Figure 7 Schematic diagram of the effect of sequence length on STIRSAN in the TKY dataset of an embodiment of the present invention;
[0036] Figure 8.a It is a similarity heat map represented by the Month cycle in the embodiments of the present invention;
[0037] Figure 8.b It is a similarity heat map represented by the Date cycle in the embodiments of the present invention;
[0038] Figure 8.c It is a similarity heat map represented by the DayofWeek cycle in the embodiments of the present invention;
[0039] Figure 8.d It is a similarity heat map represented by the Hour cycle in the embodiments of the present invention;
[0040] Figure 9.a It is a heat map of Example 1 of geographical distance and spatial relationship in the embodiments of the present invention;
[0041] Figure 9.b It is a heat map of Example 2 of geographical distance and spatial relationship in the embodiments of the present invention;
[0042] Figure 9.c It is a heat map of Example 3 of geographical distance and spatial relationship in the embodiments of the present invention;
[0043] Figure 9.d It is a heat map of Example 4 of geographical distance and spatial relationship in the embodiments of the present invention. Detailed implementation manners
[0044] The following introduces multiple preferred embodiments of the present invention with reference to the accompanying drawings of the specification, making its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.
[0045] In the embodiments of the present invention, the variables and their mathematical expressions are defined as follows: The set of users in the check-in data is defined by U = u 1 , u 2 ,... u ||}, where u represents one of the elements, and the total number of users is |U|; The set of points of interest in the check-in data is defined by L = l 1 , l 2 ,... l ||}, and the total number of check-in points of interest is |L|; The geographical coordinates of the point of interest l k are represented by p k = (lon k ), at k ), that is, longitude and latitude. The k-th check-in of the user u is denoted as indicating that the user u visits the location at time t k . Denote the trajectory of each user as Clip the user trajectory to a fixed length where N is the maximum length of the specified trajectory. If N, only consider the most recent N check-ins; if N, pad with 0s on the left side of the sequence until the length of the sequence is N. Additionally, denote the time interval between the i-th and j-th check-ins as ΔT ij =|t i -t j |, and denote the geographical distance as
[0046] ΔD ij =Haversine(p i ,p j ), where the Haversine formula is as follows:
[0047]
[0048] where R represents the radius of the earth, which is 6371 km. Therefore, the geographical interval ΔD between any two pairs within the check-in sequence is expressed as:
[0049]
[0050] The next POI recommendation refers to predicting the POI that the user will visit at the next moment given the check-in record tra(u) of user u.
[0051] The embodiments of the present invention provide a method for recommending the next POI based on spatio-temporal information representation, including the following steps:
[0052] Personalize the periodic information of the month, week, day, and hour granularities of the check-in timestamps. For the periodic representation of any one of the four granularities, use the attention mechanism to adaptively combine the representation of the time period where the check-in moment t is located with the representation of the adjacent time period to obtain the personalized representation of the time period where the granularity is located for user u;
[0053] Using the personalized representation of the time period where the granularity is located for user u, use the attention mechanism to calculate the time personalized representation of the moment t for the user u;
[0054] Use the time encoding function to map the check-in timestamp to the vector space, and calculate the time interval with the time encoding representation Φ(t) to measure the connection between two timestamps;
[0055] According to the time personalized representation of the user u, the time encoding representation, and the embedded representation of the POI where the user u checks in, calculate the embedded representation S of the user check-in sequence u ;
[0056] Combine causal convolution with Transformer to enhance the local information of the user check-in sequence. Taking the embedded representation of the check-in points of interest of the user u as the input, obtain the enhanced embedded representation S' after causal convolution u ;
[0057] Calculate the embedded representation of the geographical distance between check-in points of interest to obtain the embedded representation E of the spatial relationship Δ ;
[0058] Combine the said S u 、the said S' u 、the said E Δ to calculate the new check-in sequence representation Z of the user u u ;
[0059] Utilize the said Z u to calculate the preference of the user u for the point of interest at the said moment t
[0060] In the embodiments of the present invention, it is utilized that people's travel shows multi-granularity periodicity, considering the personalized influence of multi-granularity period information on users, designing four kinds of personalized granularity period representations, and combining the encoded representation of time and the embedded representation of geographical distance. Respectively, based on the Bochner theory and the AutoDis method, the time interval and geographical distance between check-ins are encoded and embedded, overcoming the deficiencies of existing work, and considering the representations of the time interval and geographical distance between check-ins when calculating self-attention to capture spatio-temporal effects. In addition, this method uses causal convolution to enhance the Transformer's perception of local sequences, and uses causal convolution for local information enhancement, improving the performance of the next point of interest recommendation
[0061] The following introduces this embodiment in conjunction with the accompanying drawings Figure 1 It is a schematic diagram of the check-in activities of a single user in the embodiment dataset. Where p 1 , p 2 , p 3 ,..., p i is the sequence of points of interest; Δt 1 , Δt 2 , …, Δt i-1 is the time interval between two adjacent check-in records; Δd 1 , Δd 2 , …, Δd i-1 is the geographical distance between two adjacent points of interest. The embodiments of the present invention only consider using the user's recent N check-in data to model and learn the spatio-temporal information representation, and predict the next point of interest
[0062] Figure 2Schematic diagram of a preferred embodiment model structure of the present invention. The markings in the figure are defined as follows: DistanceInterval Embedding: Geographical distance embedded representation; Inputs: Input; Input Embedding: Embedded representation of the check-in point of interest; Causal Convolution Layer: Causal convolution layer; User embedding: User embedded representation; Time Interval Encoding: Time encoding representation; PMGP: Personalized representation encoding of user u at time t; Positional Encoding: Position encoding representation; Self-Attention Layer: Self-attention layer; Feed-Forward Network: Feed-forward network layer; Dropout: Model generalization; LayerNorm Layer: Layer normalization. Among them, Dropout means that during forward propagation, the activation value of a certain neuron stops working with a certain probability p, which can make the model more general and alleviate the occurrence of overfitting.
[0063] In the embodiment of the present invention, the original data is the check-in record of users, including information such as user ID, point of interest ID, type of point of interest, longitude and latitude of the point of interest, etc. The check-in record of users can be regarded as a sequence of points of interest arranged in time, and the next point of interest recommendation task can essentially be regarded as a sequence prediction problem. The original data set is grouped according to user ID and sorted in the order of check-in time, generating the check-in sequence data of each user. For each user, its [2,N]th check-in is used as the test set, and its [2,N - 1]th check-in is used as the input sequence to predict the check-in point of interest at the Nth time.
[0064] For user u, extract the sub-attributes of its check-in time: Month, DayOfWeek, Date, Hour. Obtain the embedded representation e of user u u , and randomly initialize the embedded representations of the four granularities of month, week, day, and hour to obtain the personalized four-granularity periodic information representation of user u.
[0065] Refer to Figures 3.a to 3.d It can be seen that the categories of the check-in points of interest shown in adjacent time periods are similar, so the representations of adjacent time periods should be relatively similar. Therefore, in the embodiment of the present invention, for granularity s, the attention mechanism is used to adaptively combine the representation of the time period where time t is located with the representations of adjacent time periods to obtain the personalized representation of the time period where the granularity is located for user u.
[0066] To obtain the final representation of user u at time t, that is, to obtain the representation of the personalized multi-granularity period, for the period representations of the four granularities, the attention mechanism is used to calculate the time-personalized representation of time t for user u.
[0067] Since time intervals play an important role in expressing time effects and revealing sequence patterns, the embodiments of the present invention consider the representation of time intervals, and calculate the time intervals with the time encoding representation Φ(t) to measure the connection between two timestamps.
[0068] Since Transformer is naturally good at capturing long-term and global dependencies in sequences, but Transformer is not good at extracting fine-grained local information, while convolutional operations can well explore local information. Therefore, in order to overcome the drawbacks of Transformer, STIRSAN combines causal convolution with Transformer to enhance the local information in the sequence. Different from traditional convolutional operations, causal convolution aims to ensure that there is no information leakage from future moments during the calculation process. To implement causal convolution, it is necessary to pad (kernel size - 1) zeros at the beginning of the sequence and truncate the redundant output at the end of the sequence to ensure that the sequence lengths before and after convolution are the same. In this way, the input of the convolution at time t only contains the information of itself and previous moments, and does not contain the information of time t + 1 and later. Figure 4 Schematic diagram of causal convolution with a kernel size of 3, and the result after causal convolution is Input of causal convolution Is the embedded representation of each check-in interest point of user u.
[0069] The embedded representation of geographical distance is another important factor in the problem of next interest point recommendation. For the embedded representation of continuous numerical features, a simple solution is to treat the numerical features as categorical features and assign an independent embedding vector to each numerical feature. However, this method has serious defects, such as a large number of parameters and insufficient training of low-frequency features. In order to reduce the model parameters, the domain embedding method shares one to three embedded representations for all feature values of the same type of feature, and obtains the representations of each feature value through transformations such as multiplying with the feature value and linear interpolation method. However, the capacity of this type of method model is relatively low, resulting in a decline in performance. Therefore, the present invention designs a meta-embedding H ∈ R K×d Shared by all geographical distances. Each meta-embedding h v ∈ R d Can be regarded as a subspace in the latent space to improve the expressive power and capacity of the model. In order to capture the complex connection between geographical distance and meta-embedding, a differentiable automatic discretization module is designed, and through the weighted average method, the geographical distance ΔD between the i-th check-in and the j-th check-in interest point is obtained.ij Embedded representation The weighted average method makes the relevant meta-embeddings more conducive to providing rich information, while the irrelevant meta-embeddings are largely ignored, thereby obtaining the embedded representation E of the geographical distance ΔD between each check-in of user u ΔD and further obtaining the embedded representation E of the spatial relationship Δ .
[0070] STIRSAN calculates self-attention to assign different weights to each check-in in the trajectory to aggregate the representations of historical check-ins. The input of the spatio-temporal aware self-attention is the embedded representation S of the user check-in sequence u the embedded representation S' after enhancing the local information causal convolution of the user check-in sequence u the embedded representation E of the spatial relationship Δ During the process of calculating the attention, in addition to calculating the similarity between the check-in points of interest, the influence of the time interval and geographical distance between each check-in is also considered, obtaining the new check-in sequence representation Z of user u u .
[0071] Using the new check-in sequence representation of user u obtained in the above steps, the latent factor model is used to calculate the preference of user u for the point of interest at the moment t; model training is carried out to learn the model parameters and predict the next point of interest
[0072] Let S = {Month, DayOfWeek, Date, Hour} represent the set of four granularity period information
[0073] s ∈ S, s represents an element in the set, that is, one of the four granularities. In a preferred embodiment of the present invention, for the granularity s, the embedded representation of user u at the moment t is where, when s represents the month, it is expressed as
[0074]
[0075] where, E Month (t) ∈ R d represents the embedded representation of the month at the moment t, R represents the real number field, d is the dimension of the embedded representation, ° represents the Hadamard product, W Month ∈ R d×d is the weight coefficient of the linear transformation, randomly initialized and continuously optimized during iteration, W MonthThe embedded representation of the user can be correspondingly transformed for personalizing the month representation, and in the same way, the personalized embedded representations of the personalized weekly, daily, and hourly granularity period information can be obtained.
[0076] For the granularity s, the personalized representation of the adjacent m time periods of the moment t for the user u is Then there is the following relational expression:
[0077]
[0078] Wherein,
[0079]
[0080] Wherein, is the embedded representation of the adjacent m time periods of the granularity s at the moment t for the user u. When m < 0, it represents the -m-th time period before the time period where the moment t is located; when m > 0, it represents the m-th time period after the time period where the moment t is located; when s represents the month, t s represents the month where the moment t is located, and f is 12; when s represents the week, t s represents the week where the moment t is located, and f is 7; when s represents the day, t s represents the day where the moment t is located, and f is 31; when s represents the hour, t s represents the hour where the moment t is located, and f is 24. However, different from the calculation methods of the month, week, and day, for the hour, since a day starts from 0 o'clock, so when s represents the hour, Δ s (t, m) = (t s + m) mod f, and mod is the modulo operator.
[0081] For the granularity s, the attention mechanism is adopted to adaptively combine the representation of the time period where the moment t is located with the representations of the adjacent m time periods. The personalized representation of the time period where the granularity s is located for the user u is:
[0082]
[0083] Wherein, corresponding to the granularity s, is the personalized representation of the time period where the moment t is located for the user u. Corresponding to the four granularities of the month, week, day, and hour, the personalized representations for the user u are respectively obtained as: is the personalized representation of the adjacent m time periods of the moment t for the user u: the attention coefficient α t,m represents the similarity between and
[0084] Preferably, the attention coefficient α t,m The calculation formula is:
[0085]
[0086] where r s represents the window size of adjacent time periods of the granularity s, and r s is set as: s represents month, r s = 2; s represents week, r s = 1; s represents day, r s = 6, s represents hour, r s = 5, exp represents the exponential function with the natural constant e as the base, and n represents the value within the window of adjacent time periods of the granularity s.
[0087] In another preferred embodiment of the present invention, the time personalized representation T u (t) ∈ R d×1 for the moment t with respect to the user u is calculated by using the attention mechanism as:
[0088]
[0089] where represents the periodic representation attention coefficient of the user u for the granularity s, and W u ∈ R d×d represents the transformation of the user representation, and e u ∈ R d is the embedded representation of the user u;
[0090] T u (t) ∈ R d×1 As the final time personalized representation of the moment t with respect to the user u, it contains the representation of personalized multi-granularity periodic information.
[0091] The time interval between user check-ins reflects the degree of association between two check-ins and plays an important role in expressing the time effect and revealing the sequence pattern. Reasonably using the time information can better mine the movement rules of users. The timestamp reflects the absolute time when the user visits the point of interest. It is necessary to process the timestamp information to express the time interval. In another preferred embodiment of the present invention, the time encoding function is used to map the timestamp to a vector, i.e., Φ: t → R d , then the time interval can be expressed as the dot product of the corresponding time encoding representations:
[0092] ψ(t 1 - t 2 ) = K(t 1 , t 2 ): = <Φ(t 1), Φ(t 2 ) > (1)
[0093] Among them, K is the time kernel function, and <,> represents the dot product operation. ψ(t 1 -t 2 ) measures the connection between two timestamps. Since this time encoding function directly encodes the timestamps, it can be generalized to any timestamp, and thus the representation of any time interval can be obtained:
[0094]
[0095] Based on the Bochner theory and Monte Carlo integration, Φ(t) in Equation (1) can be realized by Φ d (t). Among them, ω = [ω 1 ,..., ω d T are the parameters to be learned by the model. Therefore, the encoded representation of time t is Φ d (t), and the time interval is obtained from the vector inner product shown in Equation (1).
[0096] In another preferred embodiment of the present invention, using the time personalization representation of the user u, the time encoding representation, and the embedded representation of each check-in point of interest of the user u, the embedded representation of the check-in sequence of the user u is calculated as:
[0097]
[0098] Among them, N is the number of check-ins of the user u, is the representation of the user u checking in at time t k :
[0099]
[0100] Among them,
[0101] is the embedded representation of the point of interest where the user u checks in at time t k , T (t u ) is the time personalization representation at time t k , which integrates the periodic information of the four granularities, Φ(t k ) is the time encoding representation of the check-in time at time t k , e k is the encoding representation of the check-in location at time t pk , e k ∈R u is the embedded representation of the user u. d
[0102] In another preferred embodiment of the present invention, a differentiable automatic discretization module is designed to capture the complex relationship between geographical distance and meta-embedding:
[0103] γ ij = ReLU(W i ΔD ij )
[0104]
[0105] where V is the number of meta-embeddings, and ΔD ij is the geographical distance between the i-th check-in and the j-th check-in, and W 1 ∈ R V ×1 and W 2 ∈ R V×V are parameters to be learned by the model. α controls the proportion of residual connections. γ ij represents the result obtained after being transformed by a neural network layer, represents the correlation between the geographical distance ΔD ij and the meta-embedding H, where represents the correlation between the geographical distance ΔD ij and the v-th meta-embedding h v .
[0106] Therefore, the embedded representation ij of the geographical distance ΔD between the i-th check-in and the j-th check-in point of interest is:
[0107]
[0108] By means of weighted average, the embedded representation ij of the geographical distance ΔD between the i-th check-in and the j-th check-in point of interest is obtained. Then, the embedded representation of the geographical distance ΔD between each check-in of user u is E ΔD ∈ R N ×N×d . Then, the embedded representation E Δ ∈ R N×N of the spatial relationship is:
[0109] E Δ = E ΔD W D
[0110] where the mapping matrix W D ∈ R d×1 is used to transform the dimension of E ΔD .
[0111] The weighted average method makes the relevant meta-embeddings more conducive to providing rich information, while the irrelevant meta-embeddings will be largely ignored.
[0112] In another preferred embodiment of the present invention, the embedded representation of the user check-in sequence is S u The embedded representation of the user check-in sequence after causal convolution enhancement of local information is S' u The embedded representation E of the spatial relationship Δ , calculate the new check-in sequence representation Z of the user u u :
[0113]
[0114]
[0115] Where Mapping matrix W Q 、W K 、W V ∈R d×d , Denotes the Hadamard product. To ensure causality when calculating self-attention, that is, when predicting the point of interest at time t, information at time t+1 and later times will not be used, a lower triangular masking matrix M with elements of 1 is introduced, and T represents the transpose operation.
[0116] In another preferred embodiment of the present invention, a feed-forward neural network (FFN) is used to introduce non-linearity for the new check-in sequence representation Z of the user u u :
[0117]
[0118] Where The weight parameters of the neural network, Is the threshold parameter of the neural network, and f1 and f2 represent the first layer network and the second layer network of the FFN.
[0119] To accelerate the training process and prevent phenomena such as overfitting and gradient disappearance during training, techniques such as LayerNorm, Dropout, and residual connections are introduced. The optimized representation of the check-in sequence of the user u Is:
[0120]
[0121] Where N is the number of check-ins of the user u, Is time t kThe optimized check-in representation of user u, where k is the k-th check-in, and t k is the time of the k-th check-in.
[0122] Preferably, a latent factor model is used to calculate the preference of user u for a point of interest k at time t where
[0123]
[0124] Among them, is the representation of the point of interest and is the optimized check-in representation of user u at time t k-1 where t k-1 is the time of the (k - 1)-th check-in.
[0125] After model training to predict the next point of interest and obtaining the preferences of user u for each point of interest at time t, the points of interest are output in descending order of preference values, which can be used to recommend the next point of interest that the user is going to visit or predict each point of interest.
[0126] To learn the parameters of the model, the BPR loss function is adopted:
[0127]
[0128] where the training set sample D = (u, l i , l j , t) | u, l i , t) ∈ I u , l j ∈ I \ I u where the positive sample (u, l i , t) represents the point of interest l i checked in by user u at time t j , and the negative sample (u, l contains all the learnable parameters of the model, and the penalty term is used to prevent the model from overfitting. σ is the sigmoid function.
[0129] Next, for the STIRSAN model established in the embodiments of the present invention, two evaluation metrics, the recall rate Recall@K and the mean reciprocal rank (MRR), are used to evaluate the performance of the model.
[0130] (1) Recall rate Recall@K
[0131] Recall refers to the ratio of the positive samples predicted by the model to all positive samples. In the next POI recommendation, it refers to the ratio of the labels in the top K after sorting the predicted values of each POI by the model in descending order. The calculation method of Recall@K is as follows:
[0132]
[0133] where K ∈ {1, 5, 10, 20}, and S label respectively represent the top K POIs recommended by the model to the user and the POIs actually visited by the user, that is, the labels. Obviously, in the next POI recommendation task, there is only one POI visited by the user at the next time step, that is, |S label | = 1.
[0134] (2) MRR
[0135] MRR is the average of the reciprocals of the ranks of positive samples among all samples, which reflects the overall ranking ability of the model. The calculation method of MRR is as follows:
[0136]
[0137] where rank u represents the rank of the label of user u in the model recommendation list. The higher the rank of the label in the recommendation list, the higher the MRR value and the better the model performance.
[0138] In the embodiments of the present invention, the STIRSAN model is evaluated in two real-world datasets, Foursquare-NYC and Foursquare-TKY. The datasets are preprocessed by deleting inactive users with less than 5 check-ins and POIs with less than 5 check-ins. The dataset statistics are as follows:
[0139] Table 1 Dataset Statistics
[0140] dataset #User #POI #Check-in Foursquare-NYC 1083 9989 179468 Foursquare-TKY 2293 15177 494807
[0141] In the experiment, only the most recent N visits of each user are used. The dataset is divided into a training set and a test set. Among them, the [1, N - 1] check-ins of each user are used as the training set, and the [2, N] check-ins are used as the test set. Among them, the [2, N - 1] check-ins are used as the input sequence to predict the check-in POI at the Nth time.
[0142] Experimental Results
[0143] The present invention mainly uses spatio-temporal information representation for the recommendation of the next point of interest. There are mainly two experiments in this experimental part: (1) A comparative experiment between the model of the embodiment of the present invention and 5 baseline models. (2) Design an ablation experiment to verify the effectiveness of each module of the model of the embodiment of the present invention. (3) Design an experiment to verify the robustness and interpretability of the model.
[0144] In the comparative experiment of the performance of the next point of interest recommendation, we compare the STIRSAN model established in the embodiment of the present invention with the following models:
[0145] (1) TMCA model: Adopts an encoder-decoder structure based on LSTM, and proposes two attention mechanisms to adaptively select relevant historical check-ins and context factors.
[0146] (2) DeepMove model: Proposes to use the attention mechanism to obtain relevant information from historical check-ins to model long-term user preferences, and uses RNN to model short-term preferences.
[0147] (3) LSTPM model: Models long-term preferences by using the temporal and spatial correlations between the current trajectory and historical trajectories, and uses geo-dilated RNN to capture the geographical connections between non-consecutive check-ins when modeling short-term preferences.
[0148] (4) STAN model: A next point of interest recommendation model based on Transformer, which considers the time interval and geographical distance between non-consecutive check-ins, and uses linear interpolation to learn the representations of different time intervals and geographical distances.
[0149] (5) TiSASRec model: A sequence recommendation model based on Transformer, which performs personalized processing on the time intervals of different users to obtain relative time intervals, and considers the influence of different time intervals when calculating self-attention.
[0150] The hyperparameters of the STIRSAN model are set as: r M = 2, r W = 1, r D = 6, r H = 5, α = 0.1, the number K of meta-embedding representations is set to 30, and the convolutional kernel size of causal convolution is set to 5. The input sequence length is 100, the dimension of the embedded representation is 100, the hidden layer dimension is also set to 100, the number of layers of CAC-Transfomer is 2, and λ is 5e-5. The number of parameters of each model is shown in Table 2.
[0151] Table 2 Comparison of the number of parameters of each model
[0152] Models TMCA DeepMove LSTPM STAN TiSASRec STIRSAN Number of parameters 2,420,753 4,207,563 3,257,863 1,795,000 1,172,400 1,268,492
[0153] In the NYC and TKY datasets, the experimental results of STIRSAN and the baseline model in the next point-of-interest recommendation task are shown in Tables 3 and 4. The data in bold in the tables are the highest experimental results.
[0154] Table 3 Comparative Experiment on Recommendation Performance on NYC Dataset
[0155]
[0156] Table 4 Comparative Experiment on Recommendation Performance on TKY Dataset
[0157]
[0158] As can be seen from Tables 3 and 4, the performance of the STIRSAN model in the embodiments of the present invention is significantly higher than that of the baseline model. In the NYC dataset, STIRSAN is higher than the baseline model by 1.7%, 4.07%, 3.32%, 0.56% and 3.14% respectively in the five evaluation metrics Recall@1, Recall@5, Recall@10, Recall@20 and MRR; in the TKY dataset, STIRSAN is higher than the baseline model by 2.48%, 2.23% and 1.11% respectively in the Recall@5, Recall@10 and MRR metrics, while it is similar to the optimal baseline model in Recall@1 and Recall@20, with differences of 0.37% and 0.92% respectively.
[0159] Combined with Table 2, the number of parameters of the STIRSAN model is significantly lower than that of TMCA, DeepMove, LSTPM and STAN, and is comparable to that of the TiSASRec model. STIRSAN achieves significantly superior performance with fewer parameters, indicating that the improvement in the performance of the STIRSAN model is achieved by mining the internal laws of check-in data rather than fitting excessive parameters.
[0160] To verify the effectiveness of each module, the embodiments of the present invention set the following variants of the STIRSAN model for ablation experiments:
[0161] (1) STIRSAN w / o CAC: A variant of the STIRSAN model, that is, the causal convolution for enhancing local sequence information is removed;
[0162] (2) STIRSAN w / o PMGP: A variant of the STIRSAN model, that is, the personalized multi-granularity periodic representation is removed;
[0163] (3) STIRSAN w / o TIE: A variant of the STIRSAN model, that is, the influence of the time interval between check-ins is removed;
[0164] (4)STIRSAN w / o DIE: A variant of the STIRSAN model, i.e., the embedded representation without the geographical distance between check-ins.
[0165] The ablation experiment results are shown in Tables 5 and 6.
[0166] Table 5 Ablation experiment results on the NYC dataset
[0167] Evaluation metrics R@1 R@5 R@10 R@20 MRR STIRSAN 0.1837 0.4414 0.5272 0.5753 0.2954 STIRSAN w / o PMGP 0.1801 0.4201 0.5115 0.5734 0.2899 STIRSAN w / o TIE 0.1810 0.4146 0.4940 0.5457 0.2881 STIRSAN w / o DIE 0.1837 0.4183 0.5032 0.5688 0.2906 STIRSAN w / o CAC 0.1625 0.4090 0.4903 0.5531 0.2722
[0168] Table 6 Ablation experiment results on the TKY dataset
[0169] Evaluation metrics R@1 R@5 R@10 R@20 MRR STIRSAN 0.1635 0.3807 0.4732 0.5608 0.2653 STIRSAN w / o PMGP 0.1618 0.3742 0.4671 0.5517 0.2625 STIRSAN w / o TIE 0.1605 0.3777 0.4758 0.5560 0.2611 STIRSAN w / o DIE 0.1622 0.3663 0.4640 0.5578 0.2648 STIRSAN w / o CAC 0.1417 0.3598 0.4605 0.5451 0.2443
[0170] Tables 5 and 6 list the experimental results of each variant model of STIRSAN. It can be seen that there is a relatively consistent trend on the two datasets of NYC and TKY, that is, each module such as causal convolution, multi-granularity periodic information representation learning, and the representation of time intervals and geographical distances contributes to the improvement of the STIRSAN model performance.
[0171] Using causal convolution to perform local sequence enhancement on Transformer results in the greatest improvement in model performance, with Recall@1 and MRR increasing by 2.12% and 2.22% respectively for the NYC dataset; for the TKY dataset, they increase by 2.18% and 2.1% respectively. Combining Tables 3 and 4, it can be seen that for the STIRSAN w / oCAC without causal convolution in the two datasets, Recall@10 and Recall@20 are higher than those of the LSTM-based DeepMove and LSTPM, but the Recall@1, Recall@5, and MRR metrics are lower than those of the LSTM-based comparison models. In addition to the proposed STIRSAN in this paper, the TiSASRec model, which is also based on Transformer, also shows the same phenomenon, that is, Recall@10 and Recall@20 are significantly higher than those of the LSTM-based comparison models such as DeepMove and LSTPM, but for the metrics Recall@1 and Recall@5 that can better illustrate the recommendation accuracy, the Transformer-based models are inferior to the LSTM-based models. This indicates that Transformer is good at capturing long-term dependencies and is insensitive to local information. For the next point-of-interest recommendation task, the user's recent access, that is, short-term preferences, is crucial for predicting the next point of interest. Therefore, in terms of accurate recommendation, the Transformer-based models are inferior to the LSTM-based models. Therefore, STIRSAN combines causal convolution with Transformer to strengthen the local information in the check-in sequence, thereby further mining the user's short-term preferences, and the experimental results also prove that the Transformer enhanced by causal convolution greatly improves the recommendation accuracy.
[0172] Next, to verify the robustness of the model, the embodiments of the present invention compare SIRSAN with the baseline model at different dimensions and sequence lengths respectively. And the periodic representation and spatial relationship learned by the STIRSAN model are visualized to prove the interpretability of the model.
[0173] (1) Influence of dimension on experimental results
[0174] Figures 5.a to 5.b Intuitively shows the trends of Recall@10 and Recall@20 of each model with the change of dimension on the NYC dataset; Figure 6 Shows the performance of the STIRSAN model in the TKY dataset under different vector dimensions. Generally speaking, the results of STIRSAN in the two datasets change relatively smoothly with the change of dimension, and are higher than the baseline model in most cases, indicating that the STIRSAN model is relatively robust to the change of vector dimension.
[0175] (2) Influence of sequence length on experimental results
[0176] Figure 7 Intuitively shows the influence of the change in sequence length in the TKY dataset on the STIRSAN model. As the sequence length decreases, the performance of STIRSAN drops, but the change is relatively gentle. So the STIRSAN model is relatively robust to sequence length.
[0177] (3) Visualization of the relationships between different granularity periods
[0178] Figures 8.a to 8.d Respectively show the heatmaps of the cosine similarity of the representations of different granularity periods learned by the STIRSAN model. From the figures, it can be observed that the representations of adjacent time periods in each period are relatively similar. Taking Figure 8.d as an example, it can be clearly seen that 24 hours is roughly divided into 4 - 14 o'clock and 15 - 3 o'clock, and it can be further divided into 8 - 11 o'clock, 11 - 14 o'clock, 14 - 18 o'clock, 19 - 23 o'clock, etc. This is also relatively consistent with the division of time in daily life.
[0179] (4) Visualization of spatial relationship E Δ of
[0180] Figures 9.a to 9.d Shows the heatmap of the geographical distance between the check - ins of four randomly selected users and the heatmap of the spatial relationship E Δ learned by the STIRSAN model. It can be clearly seen from the figure that the texture of the geographical distance heatmap is close to that of the heatmap of the spatial relationship E Δ learned by the model, while the numerical values show an opposite relationship. That is, Figure 9.a and Figure 9.b the greater the geographical distance, the smaller the corresponding E Δ value; the smaller the geographical distance, the greater the corresponding E Δ value. This indicates that the model learns that the association between two check - ins with a smaller geographical distance is greater, and vice versa.
[0181] In summary, the next interest point recommendation model based on spatio - temporal information representation proposed in the embodiments of the present invention is superior to other comparative experiments in terms of recommendation performance, thus proving the effectiveness of the embodiments of the present invention and enabling it to be applied to the recommendation task of the next interest point; in addition, through ablation experiments, the effectiveness of the multi - granularity period representation, the representation of the time interval between check - ins, the representation of geographical distance, and the causal convolution for enhancing local sequence information proposed in the embodiments of the present invention is verified.
[0182] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field according to the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A method for recommending the next point of interest based on spatio-temporal information representation, characterized in that, it includes: Periodic information of four granularities of month, week, day, and hour of personalized check-in timestamps. For the periodic representation of any one of the four granularities, the attention mechanism is used to adaptively combine the representation of the time period where the check-in moment t is located with the representation of the adjacent time period to obtain the personalized representation of the time period of the granularity for the user u; Using the personalized representation of the time period of the granularity for the user u, the attention mechanism is used to calculate the time personalized representation of the moment t for the user u; The time encoding function is used to map the check-in timestamp to the vector space, and the time interval is calculated with the time encoding representation Φ(t) to measure the connection between two timestamps; Calculate the embedded representation S of the user check-in sequence based on the time-personalized representation for user u at the given time t, the time-encoded representation, and the embedded representation of the POI where user u checks in u ; Combine causal convolution and Transformer to enhance the local information of the user check-in sequence, and take the embedded representation of the check-in points of interest of the user u as the input to obtain the enhanced embedded representation S' after causal convolution u ; Calculate the embedded representation of the geographical distance of the check-in interest points to obtain the embedded representation E of the spatial relationship Δ ; Combined with the said S u and the said S' u and the said E Δ , calculate the new check-in sequence representation Z u ; Using the said Z u , calculate the preference of user u for the point of interest at the said moment t.
2. The method for recommending the next point of interest based on spatio-temporal information representation according to claim 1, characterized in that, The personalized representation of the user u for the time period in which the granularity lies is as follows: where s ∈ S, S = {Month, DayOfWeek, Date, Hour} represents a set of four granularity period information, s represents an element in the set, that is, one of the four granularities, R represents the real number field, d is the dimension of the embedded representation, and for the granularity s, is the personalized representation of the adjacent m time periods of the granularity for the user u: when m < 0, it represents the -m-th granularity period before the time period where the moment t is located; when m > 0, it represents the m-th granularity period after the time period where the moment t is located; the attention coefficient α t,m represents the similarity with , and r s represents the window size of the adjacent time period of the granularity s.
3. The method for recommending the next point of interest based on spatio-temporal information representation according to claim 2, characterized in that, The personalized representation of user u for the adjacent m time periods of the granularity is as follows: wherein, Among them, is the embedded representation of user u at granularity s at time t. is the embedded representation of user u at m adjacent time periods of granularity s at time t, Δ s (t, m) represents the m-th granularity time period adjacent to time t; when s represents month, t s represents the month in which time t is located, and f is 12; when s represents week, t s represents the week in which time t is located, and f is 7; when s represents day, t s represents the day in which time t is located, and f is 31; when s represents hour, t s represents the hour in which time t is located, and f is 24. However, different from the calculation methods of month, week, and day, for hours, since a day starts from 0 o'clock, so when s represents hour, Δ s (t, m) = (t s + m) mod f.
4. The method for recommending the next point of interest based on spatio-temporal information representation according to claim 3, characterized in that, The attention coefficient α t,m The calculation formula is as follows: Among them, r s represents the window size of adjacent time periods of the granularity s.
5. The method for recommending the next point of interest based on spatio-temporal information representation according to claim 1, characterized in that, The time personalization representation T of the user u at the moment t u (t) ∈ R d×1 is as follows: Among them, s ∈ S, where S = {Month, DayOfWeek, Date, Hour} represents a set of four types of granularity period information, and s represents an element in the set, that is, one of the four types of granularity. represents the periodic representation attention coefficient of user u for granularity s. represents the personalized representation of the time period where the granularity s is located for user u, W u ∈ R d×d represents the transformation of the user representation, e u ∈ R d is the embedded representation of user u.
6. The method for recommending the next point of interest based on spatio-temporal information representation according to claim 1, characterized in that, The embedded representation of the check-in sequence of the user u is: where N is the number of check-ins of user u, is the representation of user u's check-in at time k t: wherein, The embedded representation of the point of interest signed in by the user u at time t k , and T is the time personalization representation at time t u (t k ), which incorporates the periodic information of the four granularities. Φ(t k ) is the encoded representation of the check-in time at time t, and e k is the encoded representation of the check-in location at time t. e k ∈R pk is the embedded representation of user u. k is the encoded representation of the check-in location at time t. u ∈R d is the embedded representation of user u.
7. The method for recommending the next point of interest based on spatio-temporal information representation according to claim 1, characterized in that, The embedded representation E of the spatial relationship Δ ∈R N×N is as follows: E Δ = E ΔD W D where E ΔD ∈R N×N×d is the embedded representation of the geographical distance between each pair of check-ins of user u, and the mapping matrix W D ∈R d×1 is used to transform the dimension of E ΔD , and N is the number of check-ins of user u.
8. The method for recommending the next point of interest based on spatio-temporal information representation according to claim 1, characterized in that, The new check-in sequence representation Z of the user u u : Among them, mapping matrix W Q 、W K 、W V ∈R d×d , o represents the Hadamard product, M is a lower triangular masking matrix with elements all being 1, and T represents the transpose operation.
9. The method for recommending the next point of interest based on spatio-temporal information representation according to any one of claims 1 or 8, characterized in that, The feed-forward neural network (FFN) is used to introduce non-linearity into the model: Among them, the weight parameters of the neural network, are the threshold parameters of the neural network, where f1 and f2 represent the first layer network and the second layer network of the FFN; The optimized representation of the check-in sequence of the user u is as follows: Where N is the number of check-ins of user u, Indicates that at time t k The optimized sign-in representation of user u, 10. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of the foregoing claims 1-9.
Citation Information
Patent Citations
Continuous interest point recommendation method based on check-in time interval mode
CN109492166A
Magellan: a context-aware itinerary recommendation system built only using card-transaction data
US20200387988A1