Interest point recommendation method based on dynamic hierarchical attention fusion space-time network

By integrating dynamic hierarchical attention with spatiotemporal networks, we solve the problems of inaccurate spatiotemporal modeling and shallow semantic understanding in POI recommendation, explicitly separate the long-term and short-term dynamics of user behavior, and achieve accurate prediction and personalized recommendation of the next POI.

CN120780923APending Publication Date: 2025-10-14BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510863424.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing POI recommendation technologies suffer from inaccurate spatiotemporal modeling, weak semantic understanding of POIs, and improper dynamic handling of user mixed behaviors, which results in limited recommendation performance.

Method used

A method based on dynamic hierarchical attention fusion spatiotemporal network is adopted. Continuous spatial influence modeling and temporal association are performed through an adaptive spatiotemporal dual-graph convolutional network. A multi-level semantic hierarchy is constructed by combining a bidirectional hierarchical graph attention network. The attention-frequency fusion Transformer is used to process user behavior sequences and explicitly separate long-term and short-term dynamic signals.

Benefits of technology

It achieves accurate prediction of user points of interest, improves the accuracy and robustness of recommendations, and can fully capture complex spatiotemporal, semantic and user behavior dynamic information, provide personalized recommendations and optimize business intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780923A_ABST
    Figure CN120780923A_ABST
Patent Text Reader

Abstract

The invention provides an interest point recommendation method based on a dynamic hierarchical attention fusion space-time network, and the method comprises the following steps: 1, obtaining user historical sign-in and interest point feature data, and carrying out the preprocessing of the user data; 2, learning robust spatio-temporal information representation of the interest points through an adaptive spatio-temporal double-graph convolutional network module; 3, constructing a'interest point-functional unit-regional unit 'multi-level semantic hierarchical structure, and learning hierarchical semantic information representation of interest points by using a bidirectional hierarchical graph attention network module; 4, after fusing the two interest point representations, coding a user track and inputting the user track into an attention-frequency fusion Transform module, so as to obtain a multi-level semantic hierarchical structure; the module separates long and short term dynamic signals in a user behavior sequence through frequency analysis and dynamically fuses the long and short term dynamic signals with a self-attention mechanism to predict a next point of interest; and 5, training by adopting a combined loss function. According to the method, the complex characteristics of the user and the POI are effectively modeled, and the accuracy and the capability of capturing the dynamic preference of the user are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of recommendation systems and data mining, and particularly relates to a technology for predicting next interest point based on user spatio-temporal trajectory data. BACKGROUND

[0002] With the popularity of intelligent mobile devices and the advancement of positioning technology, location-based services (LBS) have been deeply integrated into various fields of modern society. The spatio-temporal check-in data generated by users through various applications constitutes valuable digital trajectories. These multi-dimensional data not only record the geographical movement characteristics of users, but also imply rich individual behavior patterns and interest preference information. In this context, as a core application of spatio-temporal data mining, next interest point recommendation predicts the next destination of users through their historical visit sequences, which is of great significance to improving personalized experience, optimizing urban planning, and empowering business intelligence.

[0003] In the development of interest point recommendation technology, traditional methods have limitations in handling complex user data and context associations. In recent years, deep learning technology, especially graph neural networks, has shown significant potential and received extensive attention in the field of next interest point recommendation due to its strong relationship modeling capability. However, despite the progress made by graph neural network (GNN) based methods, there are still several key limitations in capturing the complex intrinsic characteristics of check-in data:

[0004] Firstly, existing GNN-based methods fail to accurately model the continuous decay characteristics of spatial influence and robust temporal correlation when constructing the interest point relationship graph. In the spatial dimension, these methods generally use fixed distance thresholds or k-neighbor strategies to generate binary adjacency relationships. This processing approach ignores the natural law that geographic influence usually decays continuously and smoothly with distance, which may lead to the loss of key spatial information. In the time dimension, the coarse-grained time aggregation strategy adopted by some methods may mask subtle but important periodic changes in user behavior, and simple time correlation methods are difficult to effectively distinguish between real user behavior patterns and accidental noise in the data, making the learned temporal relationship representation unreliable and inaccurate.

[0005] Second, current methods fail to fully capture the inherent complexity of POIs when understanding and representing them, particularly with regard to their semantic hierarchical structure and multiple affiliations. First, models generally treat POIs as flat, independent nodes, ignoring their inherent hierarchical structure. A POI may belong to a specific functional unit, which in turn belongs to a broader regional context, i.e., "POI-Functional Unit-Region Unit." This simplification of the POI hierarchical structure not only limits the model's representational capabilities but also exacerbates the over-smoothing of node features common during the learning process of graph neural networks. Furthermore, POIs in the real world often have multiple semantic affiliations; a single POI may simultaneously serve multiple functions, such as entertainment, cultural landmarks, and social spaces. Most current models assume a single semantic affiliation for POIs, making it difficult to effectively capture the differentiated representations of POI nodes across different semantic dimensions and their multimodal semantic features. This limits their performance in downstream recommendation tasks when handling complex scenarios.

[0006] Third, existing mainstream sequence models struggle to effectively handle the complex interweaving of long-term stable habits (such as daily commuting) and short-term, immediate intentions (such as temporary travel) within user behavior sequences. Advanced sequence models such as the Transformer excel at capturing long-range dependencies, but their mechanisms typically focus on learning holistic temporal patterns, making it difficult to explicitly distinguish and effectively integrate long-term preferences and short-term dynamics in user behavior. This can smooth out or overlook key short-term signals reflecting a user's immediate intentions, reducing the performance of next point of interest recommendations. Summary of the Invention

[0007] To address the problems of inaccurate spatiotemporal relationship modeling of points of interest, insufficient semantic understanding of points of interest, and improper dynamic processing of mixed user behaviors in the background art, this paper proposes a point of interest recommendation method based on a dynamic hierarchical attention fusion spatiotemporal network. The present invention adopts the following technical solutions:

[0008] A method for recommending points of interest based on a dynamic hierarchical attention fusion spatiotemporal network, comprising the following steps:

[0009] S1. Input the user's historical check-in data, where each check-in record contains the user ID, point of interest ID and check-in timestamp;

[0010] S2. Sort the user's historical check-in records by access time, perform data preprocessing, obtain the user check-in sequence, and divide the user's track into preset time intervals of 24 hours;

[0011] S3. Extract feature data of each point of interest, wherein the feature data includes the category of the point of interest, the historical visit frequency, and the geographic longitude and latitude coordinates;

[0012] S4, input the feature data of the interest points into an adaptive spatio-temporal double-graph convolution network module, which models the continuous attenuation of the spatial influence between interest points through a parameterized Gaussian kernel, and combines a filtered time cosine similarity calculation to enhance the robust capture of the temporal relationship between interest points, thereby learning and outputting a robust spatio-temporal information representation of each interest point;

[0013] S5, input the feature data of the interest points into a bidirectional hierarchical graph attention network module, which constructs a multi-level semantic hierarchy of "interest point - functional unit - regional unit", and processes the multiple semantic membership characteristics of the interest points using Gaussian mixture clustering, and then captures deep cross-scale semantic information through a bidirectional information propagation graph attention mechanism, thereby learning and outputting a hierarchical semantic information representation of each interest point;

[0014] S6, fuse the robust spatio-temporal information representation of the interest points obtained in step S4 with the hierarchical semantic information representation of the interest points obtained in step S5 to generate a comprehensive feature representation of the interest points;

[0015] S7, encode the user trajectory obtained in step S2 to form a vectorized representation of the user trajectory;

[0016] S8, input the vectorized representation of the user trajectory into an attention-frequency fusion Transformer, which uses frequency analysis to explicitly separate long-term and short-term dynamic signals in the user behavior sequence, and fuses with a self-attention mechanism, thereby predicting the probability of the user visiting each candidate interest point in the next step after the current trajectory;

[0017] S9, use a preset loss function as the optimization objective to train the overall model containing steps S1 to S8 until the model converges, obtaining the final interest point recommendation model

[0018] As a preferred, the user check-in record in S1 is in the following form:

[0019] q = (u, p, t)

[0020] Where u is the check-in user ID, p is the visited interest point ID, and t is the timestamp of the record.

[0021] As a preferred, the data processing in S2 is as follows:

[0022] First, exclude small interest points with less than 10 check-in records, and filter out users with less than 10 check-in records, and obtain the filtered historical check-in records, and count the visit frequency of each interest point, and sort the user's historical check-in records according to the visit time to obtain the user's check-in sequence, and divide the user's check-in sequence into trajectories in units of 24h

[0023] As preferred, the point of interest feature data extracted in S3 is as follows:

[0024] f = (lat, lon, cat, cnt)

[0025] Where lat is the latitude of the point of interest, lon is the longitude, cat is the category, and cnt is the visit frequency of the point of interest. All points of interest constitute an initial feature embedding matrix X of the point of interest.

[0026] As preferred, the process of outputting robust spatio-temporal information representation in S4 through an adaptive spatio-temporal dual graph convolution network is as follows:

[0027] a. Points of interest that are spatially adjacent usually have a stronger influence relationship. In order to quantify this influence, the distance D ij between the point of interest pair (v i , v j ) is calculated. The latitude and longitude coordinates of the point of interest pair are (lat i , lon i ) and (lat j , lon j ), respectively, and E is the radius of the earth, which is calculated by the Haversine formula:

[0028]

[0029] The distance D ij is converted into a spatial weight A space (i,j) using a parameterized Gaussian kernel function to model the continuous decay of influence:

[0030]

[0031] Where σ is the bandwidth parameter of the Gaussian kernel, and its value is the median distance between points of interest in the data set.

[0032] b. Construct a time correlation matrix between points of interest, define the time series T i = {t1, t2, …, t n} as the historical visit timestamp set of POI i , divide the time slots into 30-minute intervals for a week, count the visit frequency of each point of interest in each time slot, obtain the initial activity vector, focus on the relative visit pattern, normalize the original activity vector by L2 norm, obtain the point of interest time activity distribution vector, and calculate the time correlation degree between points of interest using cosine similarity:

[0033]

[0034] The original cosine similarity is easily disturbed by low active POIs or accidental co-occurrence factors. To filter out such false associations, a reliability indicator function I reliability :

[0035]

[0036] Here n i , n j represent the number of active time slots of POI i and POI j , respectively, and n ij is the number of common active slots, and the values of a and b are 3 and 5, respectively. This process aims to filter accidental co-occurrence phenomena and improve the reliability of the constructed time association. The final time association matrix is:

[0037] A time (i,j)=cos(i,j)·I reliability

[0038] The spatial weight matrix of the interest point and the time association matrix are normalized as follows to obtain and

[0039]

[0040] where D is the degree matrix, and I N is the identity matrix.

[0041] c. Through the L1=3 layer of the two-channel graph convolution message passing, its propagation mechanism is as follows:

[0042]

[0043] where γ space and γ time are learnable scaling parameters, and are learnable weight matrices of specific modalities, and are bias vectors, and the above parameters are iteratively trained through back propagation and optimizer, and finally converged to obtain. LayerNorm is a layer normalization operation.

[0044] The features learned by the spatial path and the time path are fused to obtain the fused features

[0045]

[0046] where α (l)The learnable fusion coefficient is used to adjust the relative contribution in the fusion process, and is a leaky ReLU activation function with a negative slope parameter of 0.2. Through 3-layer propagation, the spatiotemporal semantic representation of the interest point is obtained

[0047]

[0048] As preferred, the hierarchical semantic information flow output by the bidirectional hierarchical graph attention network in step S5 is as follows:

[0049] a. Construct a hierarchical heterogeneous graph. The probability soft clustering feature of the Gaussian Mixture Model (GMM) is used to construct a three-level hierarchical heterogeneous graph composed of “interest point-function unit-region unit” from bottom to top.

[0050] The first GMM model is fitted by the EM algorithm, and the interest points are divided into K1 medium-grained function units Each interest point node v i , whose feature is x i (the i-th row of the initial feature embedding matrix X of the interest point) belongs to the k-th function unit The conditional probability is:

[0051]

[0052] Where π k , μ k , Σ k represent the mixing weight, mean vector and covariance matrix of the k-th Gaussian distribution, respectively, and π k , μ k , Σ k are the model parameters learned and converged iteratively based on the feature x i of the interest point by the EM algorithm. The probability density value of the i-th interest point under the k-th Gaussian distribution is represented. Based on this posterior probability distribution, the initial feature of each function unit is obtained by weighting the feature x i associated with it:

[0053]

[0054] is the feature of the k-th function unit, and N is the number of interest points in the data set. The embeddings of all function units together constitute the feature matrix of the function unit layer pc1 , and the probability soft connection matrix A between the interest point and the function unit is established, whose elements areBoolean matrix A′ pc1 is constructed by setting the non-zero positions of A pc1 to 1 and the other positions to 0. To explicitly model the semantic associations of the same level, a pair-wise similarity matrix S is computed based on the initial feature embeddings of the functional units c1 and a sparse intra-level adjacency matrix A cc1 is constructed based on a threshold τ = 70%.

[0055]

[0056] Similarly, the set of functional units V c1 is further aggregated into K2 coarse-grained region units, the conditional probability that the k-th functional unit belongs to the m-th region unit is computed, and the feature of the region unit is generated based on this.

[0057]

[0058] where is the initial feature of the k-th functional unit and the embedding of all region units is stacked together to get and a probability soft connection matrix A c1c2 between functional units and region units is established, whose element f is the conditional probability that the k-th functional unit belongs to the m-th region unit. Boolean matrix A′ c1c2 is constructed by setting the non-zero positions of A c1c2 to 1 and the other positions to 0. The similarity matrix S c2 of the center node is computed, and a sparse adjacency matrix A cc2 is constructed based on a threshold τ = 80%.

[0059]

[0060] A pp is constructed in the same way as A cc1 and A cc2 , and the initial feature embedding of the interest point is Through the above process, a hierarchical semantic unit heterogeneous graph containing {V p , V c1 , V c2} three-level nodes and {A pp , A′ pc1 , A cc1 , A′ c1c2 , A cc2} five types of connections is finally constructed.

[0061] b. The graph attention mechanism with bidirectional information propagation is used to capture and fuse semantic information of different levels and granularities. The initial features of the network are and

[0062] For any node at level s∈{p,c1,c2}, given its input feature at layer l is Intra-layer message passing passes through the graph attention network and utilizes the intra-layer index edge_index corresponding to the layer s s Execute. The edge index edge_index in the layer s It is derived from the adjacency matrix of the corresponding level s. When the level s = p, the corresponding matrix is ​​A pp , when s=c1, the corresponding matrix is ​​A cc1 , when s=c2, the corresponding matrix is ​​A cc2 This process enhances the consistency of representation of nodes at the same level and the ability to perceive local structures. The specific calculation method of intra-layer message transmission is defined as follows:

[0063]

[0064] Among them, GATConv intra It is a standard graph attention convolution module used to process information interaction within the same node set. The feature representation of the node at level s is updated after the intra-layer message passing operation at layer l;

[0065] Cross-level bottom-up information propagation is along the heterogeneous edge index edge_index up Perform cross-layer attention aggregation. The inter-layer index edge index edge_index up It is derived from the adjacency matrix of the corresponding level. When the level s=p, the corresponding matrix is ​​A′ pc1 , when the level s=c1, the corresponding matrix is ​​A′ c1c2 This pathway is responsible for effectively aggregating and abstracting the fine-grained features of the low-level s into the node representation of the high-level s+1 to which it belongs, realizing the step-by-step abstraction of semantics. The process is defined as:

[0066]

[0067] Among them, GATConv up It is a standard graph attention convolution module used to achieve bottom-up information transfer. The feature representation of the node at level s+1 is updated after the bottom-up message passing operation at level l;

[0068] Top-down information propagation utilizes reverse heterogeneous edge index edge_index down with semantic guidance, the inter-layer index edge_index down is derived from the transpose of the adjacency matrix of the corresponding hierarchy, when hierarchy s = c1, the corresponding matrix is when hierarchy s = c2, the corresponding matrix is This process propagates and refines the macro semantic context information of higher hierarchy to lower hierarchy, providing global perspective and constraint for fine-grained node features, realizing the context awareness and adaptation of features, this process is defined as:

[0069]

[0070] where CATConv down is the standard graph attention convolution module, which is used to realize top-down information transmission. denotes the feature representation of the node at hierarchy s after the l-th layer of top-down message passing operation

[0071] is the feature vector captured by the three different paths, and the corresponding feature vectors are spliced, for the boundary layer, the missing path features are filled with zero vectors, to obtain the spliced feature matrix

[0072]

[0073] where || denotes the splicing operation, and the linear transformation layer is projected to the target dimension to obtain the fusion feature matrix

[0074]

[0075] where W f is the linear transformation matrix, is the bias vector. W f and are iteratively trained by back propagation and optimizer, and finally converged.

[0076] After feature fusion, batch normalization and residual connection are used for update:

[0077]

[0078] where σ is the Relu activation function, is the fusion feature matrix output by the last layer network, and BatchNorm is the batch normalization operation.

[0079] After L2=2 layer message passing, the final hidden representation of each level node is obtained in is the hierarchical semantic information representation of interest points

[0080] In order to further optimize the learned hierarchical semantic representation, an auxiliary loss function L is introduced aux , which is composed of the layer contrast loss L contrastive and the level consistency loss L consistency constitute.

[0081] The hierarchical contrast loss aims to enhance the cross-level semantic discrimination of node embeddings at different levels. Take the contrast loss from the interest point layer to the functional unit layer as an example. First, based on the soft assignment probability of GMM, for each POI i Determine hard-allocated functional units POI i The click similarity between the embedding of and the embedding of functional unit k is Contrastive loss uses infoNCE form

[0082]

[0083] Where N is the number of points of interest, K1 is the number of functional units, and τ is the temperature coefficient and is set to 0.1. Using the same processing method as above, the contrast loss from functional units to regional units can be defined as The total hierarchical contrast loss is the weighted sum of the losses at each level.

[0084]

[0085] where λ pc1 and λ c1c2 The values ​​of are all set to 0.5.

[0086] The layer consistency loss aims to use the soft assignment information of GMM to ensure the semantic coherence between layers. Take the consistency loss from interest point layer to functional unit layer as an example: the expected functional unit layer embedding of interest point i is All functional units to which it is soft-assigned are embedded By probability A pc1 The weighted average of [i,k] is:

[0087]

[0088] Where K1 is the number of functional units. The consistency loss is achieved by minimizing the interest point embedding The mean squared error between the embedding and its expected value is:

[0089]

[0090] where N is the number of interest points. Similarly, the consistency loss from the functional unit layer to the region unit layer can be defined as The total hierarchical consistency loss is the weighted sum of each hierarchical loss

[0091]

[0092] where λ′ pc1 and λ′ c1c2 are set to 0.1, the total loss of this part is

[0093]

[0094] As a preferred, S6 fuses the robust spatio-temporal information representation of interest points with the hierarchical semantic information representation of interest points, in the following way:

[0095] The hierarchical view embedding and the spatio-temporal view embedding are interacted in a non-linear way:

[0096]

[0097] where W1, W2 are learnable shared weight matrices, and α is a learnable weighting coefficient, whose range is [0, 1], all the above parameters are iteratively trained through backpropagation and optimizer, and finally converged. σ is the ReLU activation function. The enhanced views are spliced, and an attention score is calculated through a learnable linear layer, and these scores are converted into an attention weight matrix A using the Softmax function:

[0098]

[0099] where W a is a learnable linear transformation matrix, which is also obtained by training convergence. The final fused representation e p is:

[0100]

[0101] where A :,0 and A :,1 are the first column and the second column of the attention weight matrix A respectively, and ⊙ represents the element-wise multiplication operation with broadcast mechanism.

[0102] As a preferred, S7 user trajectory vectorization representation is as follows:

[0103] A single check-in record in the user trajectory is embedded and represented as:

[0104]

[0105] where, e represents the embedding representation of check-in record i u , e c , e t and represent the embedding representation of user u, interest point category c, check-in time t and POI i , respectively, is the concatenation of embedding dimensions. The embedding of the trajectory is composed of all the check-in points of the trajectory

[0106] As a preference, S8 inputs the vectorized representation of the user trajectory into an attention-frequency fusion Transformer module, which is as follows:

[0107] The trajectory embedding S is input into a self-attention encoder to obtain the features in the time domain in the trajectory. For the input S (l-1) of the l-th layer, the query, key and value matrices of each attention head can be obtained by linear transformation:

[0108]

[0109] where l is the layer index, h is the attention head index, W Q ,W K ,W V are the weight matrices of the query, key and value. Then the attention output Attn (l,h) (S (l-1) ) of each attention head is obtained, and the outputs of all attention heads are spliced to obtain the final output MultiHead (l) (S (l-1) )

[0110]

[0111] MultiHead (l) (S (l-1) ) = Concat(Attn (l,1) (S (l-1) ),...,Attn (l,H) (S (l-1) ))

[0112] where Contact is the splicing operation, M is the mask matrix, and d h is the dimension of the vector in a single attention head, which is set to 64, and the number of attention heads is set to 2. The time domain feature thereof is defined as

[0113] The trajectory is input into a frequency analysis encoder to obtain the features in the frequency domain thereof:

[0114] First, the input S (l-1)Performing a Fast Fourier Transform (FFT) converts the input sequence into the frequency domain and obtains its frequency representation:

[0115]

[0116] Separating long-term patterns in the sequence by a low-pass filter using a pre-set cutoff frequency index e and setting e to 3, and zeroing all components in the frequency representation with frequency indices greater than or equal to e results in the low-frequency components

[0117]

[0118] Reconstructing back into the time domain by an Inverse Fast Fourier Transform (IFFT) results in the low-frequency component (long-term pattern) representation of the sequence and subtracting the low-frequency component from the original time-domain input computes the high-frequency component (short-term dynamics)

[0119]

[0120] To enable the model to dynamically adjust the importance of long-term and short-term patterns, a learnable per-channel parameter vector β is introduced (l) to modulate the high-frequency component per channel the recombined frequency-domain enhanced feature is computed as follows:

[0121]

[0122] where denotes an element-wise multiplication operation with broadcasting mechanism. Fusing the time-domain and frequency-domain features, the time-domain features captured by the self-attention mechanism are fused with the frequency-domain enhanced features, and the contributions of the two features are balanced by a learnable gating coefficient a: (l)

[0123]

[0124] Subsequently, it enters a fully connected layer (FFN) for processing:

[0125]

[0126] After stacking L3=2 encoder layers, the model can capture the deep contextual information of the input user trajectory sequence and the multi-frequency dynamic characteristics of the final output

[0127] ​Based on the encoder output, three parallel lightweight prediction heads are constructed, respectively for predicting the next interest point, the next visit time offset, and the corresponding interest point category:

[0128]

[0129] where W poi , W time , W cat are linear transformation weight matrices for specific tasks, b poi , b time , b cat are bias vectors for corresponding tasks, and the above parameters are iteratively trained through backpropagation and an optimizer and finally converged. The last row of the final prediction is respectively denoted as

[0130] As a preferred, the preset loss function of step 9 is as follows:

[0131]

[0132] wherein, respectively correspond to the loss of the three sub-tasks of predicting the next interest point and the corresponding interest point category, and the next visit time offset, is the auxiliary loss function L aux in step S5. N' is the number of trajectories in the data set, and for the jth trajectory, and respectively represent the probability value predicted by the model for the real next interest point and the category to which it belongs. and respectively represent the next visit time offset and the real offset time predicted by the model for the trajectory.

[0133] In the model training stage, Adam is used as the optimizer, and all training parameters in the network are iteratively updated through the backpropagation mechanism. The goal is to minimize the total loss function If the index is not optimized for 10 rounds, the training is terminated in advance, and the model parameters at the optimal moment of the new index are saved as the final model. At the same time, the upper limit of the training rounds can be set to 150 rounds to prevent the training time from being too long.

[0134] In the model application stage, the next visited interest point of the user is predicted according to the input historical trajectory of the user.

[0135] Moreover, those skilled in the art should understand that the method proposed in the present application has geographical area adaptability, and if the present application is deployed in a new geographical area, it needs to be retrained using the data set of the area.

[0136] Compared with the prior art, the present application has the following technical effects:

[0137] In view of the limitations of the existing point of interest recommendation, such as inaccurate spatio-temporal modeling, one-sided point of interest semantic understanding and improper dynamic processing of long-term and short-term user behaviors, the present application proposes a next point of interest recommendation method based on a dynamic hierarchical attention fusion spatio-temporal network.

[0138] The present application uses a self-adaptive spatio-temporal double graph convolution network, uses a parameterized Gaussian kernel to model continuous spatial influence, and combines filtered time cosine similarity to strengthen the robustness of time correlation, to obtain accurate point of interest spatio-temporal representation, overcoming the shortcomings of traditional spatio-temporal modeling. In addition, a bidirectional hierarchical graph attention network is used to construct a multi-level semantic hierarchy of "point of interest-function-region", and a Gaussian mixture clustering is used to handle the multiple membership problems, and a bidirectional graph attention mechanism is used to capture cross-scale semantics, to generate rich hierarchical semantic representation of the point of interest, solving the one-sided problem of point of interest semantic understanding. Further, the point of interest spatio-temporal information representation and hierarchical semantic representation obtained in the foregoing are effectively fused to obtain comprehensive features of the point of interest, to support subsequent user behavior sequence modeling. Finally, an attention-frequency fusion module is used to explicitly separate long-term and short-term dynamics in the user historical trajectory based on the comprehensive features of the point of interest through frequency analysis technology, and a frequency-aware attention mechanism is used to fuse and sequence learn these dynamic signals, thereby effectively processing the challenge of mixed behavior patterns of the user, and realizing accurate prediction of the next point of interest.

[0139] The present application can comprehensively capture complex spatio-temporal, semantic and user behavior dynamic information, significantly improving the accuracy, robustness and depth of understanding of user intentions of next point of interest recommendation. The present method aims to provide more accurate personalized recommendation for users and improve travel experience, while helping businesses and platforms optimize services and realize fine operation and business intelligence. BRIEF DESCRIPTION OF DRAWINGS

[0140] Figure 1 The flowchart of the point of interest recommendation method based on the dynamic hierarchical attention fusion spatio-temporal network of the present application;

[0141] Figure 2 The structural diagram of the point of interest recommendation method based on the dynamic hierarchical attention fusion spatio-temporal network of the present application;

[0142] Figure 3 The frequency analysis structure diagram in the present application; DETAILED DESCRIPTION

[0143] The specific embodiments of the present application will be further described in detail below with reference to the accompanying drawings and examples. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments.

[0144] As shown in Figure 1 and Figure 2 The core of the technical route of the present application is: first, the user historical data is processed respectively to form the trajectory, and the feature information of the interest point is obtained; then, the spatio-temporal representation and the hierarchical representation containing multi-level semantics of the interest point are learned in parallel or sequentially, and the two representations are fused into the comprehensive features of the interest point; subsequently, the user trajectory is vectorized, and the trajectory sequence is processed by the attention-frequency fusion Transformer model to predict the next interest point; finally, the whole recommendation model is trained and optimized for the recommendation application.

[0145] In this example, the Foursquare-NYC check-in dataset is taken as an example, and the interest point recommendation method based on the dynamic hierarchical attention fusion spatio-temporal network of the present application is used to recommend the user check-in location. The following will introduce this example from three aspects of data processing and division, model construction and training, and recommendation result.

[0146] 1) Data processing and division

[0147] Step 1.1: The data in the dataset is processed as check-in records. In this embodiment, the Foursquare-NYC check-in dataset collects the daily check-in behavior of users in New York City, and each check-in record includes user name, check-in timestamp, check-in location, location category and geographic location information such as latitude and longitude coordinates.

[0148] In this embodiment, the set of check-in users is represented as U={u1,u2,…,u M}, M is the size of the user set, representing the number of check-in users; the set of interest points is represented as P={p1,p2,…,p N}, N is the size of the interest point set, representing the number of interest points; the set of interest point categories is represented as C={c1,c2,…,c O}, O is the size of the interest category set, representing the number of interest point categories;

[0149] The features f of the interest points in the dataset and the check-in records q are represented as:

[0150] f=(lat,lon,cat,cnt)

[0151] q=(u,p,t)

[0152] where lat is the latitude of the point of interest, lon is the longitude of the point of interest, cat is the category, and cnt is the visit frequency of the point of interest. u is the check-in user ID, p is the visited point of interest ID, and t is the timestamp of the record.

[0153] All points of interest in the point of interest data set constitute an initial feature embedding matrix of points of interest X.

[0154] Step 1.2: Exclude the small number of points of interest with less than 10 check-in records from the user historical check-in trajectory processed in step 1.1, filter out users with less than 10 check-in histories, and sort the check-in timestamps to obtain a user check-in sequence. The user's check-in sequence is divided into trajectories in units of 24 hours. where k represents the length of the sequence, and the trajectory is the unit for model training.

[0155] The data in the Foursquare-NYC check-in data set is filtered according to the above steps in this embodiment. After processing the data set, there are 1075 users, 5099 points of interest, 318 categories, 104074 check-in records, and 14160 trajectories.

[0156] Step 1.3: Construct the spatial weight and time correlation matrix for the points of interest obtained in step 1.1. First, calculate the distance D i between the point of interest pair (v j , v ij ), where the latitude and longitude coordinates of the point of interest pair are (lat i , lon i ) and (lat j , lon j ), and E is the radius of the earth, which is calculated by the Haversine formula:

[0157]

[0158] The distance D ij is converted into a spatial weight A space (i,j) using a parameterized Gaussian kernel function to model the continuous decay of influence:

[0159]

[0160] where σ is the bandwidth parameter of the Gaussian kernel, and in this embodiment, its value is the median distance between points of interest in the data set.

[0161] Define the time sequence T i = {t1, t2, …, t n} as the POI iThe historical access time stamp set of the POI is divided into time slots with 30 minutes interval for a week, the access frequency of each POI in each time slot is counted, and an initial activity vector is obtained, for focusing on the relative access mode, the L2 norm of the original activity vector is normalized to obtain a POI time activity distribution vector, and the cosine similarity is used to calculate the time correlation degree between POIs:

[0162]

[0163] The original cosine similarity is susceptible to low-activity POIs or accidental co-occurrence factors. In order to filter out such false correlations, a reliability indicator function I is introduced reliability

[0164]

[0165] Here n i , n j represent the number of active time slots of POI i and POI j , respectively, and n ij is the number of common active slots. In this embodiment, the values of a and b are 3 and 5, respectively. This process aims to filter accidental co-occurrence phenomena and improve the reliability of the constructed time correlation. The final time correlation matrix is:

[0166] A time (i,j)=cos(i,j)·I reliability

[0167] Step 1.4: Construct a hierarchical heterogeneous graph. Using the probability soft clustering characteristics of GMM, a three-level hierarchical heterogeneous graph composed of "POI node - functional unit - regional unit" is constructed from bottom to top based on the POIs obtained in step 1.1.

[0168] The filtered POI node set V p ={v1,v2,...,v N} and its initial feature matrix X are fitted with a first-level GMM model through the EM algorithm, and the POIs are divided into K1=40 medium-grained functional units Each POI node v i has a feature x i (the i-th row of the initial feature embedding matrix X of the POI) belonging to the k-th functional unit with the conditional probability:

[0169]

[0170] where π k , μ k , Σ kRepresent the mixture weight, mean vector and covariance matrix of the kth Gaussian distribution, π k , μ k ,Σ k It is based on the feature x of the interest point through the EM algorithm. i Iteratively learn and converge the obtained model parameters. Represents the probability density value of the i-th interest point under the k-th Gaussian distribution. Based on this posterior probability distribution, each functional unit The initial characteristics By the interest point feature x associated with it i Weighted addition:

[0171]

[0172] is the kth functional unit feature, and N is the number of interest points in the dataset. The embeddings of all functional units together constitute the feature matrix of the functional unit layer And establish the probability soft connection matrix A of interest points-functional units pc1 , whose elements are Construct Boolean matrix A′ pc1 , which will A pc1 The positions with non-zero probability values ​​in the matrix are assigned 1, and the other positions are assigned 0. To explicitly model the semantic associations at the same level, the pairwise similarity matrix is ​​calculated based on the initial features of the functional units. And based on the threshold τ c1 =70% Construct the adjacency matrix A within the sparse layer cc1 :

[0173]

[0174] Similarly, the functional unit set V c1 Further aggregated into K2 coarse-grained regional units, Count the kth functional unit Belongs to the mth regional unit The conditional probability of Based on this, regional units are generated Features

[0175]

[0176] in, is the kth functional unit The initial features of all regional units are stacked together to obtain And establish the probability soft connection matrix A of functional unit-regional unit c1c2 , whose elements are Construct Boolean matrix A′c1c2 which assigns 1 to non-zero positions of the probability values in the matrix and 0 to other positions. The similarity matrix of the center node is calculated, and based on the threshold τ c1c2 c2 = 80% and constructs the sparse adjacency matrix A cc2

[0177]

[0178] A pp is constructed in the same way as A cc1 and A cc2 , and the initial feature embedding of the interest point is Through the above process, a hierarchical semantic unit heterogeneous graph containing {V p , V c1 , V c2} three-level nodes and {A pp , A' pc1 , A cc1 , A' c1c2 , A cc2} five types of connections is finally constructed.

[0179] 2) Model establishment and training

[0180] Step 2.1: The spatial weight matrix and the time correlation matrix of the interest point are respectively normalized as follows to obtain and

[0181]

[0182] where D is the degree matrix, I N is the unit matrix, and message passing is performed through a two-channel graph convolution of L1=3 layers, and the propagation mechanism is as follows:

[0183]

[0184] where γ space and γ time are learnable scaling parameters, and are learnable weight matrices of specific modalities, and are bias vectors, and the above parameters are iteratively trained through back propagation and an optimizer, and are finally converged to obtain. LayerNorm is a layer normalization operation.

[0185] The features learned by the spatial path and the time path are fused to obtain the fused features

[0186] ​

[0187] where α (l) is a learnable fusion coefficient used to adjust the relative contribution of the fusion process, σ is the activation function of the leaky rectified linear unit (Leaky ReLU), and the negative slope parameter is set to 0.2. Through three layers of propagation, the spatiotemporal semantic representation of the interest point is obtained.

[0188]

[0189] Step 2.2: The basic feature matrix of the interest point The same as that obtained in step 1.4 and As the initial feature of message passing, a graph attention mechanism with bidirectional information propagation is adopted to capture and fuse semantic information of different levels and granularities.

[0190] For any node at level s∈{p,c1,c2}, given its input feature at layer l is Intra-layer message passing passes through the graph attention network and utilizes the intra-layer index edge_index corresponding to the layer s s Execute. The edge index edge_index in the layer s It is derived from the adjacency matrix of the corresponding level s. When the level s = p, the corresponding matrix is ​​A pp , when s=c1, the corresponding matrix is ​​A cc1 , when s=c2, the corresponding matrix is ​​A cc2 This process enhances the consistency of representation of nodes at the same level and the ability to perceive local structures. The specific calculation method of intra-layer message transmission is defined as follows:

[0191]

[0192] Among them, GATConv intra It is a standard graph attention convolution module used to process information interaction within the same node set. It represents the updated feature representation of the node at level s after the intra-layer message passing operation at layer l.

[0193] Cross-level bottom-up information propagation is along the heterogeneous edge index edge_index up Perform cross-layer attention aggregation. The inter-layer index edge index edge_index up It is derived from the adjacency matrix of the corresponding level. When the level s=p, the corresponding matrix is ​​A′ pc1 , when the level s=c1, the corresponding matrix is ​​A′ c1c2The path is responsible for effectively aggregating and abstracting the low-level s fine-grained features into the node representation of the high-level s+1 to which it belongs, realizing the semantic abstraction step by step. The process is defined as:

[0194]

[0195] where GATConv up is the standard graph attention convolution module, which is used to realize the bottom-up information transmission. represents the feature representation of the node of level s+1 updated after the self-bottom message passing operation in the lth layer.

[0196] The top-down information propagation uses the reverse heterogeneous edge index edge_index down under the guidance of semantic, the interlayer index edge_index down is derived from the adjacency matrix of the corresponding level, when level s=c1, the corresponding matrix is When level s=c2, the corresponding matrix is The process propagates and refines the macro semantic context information of the higher level to the lower level, providing global perspective and constraint for fine-grained node features, realizing the context awareness and adaptation of features. The process is defined as:

[0197]

[0198] where GATConv down is the standard graph attention convolution module, which is used to realize the bottom-up information transmission. represents the feature representation of the node of level s updated after the self-bottom message passing operation in the lth layer.

[0199] is the feature vector captured by the three different paths above. For the boundary layer, the missing path features are filled with zero vectors, and the spliced feature matrix is obtained

[0200]

[0201] where || represents the splicing operation, and the fusion feature matrix is obtained by projecting to the target dimension through the linear transformation layer

[0202]

[0203] where W f is the linear transformation matrix, is the bias vector, W f and Both are iteratively trained by backpropagation and optimizer, and finally converged.

[0204] After feature fusion, update using batch normalization and residual connection:

[0205]

[0206] where σ is the Relu activation function, is the fused feature matrix of the previous layer network output, and BatchNorm is the batch normalization operation

[0207] After L2=2 layer message passing, the final hidden representation of each level node is obtained where is the hierarchical semantic information representation of the interest point

[0208] To further optimize the learned hierarchical semantic representation, an auxiliary loss function L is introduced aux which consists of hierarchical contrastive loss L contrastive and hierarchical consistency loss L consistency .

[0209] The hierarchical contrastive loss aims to enhance the cross-level semantic discriminativeness of different level node embeddings. Taking the contrastive loss between the POI layer and the functional unit layer as an example First, based on the soft assignment probability of GMM, the hard assigned functional unit k for each POI i is determined i The click similarity between the embedding of POI i and the embedding of functional unit k is i The contrastive loss adopts the infoNCE form

[0210]

[0211] where N is the number of interest points, K1 is the number of functional units, and τ is the temperature coefficient and is set to 0.1. Using the same processing method as above, the contrastive loss between the functional unit layer and the regional unit layer can be defined The total hierarchical contrastive loss is the weighted sum of the loss of each level.

[0212]

[0213] where λ pc1 and λ c1c2 are both set to 0.5.

[0214] The hierarchical consistency loss aims to ensure the semantic coherence between levels using the soft assignment information of GMM. Taking the consistency loss between the POI layer and the functional unit layer as an example: the expected functional unit layer embedding of interest point i is​​ All functional units to which it is soft-assigned are embedded By probability A pc1 The weighted average of [i,k] is:

[0215]

[0216] Where K1 is the number of functional units. The consistency loss is achieved by minimizing the interest point embedding The mean squared error between the embedding and its expected value is:

[0217]

[0218] Where N is the number of interest points. Similarly, the consistency loss from the functional unit layer to the regional unit layer can be defined as The total level consistency loss is the weighted sum of the losses at each level

[0219]

[0220] where λ′ pc1 and λ′ c1c2 The values ​​of are all set to 0.1, and the total loss of this part is

[0221]

[0222] Step 2.3: Convert the The same as that obtained in step 2.2 To perform fusion, first perform nonlinear interaction:

[0223]

[0224] Where W1 and W2 are learnable shared weight matrices, and α is a learnable weighting coefficient in the range [0, 1]. These parameters are iteratively trained and converged through backpropagation and an optimizer. σ is the ReLU activation function. The enhanced views are concatenated, and attention scores are calculated through a learnable linear layer. These scores are then converted into the attention weight matrix A using the Softmax function:

[0225]

[0226] Where W a The learnable linear change matrix is ​​obtained by training convergence. The final fusion representation e p for:

[0227]

[0228] Among them A :,0 and A :,1are the first and second columns of the attention weight matrix A, respectively, and denotes an element-wise multiplication operation with broadcasting mechanism.

[0229] Step 2.4: The single check-in record in the user trajectory is embedded and represented as:

[0230]

[0231] where, denotes the embedding representation of check-in record i, e u , e c , e t and denote the embedding representation of user u, interest point category c, check-in time t and POI i , respectively, is the concatenation of embedding dimensions. The embedding of the trajectory is composed of all the check-in points of the trajectory

[0232] In this embodiment, the embedding dimension of e u and is 128, the embedding dimension of e c and e t is set to 32, and the dimension of e

[0233] Step 2.5: input the vectorized representation of the user trajectory in step 5.4 to the attention-frequency fusion Transformer module,

[0234] The trajectory embedding S is input to a self-attention encoder to obtain the features of the time domain in the trajectory. For the input S (l-1) of the l-th layer, the query, key and value matrices of each attention head can be obtained by linear transformation:

[0235]

[0236] where l is the layer index, h is the attention head index, W Q , W K , W V are the weight matrices of the query, key and value. Then the attention output Attn (l,h) (S (l-1) ) of each attention head is obtained, and the outputs of all attention heads are concatenated to obtain the final output MultiHead (l) (S (l-1) )

[0237]

[0238] MultiHead (l) (S (l-1))=Concat(Attn (l,1) (S (l-1) ),...,Attn (l,H) (S (l-1) ))

[0239] Where Contact is the concatenation operation, M is the mask matrix, and d h The dimension of the vector in a single attention head is set to 64, and the number of attention heads is set to 2. Its time domain feature is defined as

[0240] The frequency analysis process is as follows Figure 3 As shown, the trajectory is input into the frequency analysis encoder to obtain its frequency domain features. The specific process is as follows:

[0241] First, the input S of the lth layer (l-1) Perform a fast Fourier transform (FFT) to convert the input sequence to the frequency domain and obtain its frequency domain representation:

[0242]

[0243] The long-term pattern in the sequence is separated by a low-pass filter, using a preset cutoff frequency index e, and setting e to 3, and the frequency representation The frequency index in will set all components greater than or equal to e to zero to obtain low-frequency components

[0244]

[0245] Through the inverse fast Fourier transform IFFT Reconstruct back to the time domain to obtain the low-frequency component (long-term pattern) representation of the sequence And subtract the low frequency components from the original time domain input Calculate high-frequency components (short-term dynamics)

[0246]

[0247] In order to enable the model to dynamically adjust the importance of long-term and short-term patterns, a learnable channel-by-channel parameter vector β is introduced (l) To modulate the high frequency components channel by channel Frequency domain enhancement features after reorganization The calculation is as follows:

[0248]

[0249] Where ⊙ represents the element-by-element multiplication operation with broadcast mechanism. The time domain and frequency domain features are fused and the time domain features captured by the self-attention mechanism are with frequency domain enhanced features fusion is performed by a learnable gating coefficient a (l) to balance the contribution of the two features:

[0250]

[0251] followed by a fully connected layer (FFN) for processing:

[0252]

[0253] After stacking L3=2 encoder layers, the model is able to capture the deep contextual information of the input user trajectory sequence and the final output of the multi-frequency dynamic characteristics

[0254] Based on the encoder output, three parallel lightweight prediction heads are constructed to predict the next interest point, the next visit time offset, and the corresponding interest point category, respectively:

[0255]

[0256] where W poi , W time , W cat are linear transformation weight matrices specific to the task, and b poi , b time , b cat are bias vectors corresponding to the task, respectively. The above parameters are iteratively trained through backpropagation and an optimizer and are finally converged. For the output of the model When calculating the loss function, the last row related to the final prediction is respectively recorded as

[0257] Step 2.6: Train the model using the preset loss function:

[0258]

[0259] In this embodiment, N' = 14160, which is the number of trajectories in the data set. corresponding to the prediction of the next interest point and the corresponding interest point category, and the next visit time offset, the loss of the three sub-tasks, is the auxiliary loss function L aux in step S5. For the jth trajectory, and represent the probability values predicted by the model for the real next interest point and the category to which it belongs, respectively. and respectively represent the next visit time offset predicted by the model for the trajectory and the real offset time.

[0260] In the model training stage, Adam is used as the optimizer to minimize the total loss function For the target, all trainable parameters of the network are iteratively updated by backpropagation, if the index is not optimized for 10 consecutive rounds, the model is considered to be converged, the training is terminated in advance within the upper limit of 150 rounds, and the model parameters at the moment of optimal performance are saved, thereby obtaining the final model.

[0261] 3) recommended results

[0262] In the model application stage, the final model obtained in step 2.6 is used to predict the next interest point, the model accepts the user historical trajectory sequence, the model outputs the candidate interest point through one forward propagation, and the accuracy@k(Acc@k) and the mean reciprocal rank(MRR) are used to evaluate the performance of the model:

[0263]

[0264] In the formula, m is the number of samples (trajectories) in the data set, rank q is the rank, k takes 1, 5, 10, 20, which represents the real next interest point in the ordered list recommended by the model. If the interest point output by the model is not greater than k, the value of the formula is 1, otherwise 0.

[0265] Among the two recommendation indicators, the accuracy@k(Acc@k) calculates the proportion of samples whose real next interest point ranking is in the top k. The mean reciprocal rank(MRR) measures the average ability of the model to place the correct answer in the front position of the list by calculating the average of the reciprocal of its real ranking. The higher the above two indicators, the better the performance.

[0266] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any skilled person in the art should understand that within the scope of the spirit and essence of the present application, certain modifications and changes can be made to the present application, but should be covered within the protection scope of the present application.

Claims

1. A point of interest recommendation method based on dynamic hierarchical attention fusion spatiotemporal network, characterized by: The following steps are involved: S1. Input the user's historical check-in data, where each check-in record contains the user ID, point of interest ID and check-in timestamp; S2. Sort the user's historical check-in records by access time, perform data preprocessing, obtain the user check-in sequence, and divide the user's track into preset time intervals of 24 hours; S3. Extract feature data of each point of interest, wherein the feature data includes the category of the point of interest, the historical visit frequency, and the geographic longitude and latitude coordinates; S4. Inputting the feature data of the points of interest into an adaptive spatiotemporal dual-graph convolutional network module, which continuously attenuates the spatial influence between points of interest by using a parameterized Gaussian kernel, and combines filtered temporal cosine similarity calculation to enhance the robust capture of the temporal relationship between points of interest, thereby learning and outputting a robust spatiotemporal information representation of each point of interest; S5. Input the feature data of the POI into a bidirectional hierarchical graph attention network module, which constructs a multi-level semantic hierarchy of "POI-functional unit-region unit" and uses Gaussian mixture clustering to process the multiple semantic affiliations of POIs. Then, a bidirectional information propagation graph attention mechanism is used to capture deep cross-scale semantic information, thereby learning and outputting a hierarchical semantic information representation of each POI. S6, fusing the robust spatiotemporal information representation of the interest point obtained in step S4 with the hierarchical semantic information representation of the interest point obtained in step S5 to generate a comprehensive feature representation of the interest point; S7, encoding the user trajectory obtained in step S2 to form a vectorized representation of the user trajectory; S8. Input the vectorized representation of the user trajectory into the attention-frequency fusion transformer. This module uses frequency analysis to explicitly separate long-term and short-term dynamic signals in the user behavior sequence and integrates them with the self-attention mechanism to predict the probability of the user visiting each candidate point of interest next after the current trajectory. S9. Using a preset loss function as the optimization target, the overall model including steps S1 to S8 is trained until the model converges to obtain the final POI recommendation model.

2. The method according to claim 1, characterized in that In step S2, data preprocessing includes: First, we exclude points of interest with fewer than 10 check-in records, and filter out users with fewer than a preset threshold of 10 check-in records. We then calculate the frequency of visits to each point of interest.

3. The method according to claim 1, characterized in that In step S3, the feature data of the interest points are further used to generate an initial embedding representation of each interest point.

4. The method according to claim 1 or 3, characterized in that In step S4, the specific processing of the adaptive spatiotemporal dual-graph convolutional network module includes: a. Construct the spatial weight matrix A between interest points space , where the spatial weights between interest points are based on their geographical distances and are calculated using a parameterized Gaussian kernel whose bandwidth parameter is determined by the median of the statistical values ​​of the distances between interest points in the dataset; b. Construct the temporal correlation matrix A between points of interest time First, a week is divided into time slots according to the preset time unit of 30 minutes. The visit frequency of each interest point in each time slot is counted to form an initial activity vector. The initial activity vector is normalized by L2 norm to obtain the time activity distribution vector. Then, the cosine similarity of the time activity distribution vectors between interest points is calculated and filtered with the reliability indicator function to obtain the time correlation degree; c. Use dual-channel graph convolution to perform message passing on the spatial weight matrix and temporal correlation matrix of the interest point respectively, and adaptively fuse the output features of the two channels to obtain the robust spatiotemporal information representation of the interest point 5. The method according to claim 4, characterized in that In step S4 a, any two points of interest (v i ,v j ) between the geographical distance D ij Calculated by the spherical cosine law, the spatial weight is calculated by the formula Calculate, where σ is the bandwidth parameter of the Gaussian kernel.

6. The method according to claim 4, characterized in that In step S4 b, the reliability indicator function I reliability Used to determine whether the co-occurrence between interest points is accidental co-occurrence. i and POI j The number of common active time slots is greater than the preset threshold b, and POI i and POI j When the total number of active time slots is greater than the preset threshold a, I reliability is 1 if the value is set, otherwise it is 0.

7. The method according to claim 1, characterized in that In step S5, the specific processing of the bidirectional layered graph attention network module includes: a. Using Gaussian mixture models (GMM) for probabilistic soft clustering, a three-level hierarchical heterogeneous graph is constructed from the bottom up. This heterogeneous graph consists of a point of interest node layer, a functional unit layer, and a regional unit layer. First, the point of interest nodes are clustered into several functional units using the first-stage GMM. The conditional probability of the point of interest node belonging to each functional unit is calculated. The features of each functional unit are then weighted based on these conditional probabilities and the initial features of the point of interest nodes. Subsequently, the functional units are clustered into several regional units using the second-stage GMM, and the features of each regional unit are derived in a similar manner. b. Adopting a bidirectional heterogeneous attention propagation mechanism on the hierarchical heterogeneous graph, including intra-layer attention propagation, bottom-up information aggregation from low-level to high-level, and top-down information guidance from high-level to low-level, so as to learn the hierarchical semantic information representation of each point of interest.

8. The method according to claim 1, characterized in that In step S6, the fusion is performed in the following manner: the robust spatiotemporal information representation and hierarchical semantic information representation Perform feature interaction and attention weighted fusion, specifically, after performing linear transformation and nonlinear activation function processing on the two through a learnable weight matrix, calculate the attention weight, and perform weighted summation of the two based on the attention weight to obtain the comprehensive feature representation of the interest point e p .

9. The method according to claim 1, characterized in that In step S7, each check-in record in the user trajectory is encoded. Specifically, the corresponding POI category embedding, POI feature embedding, check-in timestamp embedding, and user ID embedding are concatenated to form a vectorized representation of the check-in record. The vectorized representation of the user trajectory is composed of the sequence of vectorized representations of each check-in record contained in it, S. In step S8, the specific processing of the attention-frequency fusion Transformer module includes: a. Input the vectorized representation S of the user trajectory into the self-attention encoding path and the frequency analysis encoding path in parallel; b. In the self-attention encoding path, a multi-head self-attention mechanism is used to process the user trajectory sequence to obtain time domain features; c. In the frequency analysis coding path, first perform fast Fourier transform on the user trajectory sequence to obtain the frequency domain representation A low-pass filter is then used to isolate the long-term patterns The short-term pattern is obtained by subtracting the long-term pattern from the original frequency domain signal. Then, the frequency enhancement feature is obtained by weighted fusion of long-term and short-term patterns through learnable parameters. Where l is the network layer index; d. Fusion of the time domain features and the frequency enhancement features through an adaptive gating mechanism to obtain a hybrid feature And processed by the fully connected layer; e. Based on the final output of the attention-frequency fusion Transformer module, a parallel prediction head is constructed to predict the probability of the next POI, the time offset of the next visit, and the corresponding POI category.

10. The method according to claim 1, characterized in that In step S9, the preset loss function is a combined loss function, including the cross entropy loss for interest point prediction Cross entropy loss for interest point category prediction and the mean squared error loss for time offset prediction and the auxiliary loss used to guide the learning of hierarchical semantic information representation in step S5 Total loss L total is the weighted sum of these losses; the auxiliary loss Includes: Hierarchical contrast loss calculated based on hard-assigned correspondences To enhance semantic discrimination across layers; and layer consistency loss based on soft assignment probability calculation To ensure semantic coherence between levels.

Citation Information

Cited By

  • Recommendation method and device, electronic equipment and storage medium

    CN121256151A