A method for recommending points of interest based on the prior of population movement patterns

By constructing a spatio-temporal graph and graph neural network to extract the population movement mode, and combining the attention mechanism to improve the feature intersection method of the Wide&Deep model, the problem that the recommendation results in the point of interest recommendation system do not conform to the travel rules, and the recommendation effect is achieved that is more in line with human travel intentions.

CN115357786BActive Publication Date: 2025-07-08ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210952103.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-07-08
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

The existing point-of-interest recommendation system cannot learn the travel rules of the people in the city in a timely manner, resulting in the recommendation results that do not conform to common sense, such as recommending cafes in commercial areas on non-working nights.

Method used

By constructing spatiotemporal graph and graph neural network algorithms, the crowd movement mode is extracted, combined with the attention mechanism, a point of interest recommendation method based on the crowd movement mode is designed, and the graph neural network is used as the extractor of the crowd movement mode, and the attention mechanism is introduced to capture the spatiotemporal information of the urban traffic mode, and the feature crossover method of the Wide&Deep model is improved.

Benefits of technology

The recommendation results are more in line with human travel intentions, and solve the problem of choice difficulties brought by promoting information on social networks to people's travel. The experimental results show that they are better than the baseline algorithm model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115357786B_ABST
    Figure CN115357786B_ABST
Patent Text Reader

Abstract

A method for recommending points of interest based on crowd mobility pattern priors includes: (1) obtaining urban traffic data, POI data and user check-in data in the same time and space, and performing data processing; (2) extracting crowd mobility patterns in the city, and designing downstream tasks to pre-train the crowd mobility pattern extraction module; (3) adding the crowd mobility pattern prior features extracted in step (2) to the deep model; (4) changing the Wide part of the original Wide & Deep model to a Cross network, and adding a feature cross layer to the linear model part; (5) connecting the Deep network in step (3) and the Cross network output in step (4), inputting the spliced ​​vector into the final single-layer perceptron, fitting the final target, and outputting the predicted score p of user j checking in at point of interest i. The recommendation result of the present invention is more in line with human travel intentions, and solves the problem of difficulty in choosing when people travel due to a large amount of promotional information flooding social networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of recommendation systems, and in particular to a method for recommending points of interest based on prior knowledge of population movement patterns. Background Art

[0002] Due to the popularity of mobile positioning devices (smartphones), location-based social networks (LBSNs) have developed rapidly. Users can share location data by checking in at points of interest (POIs) on social networks. A large amount of POI location data can reflect the preferences of the user group. POI recommendation is an important way to assist users in exploring the surrounding environment to improve the user experience. The POI recommendation method provides personalized recommendation services by learning user historical data.

[0003] The development of POI recommendation systems can be simply divided into two paths. One path is the continuous innovation of data fusion methods: from geographical location information, to user social relationships, and then to the spatio-temporal patterns of cities, more and more key elements have been discovered by researchers and used as new features of recommendation systems. The other path is closely related to the evolution of recommendation technologies themselves, from classical machine learning methods, to later graph learning technologies, and then to the recent introduction of the attention mechanism. Any technical method that plays a key role in recommendation systems, or even in the field of deep learning, may achieve unexpected results in POI recommendation as long as it is used properly. From the perspective of the evolution of data usage, early POI recommendation methods used the check-in frequency of users for recommendation. In the data fusion method, geographical data, social data, and time data were introduced successively.

[0004] However, the current POI recommendation system still has the following defects: that is, the recommendation target tends to fit the historical check-in behavior of individual users. Although this approach can achieve good indicators in offline experiments, since the feedback of the POI recommendation system is not timely, it is impossible to learn the travel patterns of the population in the city, and finally obtain recommendation results that do not conform to common sense. For example, recommending a coffee shop in the business district on a non-working day evening. Summary of the Invention

[0005] Aiming at the above-mentioned disadvantages of the prior art, the present invention proposes a method for recommending points of interest based on prior knowledge of population movement patterns. Based on multi-source heterogeneous data in the city, by means of spatio-temporal graph construction, graph neural network algorithms, and attention mechanisms, etc., explore methods for extracting population movement patterns, and design a method for recommending points of interest based on population movement patterns accordingly.

[0006] The present invention achieves the above object through the following technical solutions: A method for recommending points of interest based on prior crowd movement patterns, comprising the following steps:

[0007] (1) Obtain urban traffic data, POI data, and user check-in data in the same time and space, and perform data processing.

[0008] (2) Extract the crowd movement patterns in the city, and design a downstream task to pre-train the crowd movement pattern extraction module.

[0009] (3) Add the prior features of the crowd movement patterns extracted in step (2) to the deep model.

[0010] (4) Change the Wide part of the original Wide&Deep model to a Cross network, and add a feature crossing layer to the linear model part.

[0011] (5) Connect the output of the Deep network in step (3) and the output of the Cross network in step (4), input the concatenated vector into the final single-layer perceptron, fit the final target, and output the predicted score p of user j checking in at point of interest i.

[0012] Among them, step (1) specifically includes the following steps:

[0013] 11). Obtain the source data and perform data preprocessing, such as Mercator projection of GPS coordinates.

[0014] 12). Use the urban traffic dataset and the point of interest dataset to construct a spatio-temporal graph. Divide the selected range grid of the city into r×r regions, map each trajectory point and POI point into the divided grid, collect the traffic flow data of the city at a certain time interval, and perform temporal aggregation on the data within each interval to construct the crowd movement flow into a spatio-temporal graph G t (V t ,E t ,A t ).

[0015] Among them, |V t | = r×r, indicating that all the regions divided in the city constitute the nodes of the spatio-temporal graph E t = {(u, v) u, v ∈ V t}, representing the crowd movement relationship between any two regions in the city. Use w u,v to represent the weight of the edge (u, v), that is, the number of people flowing from region u to region v within the time slice t. A t is the node attribute matrix, representing the distribution of points of interest within the range of each node.

[0016] 13). Generate point - of - interest feature vectors to represent the distribution of functional attributes. Suppose there are k different types of point - of - interest types, and for each region i, construct where represents the number of the j - th type of point of interest in region i.

[0017] 14). Process user check - in data. Process categorical features and numerical features separately, including converting categorical features to numerical features using One - hot encoding, processing numerical features by normalization, and converting unevenly distributed numerical features to categorical features.

[0018] Among them, the crowd movement pattern extraction framework in step (2) has three main parts: a crowd movement pattern extraction module, an encoder - decoder module based on a multi - head attention mechanism, and an up - sampling module adapted to downstream tasks. Specifically, it includes the following steps:

[0019] 21). Extract crowd movement pattern features from the crowd movement pattern extraction module. First, obtain the input of this module from the spatio - temporal graph constructed in step 12): node attributes and adjacency matrix sequences. The input passes through a graph neural network and a feed - forward network module to obtain the features of the crowd movement pattern.

[0020] The specific steps for constructing the graph neural network are as follows:

[0021] First, obtain the Laplacian matrix L = D - A of graph G, and perform eigenvalue decomposition on L to get L = UΛU T . Where D is the degree matrix of graph G, D = diag(d1, d2, …, d n ), d i is the degree of node i in graph G, and A is the adjacency matrix of graph G.

[0022] Secondly, assume that the functional attribute distribution vector obtained in 13) is x, and its Fourier transform The corresponding inverse transform Get the convolution form of the graph signal x and the convolution kernel h on graph G Replace the above with where is a learnable parameter.

[0023] Finally, obtain the graph neural network framework as:

[0024]

[0025] where is the urban functional attribute matrix, is the traffic adjacency matrix, is The corresponding degree matrix, are the network parameters to be learned. XW performs a linear transformation on the feature vectors of the nodes, propagates the transformed node features to the neighbors, and normalizes the features received by the nodes. σ is a non-linear activation function, and ReLU is used in this invention. is the output result matrix of the graph convolutional neural network, representing the node features after information transmission.

[0026] After training a graph neural network, share the parameter weights of the graph convolutional neural network belonging to each time slice, and perform sum and average processing on the backpropagation gradients of the inputs of different time slices.

[0027] Finally, introduce the feed-forward network part. First, flatten the node features of a single time slice, and then input them into a multi-layer perceptron for dimensionality reduction operation to obtain the characterization vector of the urban population movement pattern of a certain time slice.

[0028] 22). Introduce the multi-head attention mechanism. Use the sequence of characterization vectors of the population movement pattern as the input characterization, and use the target data characterization of the downstream task as the output characterization. The attention mechanism can be expressed by the following formula:

[0029]

[0030] Suppose the original attention module is replicated h times, that is, h-head attention. Input the same data into the multi-head attention module to obtain the results of h single-head attentions. Concatenate the multi-head results and perform a linear layer transformation to obtain the fused feature output. The multi-head attention formula is as follows:

[0031]

[0032] 23). Upsampling module: Use a deconvolutional neural network to restore the predicted characterization vector to the same structure as the real data.

[0033] Among them, step (3) specifically includes the following steps:

[0034] 31). Convert the discrete feature characterization in step 14) into a dense feature:

[0035]

[0036] where f s k is the discrete feature, is the dense feature.

[0037] 32). Add the prior features of the population movement pattern extracted in step (2), and concatenate all dense features:

[0038]

[0039] Among them, f h is the characteristic of the urban population movement pattern.

[0040] 33). Use the Deep network to fuse the dense features obtained in 32): h0 = ReLU(W0f d + b0)

[0041] 34). Pass the fusion result in step 33) through two layers of fully connected layers (using ReLU as the activation function) to obtain the dense feature h. The formula is:

[0042] h l+1 = ReLU(W l h l + b l ) 6)

[0043] Among them, the specific steps of step (4) include:

[0044] Change the Wide part of the original Wide&Deep model to a Cross network, and use multiple Cross layers to perform feature crossing to obtain a high-order representation x of the discrete features. If the output vector of the l-th Cross layer is x l , then the output of the l + 1-th layer is

[0045]

[0046] Among them, the specific steps of step (5) include the following steps:

[0047] 51). Connect the output of the Deep network in step (3) and the output of the Cross network in step (4) to obtain the concatenated vector x c = Cat(h, x).

[0048] 52). Input the concatenated vector obtained in 51) into the last single-layer perceptron, and output the predicted score p of user j's check-in at point of interest i. The formula is as follows:

[0049] p = sigmoid(Wx c + b) 8)

[0050] The present invention includes: 1) Data acquisition and processing: Using urban traffic data, POI data, and user check-in data in the same time and space, constructing a spatio-temporal graph using urban traffic and point-of-interest datasets, and processing the features of the original data. 2) Proposing a framework for extracting crowd movement patterns: Using a graph neural network as an extractor for crowd movement patterns, introducing an attention mechanism to capture the spatio-temporal information of urban traffic patterns. Formulating downstream tasks, designing an upsampling module to restore the representation vector to the task target, and realizing end-to-end framework learning and training to complete the pre-training of the crowd movement pattern extractor. 3) Introducing prior knowledge of human patterns: Adding the prior features of the crowd movement patterns extracted in step 2) to the deep model, making the recommendation results more in line with the travel rules of humans in the city. 4) Feature crossing: Changing the Wide part to a Cross network and adding a feature crossing layer to the linear model part. 5) Output result: Connecting the network outputs of steps 3) and 4), inputting the concatenated vector into the final single-layer perceptron, and outputting the predicted score of the user's check-in at the point of interest. The experiment on point-of-interest recommendation conducted with New York as an example shows that the present invention has excellent performance in dealing with this problem.

[0051] The innovation of the present invention lies in:

[0052] (1) Proposing an architecture for extracting crowd movement patterns. This architecture uses the prediction of downstream tasks to pre-train the crowd movement pattern extraction module, enabling the crowd movement pattern extraction module to obtain the ability to extract spatio-temporal patterns under the data distribution of a specific city.

[0053] (2) Making improvements to the Wide&Deep model. Adding a feature crossing layer to the linear model part and adding prior features of crowd movement patterns to the deep model.

[0054] The advantages of the present invention are:

[0055] Aiming at the problem that the current point-of-interest recommendation results deviate from travel rules, the present invention proposes a method for point-of-interest recommendation based on prior knowledge of crowd movement patterns. Making full use of the prior knowledge of crowd movement patterns, using a graph neural network as an extractor for crowd movement patterns, introducing an attention mechanism to capture the spatio-temporal information of urban traffic patterns, and improving the feature crossing method of the breadth model. The comparative experimental results show that the algorithm proposed by the present invention performs better than the baseline algorithm model, verifying the effectiveness of the prior knowledge of crowd movement patterns in the point-of-interest recommendation system. The recommendation results of the present invention are more in line with human travel intentions, solving the problem of choice difficulties brought to people's travel by a large number of promotional messages flooding social networks. Description of the Drawings

[0056] Figure 1 It is the overall framework diagram of the present invention.

[0057] Figure 2It is the spatio-temporal graph structure diagram of the present invention.

[0058] Figure 3 It is the diagram of the crowd movement pattern extraction module of the present invention.

[0059] Figure 4 It is the diagram of the multi-head attention mechanism of the present invention.

[0060] Figure 5 It is the diagram of the upsampling module adapted to downstream tasks in the present invention.

[0061] Figures 6(a) and 6(b) are the comparison diagrams of the experimental results of the present invention and other methods in the examples of the present invention. Among them, Figure 6(a) is the comparison diagram of the ROC curves of each model, and Figure 6(b) is the comparison diagram of the PR curves of each model. Detailed implementation manners

[0062] The present invention will be further described below in combination with the travel data and check-in data of New York City for point-of-interest recommendation examples.

[0063] The overall framework of the point-of-interest recommendation method based on the prior of crowd movement patterns in this example is as Figure 1 shown, and specifically includes the following steps:

[0064] (1) Obtain urban traffic data, POI data, and user check-in data in the same spatio-temporal range, and perform data processing:

[0065] a). Obtain the Yellow taxi data, POI data, and check-in data of New York City in May and June 2012. Among them, the Yellow taxi data is shown in Table 1 below, the POI data is shown in Table 2 below, and the check-in data is shown in Table 3 below.

[0066] Table 1

[0067]

[0068] Table 2

[0069]

[0070] Table 3

[0071]

[0072]

[0073] Use the Mercator projection to transform the longitude and latitude in the dataset into northeast coordinates in meters.

[0074] b). Use the New York traffic dataset and the point-of-interest dataset to construct a spatio-temporal graph, and the spatio-temporal graph structure is as Figure 2As shown in the figure. First, the New York dataset is divided into 16*16 grids, and the starting and ending points of taxi trajectory data are mapped to two rectangular areas in the 16*16 matrix. Secondly, the data is divided at 1-hour intervals, and a (16*16)*(16*16) matrix can represent the crowd movement pattern within the city scope in a certain time slice. Finally, the data for this one hour is aggregated in time, and a slice G of the spatio-temporal graph can be obtained from the aggregated data for this one hour. t Construct the crowd movement flow as the spatio-temporal graph G t (V t , E t , A t ).

[0075] Among them, |V t | = r×r, indicating that all the divided areas in the city constitute the nodes of the spatio-temporal graph E t = {(u, v) u, v ∈ V t}, representing the crowd movement relationship between any two areas in the city. Use w u,v to represent the weight of the edge (u, v), that is, the number of people flowing from area u to area v within the time slice t. A t is the node attribute matrix, representing the distribution of points of interest within each node range.

[0076] c). Generate the feature vector of the point of interest, showing the distribution of functional attributes. The present invention introduces the point of interest distribution vector. Specifically, there are k different types of points of interest. For each area i, construct where represents the number of the jth type of point of interest in area i.

[0077] d). Process the user check-in data. Process the categorical features and numerical features separately. Specifically, use One-hot encoding to convert the categorical features and ID features into numerical features; adopt the normalization method to process the numerical features; for the unevenly distributed numerical features, such as in the dataset used in the present invention, the check-in data often occurs between 8 am and 11 pm, and there is little data in the early morning. The present invention adopts the bucketing strategy, sorts the feature values, and then finds the quantiles according to the number of buckets to bucket the samples. Finally, use the bucket ID as the feature value.

[0078] For the geographical location features (longitude, latitude) of the points of interest, the present invention maps each check-in data to a unique city block. For the check-in time information, the present invention aggregates by hour to obtain the check-in data for 1464 time slices. Arrange the check-in data in chronological order, take the first 70% of the data as the training set, and exclude the points of interest that the user has already checked in from the remaining 30% of the data to obtain the test set.

[0079] (2) Extract the crowd movement patterns in the city and design downstream tasks to pre-train the crowd movement pattern extraction module:

[0080] a). Extract the crowd movement pattern features from the crowd movement pattern extraction module, and the crowd movement pattern extraction module is as Figure 3 shown. First, obtain the input of this module from the spatio-temporal graph constructed in step (1) b): the node attributes and the adjacency matrix sequence. The input passes through the graph neural network and the feed-forward network module to obtain the features of the crowd movement pattern.

[0081] The specific steps for constructing the graph neural network are as follows:

[0082] First, obtain the Laplacian matrix L = D - A of graph G, and perform eigenvalue decomposition on L to get L = UAU T . Where D is the degree matrix of graph G, D = diag(d1, d2,..., d n ), d i is the degree of node i in graph G, and A is the adjacency matrix of graph G.

[0083] Secondly, assume that the point-of-interest feature vector obtained from (1.c) is x, and its Fourier transform The corresponding inverse transform to obtain the convolution form of graph signal x and convolution kernel h on graph G Replace the above with where are learnable parameters.

[0084] Finally, obtain the graph neural network framework as:

[0085]

[0086] where is the urban functional attribute matrix, is the traffic adjacency matrix, is the corresponding degree matrix, is the network parameter to be learned. XW linearly transforms the feature vector of the node, propagates the transformed node features to the neighbors, normalizes the features received by the node. σ is a non-linear activation function, and ReLU is used in the present invention. is the output result matrix of the graph convolutional neural network, representing the node features after information transmission.

[0087] The parameter weights of the graph convolutional neural network belonging to each time slice are shared. The present invention only trains and maintains one graph neural network, and for the backpropagation gradients of the inputs of different time slices, sum-average processing is adopted.

[0088] Finally, the feed-forward network part is introduced. First, the node features of a single time slice are flattened, and then input into a multi-layer perceptron for dimensionality reduction operation to obtain the representation vector of the urban population movement pattern of a certain time slice.

[0089] b). Introduce the multi-head attention mechanism. The framework of the head attention mechanism is as Figure 4 shown. This module uses the sequence of representation vectors of the population movement pattern as the input representation and the target data representation of the downstream task as the output representation. The attention mechanism can be expressed by the following formula:

[0090]

[0091] Suppose the original attention module is replicated h times, that is, h-head attention. The same data is input into the multi-head attention module to obtain the results of h single-head attentions. The multi-head results are concatenated and transformed through a linear layer to obtain the fused feature output. The formula for multi-head attention is as follows:

[0092]

[0093] c). Upsampling module: Use a deconvolutional neural network to restore the predicted representation vector to the same structure as the real data. The upsampling module adapted to the downstream task is as Figure 5 shown.

[0094] Generate the urban traffic congestion situation map from the predicted representation vector and record Calculate the average vehicle speed of r×r nodes within the time period. The average speed of nodes within each time slice can be regarded as a map snapshot of the urban traffic congestion state. Denote the urban traffic congestion state snapshot corresponding to time slice t as Thus, the urban traffic congestion state evaluation task is described formally. Given the sequence of traffic flow spatio-temporal graphs G seq ={G1, G2,..., G t}, the functional distribution corresponding to each region The urban traffic congestion state sequence s seq ={s1, s2,..., s t-1} is used to predict the urban traffic congestion state S t at time slice t. Let the result output by the overall framework be It is expected that is as close as possible to S t Then the optimization objective of the overall framework is:

[0095]

[0096] (3) Incorporate the prior features of the crowd movement pattern extracted in step (2) into the deep model:

[0097] a). Convert the discrete feature representation in d) of step (1) into a dense feature, and obtain the representation vectors of the category and ID features through the Embedding layer:

[0098]

[0099] where f s k is the discrete feature, is the dense feature.

[0100] b). Incorporate the prior features of the crowd movement pattern extracted in step (2), and concatenate all dense features, covering the geographical location features, time features, interest point category features, ID category features of users and interest points, etc.:

[0101]

[0102] where f h is the urban crowd movement pattern feature.

[0103] c). Use the Deep network to fuse the dense features obtained in b): h0 = ReLU(W0f d + b0)

[0104] d). Pass the fusion result in c) through two fully connected layers (using ReLU as the activation function) to obtain the dense feature h, and the formula is:

[0105] h l+1 = ReLU(W l h l + b l ) 6)

[0106] (4) Change the Wide part of the original Wide&Deep model to a Cross network, and add a feature crossing layer to the linear model part.

[0107] a). Change the Wide part of the original Wide&Deep model to a Cross network. The inputs of the Cross model are the user ID, interest point ID, and interest point category. Use multiple cross layers to perform feature crossing to obtain the high-order representation x of the discrete feature. If the output vector of the l-th cross layer is x l , then the output of the l+1-th layer is

[0108]

[0109] (5) Connect the output of the Deep network in step (3) and the output of the Cross network in step (4), input the concatenated vector into the final single-layer perceptron, fit the final target, and output the predicted score p of user j's check-in at point of interest i.

[0110] a). Connect the output of the Deep network in step (3) and the output of the Cross network in step (4) to obtain the concatenated vector x c = Cat(h, x).

[0111] b). Input the concatenated vector obtained in step a) into the final single-layer perceptron, and output the predicted score p of user j's check-in at point of interest i. The formula is as follows:

[0112] p = sigmoid(Wx c + b) 8)

[0113] (*) By comparing with other baseline models, including Deep Factorization Machine (DeepFM), Embedding Multi-Layer Perceptron (EmbeddingMLP), Neural Collaborative Filtering (NeuralCF), Wide&Deep model, and Two-Tower model. In the comparative experiment of the point of interest recommendation model, accuracy (ACC) and two AUC values (ROC and PR) are used as the experimental indicators. Figure 6 is a comparison chart of the ROC and PR curves of the experimental results of the recommendation method. The results show that the HMRec method proposed by the present invention has achieved results exceeding all baseline models in terms of both accuracy and Receiver Operating Characteristic (ROC) curve, which proves that the urban spatio-temporal attribute of the population movement pattern plays a certain role in point of interest recommendation.

Claims

1. A method for recommending points of interest based on the prior of population movement patterns, comprising the following steps: (1) Obtain urban traffic data, POI data, and user check-in data in the same time and space, and perform data processing; (2) Extract the population movement patterns in the city, and design a downstream task to pre-train the population movement pattern extraction module; (3) Add the prior features of the population movement patterns extracted in step (2) to the deep model; specifically including: 31). Convert the discrete feature representation in step 14) into a dense feature: Among them, is a discrete feature, is a dense feature; 32). Add the prior features of the population movement patterns extracted in step (2), and concatenate all dense features: Among them, f h is the characteristic of the urban population movement pattern; 33). Fuse the dense features obtained in step 32) using a Deep network: h0 = ReLU(W0f d + b0) 34). Pass the fusion result in step 33) through two layers of fully connected layers (using ReLU as the activation function) to obtain a dense feature h, and the formula is: h l+1 = ReLU(W l h l + b l ) (4) Change the Wide part of the original Wide&Deep model to a Cross network, and add a feature crossing layer to the linear model part; specifically including: Change the Wide part of the original Wide&Deep model to a Cross network, and use multiple Cross layers to perform feature crossing to obtain a high-order representation x of discrete features; if the output vector of the l-th Cross layer is x l , then the output of the l+1-th layer is (5) Connect the outputs of the Deep network in step (3) and the Cross network in step (4), input the concatenated vector into the final single-layer perceptron, fit the final target, and output the predicted score of user j's check-in at point of interest i.

2. The method for recommending points of interest based on the prior of crowd movement patterns according to claim 1, characterized in that Step (1) specifically includes: 11). Obtain the source data and perform data preprocessing, such as Mercator projection of GPS coordinates; 12). Construct a spatio-temporal graph using the urban traffic dataset and the point of interest dataset; divide the selected range grid of the city into r×r regions, map each trajectory point and POI point into the divided grid, collect the traffic flow data of the city at a certain time interval, and perform temporal aggregation on the data within each interval to construct the spatio-temporal graph G t (V t ,E t ,A t ); Among them, |V t | = r × r, indicating that all regions divided in the city constitute the nodes E of the spatio-temporal graph t = {(u, v)|u, v ∈ V t}, representing the population movement relationship between any two regions in the city. Use w u,v to represent the weight of the edge (u, v), that is, the number of people flowing from region u to region v within the time slice t; A t is the node attribute matrix, representing the distribution of points of interest within the range of each node; 13). Generate point-of-interest feature vectors to represent the distribution of functional attributes; assume there are k different types of points of interest, and for each region i, construct [f i 1 , f i 2 ,..., f i k T , where represents the number of the j-th type of point of interest in region i;​ 14). Process the user check-in data; process the categorical features and numerical features separately, including converting the categorical features to numerical features using One-hot encoding, processing the numerical features in a normalized manner, and converting the unevenly distributed numerical features to categorical features.

3. The method for recommending points of interest based on the prior of crowd movement patterns according to claim 1, characterized in that, The population movement pattern extraction framework in step (2) has three main parts: a population movement pattern extraction module, an encoder-decoder module based on a multi-head attention mechanism, and an upsampling module adapted to the downstream task; specifically including: 21). Extract the population movement pattern features from the population movement pattern extraction module; first obtain the input of this module from the spatio-temporal graph constructed in step 12): the node attributes and the adjacency matrix sequence; the input passes through the graph neural network and the feed-forward network module to obtain the features of the population movement patterns; The specific steps for constructing the graph neural network are as follows: First, obtain the Laplacian matrix L = D - A of graph G, and perform eigenvalue decomposition on L to get L = UΛU T ; where D is the degree matrix of graph G, D = diag(d1, d2, …, d n ), d i is the degree of node i in graph G, and A is the adjacency matrix of graph G; Secondly, assume that the functional attribute distribution vector obtained from (13) is x, and its Fourier transform the corresponding inverse transform gives the convolution form of the graph signal x and the convolution kernel h on the graph G Replace the above-mentioned with where are learnable parameters; Finally, the obtained graph neural network framework is: Among them is the urban functional attribute matrix, is the traffic adjacency matrix, is the corresponding degree matrix, is the network parameter to be learned, and σ is the non-linear activation function; After training a graph neural network, share the parameter weights of the graph convolutional neural network belonging to each time slice, and perform sum-average processing on the backpropagation gradients of the inputs of different time slices; Finally, introduce the feed-forward network part; first flatten the node features of a single time slice, and then input a multi-layer perceptron for dimensionality reduction operation to obtain the representation vector of the urban population movement patterns of a certain time slice; 22). Introduce the multi-head attention mechanism; use the sequence of representation vectors of the population movement patterns as the input representation, and use the target data representation of the downstream task as the output representation; the attention mechanism can be expressed by the following formula: Suppose the original attention module is replicated h times, that is, h-head attention; input the same data into the multi-head attention module to obtain the results of h single-head attentions; concatenate the multi-head results and pass through a linear layer conversion to obtain a fused feature output; the multi-head attention formula is as follows: 23). Upsampling module: Use a deconvolution neural network to restore the predicted feature vector to the same structure as the real data.

4. The method for recommending points of interest based on the prior of crowd movement patterns according to claim 1, wherein Step (5) specifically includes: 51). Connect the output of step (3) Deep network and the output of step (4) Cross network to obtain the concatenated vector x c = Cat(h, x); 52). Input the concatenated vector obtained in step 51) into the final single-layer perceptron, and output the predicted score p of user j's check-in at point of interest i. The formula is as follows: p = sigmoid(Wx c + b)8).