A Location Prediction Method and System for POI Missing

By combining the Bi-RNN and Softmax functions with the parameterization of Havercosin function and distance perception coefficient, the problem of low location prediction accuracy caused by POI loss is solved, the accuracy of the POI recommendation system is improved, and it is applied to recommendation systems, missing population analysis and personalized services.

CN115423166BActive Publication Date: 2025-07-08WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211033841.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-26
Publication Date
2025-07-08
Estimated Expiration
2042-08-26

AI Technical Summary

Technical Problem

In the prior art, the location prediction accuracy caused by the loss of POI in the absence of POI affects the data accuracy and service quality recommended by POI.

Method used

By using Bi-RNN-based method, the time interval and GPS distance are parameterized by obtaining user check-in data, using the Havercosin function and distance perception coefficient, combined with the Softmax function to predict missing locations, locations and categories, and considering the personalized eigenvectors and eigenvectors of time slots to improve prediction accuracy.

Benefits of technology

It improves the accuracy of location prediction of missing POIs, enhances the service quality of POI recommendation system, and applies it to recommendation systems, missing population analysis and personalized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115423166B_ABST
    Figure CN115423166B_ABST
Patent Text Reader

Abstract

The present invention discloses a location prediction method and system for POI missing, which relates to the field of trajectory data mining. The method comprises the following steps: obtaining valid check-in data of users based on a location social network service platform; given the time t when each user's location is missing, extracting the geographical location sequence and the corresponding hidden state before and after the time t respectively; allocating a personalized feature vector and a feature vector corresponding to the time when the POI is missing to the user and splicing them into the final feature vector; respectively calculating the x input The missing locations and corresponding location categories are predicted, and the top missing locations and location categories are selected in descending order of probability as the prediction data for the corresponding missing POI. The present invention greatly improves the prediction accuracy of the missing POI, thereby improving the convenience of application in related fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of trajectory data mining, and particularly to a method and system for predicting a location where a POI (Point of Interest, which is usually used to represent a social location with a geographical location in the real world, such as a shopping mall or a restaurant, and is associated with corresponding GPS coordinates) is missing. Background Art

[0002] The popularity of LBSNs (Location-Based Social Networks) has attracted more and more users to share their daily lives through location check-ins. The check-in behavior mainly consists of POIs, check-in times, and certain comments. With the continuous development of Internet technology and the increasing demand for personalization, LBSNs data has been effectively used in various applications, such as POI recommendation or prediction based on LBSNs data.

[0003] However, the quality problems of LBSNs data (such as missing check-in POIs and data sparsity) always limit the effectiveness of the above research and applications. See Figure 1 As shown, when a device fails or a user cheats, there is no corresponding semantic POI location check-in information for the spatial location (i.e., latitude and longitude) where a given check-in occurs, that is, the POI location is missing. Currently, research based on POI locations mainly focuses on recommending or predicting POIs that a user may go to in the future, assuming that all user check-ins specify clear POIs, without considering the situation of missing POIs. However, according to reports, more than half of the user check-ins are missing in the LBSN platforms of Twitter and Foursquare. Both POI missing and POI data sparsity will reduce the data accuracy of POI recommendation and affect the service quality of POI recommendation. Summary of the Invention

[0004] Aiming at the defects existing in the prior art, the technical problem solved by the present invention is: how to predict the missing location of a user when a POI is missing.

[0005] To achieve the above object, the present invention provides a method for predicting a location where a POI is missing, including the following steps:

[0006] S1: Obtain the valid data of user check-ins on a location-based social network service platform;

[0007] S2: Based on the data set in S1, given the time t when the location of each user is missing, respectively extract the geographical location sequences and location check-in sequences of the previous n times before the time t and the geographical location sequences and location check-in sequences and are input into the Bi - RNN to obtain the hidden state of the user's location check - in sequence and correspond to ; and correspond to, and the hidden state of the location category sequence and correspond to ; and correspond to;

[0008] S3: According to each user's and assign a trainable personalized feature vector e u to each user; Take the specified number of days X as a cycle, divide each day into several time periods Y, and divide the continuous time into Z time slots, Z = X·Y; Assign a trainable feature vector to each time slot, and determine the feature vector e corresponding to the moment when the POI is missing τ ;

[0009] S4: Concatenate e u and e τ to form the final feature vector x input , that is as the representation of user u at time t for the missing location, and respectively predict the missing location and the corresponding location category for x input through the Softmax function, to obtain several predicted missing locations and the probability of each location, as well as several predicted location categories and the probability of each location category;

[0010] S5: Select the top H missing locations and location categories in descending order of probability as the predicted data for the corresponding missing POIs.

[0011] Based on the above technical solution, the acquisition processes of and in S2 include:

[0012] Use the Havercosin function as a periodic function to parameterize the time interval ΔT between two adjacent check - ins ij :

[0013]

[0014] where ΔT ij represents the interval time between the user going to location i and location j, and calculate the time - period coefficient λ according to ΔT ij ​c (ΔT ij ),the calculation formula is: where β represents the time decay coefficient;

[0015] Given the subsequences of location categories of the previous and subsequent items of the user's check-in:

[0016]

[0017] where the superscript u of c represents the same user identifier, and the subscript ti represents the moment of the missing POI; respectively input and into Bi-RNN for training to obtain the hidden state corresponding to each input, and the training formula is:

[0018]

[0019]

[0020] According to λ c (ΔT ij ) to perform a normalization operation on the hidden state, and respectively obtain the hidden states and The calculation formula is:

[0021]

[0022]

[0023]

[0024] On the basis of the above technical solution, the acquisition processes of and in S2 include: calculating the distance perception coefficient ω ij according to the GPS coordinate distance ΔD p (ΔD ij ) between location i and location j, and the calculation formula is: where e is the natural constant and α is the coefficient that controls the attenuation of the weight with the increase of the distance;

[0025] Given the subsequences of the previous and subsequent locations of the user's check-in:

[0026] where the superscript u of p represents the same user identifier, and the subscript ti represents the moment of the missing POI;

[0027] Respectively input and into Bi-RNN for training to obtain the hidden state corresponding to each input, and the training formula is:

[0028]

[0029]

[0030] According to the distance perception coefficient ω p (ΔD ij ) perform a normalization operation on the hidden state to obtain the hidden states and The calculation formula is:

[0031]

[0032]

[0033]

[0034] On the basis of the above technical solution, in S4, the Softmax function is used to predict the missing location and the corresponding location category for x input The calculation formula is:

[0035]

[0036]

[0037] where, W p and W c are respectively trainable weight parameters, and b p and b c are respectively trainable bias parameters; represents: the probability of each predicted location, and M is the number of all locations; represents the probability of each predicted location category, and Q is the number of semantic categories of all locations.

[0038] On the basis of the above technical solution, the specific process of S1 includes: cleaning the user check-in data, retaining the user check-in data with more than 10 POI check-ins as valid data; unifying the format of the user check-in valid data; the value range of H described in S5 is 5 to 15.

[0039] The location prediction system for missing POIs provided by the present invention includes a user data acquisition module, a user hidden state training module, a user vector allocation module, a missing POI location prediction module, and a predicted data recommendation module;

[0040] The user data acquisition module is used to: acquire the user check-in valid data based on the location-based social network service platform;

[0041] The user hidden state training module is used for: based on the data set obtained by the user data acquisition module, given the time t when the location of each user is missing, respectively extract the geographical location sequences of the previous n times before the time t and the location check-in sequences as well as the geographical location sequences of the next n times after the time t and the location check-in sequences respectively input and into the Bi-RNN to obtain the hidden state of the location check-in sequence of the user and corresponding to respectively, and corresponding to and corresponding to respectively, and corresponding;

[0042] The user vector allocation module is used for: according to the and of each user, allocate a trainable personalized feature vector e to each user u ; take a specified number of days X as a cycle, divide each day into several time periods Y, and divide the continuous time into Z time slots, Z = X·Y; allocate a trainable feature vector to each time slot, and determine the feature vector e corresponding to the moment when the POI is missing τ ;

[0043] The missing POI location prediction module is used for: concatenate e u and e τ into the final feature vector x input , that is as the representation of the missing location of user u at time t, and respectively predict the missing location and the corresponding location category of x input through the Softmax function to obtain several predicted missing locations and the probability of each location, as well as several predicted location categories and the probability of each location category;

[0044] The prediction data recommendation module is used for: select the top H missing locations and location categories in descending order of probability as the prediction data for the corresponding missing POI.

[0045] Based on the above technical solution, the process by which the user hidden state training module obtains and includes:

[0046] Use the Havercosin function as a periodic function to parameterize the time interval ΔT between two consecutive check-ins ij :

[0047]

[0048] where ΔT ij represents the interval time between the user going to location i and location j. According to ΔT ij calculate the time period coefficient λ c (ΔT ij ), and the calculation formula is: where β represents the time decay coefficient;

[0049] Given the location category subsequences of the previous and subsequent items of the user's check-in:

[0050]

[0051] where the superscript u of c represents the same user identifier, and the subscript ti represents the moment of the missing POI; respectively input and into the Bi-RNN for training to obtain the hidden state corresponding to each input, and the training formula is:

[0052]

[0053]

[0054] According to λ c (ΔT ij ) perform a normalization operation on the hidden state to obtain the hidden states and respectively, and the calculation formula is:

[0055]

[0056]

[0057]

[0058] On the basis of the above technical solution, in the user hidden state training module, the acquisition process of and includes: According to the GPS coordinate distance ΔD between location i and location j ij , calculate the distance perception coefficient ω p (ΔD ij ), and the calculation formula is: where e is the natural constant, and α is the coefficient that controls the attenuation of the weight as the distance increases;

[0059] Given the subsequences of the previous and subsequent locations of user check-ins:

[0060] where the superscript u of p represents the same user identifier, and the subscript ti represents the moment of the missing POI;

[0061] Input and into the Bi-RNN for training respectively to obtain the hidden state corresponding to each input. The training formula is:

[0062]

[0063]

[0064] Perform a normalization operation on the hidden state according to the distance perception coefficient ω p (ΔD ij ) to obtain the hidden states and respectively. The calculation formula is:

[0065]

[0066]

[0067]

[0068] Based on the above technical solution, in the missing POI location prediction module, the Softmax function is used to predict the missing location and the corresponding location category of x input respectively. The calculation formula is:

[0069]

[0070]

[0071] where W p and W c are trainable weight parameters respectively, and b p and b c are trainable bias parameters respectively; represents the probability of each predicted location, and M is the number of all locations; represents the probability of each predicted location category, and Q is the number of semantic categories of all locations.

[0072] Based on the above technical solution, the working process of the user data acquisition module includes: cleaning the user check-in data, and retaining the user check-in data with the number of POI check-ins more than 10 times as valid data; unifying the format of the user check-in valid data; the value range of H in the prediction data recommendation module is 5 to 15.

[0073] Compared with the prior art, the advantages of the present invention are as follows:

[0074] Based on the prior art, after studying the missing POIs, the present invention independently developed a method for predicting the location and category of missing POIs. This method is based on bidirectional time and space and fully considers personalized preference perception, resulting in a significant improvement in prediction accuracy, thereby facilitating the application in related fields (such as recommendation systems, missing person analysis, disease tracing, personalized services, etc.). BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0076] Figure 1 It is a schematic diagram of missing POIs at a certain moment in the prior art;

[0077] Figure 2 It is a flowchart of the method for predicting the location of missing POIs in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0078] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.

[0079] The flowchart shown in the drawings is only an example, and does not necessarily include all contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.

[0080] First, introduce the research principle of this technology:

[0081] The present invention mainly focuses on the identification of missing POIs for users, that is, to identify where a user has visited at a specific time and spatial location in the past. Its significance lies in: on the one hand, considering the candidate distribution of the missing POIs, it helps to enrich the user check-in data, so as to better understand and model the user's mobility behavior and improve the service quality of POI recommendation; on the other hand, almost all POI recommendation or prediction tasks have the problem of data sparsity, and the identification of missing POIs can alleviate the data sparsity problem; this task can also be used for social welfare, such as suspect tracking or missing person search and analysis. In addition, by determining which places a patient has visited during the period of being out of contact in the past, we can trace the people who may have been in the same place as him / her to help avoid the risk of cross-infection.

[0082] In existing research, great progress has been made in the problems of POI recommendation and prediction. FPMC (Factorized Personalized Markov Chain) is a widely used method optimized by BPR (Bayesian Personalized Ranking), which embeds user preferences and personalized Markov chains for POI recommendation. Based on the extended RNN (Recurrent Neural Network), STRNN (Spatio-Temporal Recurrent Neural Network) is proposed, which models the local spatio-temporal environment of each layer with time and distance transformation matrices of time intervals and geographical distances respectively. In addition, we also adopt the HME (Hyperbolic Metric Embedding) method to replace the traditional Euclidean metric space to learn how to capture the underlying hierarchical structure and learn complex behavior patterns.

[0083] Although these methods have achieved satisfactory results, there are still some drawbacks. First, for the impacts of time periodicity, geographical proximity, user preferences, and sequential transition patterns, existing methods rarely consider all these influencing factors. Second, the preference of users for semantic categories can be regarded as a kind of coarse-grained preference, while the preference for specific POIs is a relatively fine-grained preference, and rich transfer data is definitely needed. However, most current methods directly model the POI location-level user preferences and sequential transitions on sparse data without considering the hierarchical structure in user mobility; finally, as a sequence prediction problem, the RNN network model is naturally used for time series data mining; then, due to the data sparsity problem, existing methods parameterize the spatio-temporal context as a matrix and incorporate it into each RNN layer, which may oversimplify the temporal features and spatial laws of mobility behavior and cannot measure the dynamic importance of relevant historical items; finally, the current mainstream POI location prediction methods mainly focus on the recommendation or prediction of the next moment's POI, rather than being specifically designed for the identification of missing POIs, because such problems require using relevant context information before and after the given query time, rather than recommending or predicting from a single perspective.

[0084] On this basis, seeFigure 2 As shown in the figure, the method for predicting the location where POI is missing according to the present invention includes the following steps:

[0085] S1: Obtain the valid user check-in data based on the location-based social network service platform (such as WeChat, Foursquare, Facebook, etc.). The specific process is as follows: Clean the user check-in data, and retain the user check-in data with the number of POI check-ins more than 10 times as the valid data; unify the format of the valid user check-in data to ensure data consistency.

[0086] S2: Based on the data set in S1, given the time t when each user has a location missing, extract the geographical location sequence of the previous n times before the time t and the location check-in sequence as well as the geographical location sequence of the next n times after the time t and the location check-in sequence Respectively input and into the Bi-RNN (Bi-direction Recurrent Neural Network, a two-way and two-grained recurrent neural network) to obtain the hidden state of the user's location check-in sequence (corresponding to ) and (corresponding to ), and the hidden state of the location category sequence (corresponding to ) and (corresponding to ).

[0087] Preferably, in S2 and are obtained based on the following principles: From a time perspective, the user's movement behavior shows strong periodicity and accompaniment. Intuitively, people tend to determine the category of the location they are going to at a specific periodic time and then determine a specific location, which naturally reflects the user's long-term preferences. Secondly, the user's movement behavior shows an obvious semantic category conversion rule. The shorter the time interval between two check-ins, the stronger the accompaniment.

[0088] On this basis, the acquisition methods of and in S2 are as follows: Model the above factors. This application uses the Havercosin function as a periodic function to parameterize the time interval ΔT between two adjacent check-ins ij , with the unit of (days), specifically:

[0089]

[0090] where ΔT ij represents the interval time between the user's travel from location i to location j. The above formula can be used to mine the hidden state of weekly preferences in trajectory movement, which helps to process sparse sequences. At the same time, the user's conversion preference between locations is not only affected by long-term periodic preferences, but also has accompaniment (such as going to have a meal after shopping and going to watch a movie after having a meal), and the strength of this accompaniment will weaken as the interval time between events becomes longer.

[0091] Therefore, while modeling the above periodic influence, it is necessary to introduce a decay factor based on the time interval, that is, calculate the time period coefficient λ ij according to ΔT c (ΔT ij ), and the calculation formula is:

[0092]

[0093] where β represents the time decay coefficient.

[0094] After that, given the subsequences of the location categories of the previous and subsequent items of the user's check-in:

[0095]

[0096] where the superscript u of c represents the same user identifier, and the subscript ti represents the moment of the missing POI;

[0097] Input and into the Bi-RNN for training respectively to obtain the hidden state corresponding to each input. The training formula is:

[0098]

[0099]

[0100] Finally, normalize the above hidden states according to λ c (ΔT ij ) to obtain the hidden states and respectively. The calculation formula is:

[0101]

[0102]

[0103]

[0104] Preferably, in S2 and The acquisition principle is as follows: Spatially, the check-ins of users in some frequently visited areas tend to be concentrated at certain locations. Most cities are divided into areas with certain implicit "functions", such as shopping, working, and resting. Therefore, the closer a user is to these areas, the more predictable their movement behavior becomes. This means that given the locations visited in the historical movement trajectory, the closer they are to the current moment, the greater their contribution to predicting the location at the next moment.

[0105] Based on this, in S2 and are obtained by spatially distance-parameterizing the relevant RNN hidden states output by some past locations, resulting in the following distance-aware weights, that is, according to the GPS coordinate distance ΔD ij , calculate the distance-aware coefficient ω p (ΔD ij ), and the calculation formula is:

[0106]

[0107] where e is the natural constant and α is the coefficient that controls the decay of the weight as the distance increases.

[0108] After that, given the previous and subsequent location subsequences of the user's check-in:

[0109]

[0110] where the superscript u of p represents the same user identifier and the subscript ti represents the moment of the missing POI;

[0111] Input and into the Bi-RNN for training respectively to obtain the hidden states corresponding to each input. The training formula is:

[0112]

[0113]

[0114] Finally, normalize the above hidden states according to the distance-aware coefficient ω p (ΔD ij ) to obtain the hidden states and The calculation formula is:

[0115]

[0116]

[0117]

[0118] S3: According to each user's and To express the personalized preferences of users, a trainable personalized feature vector e needs to be assigned to each user u , so as to achieve the differential expression of each individual user. At the same time, to characterize the personalized periodic preferences of each user, the continuous time needs to be discretized, that is, taking the specified number of days X as a cycle, and dividing each day into several time periods Y, and dividing the continuous time into Z time slots, Z = X·Y; assign a trainable feature vector to each time slot, and determine the feature vector e corresponding to the moment when the POI is missing τ .

[0119] It should be noted that considering that people's daily life is cycled by 7 days a week, so X is at least 7, and Y can be set by oneself, but it should not be less than 3. In this embodiment, X is 7 and Y is 6 (that is, every 4 hours is 1 time period).

[0120] S4: Combine e u and e τ to form the final feature vector x input , that is As the representation of user u at time t for the missing location, the Softmax function is used to predict the missing location and the corresponding location category for x input respectively (that is, identify "where did he / she go at a certain moment in the past?"), and obtain the predicted missing locations and the probabilities of each location, as well as the predicted location categories and the probabilities of each location category for each missing POI.

[0121] Preferably, the calculation formula for predicting the missing location and the corresponding location category for x input respectively through the Softmax function in S4 is:

[0122]

[0123]

[0124] where, W p and W c are trainable weight parameters respectively, b p and b c are trainable bias parameters respectively, represents: the probability of each predicted location (M is the number of all locations), represents the probability of each predicted location category (Q is the number of semantic categories of all locations).

[0125] S5: Select the top H missing location points and location categories in descending order of probability as the predicted data for the corresponding missing POIs.

[0126] The applicant ran on a computer with an Intel(R) Core(TM) i7-7700K CPU@4.20GHz and a 2080Ti GPU, using the method of this embodiment and the publicly available datasets NYC and TKY and the literature (Q. Liu, S. Wu, L. Wang, and T. Tan, “Predicting the next location: A recurrent model with spatial and temporal contexts,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 30, no. 1, 2016.), (C. Yang, L. Bai, C. Zhang, Q. Yuan, and J. Han, “Bridging collaborative filtering and semi-supervised learning: a neural approach for poi recommendation,” in Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2017, pp. 1245–1254.), (D. Xi, F. Zhuang, Y. Liu, J. Gu, H. Xiong, and Q. He, “Modelling of bidirectional spatio-temporal dependence and users’ dynamic preferences for missing poi check-in identification,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, 2019, pp. 5458–5465.) for comparison and obtained the following information:

[0127] The number of test samples (i.e., the number of all missing POIs) is 13,258. When the recommended missing location for each missing POI is 1, the correct rate is approximately 18.78% (i.e., the probability of the predicted correct location of the missing POI). When the recommended missing location for each missing POI is 5, the correct rate is approximately 41.79%. When the recommended missing location for each missing POI is 10, the correct rate is approximately 52.26%.

[0128] It can be seen from this that when the recommended missing location is 1, the correct rate is relatively low. When the recommended missing location is more than 5, the correct rate gradually increases. However, if the recommended missing location is too many, it is not conducive to later applications. Therefore, in this embodiment, the value range of H is 5 to 15.

[0129] As can be seen from the above, based on the prior art, after researching the missing POIs, the embodiment of the present invention independently developed a method for predicting the location and category of missing POIs. This method is based on bidirectional time and space and fully considers personalized preference perception, and its prediction accuracy has been greatly improved, thereby facilitating the applications in related fields (such as recommendation systems, missing person analysis, disease tracing, personalized services, etc.).

[0130] The location prediction system for missing POIs provided by the present invention includes a user data acquisition module, a user hidden state training module, a user vector allocation module, a missing POI location prediction module, and a prediction data recommendation module;

[0131] The user data acquisition module is used to: acquire the valid user check-in data based on the location-based social network service platform;

[0132] The user hidden state training module is used to: based on the data set acquired by the user data acquisition module, given the time t when each user has a missing location, respectively extract the geographical location sequences of the n times before time t and the location check-in sequences as well as the geographical location sequences of the n times after time t and the location check-in sequences respectively input and into the Bi-RNN to obtain the hidden states of the location check-in sequences of the user and corresponding to respectively, corresponding to respectively, as well as the hidden states of the location category sequences and corresponding to respectively, corresponding to respectively;

[0133] The user vector allocation module is used to: according to each user's and allocate a trainable personalized feature vector e for each user u ; taking the specified number of days X as a cycle, dividing each day into several time periods Y, and dividing the continuous time into Z time slots, where Z = X·Y; allocating a trainable feature vector for each time slot, and determining the feature vector e corresponding to the moment when the POI is missing τ ;

[0134] The missing POI location prediction module is used to: e u and e τ concatenate them into the final feature vector x input , that is as the representation of the missing location of user u at time t, and respectively predict the missing location and the corresponding location category of x input through the Softmax function, obtaining several predicted missing locations and the probabilities of each location, as well as several predicted location categories and the probabilities of each location category for each missing POI;

[0135] The prediction data recommendation module is used to: select the top H missing locations and location categories in descending order of probability as the prediction data for the corresponding missing POI.

[0136] Based on the above technical solution, the process for the user hidden state training module to obtain and includes:

[0137] Using the Havercosin function as a periodic function to parameterize the time interval ΔT between two adjacent check-ins ij :

[0138]

[0139] where ΔT ij represents the interval time between the user going to location i and location j, and calculates the time period coefficient λ ij according to ΔT c (ΔT ij ), and the calculation formula is: where β represents the time decay coefficient;

[0140] Given the subsequence of location categories of the previous and subsequent items of the user's check-in:

[0141]

[0142] where the superscript u of c represents the same user identifier, and the subscript ti represents the moment of the missing POI; respectively, and are input into the Bi - RNN for training to obtain the hidden state corresponding to each input. The training formula is:

[0143]

[0144]

[0145] According to λ c (ΔT ij ), a normalization operation is performed on the hidden state to obtain the hidden states and respectively. The calculation formula is:

[0146]

[0147]

[0148]

[0149] On the basis of the above technical solution, the acquisition process of and in the user hidden state training module includes: calculating the distance perception coefficient ω ij (ΔD p ) according to the GPS coordinate distance ΔD ij between location i and location j. The calculation formula is: where e is the natural constant and α is the coefficient that controls the attenuation of the weight with the increase of the distance;

[0150] Given the previous and subsequent location subsequences of the user's check - in:

[0151] where the superscript u of p represents the same user identifier, and the subscript ti represents the moment of the missing POI;

[0152] respectively, and are input into the Bi - RNN for training to obtain the hidden state corresponding to each input. The training formula is:

[0153]

[0154]

[0155] According to the distance perception coefficient ω p (ΔD ij ), a normalization operation is performed on the hidden state to obtain the hidden states and The calculation formula is as follows:

[0156]

[0157]

[0158]

[0159] Based on the above technical solution, in the missing POI location prediction module, the calculation formula for predicting the missing location and the corresponding location category through the Softmax function for x respectively is as follows: input The calculation formula for predicting the missing location and the corresponding location category is as follows:

[0160]

[0161]

[0162] Where, W p and W c are respectively trainable weight parameters, and b p and b c are respectively trainable bias parameters; represents: the probability of each predicted location, and M is the number of all locations; represents the probability of each predicted location category, and Q is the number of semantic categories of all locations.

[0163] Based on the above technical solution, the working process of the user data acquisition module includes: cleaning the user check-in data, retaining the user check-in data with more than 10 POI check-in times as valid data; unifying the format of the user check-in valid data; the value range of H in the predicted data recommendation module is 5 to 15.

[0164] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware and their appropriate combinations. In the hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or be implemented as hardware, or be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable storage medium, and the computer-readable storage medium can include a computer-readable storage medium (or non-transitory medium) and a communication medium (or transitory medium).

[0165] As is known to those of ordinary skill in the art, the term computer-readable storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is known to those of ordinary skill in the art, communication media typically contain computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0166] Exemplarily, the computer-readable storage medium may be an internal storage unit of the electronic device in the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device.

[0167] The above are only specific embodiments of the embodiments of the present invention, but the protection scope of the embodiments of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the embodiments of the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the embodiments of the present invention. Therefore, the protection scope of the embodiments of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for predicting the location of missing POIs, characterized in that, The method includes the following steps: S1: Obtain the valid user check-in data of the location-based social network service platform; S2: Based on the data set of S1, given the time t when the location of each user is missing, extract the geographical location sequences of the previous n times before time t and the location check-in sequences as well as the geographical location sequences of the next n times after time t and the location check-in sequences respectively input and into the Bi-RNN to obtain the hidden states of the user's location check-in sequences and corresponding to , and corresponding to, as well as the hidden states of the location category sequences and corresponding to , and corresponding; S3: According to each user's and assign a trainable personalized feature vector e to each user u ; take the specified number of days X as a cycle, divide each day into several time periods Y, divide the continuous time into Z time slots, Z = X·Y; assign a trainable feature vector to each time slot, and determine the feature vector e corresponding to the moment when the POI is missing τ ; S4: Combine e u and e τ to form the final feature vector x input , that is As the representation of user u at time t for the missing location, use the Softmax function to separately predict the missing location and the corresponding location category for x input to obtain several predicted missing locations and the probabilities of each location, as well as several predicted location categories and the probabilities of each location category for each missing POI; S5: Select the top H missing location places and place categories in descending order of probability as the predicted data for the corresponding missing POIs.

2. The method for predicting the location where a POI is missing according to claim 1, wherein As described in S2 and The acquisition process includes: Use the Havercosin function as a periodic function to parameterize the time interval ΔT between two consecutive check-ins ij : where ΔT ij represents the interval time between the user going to location i and location j. According to ΔT ij calculate the time period coefficient λ c (ΔT ij ), and the calculation formula is: where β represents the time decay coefficient; Given the subsequences of place categories of the previous and subsequent items of the user check-in: where the superscript u of c represents the same user identifier, and the subscript ti represents the time when the POI is missing; respectively input and into the Bi-RNN for training to obtain the hidden state corresponding to each input, and the training formula is: According to λ c (ΔT ij ) perform a normalization operation on the hidden state to obtain the hidden states and The calculation formula is:

3. The method for predicting the location of missing POIs according to claim 1, wherein As described in S2 and The acquisition process includes: calculating the distance perception coefficient ω ij based on the GPS coordinate distance ΔD between location i and location j p (ΔD ij ), and the calculation formula is: where e is the natural constant and α is the coefficient controlling the attenuation of the weight with the increase of distance; Given the subsequences of places of the previous and subsequent items of the user check-in: Where the superscript u of p represents the same user identifier, and the subscript ti represents the moment of the missing POI; Input and into the Bi-RNN for training respectively to obtain the hidden state corresponding to each input. The training formula is as follows:​​​​ According to the distance perception coefficient ω p (ΔD ij ) perform a normalization operation on the hidden state to obtain the hidden states and The calculation formula is as follows:

4. The method for predicting the location of missing POI according to claim 1, wherein In S4, the Softmax function is used to separately predict the missing location and the corresponding location category for x input The calculation formula for this prediction is as follows: Among them, W p and W c are respectively trainable weight parameters, and b p and b c are respectively trainable bias parameters; represents the probability of each predicted location, where M is the number of all locations; represents the probability of each predicted location category, where Q is the number of semantic categories of all locations.

5. The method for predicting the location of the missing POI according to any one of claims 1 to 4, characterized in that The specific process of S1 includes: cleaning the user check-in data, retaining the user check-in data with the number of POI check-ins above 10 as valid data; unifying the format of the valid user check-in data; the value range of H in S5 is 5 to 15.

6. A location prediction system for POI missing, characterized in that, The system includes a user data acquisition module, a user hidden state training module, a user vector allocation module, a missing POI location prediction module, and a predicted data recommendation module; The user data acquisition module is used to: obtain the valid user check-in data of the location-based social network service platform; The user hidden state training module is used to: based on the data set obtained by the user data acquisition module, given the time t when the location of each user is missing, extract the geographical location sequences of the previous n times before the time t respectively and the location check-in sequences as well as the geographical location sequences of the next n times after the time t and the location check-in sequences respectively input and into the Bi-RNN to obtain the hidden state of the user's location check-in sequence and corresponding to and corresponding to as well as the hidden state of the location category sequence and corresponding to and corresponding to ; The user vector allocation module is used to: according to and allocate a trainable personalized feature vector e to each user u ; take the specified number of days X as a cycle, divide each day into several time periods Y, divide the continuous time into Z time slots, Z = X · Y; allocate a trainable feature vector to each time slot, and determine the feature vector e corresponding to the moment when the POI is missing τ ; The missing POI location prediction module is used to: combine e u and e τ to form the final feature vector x input , that is as the representation of the missing location of user u at time t. Through the Softmax function, the missing location and the corresponding location category are predicted for x input respectively, obtaining several predicted missing locations and the probabilities of each location, as well as several predicted location categories and the probabilities of each location category for each missing POI; The predicted data recommendation module is used to: select the top H missing location places and place categories in descending order of probability as the predicted data for the corresponding missing POIs.

7. The location prediction system for POI missing as claimed in claim 6, wherein The user hidden state training module obtains and The process includes: Using the Havercosin function as a periodic function to parameterize the time interval ΔT between two consecutive check-ins ij : where ΔT ij represents the interval time between the user's going to location i and location j, and based on ΔT ij calculate the time period coefficient λ c (ΔT ij ), and the calculation formula is: where β represents the time decay coefficient; Given the subsequences of place categories of the previous and subsequent items of the user check-in: where the superscript u of c represents the same user identifier, and the subscript ti represents the time when the POI is missing; respectively input and into the Bi-RNN for training to obtain the hidden state corresponding to each input. The training formula is: According to λ c (ΔT ij ) perform a normalization operation on the hidden state to obtain the hidden states and The calculation formula is:

8. The location prediction system for POI absence as claimed in claim 6, wherein In the user hidden state training module, the and acquisition process includes: calculating the distance perception coefficient ω ij based on the GPS coordinate distance ΔD between location i and location j p (ΔD ij ), and the calculation formula is: where e is the natural constant and α is the coefficient controlling the attenuation of the weight with the increase of distance; Given the subsequences of places of the previous and subsequent items of the user check-in: Where the superscript u of p represents the same user identifier, and the subscript ti represents the time of the missing POI; Input and into the Bi-RNN for training respectively to obtain the hidden state corresponding to each input. The training formula is as follows:​​​​ According to the distance perception coefficient ω p (ΔD ij ) perform a normalization operation on the hidden state to obtain the hidden states and The calculation formula is as follows:

9. The location prediction system for POI absence as described in claim 6, wherein Missing POI location prediction module uses the Softmax function to separately predict the missing location and the corresponding location category for x input The calculation formula for predicting the missing location and the corresponding location category is as follows: Among them, W p and W c are trainable weight parameters respectively, and b p and b c are trainable bias parameters respectively; represents the probability of each predicted location, where M is the number of all locations; represents the probability of each predicted location category, where Q is the number of semantic categories of all locations.

10. The location prediction system for POI missing according to any one of claims 6 to 9, characterized in that The working process of the user data acquisition module includes: cleaning the user check-in data, retaining the user check-in data with the number of POI check-ins above 10 as valid data; unifying the format of the valid user check-in data; the value range of H in the predicted data recommendation module is 5 to 15.

Citation Information

Patent Citations

  • Missing POI track completion method based on mask and bidirectional model

    CN114116692A

  • Points of interest (POI) ranking based on mobile user related data

    US20130262479A1