Interest point sign-in sequence generation method based on diffusion model
Through space-time lossless encoding and conditional U-shaped network based on diffusion model, combined with a comparative learning strategy, the problems of different lengths and insufficient consideration of space-time correlation in the prior art are solved, and high-quality sign-in sequence generation is achieved.
Patent Information
- Application Number
- CN202411931849.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The prior art is difficult to generate a complete and high-quality point-of-interest check-in sequence, especially when processing trajectory data of different lengths, noise is easily introduced and space-time correlation is not fully considered.
Using a diffusion model-based method, the check-in sequences of different lengths are converted into check-in vectors of equal lengths through space-time lossless encoding, a spatial diffusion module and a temporal diffusion module are established, and the spatiotemporal features and correlations are captured using conditional U-shaped networks and contrast learning strategies.
Effectively handling check-in sequences of unfixed lengths significantly improves the quality and authenticity of the generated check-in sequences and can better capture spatial and temporal context information.
Smart Images

Figure CN120011656A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of spatiotemporal data mining and deep learning technology, and relates to a method for generating a check-in sequence of points of interest based on a diffusion model. Background Art
[0002] The sign-in sequence generation task belongs to the research field of "spatial-temporal data generation". The existing published technical solutions for spatial-temporal data generation include the following, and their shortcomings are also described:
[0003] (1) A next POI recommendation method based on user preference and spatiotemporal context information (CN117194763A) obtains user check-in records for preprocessing, uses low-dimensional dense vectors to embed points of interest and various auxiliary information, builds and trains a next point of interest recommendation model based on user preference and spatiotemporal context information, inputs the user's long-term check-in sequence and short-term check-in sequence, generates k points of interest, and recommends them to the user; however, this method can only generate the user's next point of interest, but not the complete point of interest check-in sequence. Although this method introduces spatiotemporal context information to improve the recommendation accuracy, it does not fully consider the spatiotemporal correlation in the point of interest check-in sequence, which may cause the recommendation result to be inconsistent with the user's actual situation.
[0004] (2) A method and system for automatically generating the activity trajectory and destination of a person (CN111985452A), including: determining the information of the person to be analyzed and the vehicle information under the name of the person to be analyzed and his / her related persons; determining the first activity trajectory of the vehicle of the person to be analyzed within the set time period based on the information of the checkpoints that the vehicle has passed through within the set time period; determining the itinerary of the person to be analyzed and the key point information in all the itineraries based on the activity trajectory; confirming the destination information of the person to be analyzed based on the key point location information. This method relies on vehicle information and existing places and equipment, which may result in the inability to accurately generate the activity trajectory and destination information of some people, and will be affected by incomplete data and privacy issues; at the same time, the method mainly determines the activity trajectory based on the information of the checkpoints that the vehicle has passed through, and may not be able to generate accurate information when there is no vehicle or the vehicle has not passed through the checkpoint.
[0005] (3) A vehicle trajectory generation method (CN110223515B) A vehicle trajectory generation method based on GAN is proposed, step 1, data processing stage, the data processing stage is to first pre-process the trajectory data and map data; step 2, data generation model stage, the data generation model stage includes generating a road segment trajectory generation model and a trip trajectory generation model; step 3, data set generation stage, the data set generation stage is to load the road segment trajectory generation model and the trip trajectory generation model to obtain trajectory data. Since the length of the check-in sequence is not fixed, this method is difficult to effectively process trajectory data of different lengths. This method requires fixed-length input data, so when processing trajectory data of different lengths, padding or truncation operations may be required. However, these operations may introduce additional noise, affecting the accuracy and quality of the generated trajectory; in addition, due to the extreme instability of GAN training, its generated results may be unrealistic or unreasonable.
[0006] In summary, the published technical solutions do not fully consider the spatiotemporal correlation in the POI check-in sequence, rely on existing venues and equipment, and are difficult to effectively process trajectory data of different lengths, resulting in the inability to generate a complete and high-quality POI check-in sequence.
[0007] The development of wireless network technology enables individual mobility events to be detected with high precision. These records are called check-in sequences of points of interest, which provide feasibility for fully understanding the mobility of urban populations. By extracting high-level semantics of individual mobility events, they can be applied to downstream tasks such as recommendation of points of interest, selection of commercial locations, and time prediction. However, directly using real-life check-in sequences to support downstream tasks will inevitably lead to privacy issues, making it difficult to obtain large-scale publicly available data on human activities.
[0008] Specific details of the data preprocessing part: Source dataset introduction:
[0009] University of Canberra Wireless Network Usage Dataset.
[0010] (1) Dataset content
[0011] This dataset records the connection location information, connection bandwidth, connection duration, and other important sign-in sequence data that can reflect user behavior when teachers and students of the University of Canberra connect to the campus free wireless network. The data is stored in the form of a CSV file, and the main data fields are as follows:
[0012] Canberra_checkins.csv
[0013] This file records each record of a user accessing an AP, including the access location, access time, and access duration.
[0014]
[0015] (2) Dataset size
[0016] The downloaded data set contains a total of 215,527 wireless network access records, which correspond to 4,441 users and 317 AP access points after classification.
[0017]
[0018] Dartmouth College Wireless Network Usage Dataset.
[0019] (1) Dataset content
[0020] This dataset records the wireless network usage of Dartmouth College students, involving 476 wireless network access points and covering 161 buildings of Dartmouth College. The data is stored in the form of a CSV file (two-dimensional table). The main data fields are as follows:
[0021] AP_Locations.csv, this file records the geographic location coordinates corresponding to each AP (wireless network access point) number.
[0022]
[0023] [ID].csv, there are several files, which record the AP access records of each user and are stored in a folder named "2001-2003", where [ID] is the ID of the corresponding user.
[0024]
[0025] (2) Dataset size. The downloaded dataset contains 623 AP access point numbers (including some AP access point data with missing values) and 6202 users, with a total of 193,958 records.
[0026]
[0027] Weeplace Dataset:
[0028] (1) Dataset content,This dataset records the check-in activities of users on location-based social platforms. The data is stored in the form of CSV files, where the main fields are as follows:
[0029] Weeplace_checkins.csv
[0030]
[0031]
[0032] (2) Dataset size. The downloaded dataset contains 1097 AP access point numbers (including some AP access point data with missing values) and 1432 users, with a total of 3,041,596 records.
[0033]
[0034] Gowalla Dataset:
[0035] (1) Dataset content. This dataset is similar to the Weeplace dataset. Both datasets record the check-in activities of users on location-based social platforms. The data is stored in CSV format. The main fields are as follows:
[0036] Gowalla_checkins.csv
[0037]
[0038] (2) Dataset size. The downloaded dataset contains 3028 AP access point numbers (including some AP access point data with missing values) and 3286 users, with a total of 2148575 records.
[0039]
[0040] Data cleaning process: Since the quality of records in the directly acquired data set varies, a series of cleaning operations are required to remove records containing missing values in the data set and screen users with high-quality check-in sequences for subsequent model training.
[0041] All datasets adopt a unified processing flow, and the common fields of the above datasets are as follows:
[0042]
[0043] Data filtering;
[0044] (1) After deleting the users that do not contain trajectory point data after filtering (i.e., deleting the empty CSV files), valid users are obtained, i.e., users with at least one trajectory point;
[0045] (2) Data with a SESSION_DURATION value less than 3 were deleted (connection duration less than 3 minutes was considered an invalid connection);
[0046] (3) Data with a field AVG_KBPS value less than 1 is deleted (a bandwidth less than 1 kbps cannot be used for normal network communication and is considered an invalid connection).
[0047] The filtered data all meet the requirement that the SESSION_DURATION value is greater than or equal to 3, and the AVG_KBPS value is greater than or equal to 1.
[0048] Data de-scrambling: The value of the LOCATION_CODE field of the continuous data with a length greater than or equal to 4 that jumps back and forth between two AP connection points is uniformly modified to the LOCATION_CODE value of the data with the largest SESSION_DURATION value.
[0049]
[0050] Data merging;
[0051] (1) Aggregation based on the day granularity:
[0052] Based on the CONNECT_DATE field, the connection data of the same day are merged into one connection data with the day as the granularity. The merged connection data includes:
[0053] The value of CONNECT_DATE is the date of that day (format: Y / m / d);
[0054] The value of SESSION_DURATION is the sum of the connection data within the same day;
[0055] The value of AVG_KBPS is the weighted sum of the connection data within the same day (the weight is the proportion of SESSION_DURATION);
[0056] The values of the remaining fields are equal to the values of the fields corresponding to the data with the largest SESSTION_DURATION value.
[0057] (2) Merge adjacent identical AP records:
[0058] The data of users' WiFi access time on the same AP within 2 hours are merged into one piece of data. The new data field values after the merger are as follows:
[0059] The value of CONNECT_DATE is the earliest date in the original data to be merged (format: Y / m / d H:M:S) year / month / day hour:minute:second;
[0060] The value of SESSION_DURATION is the sum of the values of the corresponding fields of the original data to be merged;
[0061] The value of AVG_KBPS is the weighted sum of the corresponding field values in the data to be merged (the weight is the proportion of SESSION_DURATION);
[0062] The values of the remaining fields are equal to the values of the fields corresponding to the data with the largest SESSTION_DURATION value in the original data to be merged.
[0063] Data processing results;
[0064] After preprocessing the data set, users with high check-in sequence quality (the length of check-in sequences in a day is between 10 and 25) are screened out. The data size after screening is as follows:
[0065] Summary of the invention
[0066] In view of the problems in the prior art, the present invention provides a method for generating a check-in sequence of a point of interest based on a diffusion model, which generates a sequence by learning the distribution of a real sequence or supplementing the real sequence, obtains an equivalent data analysis result through the generated sequence and supports high-level semantic modeling.
[0067] A method for generating a check-in sequence of interest points based on a diffusion model includes the following steps:
[0068] Step 1: Data preprocessing: cleaning and smoothing the real-world data set to ensure data quality and consistency.
[0069] Step 2: Perform space-time lossless coding to convert sign-in sequences of different lengths into sign-in vectors of equal length, and encode them through spatial frequency vectors and time bucket vectors to retain the information of the original sequence.
[0070] Step 3: Establish a diffusion model, construct a spatial diffusion module and a temporal diffusion module, and use the forward diffusion process and the reverse reconstruction process to capture the spatiotemporal characteristics.
[0071] Step 4: Introduce a conditional U-network into the diffusion module and propose a denoising network to capture complex spatiotemporal correlations, modeled through a self-attention mechanism.
[0072] Step 5: Using the contrastive learning strategy, ternary contrastive learning is adopted to further capture the spatiotemporal correlation of the check-in sequence and strengthen the connection between the temporal and spatial diffusion modules.
[0073] The advantages of the present invention are that it can effectively solve the problems of different input code lengths, difficulty in capturing spatiotemporal correlation, and noise interference. The method is used for generating a sign-in sequence of interest points, and has a significant effect.
[0074] Table 1 Comparison on four real-world datasets
[0075]
[0076]
[0077] Table 1 shows the performance of the proposed model on four real-world datasets.
[0078] Table 1 shows the performance comparison of the proposed model on four real-world datasets, using different metrics to evaluate the performance of the model, including JSD-all, JSD-t, JSD-r, and JSD-u.
[0079] JSD-all represents the overall Jensen-Shannon distance, which measures the similarity between the generated check-in sequence and the real data.
[0080] A lower JSD-all value indicates that the sequences generated by the model are closer to the real data and the model performs better.
[0081] JSD-t represents the Jensen-Shannon distance in the time dimension, which evaluates the difference between the time distribution of the check-in sequence generated by the model and the real data.
[0082] A lower JSD-t value indicates that the model is better at generating patterns over time.
[0083] JSD-r represents the Jensen-Shannon distance in the spatial dimension, which evaluates the difference between the spatial distribution of the sign-in sequence generated by the model and the real data. A lower JSD-r value indicates that the model has better generation effect in space.
[0084] JSD-u represents the Jensen-Shannon distance in the user dimension, which evaluates the difference between the check-in sequence generated by the model and the real data in terms of user distribution.
[0085] A lower JSD-u value indicates that the model performs better on the users.
[0086] It can be seen from the table that on the four data sets, the model of the present invention shows relatively low values in the JSD-all, JSD-t, JSD-r and JSD-u indicators, which indicates that the similarity between the generated check-in sequences and the real data is high and the model has good performance.
[0087] Compared with the prior art, the present invention has the following advantages:
[0088] Strong ability to process sequences of variable length: The technology of the present invention has stronger processing capabilities for check-in sequences of variable length. By adopting spatiotemporal lossless coding technology, the method of the present invention can effectively process check-in sequences of different lengths without introducing additional noise or losing key information.
[0089] Effectively capture spatial and temporal context information: The technology of the present invention utilizes two different diffusion modules to distribute and capture temporal and spatial context information to ensure that the generated check-in sequence is closer to the actual situation.
[0090] Effectively capture spatial and temporal correlations: The technology of the present invention adopts a contrastive learning strategy to capture the spatiotemporal correlations of the check-in sequence and strengthen the connection between the temporal and spatial diffusion modules.
[0091] The effect of decoupling strategy:
[0092] To demonstrate the necessity of adopting a separation strategy, spatial and temporal features are intentionally merged without separation. In this configuration, only a denoising network is used to generate the check-in sequence, without incorporating contrastive learning or additional conditions. The rest of the settings are kept the same as STCDM. This setting allows evaluating the effectiveness of the separation generation. Figure 5a , Figure 5b , Figure 5c , Figure 5d As shown in , when temporal and spatial features are not separated, noise may interfere with each other. In addition, temporal and spatial features show significant differences. Mixing spatial and temporal features together affects the model's ability to capture their characteristics. The final result is even worse than many baseline models. The difference is even greater when compared to the performance of Diff-Traj. This shows that separate modeling significantly affects the overall effectiveness of the model. This suggests that relying solely on a single denoising network is not enough to effectively model spatial-temporal correlation. After separation, the effective spatial-temporal correlation strategy proposed by the model can significantly improve the quality of the generated check-in sequences. This also shows that compared with CUnet, the denoising network of Diff-Traj is more suitable for this specific setting, but further improvements are difficult to achieve. The denoising network of can show better performance when combined with the decoupling framework.
[0093] Ablation experiment:
[0094] To further evaluate the effects of different components in STCDM, these four variants are compared with the STCDM model. Ablation experiments are performed and the experimental results on all datasets are analyzed.
[0095] w / o Cond\&Cont: The conditioning information is removed from the denoising network. The input and temporal embeddings are used directly to predict the noise. Contrastive learning is removed. The rest of the settings are the same as STCDM.
[0096] w / o Cond: The conditional information is removed from the denoising network. The input and temporal embedding are used directly to predict the noise. The rest of the settings are the same as STCDM.
[0097] Remove Contrastive Learning (w / o Cont): The contrastive loss is directly removed. The rest of the settings are the same as STCDM. This setting is used to evaluate the effectiveness of contrastive learning at the sequence level.
[0098] w / o CUnet: Use normal Unet as the denoising network instead of CUnet. The rest of the settings are the same as STCDM. Use this setting to evaluate the performance of CUnet.
[0099] Figure 6a , Figure 6b , Figure 6c , Figure 6d The results show that the two diffusion modules and contrastive learning design of play an important role in modeling space-time correlation. These methods effectively capture the space-time correlation from their different spatial and temporal features. At the same time, the ability to generate high-quality check-in sequences relies on the unique denoising network design. In summary, each module and block of the design contributes to improving the quality of the generated check-in sequences. The superior combination of these modules greatly enhances the performance of the model and the realism of the generated sequences.
[0100] In order to more intuitively compare the generated results, the model and Diff-Traj are visualized, focusing on two aspects of the Dartmouth dataset. First, the visualization of the number of visits to 256 POIs selected from the Dartmouth dataset is presented, providing an overview of the visit frequency distribution. Figure 7a , Figure 7b , Figure 7c and Figure 7d shown.
[0101] Figure 7a : Real-world: Sequence distribution in the real world.
[0102] Figure 7b :STCDM:The method for generating check-in sequences of interest points based on diffusion model proposed in the present invention. That is, the distribution of user sequences generated by the STCDM method.
[0103] Figure 7c : Diff-Traj: Data distribution generated by one of the baseline models named Diff-Traj
[0104] Figure 7d :Real: In the chart legend, indicates the real data distribution.
[0105] STCDM: A method for generating check-in sequences of interest points based on a diffusion model proposed in this paper.
[0106] Visit density in a day: The distribution of different models and real data in a day.
[0107] The check-in sequences generated by the model are closer to the spatial distribution observed in the real world. Secondly, the visit time distribution of a specific POI is presented. Compared with Diff-Traj, the model generates a visit time distribution that better captures the patterns observed in the real world. These cases highlight the significant advantage of the model in capturing the spatial-temporal distribution of the generated check-in sequences. They demonstrate the ability of the model to effectively capture spatial-temporal correlations and generate more realistic check-in sequences. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. As shown in the figure:
[0109] Figure 1 It is the spatiotemporal contrast diffusion model framework of the present invention.
[0110] Figure 2 This is the spatiotemporal lossless coding method of the present invention.
[0111] Figure 3 It is the conditional U-type network of the present invention.
[0112] Figure 4 This is the comparative learning process of the present invention.
[0113] Figure 5a This is a diagram of the decoupling experiment of the present invention.
[0114] Figure 5b This is the second figure of the decoupling experiment of the present invention.
[0115] Figure 5c These are three diagrams of the decoupling experiment of the present invention.
[0116] Figure 5d These are four figures of the decoupling experiment of the present invention.
[0117] Figure 6a This is one of the ablation experiment results of the present invention.
[0118] Figure 6bThis is the second ablation experiment result of the present invention.
[0119] Figure 6c This is the third ablation experiment result of the present invention.
[0120] Figure 6d This is the fourth ablation experiment result of the present invention.
[0121] Figure 7a This is one of the visual comparisons of the present invention.
[0122] Figure 7b This is the second visualization comparison of the present invention.
[0123] Figure 7c This is the third visualization comparison of the present invention.
[0124] Figure 7d This is the fourth visualization comparison of the present invention.
[0125] Explanation of foreign language in the picture:
[0126] Canberra: refers to the University of Canberra wireless network access dataset.
[0127] Dartmouth: refers to the Dartmouth College Wireless Network Access Dataset.
[0128] Weeplace: This is the name of a dataset that records wireless network access by users in a city.
[0129] Gowalla: This is a well-known location-based social network dataset.
[0130] STCDM: abbreviation of the Spatial-Temporal Conditional Diffusion Model (Spatial-Temporal Conditional Diffusion Model) proposed in the present invention.
[0131] w / o Cond&Cont: Abbreviation for "without Condition and Content", indicating a model without conditions and content.
[0132] w / o Cond: Abbreviation for "without Condition", indicating a model without conditions.
[0133] w / o Cont: Abbreviation for "without Content", indicating a model without content.
[0134] w / o STUnet: Abbreviation for "without Spatio-Temporal U-net", which means a model without spatio-temporal U-net.
[0135] w / o disentanglement: non-decoupling mode.
[0136] JSD-all: used to measure the similarity between the model-generated sequence and the real data in all aspects.
[0137] JSD-t: measures the similarity between the model-generated sequence and the real data in the time dimension.
[0138] JSD-r: Measures the similarity between the model-generated sequence and the real data in the spatial dimension.
[0139] JSD-u: Measures the similarity between the model-generated sequence and the real data based on user-related features.
[0140] Real-world: real-world data distribution.
[0141] Diff-Traj: Data distribution generated by one of the baseline models named Diff-Traj.
[0142] Visit density in a day: The distribution of different models and real data in a day. DETAILED DESCRIPTION
[0143] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0144] Example 1: Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5a , Figure 5b , Figure 5c , Figure 5d , Figure 6a , Figure 6b , Figure 6c , Figure 6d , Figure 7a , Figure 7b , Figure 7c and Figure 7d As shown in Figure 1, a method for generating check-in sequences of interest points based on a diffusion model is proposed. The model learns a parameterized denoising network and gradually denoises the noise sampled in the prior distribution based on the spatiotemporal mutual information, thereby generating a high-quality check-in sequence. The overall framework of the model is shown in Figure 1. Figure 1 shown.
[0145] To better illustrate the example, the following definitions are made based on the check-in data of points of interest:
[0146] Interest point check-in sequence c: The interest point check-in sequence c is defined as a sequence containing check-in points c = <(s1,τ1),(s2,τ2),…,(s m ,τ m ),…,(s M ,τ M )>.
[0147] Where M is the length of the check-in sequence.
[0148] s m ∈S is the index number of a POI (point of interest), and S is a POI set containing L POIs, that is, |S|=L.
[0149] m: is a non-negative integer.
[0150] s1: The index number of the first POI (point of interest).
[0151] s m : The index number of the mth POI (point of interest).
[0152] s M : The index number of the Mth POI (point of interest).
[0153] τ1: The index number of the first arrival time slot.
[0154] τ m : The index number of the mth arrival time slot.
[0155] τ M : The index number of the Mth arrival time slot.
[0156] The total check-in dataset C is a set of check-in sequences c, i.e., c∈C. Hereinafter, we refer to them as check-in sequences. Given a real-world check-in dataset, the goal of this problem is to train a diffusion model that can generate synthetic check-in datasets whose distribution is similar to that of the original check-in sequences.
[0157] The parameters are defined as follows:
[0158] L: The number of overall points of interest (POI) in the dataset.
[0159] The number of time slots.
[0160] T: total number of diffusion steps.
[0161] t: the step in the diffusion process.
[0162] (s,τ): Check-in sequence element, i.e., the index number of the arrival point of interest and the index number of the arrival time slot.
[0163] s m : The index number of POI (point of interest).
[0164] τ m : The index number of the arrival time slot, indicating a specific time period.
[0165] Time buckets for check-in sequences.
[0166] x0: A check-in vector.
[0167] The spatial frequency vector of the check-in sequence.
[0168] Time bucket vector of check-in sequences.
[0169] Generated spatial positive samples.
[0170] Generated spatial negative samples.
[0171] Generated temporal positive samples.
[0172] Generated temporal negative samples.
[0173] Loss of the spatial diffusion module.
[0174] Loss of the time diffusion module.
[0175] L C (θ s ): spatial contrast loss.
[0176] L C (θ τ ): time contrast loss.
[0177] λ: weight of contrastive loss.
[0178] Space CUnet network.
[0179] Time CUnet network.
[0180] θ s : Parameters of the spatial CUnet network.
[0181] θ τ: Parameters of the temporal CUnet network.
[0182] A: An anchor sample.
[0183] P: A positive sample.
[0184] N: a negative sample.
[0185] Standard Gaussian distribution.
[0186] ξ: asymmetric time interval.
[0187] A method for generating a sign-in sequence based on a spatiotemporal contrast diffusion model comprises the following steps:
[0188] Step 1: Data preprocessing: cleaning and smoothing the real-world data set to ensure data quality and consistency.
[0189] Step 2: Perform space-time lossless coding to convert sign-in sequences of different lengths into sign-in vectors of equal length, and encode them through spatial frequency vectors and time bucket vectors to retain the information of the original sequence.
[0190] Step 3: Establish a diffusion model, construct a spatial diffusion module and a temporal diffusion module, and use the forward diffusion process and the reverse reconstruction process to capture the spatiotemporal characteristics.
[0191] Step 4: Introduce a conditional U-network into the diffusion module and propose a denoising network to capture complex spatiotemporal correlations, modeled through a self-attention mechanism.
[0192] Step 5: Using the contrastive learning strategy, ternary contrastive learning is adopted to further capture the spatiotemporal correlation of the check-in sequence and strengthen the connection between the temporal and spatial diffusion modules.
[0193] The data preprocessing of step 1 includes the following steps: performing data preprocessing on the acquired real-world data set.
[0194] These datasets contain user visit sequences to various POIs, which fully reflect user behaviors in different environments. Multiple preprocessing operations are performed on the original check-in sequences of each dataset.
[0195] First, data cleaning was performed by eliminating outliers and removing check-in points with missing values or connection bandwidth lower than 1k bps (rate). Next, data smoothing techniques were applied by considering consecutive records with a connection duration difference of less than 1 hour as the same POI, as long as they only switched between two points of interest. Finally, adjacent records with the same point of interest and a small time interval were merged to further improve the quality of the dataset.
[0196] Step 2 performs spatiotemporal lossless coding, including the following steps: obtaining a sign-in vector of equal length,
[0197] First, split the original sign-in sequence c of length M into spatial sequences s = <s1,s2,…,s m ,…,s M > and time series τ=<τ1,τ2,…,τ m ,…,τ M >.
[0198] s1: The index number of the first POI (point of interest).
[0199] s m : The index number of the mth POI (point of interest).
[0200] s M : The index number of the Mth POI (point of interest).
[0201] τ1: The index number of the first arrival time slot.
[0202] τ m : The index number of the mth arrival time slot.
[0203] τ M : The index number of the Mth arrival time slot.
[0204] Interest point check-in sequence c: The interest point check-in sequence c is defined as a sequence of check-in points c = <(s1,τ1),(s2,τ2),…,(s m ,τ m ),…,(s M ,τ M )>.
[0205] Where M is the length of the check-in sequence.
[0206] s m ∈S is the index number of a POI (point of interest), and S is a POI set containing L POIs, that is, |S|=L.
[0207] τ m It is the index number of the arrival time slot, indicating a specific time period.
[0208] Then, the spatial sequence s and the time series τ are encoded into spatial frequency vectors of equal length and the time bucket vector
[0209] These vectors can be decoded into the original sequence without losing information.
[0210] This shows that an equivalent conversion can be made between check-in sequences and check-in vectors.
[0211] Specifically, the spatial frequency vector It is obtained by calculating the visit frequency of each POI in the sequence.
[0212] Its i-th element is defined as:
[0213] where i=0,1,…,L
[0214] Where L is the total number of POIs, M is the length of this check-in sequence, and i # represents the i-th POI, and F is an indicator function. If s m Equal to i # , then F=1, otherwise F=0.
[0215] For example, Figure 2 As shown, a sign-in sequence c generated from the Canberra dataset is
[0216] <(s0,τ0),(s1,τ1),…,(s 10 ,τ 10 )>.
[0217] s1: The index number of the first POI (point of interest).
[0218] s m : The index number of the mth POI (point of interest).
[0219] s M : The index number of the Mth POI (point of interest).
[0220] τ1: The index number of the first arrival time slot.
[0221] τ m : The index number of the mth arrival time slot.
[0222] τ M : The index number of the Mth arrival time slot.
[0223] By calculating the visit frequency of each POI in c, we know that he visited the first POI once, the second POI three times, the 314th POI five times, and the 315th POI twice.
[0224] In addition, -1 is used to indicate that the corresponding POI has not been visited.
[0225] Next, we introduce the time bucket vector First, for each POIi #Construct a corresponding time bucket list L i To record the time slots for visiting the POI:
[0226]
[0227] τ j : The index number of the j-th time slot.
[0228] j: the subscript of the interest point sequence.
[0229] s j : The index number of the POI (point of interest) of the jth point of interest.
[0230] i # : represents the i-th POI.
[0231] Represents a list of times built for each POI.
[0232] Represents the time list constructed for the i-th POI.
[0233] For example, Figure 2 As shown in Figure 2, the user visited the second POI three times in the 10th, 12th, and 22nd time slots, so
[0234] Then, a new numerical processing method is adopted, which can list each time bucket Convert to a positive integer and restore the original without losing information
[0235] Time bucket conversion: Specifically, first convert The index of each time slot in is converted to binary form. These binary numbers are padded with zeros on the left to form Number of digits.
[0236] Here, ceil(·) represents the rounding up operation. For base 2 Finally, these binary numbers are concatenated into an aggregate binary number and the aggregate binary number is converted to a positive decimal using the binary conversion method as
[0237] i: represents the index number of the i-th POI.
[0238] τ: The superscript τ indicates the time dimension.
[0239] In addition, use -1 to indicate Indicates that the corresponding POI has not been visited.
[0240] Time bucket recovery: On the other hand, restore the positive integers to the original time bucket list L i The following steps are involved:
[0241] First convert the positive decimal to binary, and pad with zeros on the left if the number of binary digits is less than υ·ceil(log2(t)). Here, υ is the number of digits to reach i # The binary is divided into v segments, each segment is converted from binary format to a decimal time slot index number, and finally, the index numbers of all time slots are combined into the original time bucket list L i .
[0242] In summary, by implementing spatiotemporal lossless coding, the actual input of the spatial diffusion and temporal diffusion modules is obtained, namely, the spatial frequency vector and the time bucket vector
[0243] Refers to the first element in the spatial frequency vector.
[0244] Refers to the Lth element in the spatial frequency vector.
[0245] Refers to the first element in the time bucket vector.
[0246] Refers to the Lth element in the time bucket vector.
[0247] This encoding enables the conversion of check-in sequences c = (s, τ) of different lengths into normalized check-in vectors of the same length The check-in vector is the actual input of the subsequent diffusion module.
[0248] L: The number of overall points of interest (POI) in the dataset.
[0249] The superscript s means: s represents the spatial dimension, so is the spatial frequency vector.
[0250] The superscript τ means: τ represents the time dimension, so is the time bucket vector.
[0251] Step 3 of establishing a spatiotemporal diffusion model includes the following steps:
[0252] The model consists of a spatial diffusion module and a temporal diffusion module. Each module uses an independent denoising network: a spatial U-network and a temporal U-network. The U-network will be explained in step 4. The outputs of these two diffusion modules serve as conditional inputs to each other in their interrelated forward diffusion process and reverse reconstruction process.
[0253] The forward diffusion process of the spatial and temporal diffusion modules is defined as follows:
[0254]
[0255] Among them, the superscripts s and τ represent the spatial component and time component of the original real data respectively. In order to simplify the explanation, both are abstracted as the original real data represented by x0.
[0256] The relevant parameters of the above formula are described as follows:
[0257] It means that under the condition of time τ, the spatial component The probability distribution of .
[0258] It means that under the condition of time τ, the time component The probability distribution of .
[0259] and Represent the initial states of the spatial component and the temporal component respectively.
[0260] and They represent the spatial component and temporal component in the t-th step diffusion process respectively.
[0261] represents the normal distribution and I represents the identity matrix.
[0262] α t It indicates the proportion of the original signal retained during the t-th step diffusion process.
[0263] Represents the number of steps from the first α1 to the tth α t The continuous multiplication of .
[0264] Assume that data x0 follows the true distribution P(x0), x1, x2…x E represents the data after x0 is denoted by different diffusion steps [1, T]. The denoising process can be described as:
[0265]
[0266] in t∈[1,T], α for different t t ∈(0,1) is predefined and increases linearly from the diffusion step number 1 to T, satisfying α1<α2<…<α E .
[0267] After reparameterization, the noise addition process is expressed as:
[0268]
[0269] p(x t |x t-1 ) means that in x t-1 Under the condition x t The true distribution (conditional probability) of .
[0270] Through mathematical derivation, we can find that:
[0271]
[0272] That is, for the noise adding process, only one step of calculation is needed to know x after any noise adding step t t .
[0273] For the reverse reconstruction process, there is no way to directly obtain p(x t-1 |x P ), but since the reconstruction process is a Markov process, we can use the Bayesian formula to know:
[0274]
[0275] p(x t-1 |x t ,x0) means that when x0 is known, t Under the condition x t-1 The true distribution (conditional probability) of t |x0) we can see that,
[0276]
[0277] So p(x t-1 |x t ) can be expressed as:
[0278]
[0279] because x E It is pure Gaussian noise, from x E Gradually add conditional denoising to generate x0. However, in the generation stage, the real ε(x t ,t) is unknown, and cannot be derived or calculated, so a neural network ε is needed to be trained θ (x t ,t,ODT) as a denoiser, where θ represents the parameters of the neural network. The explanation of ODT is as follows: O represents the departure area of the sample, D represents the destination area and T represents the departure time, which is used to fit the real ε(x t ,t), thus completing the gradual denoising of noisy data. That is:
[0280]
[0281] Represents variance.
[0282] The explanation of ODT is as follows, where O represents the departure area of the sample, D represents the destination area, and T represents the departure time.
[0283] Specifically, in the method of the present invention, the sampled noise is converted into a meaningful spatial frequency vector by a spatial and temporal diffusion module. and the time bucket vector The prior distributions of spatial and temporal noise are defined as follows:
[0284]
[0285] in In order to (or ) denoising to (or ), the spatial diffusion module (or temporal diffusion module) is used in the reverse process (or ) to capture the space-time correlation.
[0286] Formally, the reverse process of the spatial and temporal diffusion modules is defined as follows:
[0287]
[0288] The reverse transition probability at each step is defined as follows:
[0289]
[0290] θ p Represents the parameters of the diffusion model in space.
[0291] θ m represents the parameters of the diffusion model over time.
[0292] and Represents the mean term, which represents the mean of the spatial dimension Depends on the current space status and time status and the mean of the time dimension Depends on the current time state and space status
[0293] Spatial and temporal denoising networks are and In the expression and
[0294] In this way, the two diffusion modules jointly generate a check-in vector Two related parts of
[0295] They refer to the spatial part and the temporal part corresponding to the signed vector respectively.
[0296] Further, Figure 1 The overall workflow of the proposed diffusion model is demonstrated.
[0297] Step 4 of introducing a conditional U-type network into the diffusion module includes the following steps: introducing a conditional U-type network into the diffusion module.
[0298] Existing denoising networks for sequence generation lack the ability to capture the complex spatial-temporal correlations in check-in sequences.
[0299] To bridge this gap, a denoising network called conditional U-network is proposed. Figure 3 As shown in the figure, the conditional U-network adopts an architecture similar to the U-network, which consists of 2Z+1 spatiotemporal perception blocks, where Z is a hyperparameter and Z is set to 2. Each spatiotemporal perception block uses a self-attention mechanism to effectively capture spatial and temporal dependencies.
[0300] Furthermore, the fourth step is specifically as follows:
[0301] f) Figure 3 As shown on the left, the i-th block of the conditional U-type network is I (i) is the input, C (i) is the conditional information, and consider the diffusion step t. In the tth step of the reverse process, (or ) denoising (or ) step, I (0) yes (or ), while C (i) yes (or ). The output size of the i-th block and the (2Z+2-i)-th block is the same.
[0302] like Figure 2 As shown, these two blocks are connected together through a residual connection, where is a concatenation operation. A linear layer is used to perform up / down sampling operations to change the output size. The i+1th block C (i+1) The condition information is defined as follows:
[0303]
[0304] g) Next, the diffusion step embedding e is initialized using the following equation t ∈R D
[0305]
[0306] where e t Initialize the embedding vector in the diffusion step of the equation, d = [1, · · ·, D / 2], t ∈ [1, T], D is the input I in the i-th diffusion module (i) Dimension.
[0307] h), then, I (i) , C (i) Through the space-time perception block, to obtain I (i+1) The self-attention operation in the spatiotemporal-aware block involves utilizing the query (Q), key (K), and value (V). Specifically, Q represents the query matrix, K represents the key-value matrix, and V represents the value matrix. Taking the embedding in the i-th block and And apply linear projection Q (i) , K (i) , V (i) Converted into three matrices, where They represent the linear projection calculations for the query matrix, key matrix, and value matrix in the i-th block, respectively. (i) , K (i) , V (i) Denote the query matrix, key matrix, and value matrix corresponding to the i-th block respectively. Using the embedding in the i-th block and And apply linear projection Q (i) , K (i) , V (i) Convert to three matrices:
[0308]
[0309] Among them, the scaled dot product attention and the output H of the i-th block (i) The definition is as follows:
[0310]
[0311] H (i) =Attention(Q (i) ,K (i) ,V (i) 8. Attention means attention mechanism
[0312] i) Then, an up / down sampling network is used to convert the attention output H (i) Convert to I (i+1) :
[0313]
[0314] j) Finally, the output H of the last block (2Z) Feed into a fully connected layer to obtain the predicted noise.
[0315] The use of the contrastive learning strategy in step 5 includes the following steps: the contrastive learning strategy is used to further capture the spatiotemporal correlation of the sign-in sequence.
[0316] In order to further bind the spatial-temporal diffusion modules together and capture the spatiotemporal correlation of the check-in sequence, contrastive learning with triplet loss is adopted. This loss function aims to bring the positive samples closer to the anchor samples. It can be expressed as:
[0317]
[0318] Among them, A is the anchor sample, P is the positive sample, N is the negative sample, d is the cross entropy, m is the interval between positive and negative samples, S is the number of samples, and max represents the function of taking the maximum value in the sequence.
[0319] Among them, A i is the i-th anchor sample, P i is the i-th positive sample, N i is the i-th negative sample. d(A i ,P i ) represents the cross entropy between the i-th anchor sample and the positive sample, d(A i ,N i ) represents the cross entropy between the i-th anchor sample and the negative sample.
[0320] Figure 4 The overall process of contrastive learning using anchor samples, positive samples, and negative samples is explained. Contrastive learning is applied to the spatial and temporal diffusion modules respectively. To maintain brevity and clarity, the contrastive learning process is mainly explained using the spatial diffusion module. The contrastive learning process of the temporal diffusion module also follows the same method. In the method of the present invention, a real sample is As anchor samples, and under the condition The samples generated under the condition As a positive sample. To obtain a negative sample, use the negative condition generate
[0321] Generating samples in the diffusion module is computationally intensive and can cause significant delays in training time. This problem is further exacerbated by generating positive and negative samples for contrastive learning in t steps in each training iteration. To reduce the computational burden, a strategy of estimating positive and negative samples is adopted. Specifically, the spatial diffusion module is used to predict and Its equation is as follows:
[0322]
[0323] in, and They represent the positive and negative samples in the spatial dimension during the diffusion process respectively.
[0324] Similarly, for the time diffusion module, a similar strategy can be used to estimate and Once the positive and negative samples are generated, the contrastive learning loss is calculated using the formula.
[0325] It should be noted that the negative condition and is the key to generating negative samples. and Represent the negative conditions in the time dimension and the negative conditions in the space dimension respectively. Figure 4 The method shown in , generates negative conditions by randomly rearranging the set of spatial frequency and temporal bucket vectors to ensure that they do not match. For example, given Its negative condition is obtained from different randomly selected time bucket vectors. Similarly, for the time diffusion module Its negative condition It is also generated using the same method.
[0326] Example 2: Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5a , Figure 5b , Figure 5c , Figure 5d , Figure 6a , Figure 6b , Figure 6c , Figure 6d , Figure 7a , Figure 7b , Figure 7c and Figure 7dAs shown, a method for generating a check-in sequence of interest points based on a diffusion model is shown. Taking the wireless network usage dataset of the University of Canberra (Canberra) as an example, the goal of the model is to generate a check-in sequence of interest points. The following is the specific implementation process of the technology.
[0327] Step 1: Data preprocessing.
[0328] The detailed preprocessing process is shown in Section 5. The original dataset is cleaned, records containing outliers are removed, and smoothing is performed. Adjacent check-in data with the same AP are merged. Finally, users with higher check-in sequence quality are selected as the training dataset.
[0329] Step 2: Data encoding.
[0330] Split the original sign-in sequence c of length M into spatial sequences s = <s1,s2,…,s m ,…,s M > and time series τ=<τ1,τ2,…,τ m ,…,τ M >. Then, the spatial sequence s and the time series τ are encoded into spatial frequency vectors of equal length and the time bucket vector like Figure 2 shown.
[0331] Step 3: Divide the data set.
[0332] Data normalization and division into training set and test set. A normalization method based on minimum-maximum scaling is used to map the data to a specified interval between -1 and 1 to eliminate the scale differences between features. Then 80% of the samples are divided into training set and 20% into test set.
[0333] Step 4: Use the samples obtained in the second step to train the model.
[0334] First, initialize the parameters θ of the spatial and temporal conditional U-networks s and θ m Then, t is sampled from a uniform distribution and two diffusion modules are trained using Eq.
[0335] Subsequently, negative conditions for contrastive learning are generated, and then positive and negative samples are generated, and the contrastive learning loss is calculated using the formula.
[0336] Finally, the contrastive learning loss is integrated with the diffusion loss, and the parameters θ are updated accordingly. s and θ m . λ is the weight of the contrastive loss.
[0337] The sixth step algorithm is shown in the following table:
[0338] Algorithm 1: Training process
[0339] 1. Initialize spatial CUnet network parameters θ s and time CUnet network parameters θ τ .
[0340] 2. Repeat the following steps until convergence:
[0341] a. From probability distribution The sampling spatial frequency vector and probability distribution The vector of sample time buckets
[0342] b. From the standard Gaussian distribution The noise ∈ is sampled from a uniform distribution
[0343] The sampling diffusion step t in Uniform({1,…,T}).
[0344] c. Calculate spatial diffusion loss and temporal diffusion loss.
[0345] d. Construct time negative conditions and space negative conditions.
[0346] e. Generate temporal positive and negative samples and spatial positive and negative samples.
[0347] f. Calculate spatial contrast loss and temporal contrast loss.
[0348] g. Update the spatial diffusion module loss and the temporal diffusion module loss.
[0349] During testing, the JSD-all, JSD-t, JSD-r, and JSD-u evaluation indicators are calculated to evaluate the trained model and select the optimal model.
[0350] Step 5: Generate a check-in sequence using the optimized model.
[0351] Based on the trained model in the previous step, the sign-in sequence of interest points can be generated.
[0352] Algorithm 2 provides a detailed description of the sampling process.
[0353] First, sample from a Gaussian distribution and Then, these samples are transformed into and At each step, and The condition of is based on the denoised samples from the spatial and temporal diffusion modules in the previous time step t.
[0354] Experiments have found that using asymmetric time intervals in the sampling process of generating sign-in vectors can improve the sampling quality of the model, which can be expressed as Here, ξ represents a small non-negative time interval parameter. Note that the training phase remains the same and no time interval is used.
[0355] Finally, the output check-in vector x0 is post-processed by denormalizing and rounding. Then, the time bucket recovery method introduced in the second step is used to recover the time bucket of the i-th POI, where the arrival time ν of the i-th POI is equal to Then, all time buckets are expanded in the order of the index numbers of the time slots to obtain the final generated check-in sequence.
[0356] The seventh step algorithm is as follows:
[0357] Sampling process:
[0358] 1. Randomly select a sample from the set of all possible spatial frequency vectors and time bucket vectors as the initial sample, expressed as and
[0359] 2. Starting from the last time step, generate samples for each time step step by step until you reach the first time step.
[0360] a. For each time step, if it is the first time step, set the noise z to a zero vector; otherwise, sample a noise vector z from a standard Gaussian distribution.
[0361] b. Use the diffusion model and denoising network to combine the samples and noise vector of the current time step to generate the samples of the previous time step and
[0362] 3. Return the sample of the first time step generated and
[0363] like Figure 1 As shown in the figure, a space-time conditional diffusion model is proposed to generate the check-in sequence, which is divided into five parts from left to right:
[0364] 1) First, a space-time lossless encoding method is used to match the check-in sequences as appropriate inputs of the model, which converts check-in sequences of different lengths into check-in vectors of equal length.
[0365] 2) The entire model consists of a spatial diffusion module and a temporal diffusion module with independent denoising networks. During training, the denoising network is optimized through forward diffusion and reverse reconstruction processes, and during sampling, sequences are generated through the denoising network.
[0366] 3) During the reconstruction process, contrastive learning is used to capture the correlation between the spatial and temporal aspects of the check-in sequence.
[0367] like Figure 2 As shown, a check-in sequence c=<(s0,τ0),(s1,τ1),…,(s 10 ,τ 10 )>. By calculating the visit frequency of each POI in c, we know that he visited the first POI once, the second POI three times, the 314th POI five times, and the 315th POI twice. In addition, -1 is used to indicate that the corresponding POI has not been visited. The user visited the second POI three times in the 10th, 12th, and 22nd time slots, so L2 = [10, 12, 22]. Then, a new numerical processing method is adopted, which can list each time bucket l i Convert to a positive integer and restore the original L without losing information i .
[0368] like Figure 3 As shown in Figure 1, the conditional U-network adopts a similar architecture to the U-network, consisting of 2Z+1 spatiotemporal perception blocks, where Z=2. Each spatiotemporal perception block uses a self-attention mechanism to effectively capture the spatial and temporal dependencies. Figure 3 As shown on the left, the i-th block of the conditional U-type network is I (i) is the input, C (i) is the conditional information, and consider the diffusion step t. In the tth step of the reverse process, (or ) denoising (or ) step, I (0) yes (or ), while C (i) yes (or ). The output size of the i-th block and the (2Z+2-i)-th block is the same.
[0369] like Figure 2 As shown, these two blocks are connected together through a residual connection, where is a concatenation operation. A linear layer is used to perform up / down sampling operations to change the output size. The i+1th block C (i+1) The condition information is defined as follows:
[0370]
[0371] like Figure 4As shown in Figure 1, the overall process of contrastive learning using anchor samples, positive samples, and negative samples. Contrastive learning is applied to the spatial and temporal diffusion modules respectively. To keep it concise and clear, the contrastive learning process is mainly explained using the spatial diffusion module. The contrastive learning process of the temporal diffusion module also follows the same approach. In the method, a real sample is As anchor samples, and under the condition The samples generated under the condition As a positive sample. To obtain a negative sample, use the negative condition generate
[0372] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for generating a check-in sequence of points of interest based on a diffusion model, characterized in that: Contains the following steps: Step 1: Data preprocessing: cleaning and smoothing the real-world data set to ensure data quality and consistency; Step 2: Perform space-time lossless coding to convert sign-in sequences of different lengths into sign-in vectors of equal length, and encode them through spatial frequency vectors and time bucket vectors to retain the information of the original sequence; Step 3: Establish a diffusion model, construct a spatial diffusion module and a temporal diffusion module, and use the forward diffusion process and the reverse reconstruction process to capture the spatiotemporal characteristics; Step 4: Introduce a conditional U-shaped network into the diffusion module and propose a denoising network to capture complex spatiotemporal correlations, modeled through a self-attention mechanism; Step 5: Using the contrastive learning strategy, ternary contrastive learning is adopted to further capture the spatiotemporal correlation of the check-in sequence and strengthen the connection between the temporal and spatial diffusion modules.
2. According to the method for generating a check-in sequence of points of interest based on a diffusion model according to claim 1, it is characterized in that: The data preprocessing in step 1 includes the following steps:
1. Preprocess the acquired real-world datasets, which contain the user's visit sequences to various POIs and fully reflect the user behavior in different environments.
2. Perform multiple preprocessing operations on the original check-in sequence of each dataset. First, data cleaning was performed by eliminating outliers and removing check-in points with missing values or connection bandwidth lower than 1 kbps. Next, data smoothing techniques were applied by considering consecutive records with a connection duration difference of less than 1 hour as the same POI, as long as they only switched between two points of interest. Finally, adjacent records with the same point of interest and a small time interval were merged to further improve the quality of the dataset.
3. The method for generating a check-in sequence of points of interest based on a diffusion model according to claim 1, characterized in that: Step 2 performs spatiotemporal lossless coding, including the following steps: obtaining a sign-in vector of equal length, First, split the original sign-in sequence c of length M into spatial sequences s = <s1,s2,…,s m ,…,s M > and time series τ=<τ1,τ2,…,τ m ,…,τ M >, The relevant parameters are defined as follows: m: is a non-negative integer, s1: the index number of the first POI (point of interest), s m : The index number of the mth POI (point of interest), s M : The index number of the Mth POI (point of interest), τ1: the index number of the first arrival time slot, τ m : The index number of the mth arrival time slot, τ M : The index number of the Mth arrival time slot, Interest point check-in sequence c: The interest point check-in sequence c is defined as a sequence containing check-in points c = <(s1,τ1),(s2,τ2),…,(s m ,τ m ),…,(s M ,τ M )>, c is the sign-in point sequence, Where M is the length of the check-in sequence, and the number of POIs in the entire POI set is L. Then, the spatial sequence s and the time series τ are encoded into spatial frequency vectors of equal length and the time bucket vector in, is the initial state of the spatial component of the diffusion model (spatial frequency vector), is the initial state of the time component of the diffusion model (time bucket vector), These vectors can be decoded into the original sequence without losing information, This shows that an equivalent conversion can be made between check-in sequences and check-in vectors. Specifically, the spatial frequency vector It is obtained by calculating the visit frequency of each POI in the sequence. Its i-th element is defined as: where i=0,1,…,L in refers to the number of visits to the i-th POI in the spatial frequency vector, L is the total number of POIs, M is the length of this check-in sequence, and i # represents the i-th POI, s m is the index number of the mth POI (point of interest), F is an indicator function, if s m Equal to i # , then F=1, otherwise F=0, A check-in sequence c=<(s1,τ1),(s2,τ2),…,(s 10 ,τ 10 )>, s1: the index number of the first POI (point of interest), s 10 : The index number of the 10th POI (point of interest), τ1: the index number of the first arrival time slot, τ 10 : The index number of the 10th arrival time slot, By calculating the visit frequency of each POI in c, we know that the first POI is visited once, the second POI is visited three times, the 314th POI is visited five times, and the 315th POI is visited twice. In addition, -1 is used to indicate that the corresponding POI has not been visited. Next, we introduce the time bucket vector First, build a corresponding time list for each POI To record the time slots for visiting the POI: τ j : The index number of the j-th time slot, j: the subscript of the interest point sequence, s j : The index number of the POI (point of interest) of the jth point of interest, i # : represents the i-th POI, Represents a time list built for each POI, represents the time list constructed for the i-th POI, Then, a numerical processing method is used to list each time bucket Convert to a positive integer and restore the original without losing information in Refers to the i-th element in the time bucket vector, that is, the positive integer after the i-th POI visit time is encoded. Refers to the number of visits to the i-th POI in the spatial frequency vector Time bucket conversion: Specifically, first convert The index of each time slot in is converted to binary form, and these binary numbers are padded with zeros on the left to form Number of digits, Here, ceil(·) represents the rounding up operation. Indicates base 2 pair Finally, concatenate these binary numbers into an aggregate binary number and use the binary conversion method to convert the aggregate binary number into a positive decimal as i: represents the index number of the i-th POI, τ: The superscript τ represents the time dimension, In addition, use -1 to indicate Indicates that the corresponding POI has not been visited. Time bucket recovery: On the other hand, restore the positive integers to the original time bucket list L i The following steps are involved: Convert a positive decimal to binary. If the number of binary digits is less than υ·ceil(log2(T)), pad with zeros on the left. Here, v is the value of i. # The number of times (i # The binary is divided into v segments, each segment is converted from binary format to the index number of the decimal time slot, and finally, the index numbers of all time slots are combined into the original time bucket list. By implementing spatiotemporal lossless coding, the actual input of the spatial diffusion and temporal diffusion modules is obtained, namely the spatial frequency vector and the time bucket vector Refers to the first element in the spatial frequency vector, refers to the Lth element in the spatial frequency vector, Refers to the first element in the time bucket vector, refers to the Lth element in the time bucket vector, This encoding enables the conversion of check-in sequences c = (s, τ) of different lengths into normalized check-in vectors of the same length The check-in vector is the actual input of the subsequent diffusion module. L: the number of overall points of interest (POI) in the dataset, The superscript s means: s represents the spatial dimension, so is the spatial frequency vector, The superscript τ means: τ represents the time dimension, so is the time bucket vector.
4. The method for generating a check-in sequence of points of interest based on a diffusion model according to claim 1, characterized in that: Step 3 of establishing a spatiotemporal diffusion model includes the following steps: The model consists of a spatial diffusion module and a temporal diffusion module, each of which uses an independent denoising network: a spatial U-network and a temporal U-network. The U-network will be explained in step 4. The outputs of these two diffusion modules serve as conditional inputs to each other in their interrelated forward diffusion process and reverse reconstruction process. The forward diffusion process of the spatial and temporal diffusion modules is defined as follows: The superscripts s and τ represent the spatial component and temporal component of the original real data, respectively. To simplify the description, both are abstracted as the original real data represented by x0. The relevant parameters of the above formula are described as follows: It means that under the condition of time τ, the spatial component The probability distribution of It means that under the condition of time τ, the time component The probability distribution of and denote the initial states of the spatial component and the temporal component respectively, and represent the spatial component and the temporal component of the diffusion process in the t-th step, respectively. represents the normal distribution, I represents the identity matrix, α t It indicates the proportion of the original signal retained during the t-th step diffusion process, Represents the number of steps from the first α1 to the tth α t The multiplication of Assume that data x0 follows the true distribution P(x0), x1, x2…x E represents the data after x0 is denoted by different diffusion steps [1, T], where T is the total number of diffusion steps. The noise adding process can be described as: in t∈[1,T], α for different t t ∈(0,1) is predefined and increases linearly from the diffusion step number 1 to T, satisfying α1<α2<…<α T , After reparameterization, the noise addition process is expressed as distribution: p(x t |x t-1 ) means that in x t-1 Under the condition x t The true distribution (conditional probability), After mathematical derivation, That is, for the noise adding process, only one step is needed to know x after any t-th step of noise adding. t , For the reverse reconstruction process, there is no way to directly obtain p(x t-1 |x t ), but since the reconstruction process is a Markov process, we can use the Bayesian formula to know: p(x t-1 |x t ,x0) means that when x0 is known, t Under the condition x t-1 The true distribution (conditional probability) of t |x0) we can see that, μ(x t ,t) represents the estimation function of x0 in the tth step of the diffusion process, At the same time p(x t-1 |x t ) can be expressed as: because x T It is pure Gaussian noise, from x T Gradually add conditional denoising to generate x0, ε(x t ,t) represents the real Gaussian noise of step t, but the real ε(x t ,t) is unknown, and cannot be derived or calculated, so a neural network ε is needed to be trained θ (x t ,t,ODT) as a denoiser, where θ represents the parameters of the neural network. The explanation of ODT is as follows: O represents the departure area of the sample, D represents the destination area and T represents the departure time, which is used to fit the real ε(x t ,t), that is, the real Gaussian noise of the tth step, thereby completing the gradual denoising of the noisy data, that is: represents the variance, The explanation of ODT is as follows, O represents the departure area of the sample, D represents the destination area and T represents the departure time. The sampled noise is converted into a meaningful spatial frequency vector through spatial and temporal diffusion modules and the time bucket vector The prior distributions of spatial and temporal noise are defined as follows: in In order to (or ) denoising to (or ), the spatial diffusion module (or temporal diffusion module) is used in the reverse process (or ) to capture the spatial-temporal correlation, Formally, the reverse process of the spatial and temporal diffusion modules is defined as follows: The reverse transition probability at each step is defined as follows: The relevant parameters are defined as follows: θ s represents the parameters of the diffusion model in space, θ τ represents the parameters of the diffusion model over time, and Represents the mean term, which represents the mean of the spatial dimension Depends on the current space status and time status and the mean of the time dimension Depends on the current time state and space status Spatial and temporal denoising networks are and In the expression and In this way, the two diffusion modules jointly generate a check-in vector Two related parts of Sign in vector Corresponding spatial and temporal parts.
5. The method for generating a check-in sequence of points of interest based on a diffusion model according to claim 1, characterized in that: Step 4 of introducing a conditional U-type network into the diffusion module comprises the following steps: introducing a conditional U-type network into the diffusion module, The conditional U-shaped network adopts a similar architecture to the U-shaped network, which consists of 2Z+1 spatiotemporal perception blocks, where Z is a hyperparameter and Z=2. Each spatiotemporal perception block uses a self-attention mechanism to effectively capture the spatial and temporal dependencies. Furthermore, a) In the i-th block of the conditional U-type network, I (i) is the input of the ith block, C (i) is the conditional information of the ith block, and considering the diffusion step t, in the tth step of the reverse process, (or ) denoising (or ) step, I (0) yes (or ), while C (i) yes (or ), the output size of the i-th block and the (2Z+2-i)-th block is the same, and the two blocks are connected together through a residual, where is a connection operation, using a linear layer to perform up / down sampling operations to change the output size, the i+1th block C (i+1) The condition information is defined as follows: C (i+1) Indicates the condition information of the i+1th block, b) Next, the diffusion step embedding e is initialized using the following equation t ∈R D where e t Initialize the embedding vector in the diffusion step of the equation, d = [1, · · ·, D / 2], t ∈ [1, T], D is the input I in the i-th diffusion module (i) The dimension of c) Then, I (i) , C (i) Through the space-time perception block, to obtain I (i+1) , the self-attention operation in the spatiotemporal perception block involves utilizing query (Q), key (K) and value (V). Specifically, Q represents the query matrix, K represents the key-value matrix, and V represents the value matrix. The embedding in the i-th block is adopted. and And apply linear projection Q (i) , K (i) , V (i) Converted into three matrices, where They represent the linear projection calculations for the query matrix, key matrix, and value matrix in the i-th block, respectively. (i) , K (i) , V (i) Respectively represent the corresponding query matrix, key value matrix and value matrix in the i-th block: Among them, the scaled dot product attention and the output H of the i-th block (i) The definition is as follows: H (i) =Attention(Q (i) ,K (i) ,V (i) ), Attention represents the attention mechanism, d), then, an up / down sampling network is used to output the attention H (i) Convert to I (i+1) : e) Finally, the output H of the last block (2Z) Feed into a fully connected layer to obtain the predicted noise.
6. The method for generating a check-in sequence of points of interest based on a diffusion model according to claim 1, characterized in that: The contrastive learning strategy used in step 5 includes the following steps: The contrastive learning strategy is used to further capture the spatiotemporal correlation of the sign-in sequence, and the contrastive learning of the ternary loss is used. This loss function aims to bring the positive sample closer to the anchor sample, which is expressed as: Among them, A is the anchor sample, P is the positive sample, N is the negative sample, d is the cross entropy, m is the interval between positive and negative samples, S is the number of samples, and max represents the function of taking the maximum value in the sequence. Among them, A i is the i-th anchor sample, P i is the i-th positive sample, N i is the i-th negative sample, d(A i ,P i ) represents the cross entropy between the i-th anchor sample and the positive sample, d(A i ,N i ) represents the cross entropy between the i-th anchor sample and the negative sample, The overall process of contrastive learning using anchor samples, positive samples and negative samples, respectively, applies contrastive learning to the spatial and temporal diffusion modules. In order to keep it concise and clear, the spatial diffusion module is mainly used to explain the contrastive learning process. The contrastive learning process of the temporal diffusion module also follows the same method. As anchor samples, and under the condition The samples generated under the condition As a positive sample, in order to obtain a negative sample, use the negative condition generate A strategy for estimating positive and negative samples, specifically, using the spatial diffusion module to predict and Its equation is as follows: in, and Respectively represent the positive samples and negative samples in the spatial dimension during the diffusion process, Similarly, for the time diffusion module, a similar strategy can be used to estimate and Once the positive and negative samples are generated, the contrastive learning loss is calculated using the formula, Negative Condition and is the key to generating negative samples, where and denote the negative condition in the time dimension and the negative condition in the spatial dimension, respectively. The negative condition is generated by randomly rearranging the set of spatial frequency and time bucket vectors to ensure that they do not match. Given the negative condition from the spatial diffusion module Its negative condition is obtained from different randomly selected time bucket vectors. Similarly, for the time diffusion module Its negative condition It is also generated using the same method.
Citation Information
Patent Citations
A method for generating vehicle trajectories
CN110223515B
Method and system for automatically generating people movement track and foothold
CN111985452A
Next POI recommendation method based on user preference and spatio-temporal context information
CN117194763A
Virtual fitting image restoration method based on diffusion condition generation algorithm
CN116703747A
Next interest point recommendation method based on double-contrast learning
CN118364176A