A point of interest recommendation method and system based on multi-stage denoising
Patent Information
- Application Number
- CN202610942126.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-29
AI Technical Summary
然而,真实场景下的轨迹数据普遍存在偶发性签到、误操作、记录缺失等多种类型噪声,这些噪声破坏了签到序列的语义连贯性,加剧了数据稀疏性,并导致模型学习到不可靠的POI转移关系
本发明在个体轨迹层面,利用大语言模型的语义理解与推理能力,对原始签到轨迹进行语义噪声过滤,识别并剔除与用户行为逻辑不一致的偶发性签到、误操作等异常记录;同时,通过融合相似用户的协同行为模式对缺失信息进行语义补全,有效缓解了数据稀疏性问题,恢复了轨迹的语义连贯性与完整性;其次,在全局拓扑层面,基于增强后的轨迹构建兴趣点转移图,通过提取节点特征后评估边的可信度,并采用自适应剪枝策略剔除低可信的噪声连接边,阻断错误转移关系的传播路径;再将剪枝后的高可信拓扑结构进行表示学习,生成鲁棒的兴趣点拓扑嵌入。最终,实验结果表明,本发明在多个真实数据集上相较于现有方法取得了3%–6%的性能提升,能够为用户提供更精准的下一兴趣点推荐,显著提升出行规划效率与平台用户体验。
Smart Images

Figure CN122489852B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of point of interest recommendation technology, and in particular to a point of interest recommendation method and system based on multi-level denoising. Background Technology
[0002] The statements in this section are merely to provide background information related to the present invention and do not necessarily constitute prior art.
[0003] With the massive accumulation of user check-in trajectory data in location-based social networks, Point of Interest (POI) recommendation has become a key technology for improving personalized travel services and intelligent urban decision-making, such as geolocation recommendations for shops, restaurants, and attractions. However, trajectory data in real-world scenarios generally suffers from various types of noise, including sporadic check-ins, accidental operations, and missing records. This noise disrupts the semantic coherence of check-in sequences, exacerbates data sparsity, and leads to models learning unreliable POI transition relationships.
[0004] Existing denoising methods in recommender systems, whether based on graph structure and attention mechanisms such as session filtering, hierarchical sequence denoising, or multimodal joint denoising, are primarily designed for interaction noise. They struggle to adapt to the complex characteristics of POI recommendations, where noise is deeply coupled with spatiotemporal context, mobility continuity, and semantic preferences. Therefore, current technologies cannot systematically denoise at both the individual trajectory and global topology levels, addressing the multi-source, semantically, and structurally coupled nature of noise in POI recommendations. This results in semantic and structural noise in the trajectory continuously interfering with user preference modeling, making it difficult to accurately predict which shops, restaurants, or attractions a user might visit next. Consequently, users frequently receive irrelevant recommendations during travel planning, degrading the user experience. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method and system for recommending points of interest based on multi-level denoising. It uses a large language model to perform semantic denoising and completion at the individual trajectory level, and combines global topology-level credibility pruning to systematically suppress complex trajectory noise and improve the accuracy of point of interest recommendations.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solution: In a first aspect, the present invention provides a method for recommending points of interest based on multi-level denoising, comprising: Obtain the user's original check-in history and the global user set; The original check-in trajectory is subjected to semantic noise filtering to obtain a denoised trajectory, thus completing the initial denoising at the individual trajectory level. Similar user trajectories are extracted from the global user set and then semantically completed to obtain enhanced trajectories. A global interest point transfer graph is constructed based on the enhanced trajectory. After extracting node features, the edges in the graph are evaluated for credibility and adaptively pruned. The pruned high-credibility topology is then subjected to graph representation learning to obtain the interest point topology embedding, thus completing the second denoising at the global topology level. The enhanced trajectory is fused with the interest point topology embedding to obtain recommended interest points.
[0007] Furthermore, the initial denoising specifically includes: The original check-in trajectory is constructed into a semantic trajectory; The semantic trajectory and the preset denoised prompt template are input into the large language model to identify and output the noisy check-in set; The noisy check-in set is removed from the original check-in trajectory to obtain the denoised trajectory.
[0008] Furthermore, the semantic completion specifically includes: A pre-trained model is used to extract user embeddings, and the top preset number of users with the highest similarity are retrieved to obtain similar user trajectories; The denoised trajectory and similar user trajectories are input into a large language model. Reasoning is performed through chain-like thinking prompts. The generated candidate completion check-in points are then incorporated into the denoised trajectory to obtain the enhanced trajectory.
[0009] Furthermore, the similarity is measured by calculating the cosine similarity between the target user and the embedded representations of each user in the global user set.
[0010] Furthermore, a global interest point transfer graph is constructed based on the enhanced trajectories of all users. In the graph, nodes represent interest points, and directed edges represent the access transfer relationships between interest points. The global interest point transfer graph is propagated, the graph convolutional augmented representation of each interest point is extracted, and the cosine similarity between the graph convolutional augmented representations of the two endpoints of each edge is used as the structural credibility score of that edge. For each source interest point, all outgoing edges are sorted according to their structural credibility score, and only the preset number of outgoing edges with the highest scores are retained to obtain a pruned high-credibility topology.
[0011] Furthermore, the graph representation learning employs a gated graph neural network. During each update, neighborhood messages are first aggregated, and then the gated state is fused through the update gate and the reset gate. After multiple iterations, interest point topology embeddings are generated, completing secondary denoising at the global topology level.
[0012] Furthermore, the aggregation of neighborhood messages specifically involves: during each iteration update, for each target interest point, multiplying the previous hidden state of all outgoing neighbor nodes by the edge weight and summing the results to obtain the neighborhood message aggregation vector. The gated state fusion is specifically as follows: the hidden state of the current node at the previous moment and the neighborhood message aggregation vector are input into the gated loop unit. The new hidden state of the node is generated by weighted fusion by updating the gate to control the retention ratio of historical states and resetting the gate to control the introduction of new messages.
[0013] Furthermore, the interest point topology embedding, user embedding, time embedding, and category embedding of each check-in in the enhanced trajectory are spliced and fused to construct a multimodal check-in representation sequence; The multimodal check-in representation sequence is input into the Transformer encoder for temporal modeling, and the next point of interest of the user is predicted by the multilayer perceptron to obtain the recommended point of interest.
[0014] In a second aspect, the present invention provides an interest point recommendation system based on multi-level denoising, comprising: Data acquisition module: used to acquire users' original check-in history and the global user set; The first-level denoising module is used to perform semantic noise filtering on the original check-in trajectory to obtain a denoised trajectory, thus completing the initial denoising at the individual trajectory level. The trajectory enhancement module is used to extract similar user trajectories from the global user set and obtain enhanced trajectories through semantic completion. The second-level denoising module is used to construct a global interest point transfer graph based on the enhanced trajectory, extract node features, and then perform credibility evaluation and adaptive pruning on the edges in the graph; the pruned high-confidence topology structure is then subjected to graph representation learning to obtain the interest point topology embedding, thus completing the second-level denoising at the global topology level. Recommendation module: used to fuse the enhanced trajectory with the topological embedding of interest points to obtain recommended interest points.
[0015] In a third aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method as described in the first aspect.
[0016] Compared with the prior art, the beneficial effects of the present invention are: This invention, at the individual trajectory level, leverages the semantic understanding and reasoning capabilities of a large language model to filter semantic noise from the original check-in trajectory, identifying and removing abnormal records such as occasional check-ins and erroneous operations that are inconsistent with user behavior logic. Simultaneously, it semantically completes missing information by fusing collaborative behavior patterns of similar users, effectively alleviating data sparsity and restoring the semantic coherence and integrity of the trajectory. Secondly, at the global topology level, it constructs an interest point transition graph based on the enhanced trajectory, evaluates the credibility of edges after extracting node features, and employs an adaptive pruning strategy to remove low-credibility noisy connections, blocking the propagation path of erroneous transition relationships. Finally, it performs representation learning on the pruned high-credibility topology structure to generate robust interest point topology embeddings. Experimental results show that this invention achieves a 3%–6% performance improvement over existing methods on multiple real-world datasets, providing users with more accurate next interest point recommendations and significantly improving travel planning efficiency and platform user experience. Attached Figure Description
[0017] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0018] Figure 1 This is a flowchart illustrating the method architecture of Embodiment 1 of the present invention; Figure 2 This is a comparison of the impact of hyperparameter k on the New York City dataset. Figure 3 This is a comparison chart showing the impact of hyperparameter k on the Tokyo dataset. Figure 4 This is a comparison chart showing the impact of hyperparameter k on the London City dataset. Detailed Implementation
[0019] The following detailed description is exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0020] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0021] In this embodiment of the invention, the data collection and processing strictly adhere to the requirements of relevant laws and regulations, obtaining informed consent or separate consent from the data subject, and conducting subsequent data use and processing within the scope of laws, regulations, and the data subject's authorization. All data acquisition in this embodiment is based on compliance with laws and regulations and user consent, representing the lawful application of the data.
[0022] Terminology Explanation: Point of Interest (POI): such as the geographical location of shops, restaurants, attractions, etc.
[0023] Large Language Model (LLM): Possesses powerful semantic understanding and reasoning capabilities.
[0024] Graph Convolutional Network (GCN): Used to capture high-order semantic neighborhood relationships in graph structures.
[0025] Gated Graph Neural Network (GGNN): Used to capture high-order dependencies in graph structures.
[0026] The terms mentioned above are commonly used technical terms in this field, but they have specific meanings in the context of this invention. Those skilled in the art should understand these terms in conjunction with the overall technical solution of this invention. The following embodiments will provide a detailed description of the technical solution of this invention to facilitate its implementation by those skilled in the art.
[0027] Example 1 Compared to existing technologies that primarily address interactive noise and struggle to adapt to the deep coupling of noise with spatiotemporal context and semantic preferences in POI recommendations, this invention solves the problem by constructing a two-level denoising mechanism involving individual trajectories and global topology. In real-world scenarios, user check-in trajectories for points of interest such as shops, restaurants, and attractions often contain noise from sporadic check-ins or missing records. This invention first utilizes the semantic understanding and reasoning capabilities of a large language model at the individual trajectory level to identify and filter out noisy check-ins inconsistent with user behavior logic, such as users continuously visiting unrelated points of interest within a short period. Then, it integrates collaborative behavior patterns of similar users to semantically complete missing check-in information, such as inferring that a user might visit a nearby cafe after visiting a museum, thereby restoring the coherence and completeness of the trajectory. Building upon this, at the global topology level, this invention constructs a POI transition graph based on the enhanced trajectories. Through credibility assessment and adaptive pruning, it removes low-credibility edges introduced by accidental co-occurrence or popularity bias, such as eliminating false transition relationships formed by a large number of random check-ins at popular attractions. Finally, it learns robust POI embeddings on a high-credibility structure. Therefore, this invention systematically suppresses the interference of noise on user preference modeling from both semantic and structural dimensions, and effectively denoises complex trajectory noise of points of interest such as shops, restaurants, and attractions, thereby improving the accuracy of next point of interest recommendation and user experience.
[0028] Specifically, this invention proposes an LLM-based Multi-level Denoising Framework for POI Recommendations (LMD-POI) for POI recommendation tasks. This framework integrates the contextual understanding capabilities of the LLM with the adaptive optimization mechanism of graph structures. By implementing progressive denoising at both the individual trajectory level and the global topology level, it systematically improves the quality of trajectory data and enhances the robustness of representation learning. Specifically, at the individual trajectory level, this invention introduces a trajectory purification and enrichment module based on the LLM: this module utilizes the powerful contextual understanding capabilities of the LLM to identify and filter noisy check-in behaviors in the original user check-in trajectory. Simultaneously, by fusing trajectory information from similar users, it completes and enriches the target user's trajectory, thereby constructing a trajectory sequence that more accurately reflects user preferences. At the global topology level, this invention proposes a credibility-aware graph structure optimization module. This module evaluates the reliability of the global topological relationships between Points of Interest (POIs) and adaptively prunes untrusted connections to obtain a highly credible trajectory topology. The optimized topology further guides a gated graph neural network to capture higher-order dependencies, ultimately generating POI representations that are more robust to noise. To verify the effectiveness of the proposed framework, extensive comparative and ablation experiments were conducted on three real-world datasets. Experimental results show that LMD-POI consistently outperforms state-of-the-art baseline methods on multiple evaluation metrics, achieving performance improvements of 3%–6%, thus validating the effectiveness of each component in the framework.
[0029] In a typical embodiment of the present invention, such as Figure 1 As shown, a multi-level denoising-based interest point recommendation method is disclosed, the specific steps of which include: S1: Obtain the user's original check-in history and the global user set; S2: Perform semantic noise filtering on the original check-in trajectory to obtain a denoised trajectory, thus completing the initial denoising at the individual trajectory level; S3: Extract similar user trajectories from the global user set, and obtain enhanced trajectories through semantic completion; S4: Construct a global interest point transfer graph based on the enhanced trajectory, extract node features, and then perform credibility evaluation and adaptive pruning on the edges in the graph; perform graph representation learning on the pruned high-credibility topology to obtain the interest point topology embedding, and complete the second denoising at the global topology level. S5: The enhanced trajectory is fused with the interest point topology embedding to obtain recommended interest points.
[0030] The above-mentioned interest point recommendation method based on multi-level denoising will be described in detail below with reference to specific embodiments.
[0031] In step S1, obtaining the user's original check-in trajectory and the global user set includes: collecting timestamped check-in records of all users from a location-based social network platform; sorting each user's check-in records in chronological order to form the user's original check-in trajectory sequence; segmenting the original check-in trajectory using a preset time window (24 hours) and filtering out trajectories containing only a single check-in; and organizing all users' check-in records and trajectory sequences into a global user set, denoted as […]. Each user The original check-in trajectory is recorded as Each time you sign in It should include at least user identifier, point of interest identifier, and timestamp information.
[0032] In this embodiment, user check-in data from Foursquare was collected, spanning from April 2012 to February 2013, covering three cities: New York City (NYC), Tokyo (TKY), and London (Great Britain (GB)). Experiments were conducted on these three datasets. Each dataset contains not only user-POI interaction information but also rich contextual information, such as timestamps and geographic coordinates (longitude and latitude). This invention follows the standard preprocessing and data partitioning strategies commonly used in previous POI recommendation studies. Specifically, user check-in records are first sorted chronologically to form trajectory sequences, and the trajectories are segmented using a 24-hour time window. Trajectories containing only a single check-in are discarded. Subsequently, to maintain temporal consistency, each dataset is divided into a training set, a validation set, and a test set in chronological order, with a ratio of 8:1:1. The statistical information of the datasets is summarized in Table 1.
[0033] Table 1. Dataset Statistics
[0034] For ease of understanding, this embodiment first formalizes the symbolic representation used in this document. Let the set of interest points and the set of users be represented as follows: and ,in, and These represent the total number of POIs and users, respectively. For each user... Its check-in trajectory is represented as ,in Indicates user The total number of check-ins. Each check-in is defined as a triple. , indicating user In time Visited points of interest Ultimately, the set of all users' check-in records can be represented as: .
[0035] In POI recommendation tasks, users' historical check-in trajectories typically face two key challenges. First, the trajectories may contain noisy check-ins that are inconsistent with the user's true preferences or reasonable travel patterns, such as records caused by accidental operations or occasional visits. Second, due to missing or incomplete check-ins, the trajectories are often sparse and accompanied by significant information loss. Traditional sequence denoising methods mainly rely on frequency statistics or local interaction patterns to filter noise, thus having limitations in identifying semantic inconsistencies and failing to reasonably infer or complete missing user preference information. To address these issues, this invention employs a trajectory cleansing and enrichment module based on a large language model. This module leverages the powerful semantic understanding, logical reasoning, and context modeling capabilities of LLM to achieve semantic cleansing and information completion at the single trajectory level, thereby constructing a cleaner and more complete user behavior trajectory and providing high-quality input for subsequent graph-based representation learning.
[0036] Specifically, the trajectory cleanup and enrichment module comprises two stages: (1) LLM-driven data cleanup and (2) LLM-driven trajectory enrichment. In the first stage, LLM performs semantic-level noise detection and filtering on the user's original trajectory to generate a denoised trajectory set. In the second stage, based on collaborative semantic similarity retrieval and the generative reasoning capabilities of LLM, the denoised trajectory can be enhanced and completed in a context-aware manner.
[0037] In step S2, the initial noise reduction specifically includes: First, the original check-in trajectory is constructed into a semantic trajectory; Secondly, the semantic trajectory and the preset denoised prompt template are input into the large language model to identify and output the noisy check-in set; Finally, the noisy check-in set is removed from the original check-in trajectory to obtain the denoised trajectory.
[0038] The semantic trajectory is a text sequence obtained by converting the interest point identifiers, categories, and access times of each check-in in the original check-in trajectory into natural language descriptions in sequence; the preset denoising prompt template is used to guide the large language model from behavioral meaning. Figure 1 A structured instruction that performs step-by-step reasoning and noise discrimination on sign-in records from the perspectives of consistency, category rationality, and time rationality.
[0039] In this embodiment, to facilitate understanding of the denoising and enhancement operations performed by the large language model on the user trajectory, the present invention further formalizes the trajectory representation before and after data processing. Given a target user... Original check-in trajectory The trajectory obtained after LLM denoising is represented as follows: , in The set of denoised trajectories for all users is denoted as . Subsequently, LLM further performs trajectory enhancement and information completion operations on the denoised trajectory to generate an enhanced trajectory. ,in Ultimately, the enhanced trajectory set for all users is represented as: Enhanced trajectory for a given target user The goal of the POI recommendation task is to predict the target user. The next most likely point of interest to be visited in the future.
[0040] As one implementation method, the specific process of LLM-driven data cleansing is as follows: While existing sequence denoising methods can remove some anomalous noise based on interaction frequency or manually designed denoising strategies, their ability to identify semantically inconsistent noise in POI trajectories is limited, such as random check-ins that deviate from the underlying travel theme. This type of noise may not show obvious anomalies in frequency-based features, but it can disrupt the semantic coherence of the trajectory, thus interfering with user preference modeling. To address this issue, this invention introduces a large language model as a semantic inference engine, utilizing the functional semantics of POIs, temporal context, and travel logic to achieve fine-grained identification and filtering of semantic noise. The proposed process includes three stages: trajectory semantic description construction, generative noise discrimination, and denoised trajectory construction.
[0041] The specific steps include: First, the original check-in trajectory is constructed into a semantic trajectory description. The semantic trajectory description is generated by translating the interest point identifiers, interest point categories, and access times of each check-in into natural language text according to the original access order, and is used as input for noise discrimination and recognition in a large language model.
[0042] Large Language Models (LLMs) rely on natural language for reasoning and cannot directly understand discrete POI identifiers or raw timestamps. To enable LLMs to perform semantic reasoning on user trajectories, this invention converts structured check-in trajectories into natural language descriptive text containing rich contextual information. This process preserves the original access order while explicitly encoding semantic attributes such as POI identifiers, POI categories, and access times, thus forming an LLM-interpretable trajectory description, rather than a learned embedding representation.
[0043] Specifically, given the target user Original check-in trajectory Define a noise reduction tooltip constructor function. It is used to extract and organize check-in information and generate semantic trajectory descriptions. :
[0044] in, It is a semantic description of the trajectory after textualization, which is used as input for subsequent noise inference and recognition processes based on LLM.
[0045] Secondly, the semantic trajectory description and the preset denoised prompt template are input into the large language model to guide the large language model from behavioral intention. Figure 1 Each check-in is evaluated from the perspectives of consistency, category rationality, and time rationality, identifying and outputting a set of semantically anomalous noisy check-ins. This is generative noise discrimination, which introduces a Large Language Model (LLM) as a semantic discriminator to identify and extract semantically anomalous check-in records from the user's trajectory. Unlike traditional supervised classification methods, this invention's model utilizes the generative reasoning and common-sense understanding capabilities of LLM to directly infer which check-in behaviors violate the inherent logic of the trajectory, the functional semantics of POIs, or real-world travel patterns.
[0046] Specifically, the semantic trajectory description obtained in the previous step This, along with a carefully designed denoising cue template (as shown in Table 2), will be input into the LLM to guide its step-by-step reasoning. The structured description of the denoising cue template is shown in Table 2 below.
[0047] Table 2. Structured Description of Noise Reduction Cue Template
[0048] This prompt requires the LLM to evaluate each check-in from the perspectives of consistency of behavioral intent, category rationality, and temporal rationality, and to identify access records that conflict with the overall travel logic. In this way, semantic noise in the user's trajectory can be accurately located and extracted. Specifically, consistency of behavioral intent involves the LLM analyzing the functional semantic relationships between consecutive check-ins in the trajectory to determine whether the current check-in conforms to the travel purpose chain reflected in the user's historical behavior. For example, a user consecutively checking into restaurants after a gym session aligns with the behavioral intent of "eating after exercise." Category rationality involves the LLM evaluating whether the category of the current check-in matches the category distribution in the user's historical preferences and the category combination pattern of adjacent check-ins in the trajectory context, based on the functional category labels of points of interest. For example, if a user frequently visits cafes and bookstores, a sudden check-in at a car repair shop would have low category rationality. The time reasonableness is specifically defined as follows: The large language model combines the specific timestamp of the check-in with the functional time period characteristics of the point of interest to determine whether the access time conforms to the regular business hours of the point of interest or the user's personal time period behavior pattern. For example, checking in at a museum at 3 a.m. or frequently checking in at a bar on a weekday morning are both considered noisy check-ins with questionable time reasonableness.
[0049] Specifically, let Represented as user The noise check-in set identified can then be denoised using an LLM-based process as follows:
[0050] in, This represents the parameters of the LLM. If the LLM infers that there are no obvious semantic anomalies in the user trajectory, then... If empty, the trajectory remains unchanged.
[0051] Finally, the noisy check-in set is removed from the original check-in trajectory, and the remaining check-in records are retained in their original order to obtain the denoised trajectory.
[0052] Specifically, for each user LLM will start from the original trajectory A set of sign-in records with semantic anomalies was identified. Subsequently, these noisy check-ins are directly removed from the original trajectory using set difference operations, thus constructing the denoised trajectory:
[0053] The final denoised trajectory Only semantically consistent and behaviorally reasonable check-in records are retained, while those that contradict the user's potential preferences or overall travel logic are removed. The backslash here represents a set difference operation; the formula means that from the original check-in trajectory set, all elements in the noisy check-in set are removed (subtracted) to obtain the denoised trajectory set. This process effectively improves the semantic coherence and reliability of user trajectories, providing a clean and reliable starting point for subsequent accurate modeling of users' true preferences.
[0054] After data cleansing at the individual trajectory level, the semantic consistency and logical coherence of user trajectories were significantly improved. However, relying solely on denoising inevitably reduces the amount of usable interaction data, further exacerbating the data sparsity problem, especially for long-tail users with limited historical records or cold-start users. For these users, the remaining valid check-ins after denoising may not be sufficient to support the effective learning of long-term preferences.
[0055] Furthermore, real-world user behavior is often only partially observable by the platform. Users may engage in many preference-related activities without any recorded check-in history (e.g., going to the gym or coffee shop but not checking in). These unobserved but semantically significant behaviors constitute implicit positive feedback signals, and their absence leads to incomplete and biased user preference modeling.
[0056] To alleviate the problems of data sparsity and the lack of implicit behavior, this invention introduces a collaborative semantic enhancement mechanism. Its core lies in the fact that users with similar historical preferences often exhibit consistent behavioral patterns and shared travel motivations. Traditional collaborative filtering methods typically identify similar users based on the co-occurrence statistics of user-POI interactions, but this approach struggles to capture the deep semantic intent behind user behavior. In contrast, the module proposed in this invention utilizes the powerful semantic understanding and reasoning capabilities of Large Language Models (LLMs) to identify semantically similar users from a behavioral semantic perspective. By modeling at the semantic level, LLMs can identify users with similar functional preferences and activity patterns even when explicit interactions are sparse. Subsequently, reliable behaviors from these similar users are integrated into the denoised trajectory of the target user, achieving semantic completion and trajectory enhancement. This collaborative enhancement process not only enriches user trajectories with reasonable and interpretable behavioral signals but also provides a more robust and informative foundation for subsequent graph-based representation learning and POI recommendation. Specifically, this module mainly includes two stages: (1) user representation extraction and similar user retrieval, and (2) LLM-based user semantic reasoning and trajectory completion.
[0057] In step S3, similar user trajectories are extracted from the global user set, and enhanced trajectories are obtained through semantic completion. Specifically, the semantic completion process includes: A pre-trained model is used to extract user embedding representations, and the top preset number of users with the highest similarity are retrieved to obtain similar user trajectories. The denoised trajectory and similar user trajectories are input into a large language model, and semantic reasoning is performed through chain-like thinking prompts. The generated candidate completion check-in points are incorporated into the denoised trajectory to obtain the enhanced trajectory.
[0058] As one implementation method, LLM-driven trajectory enrichment involves the following process: (1) User representation extraction and similar user retrieval. In order to identify users with similar behavioral patterns at the semantic level, this invention uses a pre-trained spatio-temporal clustering model (IPCM) as a user representation extractor. IPCM jointly models personalized spatio-temporal clustering features and historical check-in trajectories, enabling it to capture users' long-term mobile preferences and periodic behavioral patterns, rather than relying solely on simple co-occurrence statistics.
[0059] First, the target user's original check-in trajectory is input into a pre-trained interest point recommendation model, encoded into a latent user embedding representation that can characterize the user's global movement preferences. Then, the cosine similarity between the target user and other users in the global user set is calculated, and the users with the highest similarity are retrieved. For each user, their denoised trajectory is used as the trajectory of similar users.
[0060] Specifically, given a user Original check-in trajectory Pre-trained model Encode the trajectory into a potential user embedding representation:
[0061] in, It represents the potential spatiotemporal behavior of users and can characterize users' global mobility preferences in a high-dimensional space.
[0062] Target users With another user Behavioral similarity between users is measured using cosine similarity. Subsequently, from the global user set... The search results for the most similar items were retrieved from the list. For each user, construct a set of similar users. :
[0063]
[0064] in, For a set of similar users; To retrieve the top with the highest similarity One user; For users and users Similarity score calculation; For users Embedded representation; For users Embedded representation; For users, the complete collection Delete user The remaining users.
[0065] The collection of denoised trajectories of these similar users This will serve as a highly reliable signal of collaborative behavior for subsequent semantic reasoning and trajectory enhancement processes.
[0066] (2) User Semantic Reasoning and Trajectory Completion Based on LLM. To achieve intelligent and interpretable semantic completion of user trajectories, this invention abandons the random-driven strategies (such as random insertion or replacement) commonly used in traditional data augmentation. Although these strategies can easily expand the data scale, they often introduce additional semantic noise because they fail to consider the inherent logic and contextual consistency of user behavior. To this end, this invention designs a generative semantic completion method based on the Chain-of-Thought (CoT) prompting strategy. This method aims to guide LLM to simulate a human-like multi-step progressive reasoning process: first, understand the historical behavior patterns of the target user; second, abstract the common behavioral characteristics of semantically similar user groups; and finally, integrate the two types of knowledge to generate a logically consistent completion result for the current trajectory. This structured reasoning path improves the transparency and reliability of the completion process, ensuring that the generated potential check-ins are consistent with the user's real preferences and real-world behavioral logic.
[0067] First, the denoised trajectory of the target user and the retrieved trajectories of similar users are used to construct an enhanced semantic description. Second, the enhanced semantic description and a preset chain-like thinking prompt template are input into a large language model, guiding the large language model to perform the following inferences in sequence: extracting user historical preferences, summarizing similar user preferences, and judging the completeness of the trajectory. When it is determined that completion is needed, the large language model generates candidate completion check-in points under the constraints of time sequence, interest point category, and behavioral logic. Finally, the generated completion check-in points are incorporated into the denoised trajectory to obtain the enhanced trajectory.
[0068] Specifically, the CoT prompt template, as shown in Table 3, guides the LLM to execute the following four core inference steps in sequence.
[0069] Table 3. Structured Description of Track Enhancement Hint Templates
[0070] 1) Extract user historical preferences. This is because it directly involves the discrete denoised trajectory. Performing high-level reasoning on top of this is quite challenging. We first suggest using LLM analysis to examine the logical relationships between POI identifiers, categories, and access times in the denoised trajectory data, generating a concise summary of the user's historical behavior patterns (e.g., "Users frequently go to the gym on weekday evenings and tend to travel along a leisurely route from a coffee shop to a bookstore on weekends"). This summary extracts the user's core activity preferences, time patterns, and potential travel motivations in natural language, thus providing a personalized contextual basis for subsequent reasoning.
[0071] 2) Summarize similar user preferences. Relying solely on individual historical data can lead to insufficient information, while similar user groups often contain transferable and relatively stable behavioral patterns. By analyzing a large amount of trajectory data within this group, LLM can identify high-frequency behavioral patterns that are potentially related to the target user's historical preferences and extract them into a group-level summary of collaborative behaviors. This step aims to tap into collective wisdom; for example, "Among users with similar behavioral patterns to the target user, the probability of visiting a light meal restaurant within one hour after engaging in exercise is significantly higher." Such summaries provide group-validated potential behavioral associations, effectively compensating for and enriching the limitations of individual experience.
[0072] 3) Trajectory Inference. After obtaining the target user's historical behavioral preferences and the collaborative behavioral characteristics of similar user groups in the first two stages, LLM performs comprehensive inference on these two types of preference information in this stage. By integrating the user's long-term historical preferences with the common behavioral patterns at the group level, the model evaluates the semantic continuity and behavioral rationality of the current user trajectory to determine whether there are potential information gaps or semantic breaks, thereby deciding whether trajectory enhancement and completion are needed.
[0073] 4) Context-Aware Generative Semantic Completion. When the previous stage determines that the current trajectory needs completion, LLM will further generate supplementary check-in points for the inserted trajectory under contextual constraints such as time sequence, POI category, and overall behavioral logic. Specifically, the model infers the most reasonable missing access to compensate for semantic breaks caused by excessively long time intervals or unnatural activity type transitions, and generates two candidate completion points for each trajectory. Each completion point contains a reasonable POI identifier, corresponding category, and estimated check-in timestamp. If the model determines that the trajectory already has satisfactory continuity and completeness, completion is not performed, and an empty result is returned.
[0074] The CoT inference process described above is as follows. This invention is for the target user. Define an enhanced suggestion constructor function Its input is the user's denoised trajectory. and a set of denoised trajectories of similar users .from and The check-in information is extracted to construct the enhanced semantic trajectory description. :
[0075] make This represents the semantically enhanced check-in set representing the current user trajectory, and models the semantic enhancement completion and inference generation process of LLM as a function. The definition is as follows:
[0076] in, This represents the parameters of the LLM. If the LLM determines that the user trajectory information is complete and no enhancement or completion is needed, then... Since it is an empty set, the trajectory remains unchanged.
[0077] For each user LLM will identify a set of semantically missing sign-in records. Used to supplement the denoised trajectory Subsequently, these predicted candidate completion check-in points are integrated into the denoised trajectory through set union operation, thereby constructing the enhanced trajectory:
[0078] The final enhanced trajectory LLM-driven semantic augmentation effectively mitigates data sparsity, making trajectories more semantically coherent and statistically robust. Simultaneously, it provides high-quality input for subsequent global topology denoising, thereby improving the model's generalization ability in scenarios involving long-tail users and sparse behavior.
[0079] After obtaining semantically cleaner and more information-rich user trajectories at the single-trajectory level, a credibility-aware graph structure optimization module is introduced. This module further models POI transition relationships from a global perspective to enhance the robustness of representation learning to structural noise. Since global POI transition graphs often contain low-credibility edges introduced by occasional visits, POI popularity bias, or data sparsity, these noisy connections may mislead the model into learning unreliable association patterns, thereby weakening the discriminative ability of POI representations. To address this issue, this invention employs a credibility-aware graph structure optimization module, specifically a credibility-aware graph structure optimization module, to identify and suppress transition relationships lacking semantic reliability.
[0080] In step S4, the secondary denoising at the global topology level specifically includes: First, a global interest point transfer graph is constructed based on the enhanced trajectories of all users. In the graph, nodes represent interest points and directed edges represent the access transfer relationships between interest points. Secondly, a graph convolutional network is used to propagate the global interest point transfer graph, extract the graph convolutional augmented representation of each interest point, and use the cosine similarity between the graph convolutional augmented representations of the two endpoints of each edge as the structural credibility score of that edge; then, for each source interest point, all its outgoing edges are sorted from high to low according to the structural credibility score, and only the top k outgoing edges with the highest scores are retained, while the remaining low credibility edges are pruned to obtain the pruned high credibility topology. Finally, the pruned high-confidence topology structure is input into the gated graph neural network. During each update, neighborhood messages are aggregated first, and then the gated state is fused through the update gate and the reset gate. The fusion weight of neighborhood information and node historical state is adaptively adjusted through the update gate and the reset gate. After multiple iterations, the topology embedding of each interest point is generated, thereby completing the secondary denoising at the global topology level.
[0081] The value of k for the first k outgoing edges is determined to be 15 through hyperparameter sensitivity analysis. The hyperparameter sensitivity analysis specifically involves adjusting the number of edges k retained during pruning within a preset range, observing the changes in model performance, and determining the optimal value of k for the recommendation effect.
[0082] In this embodiment, the credibility-aware graph structure optimization module includes three stages: global transfer network modeling and graph convolution feature extraction, edge credibility estimation and adaptive pruning, and GGNN representation learning based on high-credibility topology.
[0083] (1) Global transfer network modeling and graph convolution feature extraction. First, a directed transfer graph is constructed from the enhanced trajectories of all users, and a graph convolutional network (GCN) is used to obtain POI representations that integrate global neighborhood information. These representations will serve as a reference for subsequent edge confidence estimation.
[0084] To model potential transfer patterns and global collaborative relationships among Points of Interest (POIs) from a holistic perspective, an enhanced trajectory set based on all users is used. Construct a global POI transfer graph This serves as the structural foundation for subsequent graph representation learning.
[0085] Definition 1 (Global POI Transition Graph). The global POI transition graph is denoted as... , where the set of nodes Represents the set of all unique Points of Interest (POIs). A directed edge. This indicates that in the user's POI access history, from the access... Visit afterwards The frequency of transitions.
[0086] However, this transition graph simultaneously contains high-quality connections reflecting stable behavioral patterns, as well as low-reliability edges introduced by occasional visits, data sparsity, and popularity bias. Direct representation learning on such a graph structure is susceptible to structural noise.
[0087] To initially assess the reliability of edges and provide a basis for subsequent pruning, it is necessary to obtain vector representations that can capture the latent semantics of POIs. Therefore, in the POI transition graph... The initial embedding learning is performed using a graph convolutional network (GCN), aiming to fuse multi-hop neighborhood information and obtain a GCN-enhanced representation for each point of interest (POI). To simplify processing, the transition graph is symmetricized during GCN propagation. Although these representations are still affected by noisy edges, they better characterize the position of nodes in the global graph structure and coarse-grained semantic relationships compared to the initial features, thus providing a more informative reference for the subsequent edge confidence estimation stage.
[0088] Specifically, let This represents the initial feature matrix of the POI. It uses a matrix containing... The GCN of the first layer propagates. layer( Its propagation rules are defined as follows:
[0089] in, Indicates a self-loop Adjacency matrix, express The degree matrix, For a trainable weight matrix, The activation function is (e.g., ReLU). After... After layer propagation, the GCN-enhanced representation matrix of the POI is obtained. .in, each line Corresponding POI The vector representation that incorporates the POI transition graph Information about multi-hop neighborhoods. It is regarded as a reference representation for evaluating the reliability of edges, and the semantic similarity information it captures will be used for subsequent credibility assessment and pruning decisions.
[0090] (2) Edge credibility estimation and adaptive pruning. Using the POI representation learned by GCN, the semantic similarity between node pairs is calculated and used as an indicator of edge credibility. Based on this credibility, the POI transition graph is sparsified to retain high-credibility connections, thus obtaining the purified pruned graph structure.
[0091] In obtaining enhanced representation Next, the reliability of edges in the original transition graph is further evaluated to facilitate subsequent structural cleanup. Considering that: if an edge... For semantically consistent and repetitive transition patterns (such as functional compatibility or common route connections), the representation of its endpoint POI is... and Edges are more likely to be close to each other in the learned semantic space. Conversely, edges introduced by generalization transfers driven by incidental visit trajectories or popularity (i.e., structural noise) typically exhibit weaker representational consistency in the global graph context. Therefore, edges... Structural credibility score Defined as the cosine similarity between the GCN-enhanced representations at their two endpoints:
[0092] higher Indicate the edge Stronger representational consistency within a global graph context suggests a more stable and semantically coherent transition pattern; conversely, lower scores indicate weaker edge reliability, potentially induced by accidental co-occurrence or structural noise. Based on these reliability estimates, this invention employs an adaptive top-... Pruning strategies are used to clean up the graph structure.
[0093] Specifically, for each source POI Consider the set of all its outgoing edges. ,in express The set of out-neighbors. Based on structural credibility score. Sort these edges in descending order and keep only the top-scoring edges. The first k edges are pruned, and the remaining connections are removed. This pruning process is applied independently to all nodes, resulting in a sparser adjacency matrix. And the transfer diagram after purification .
[0094] By selectively retaining edges with higher semantic consistency in the learning representation space, the pruned graph effectively suppresses spurious transition relationships caused by randomness or popularity bias, while preserving more structurally informative transition connections. This refined graph structure provides a cleaner and more reliable foundation for subsequent POI representation learning based on GGNN.
[0095] (3) GGNN representation learning based on high-confidence topology. Gated graph neural network (GGNN) is applied to the pruned high-confidence topology. The information propagation is regulated by the gating mechanism, and a more robust POI embedding representation is finally generated.
[0096] After obtaining a high-confidence topology through adaptive graph pruning, this stage aims to learn POI representations on the pruned graph, enabling it to capture high-order transition dependencies while remaining robust to potential residual noise. Although many low-confidence edges have been removed, POI transitions are inherently directional and strictly temporally ordered, constituting a serialization process. Therefore, the model needs to effectively characterize dynamic multi-hop dependencies. While traditional multi-layer GCNs can aggregate high-order neighborhood information, they often encounter smoothing problems during deep propagation, thus weakening the discriminative power of node representations. Furthermore, their static aggregation mechanism struggles to model heterogeneous long-term and short-term dependencies in sequential decision-making.
[0097] To address these limitations, this invention employs a gated graph neural network (GGNN) as the representation learner on a high-confidence topology. The core advantage of GGNN lies in its introduction of a gating mechanism similar to recurrent neural networks (e.g., GRU), modeling node representation learning as a dynamically evolving process over discrete time steps. The update and reset gates adaptively adjust the trade-off between newly aggregated neighborhood information and node historical states based on the current context, thereby achieving dynamic control over the information flow. This mechanism not only captures multi-hop dependencies along cleaned directed edges but also effectively avoids over-smoothing, preserving node-specific semantic features and generating more discriminative representations. Therefore, GGNN is well-suited for generating the final robust topological embeddings required for recommendation tasks, maintaining global structural consistency while accommodating individual semantic differences.
[0098] To avoid over-smoothing of input features, this invention uses the initial POI feature matrix obtained in the first stage. Instead of the representation enhanced by GCN. This is used as the input to the GGNN. For each POI node... Its hidden state is initialized to Subsequently, GGNN defines the adjacency matrix based on the high-confidence pruning topology. ,implement Each iteration update consists of two core operations: Aggregate neighborhood messages: During each iteration update, for each target interest point, multiply the hidden states of all its outgoing neighbor nodes in the previous time step by the edge weights and sum them to obtain a comprehensive neighborhood message aggregation vector.
[0099] Gated state fusion: The hidden state of the current node at the previous moment and the neighborhood message aggregation vector are input into the gated loop unit. The retention ratio of historical states is controlled by updating the gate and the degree of introduction of new messages is controlled by resetting the gate. The two are weighted and fused to generate a new hidden state of the node.
[0100] In this embodiment, aggregating neighborhood messages specifically involves: at each time step In the middle, node From all outgoing neighbors within the high-confidence topology Receive information. Compute the neighborhood message aggregation vector of the node. as follows:
[0101] This aggregation vector represents the combined signal that the current node acquires from its neighborhood. This operation ensures that POI transfer information propagates only through the high-confidence edges retained after pruning, thereby effectively isolating the spread of structural noise at the model structure level.
[0102] Gated state fusion specifically involves: obtaining the neighborhood message aggregation vector Then, a Gated Recurrent Unit (GRU) is used to update the hidden state of the node. This process involves updating the gate. With Reset Door To achieve fine-grained information control and generate candidate states. :
[0103] in, This represents the sigmoid function. This represents element-wise multiplication. These are learnable parameters. Reset gate. Used to control the hidden state of the previous moment. How much information is incorporated into the calculation of the new candidate state? Simultaneously, the update gate... To what extent the current state is determined by the candidate states Update, rather than retaining, the node's new state. It is a weighted combination of the old state and the candidate state. After... After the iteration, the hidden state at the final time step is used as the robust topological embedding of the POI, i.e., the topological embedding of a single interest point is denoted as... The embeddings of all POIs together constitute the topological embedding matrix of the points of interest. .
[0104] By employing GGNN representation learning on a purified, high-confidence topology, this module achieves global topology-level denoising. Restricting information propagation to reliable POI transition relationships effectively suppresses the impact of structural noise. Simultaneously, the gating mechanism adaptively balances node-specific features with neighborhood information, mitigating over-smoothing and enhancing semantic discriminative ability. The resulting POI representation compactly encodes global collaborative behavior patterns, providing robust and information-rich input for subsequent POI recommendation stages, thereby further improving overall recommendation performance.
[0105] After semantic cleansing and enrichment at the individual trajectory level and structural graph optimization at the global topology level, two key types of information were obtained: the first is the enhanced user trajectory after semantic cleansing and completion, which can more completely and coherently reflect the user's true behavioral patterns; the second is the robust POI semantic topology embedding learned by GGNN on a high-confidence topology, which encodes reliable global transfer patterns between POIs. However, the final next step, POI prediction, not only needs to understand the user's current travel context and immediate intent, but also needs to integrate the user's long-term stable preferences and the attractiveness of POIs in specific spatiotemporal contexts, thereby achieving more accurate recommendation decisions.
[0106] To this end, this invention designs a recommendation architecture consisting of two core modules: (1) a multimodal check-in representation learning module, and (2) a Transformer-based temporal preference modeling and prediction module. The former aims to construct a rich multimodal representation for each check-in, integrating multi-source information, including POI semantics, user preferences, temporal behavior patterns, and POI functional categories. The latter utilizes the powerful trajectory modeling capabilities of Transformer to encode the user behavior trajectory composed of the multimodal check-in representation, capture complex temporal dependencies and preference evolution processes, and finally jointly predict the next visited POI and its related contextual attributes through a multi-task prediction head.
[0107] In section S5, the multimodal sign-in representation learning module is first introduced, as follows: User check-in behavior is typically influenced by a combination of factors, including the visited Points of Interest (POIs), user preferences, temporal regularity of behavior, and the functional category of the POIs. Modeling only the POI identifier is insufficient to effectively capture these multidimensional potential dependencies. To address this limitation, this invention proposes a multimodal check-in representation learning strategy that integrates four types of embedding features (point of interest embedding, user embedding, temporal embedding, and category embedding) to construct a unified representation for each check-in.
[0108] POI Embedding: Introducing a Robust POI Semantic Topological Embedding Matrix This matrix is learned by a graph denoising module at the global topology level, encapsulating the global semantic and structural information of POIs in the cleaned transition topology. The final POI embedding representation is denoted as... :
[0109] User Embedding: Definition 2 (User-POI Graph). The user-POI graph is denoted as... ,in, , and These represent the user set, the POI set, and the edge set, respectively. If the user... POI visited in its check-in history Then in the user With POI There is an edge between them .
[0110] Based on the user-POI graph Two-layer GCN ( ) Perform message propagation. Initialize user embedding. With POI embedding After concatenation, the data is input into the GCN layer. After propagation, user embeddings are calculated by aggregating messages from neighboring POIs. (The last sentence appears to be a fragment and doesn't translate directly.) The front of the layer representation The rows (corresponding to the user part) are combined with residual fusion and linear transformation to generate the final user embedding representation. :
[0111] in, User-POI Interaction Diagram The normalized adjacency matrix. To facilitate subsequent message propagation in GCN, the POI index is determined by offset. It was translated. For trainable weight matrix, This represents the activation function (ReLU).
[0112] Time embedding: Timestamps are embedded using Time2Vec encoding technology. Mapping to a vector space. This method can simultaneously capture the linear trend and periodic dynamics of time, thus obtaining a temporal embedding representation. :
[0113] Category embedding: The category information of POIs reflects the underlying semantic motivations behind user behavior. A learnable embedding layer is used. Index by category Mapped to dense vectors :
[0114] Ultimate Embedded Fusion: For Users Enhanced trajectory This invention designs an embedded fusion function. This is used to fuse interest point embeddings, user embeddings, time embeddings, and category embeddings. The fused representation... The calculation is as follows:
[0115] in, This represents the dimension of the merged user trajectory representation. and These are learnable parameters. For POI embedding representation, Embedded for users, An embedded representation of the check-in time. An embedded representation of POI category information.
[0116] Secondly, the Transformer-based temporal preference modeling and prediction module is as follows: Traditional RNN / GRU models are prone to gradient vanishing or memory decay when processing long trajectory sequences, making it difficult to effectively model long-range dependencies. Convolutional models, on the other hand, are limited by fixed receptive fields and lack the flexibility to capture variable-length contextual information. To overcome these limitations, a Transformer encoder is used to model user check-in trajectories. Thanks to its self-attention mechanism, the Transformer can establish dependencies between arbitrary locations throughout the entire trajectory, allowing the network to simultaneously characterize short-term movement patterns and long-term interest shifts, thus enabling more granular modeling of complex user travel behaviors.
[0117] In this embodiment, a set of all enhanced user trajectories is given. The system stacks the corresponding embedding representations to form the input tensor of the encoder layer. To enhance the model's understanding of time trajectories and movement patterns, several auxiliary tasks were introduced in addition to the main task of predicting the next POI. During the decoding phase, a multi-head decoder was constructed based on a multilayer perceptron (MLP) to perform multiple prediction tasks simultaneously.
[0118] in, , and These represent the weight matrices of each MLP prediction head. This represents the output of the Transformer encoder, while , and These represent the number of POIs, the number of time intervals, and the number of POI categories, respectively. (Vector) , and These represent the model's predicted probability distributions for the next POI, the next time interval, and the next POI category, respectively.
[0119] To optimize the model within a multi-task learning framework, this invention employs the cross-entropy function as the target loss, used to simultaneously predict the next POI, the next time interval, and the next class. The overall loss function is defined as the sum of the losses from each task:
[0120] Specifically, each component is defined using the cross-entropy loss form as follows:
[0121] in, , and These represent the actual labels for the next POI, the next time interval, and the function category, respectively. Furthermore, , and These represent the POI set, the time interval set, and the POI category set, respectively.
[0122] In summary, this invention first uses a large language model to perform semantic noise filtering and collaborative completion on users' historical check-in trajectories. This effectively eliminates abnormal check-ins caused by misoperation, occasional check-ins, or missing records, restoring the continuity and completeness of the trajectory, thereby alleviating the data sparsity problem and significantly improving the recommendation quality for long-tail users. Second, through global topology-level credibility assessment and adaptive pruning, it can eliminate false transition relationships formed by random check-ins or accidental co-occurrence of popular attractions, blocking the propagation path of noise in the graph structure, making the final generated interest point topology embedding more robust. It can be applied to real-world interest point recommendation scenarios such as shops, restaurants, and attractions. Furthermore, experimental results on real-world city datasets such as New York, Tokyo, and London show that this invention consistently outperforms existing state-of-the-art methods in terms of accuracy and normalized discount cumulative gain, achieving a performance improvement of 3% to 6%. It can provide users with more accurate recommendations for the next shop, restaurant, or attraction, improving travel planning efficiency and platform user experience.
[0123] To evaluate the performance of predicting the next POI, this invention employs two widely used evaluation metrics in the recommender system: accuracy (Acc) and normalized discounted cumulative gain (NDCG). The top k recommendation accuracy (Acc@k) measures whether the target user's next check-in POI appears in the top k recommendation list; if the test POI is included in the recommendations, it is considered "visited." On the other hand, the top k normalized discounted cumulative gain (NDCG@k) not only evaluates the correctness of the recommendations but also their rank position, assigning higher weights to POIs appearing at the top of the list and gradually reducing the contribution of lower-ranked POIs based on their marginal fractional utility.
[0124] Table 4 shows the performance comparison of LMD-POI with various baseline methods on the New York City (NYC), Tokyo (TKY), and London (Great Britain (GB) check-in datasets. Acc@(k) and NDCG@(k) are used. () as an evaluation indicator.
[0125] Table 4. Performance comparison of different methods on the NYC, TKY, and GB datasets.
[0126] Based on the experimental results, the following key observations can be made: First, on the NYC dataset, the proposed LMD-POI framework consistently outperforms all baseline methods across all evaluation metrics. Specifically, compared to the second-best state-of-the-art baseline model (Spatio-Temporal Clustering Model, IPCM), LMD-POI achieves relative improvements of 3.27%, 2.46%, 5.43%, and 5.27% on Acc@5, Acc@20, NDCG@5, and NDCG@20, respectively. These significant gains validate the effectiveness of the proposed multi-level denoising mechanism in improving recommendation accuracy.
[0127] Traditional sequence recommendation methods, such as self-attention-based sequential models (SASRec) and personalized long-and-short-term preference learning (PLSPL) models, perform relatively poorly. This is mainly because these models only focus on the sequential dependencies in the check-in trajectory, while ignoring key factors in POI recommendation, including spatiotemporal contextual information (such as geographical proximity and time periodicity) and the semantic attributes of POIs.
[0128] Methods incorporating spatiotemporal information and graph structures, such as Spatio-Temporal Attention Networks (STAN) and Disentangled Contrastive Hypergraph Learning (DCHL), have achieved significant improvements over pure sequence models. DCHL, in particular, captures higher-order relationships through hypergraph modeling, thus achieving competitive performance. However, these methods typically assume the original interaction data is completely reliable and lack mechanisms for explicitly handling noisy observations. Therefore, their performance is often limited when applied to real-world trajectory data containing random check-ins or erroneous records.
[0129] More recent baseline methods, such as the multi-modal content-aware framework for point-of-interest (POI) recommendation (MMPOI) and the spatio-temporal clustering model (IPCM), have further improved performance through multimodal fusion or clustering strategies. The spatio-temporal clustering model (IPCM), in particular, has achieved highly competitive results using a personalized spatio-temporal clustering mechanism. However, LMD-POI continues to consistently outperform these strong baseline methods with a stable advantage.
[0130] The superiority of LMD-POI primarily stems from its novel two-layer denoising design. At the individual trajectory level, LMD-POI fully leverages the reasoning capabilities of large language models to identify and filter semantically inconsistent noisy check-ins, while simultaneously performing context-aware preference completion based on similar user behaviors. This process not only alleviates the data sparsity problem but also enhances the semantic coherence of trajectories, providing high-quality input for subsequent learning.
[0131] At the global topology level, LMD-POI performs confidence-based adaptive pruning on the global POI transition graph and learns robust POI embeddings through GGNN on the cleaned high-confidence topology, thereby effectively blocking the propagation of structural noise.
[0132] On the TKY and GB datasets, LMD-POI also demonstrates a consistent advantage, with average performance improvements comparable to those observed on the NYC dataset (approximately 3%–6%). These results further validate the strong generalization ability and robustness of the proposed framework.
[0133] In this invention, ablation experiments were also conducted, and six variants of LMD-POI were designed, as shown in Table 5, to demonstrate the importance of key components in the model of this invention: Table 5. Ablation experiments of different components on the NYC, TKY, and GB datasets.
[0134] Wherein, "w / o" means without, that is, removing the corresponding module or feature; -w / o LLM means removing the large language model module, that is, disabling the trajectory denoising and trajectory enhancement process driven by the large language model, and directly inputting the original user trajectory into the subsequent recommendation model. -w / o LLM Augment means removing the large language model trajectory enhancement module, that is, retaining the large language model denoising function, but no longer performing semantic completion for missing check-ins based on similar user behavior. -w / o Denoising Graph means removing the denoising graph structure optimization module, that is, not performing graph pruning based on edge confidence and subsequent high-confidence topology learning, but directly performing representation learning on the original global interest point transfer graph. -w / o User means removing user embedding, ignoring user personalized preference features; -w / o Time means removing time embedding and time prediction loss, ignoring the time features of check-in behavior; -w / o Cat means removing interest point category embedding and category prediction loss, used to verify the impact of category information on user preference modeling and interest point recommendation.
[0135] In this embodiment, -w / o LLM: This variant removes the entire LLM module. Specifically, both LLM-driven denoising and enhancement processes are disabled, and the original user trajectory is directly input into the subsequent model architecture.
[0136] -w / o LLM Augment: In this configuration, the LLM-based denoising function is retained, but the enhancement module is removed. The model only uses the LLM-cleaned trajectories and no longer performs semantic completion for missing check-ins based on similar user behavior.
[0137] -w / o Denoising Graph: This variant removes the confidence-based graph denoising module at the global topology level. Specifically, this invention omits the edge pruning process and subsequent GGNN operations on the POI transition graph, instead directly employing standard GCN for message passing and representation learning on the constructed global transition graph.
[0138] -w / o User: Remove user embeddings and ignore learned user preference features.
[0139] -w / o Time: Removes time embedding and time prediction loss, ignoring the time features of check-in behavior.
[0140] -w / o Cat: Removes POI category embeddings and category prediction loss, avoids the model considering category information, and tests the importance of category features in capturing user preferences and points of interest features.
[0141] First, trajectory cleansing and enrichment modules based on a large language model are crucial for performance improvement. The performance degradation observed in the -w / o LLM variant directly indicates that semantic noise inherent in the original trajectories severely interferes with the accurate learning of user preferences. In contrast, although -w / o LLM Augment outperforms -w / o LLM, it still lags significantly behind the full model. This demonstrates that noise filtering alone is insufficient; combining similar user behaviors with LLM inference capabilities for trajectory completion is essential for mitigating data sparsity and is a key factor in improving recommendation accuracy.
[0142] Secondly, the credibility-aware graph structure optimization module also makes a significant contribution. The performance degradation of the -w / o Denoising Graph variant indicates that directly learning representations on the original transition graph containing a large number of low-credibility transition edges leads to the embedding of POIs with aggregated noise information, resulting in over-smoothing and semantic drift. In contrast, the module proposed in this invention can effectively cleanse the graph structure and promote the learning of more robust POI representations.
[0143] Finally, the fusion of multimodal contextual information is essential. Among the three feature ablation variants, -w / o User exhibited the most severe performance degradation, highlighting its central role in modeling personalized user preferences in POI recommendation. Meanwhile, the significant performance decline of -w / o Time and -w / o Cat further validates the indispensable role of temporal context and POI feature category information in capturing user behavioral motivations and periodic patterns.
[0144] To further explore the impact of the key hyperparameter k in the global topology level on representation learning performance of the credibility-aware graph structure optimization module, where k controls the number of top-k high-credibility edges retained after pruning, this invention also conducted comprehensive hyperparameter sensitivity analysis experiments on three benchmark datasets. During the experiments, the value range of k was set to {5, 10, 15, 20, 25}, and NDCG@20 and Acc@20 were used as evaluation metrics.
[0145] like Figure 2 The figure shown is a comparison of the impact of the hyperparameter k on the New York City dataset (NYC) of this invention. From... Figure 2 It can be seen that as the value of k increases, the model's performance on Acc@20 and NDCG@20 exhibits an inverted U-shaped trend of first rising rapidly and then slowly declining. When k=5, due to excessive pruning, a large number of potential semantic transition edges are removed, resulting in significantly lower performance. When k increases to 15, both metrics reach their peak, with NDCG@20 showing an improvement of about 7% compared to k=5, indicating that this pruning intensity is most suitable for the graph structure characteristics of the NYC dataset. When k exceeds 15, the performance begins to gradually decline, but the decline is relatively gentle, indicating that the NYC dataset has good differentiation between high-confidence edges and noisy edges, and retaining slightly more edges will not cause severe interference.
[0146] like Figure 3 The figure shows a comparison of the impact of hyperparameter k on the Tokyo City (TKY) dataset. The TKY dataset has the largest number of users and check-ins, resulting in richer trajectories. The characteristics shown in the figure differ slightly from those of the NYC dataset: when k increases from 5 to 15, both Acc@20 and NDCG@20 significantly improve, with the peak also occurring at k=15; however, when k continues to increase above 20, the performance degradation rate is significantly faster than on the NYC dataset. This is mainly because the TKY dataset has a denser distribution of interest points and a greater number of low-confidence edges. Excessive edge retention introduces more structural noise, thus degrading representation learning quality more quickly. This characteristic illustrates that accurate top-15 pruning is particularly crucial on dense datasets.
[0147] like Figure 4 The figure shows a comparison of the impact of the hyperparameter k on the London City (GB) dataset. The GB dataset has a relatively low overall check-in density and high trajectory sparsity. As can be seen from the figure, the performance loss is most severe when k=5, because sparse data inherently lacks sufficient connections, and excessive pruning leads to information gaps. As k increases to 15, the performance steadily recovers to its optimal level. When k exceeds 15, the performance decline trend is between that of NYC and TKY, but the overall fluctuation is small. This indicates that on sparse datasets, top-15 pruning can both supplement necessary topological information and effectively suppress noise, demonstrating good robustness.
[0148] comprehensive Figures 2-4 It can be seen that although the three datasets differ in data density and noise level, the optimal k value remains stable at 15. This result verifies the generalization ability of the hyperparameter setting proposed in this invention: k=15 achieves the best trade-off between structural integrity and denoising effectiveness, and is applicable to interest point recommendation scenarios with different city scales and sparsity levels. Therefore, in this embodiment, k=15 is fixed in all experiments to ensure the best trade-off between global structural integrity and denoising effectiveness.
[0149] Understandably, based on the above embodiments, the present invention may also have a variety of alternative implementation methods.
[0150] As an optional embodiment, this embodiment uses a gated graph neural network as the core for feature extraction at the global topology level, but the invention is not limited to this. In practical applications, graph convolutional networks, graph attention networks, or graph Transformers can also be used as alternatives. These structures can also perform message propagation and node feature updates on the purified high-confidence topology structure of this invention.
[0151] As an optional embodiment: This embodiment uses a large language model to drive the denoising module, but the present invention is not limited to the mainstream large-scale pre-trained model. For terminal devices with limited computing resources, a lightweight language model after knowledge distillation or a logic verification engine based on semantic knowledge graph can also be used as an alternative to achieve semantic noise recognition and trajectory completion functions.
[0152] As an optional embodiment: This embodiment employs a top-k hard pruning strategy based on credibility scores for graph structure optimization, but the invention is not limited to this. Alternatively, an attention mechanism can be introduced to assign continuous weights to each edge, achieving soft denoising through soft masking; or a graph smoothing technique based on Laplace constraints can be used to optimize the topology, achieving the same effect of suppressing structural noise.
[0153] As an optional embodiment: This embodiment has good scalability. In other implementations, auxiliary information such as the user's historical consumption preferences, the heat map distribution of people at the check-in point, or the user's social network can be added to the input to further enrich the multimodal check-in representation and improve the accuracy of the final point of interest prediction.
[0154] Those skilled in the art should understand that the above alternative embodiments are all within the scope of protection covered by this invention.
[0155] Example 2 In a typical embodiment of the present invention, an interest point recommendation system based on multi-level denoising is provided, comprising: Data acquisition module: used to acquire users' original check-in history and the global user set; The first-level denoising module is used to perform semantic noise filtering on the original check-in trajectory to obtain a denoised trajectory, thus completing the initial denoising at the individual trajectory level. The trajectory enhancement module is used to extract similar user trajectories from the global user set and obtain enhanced trajectories through semantic completion. The second-level denoising module is used to construct a global interest point transfer graph based on the enhanced trajectory, extract node features, and then perform credibility evaluation and adaptive pruning on the edges in the graph; the pruned high-confidence topology structure is then subjected to graph representation learning to obtain the interest point topology embedding, thus completing the second-level denoising at the global topology level. Recommendation module: used to fuse the enhanced trajectory with the topological embedding of interest points to obtain recommended interest points.
[0156] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the method of Embodiment 1.
[0157] The steps and methods involved in Embodiment 3 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0158] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0159] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0160] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for recommending points of interest based on multi-level denoising, characterized in that, include: Obtain the user's original check-in history and the global user set; The original check-in trajectory is subjected to semantic noise filtering to obtain a denoised trajectory, thus completing the initial denoising at the individual trajectory level; including constructing the original check-in trajectory into a semantic trajectory, inputting the semantic trajectory and a preset denoised prompt template into a large language model, identifying and outputting a set of noisy check-ins; The process involves extracting similar user trajectories from the global user set and obtaining enhanced trajectories through semantic completion; this includes inputting the denoised trajectory and similar user trajectories into a large language model, performing reasoning through chain-like thinking prompts, and incorporating the generated candidate completion check-in points into the denoised trajectory to obtain the enhanced trajectory. Based on the enhanced trajectories, a global interest point transfer graph is constructed. After extracting node features, the edges in the graph are evaluated for credibility and adaptively pruned. This includes: constructing a global interest point transfer graph based on the enhanced trajectories of all users, where nodes represent interest points and directed edges represent access and transfer relationships between interest points; propagating the global interest point transfer graph, extracting the graph convolutional enhanced representation of each interest point, and using the cosine similarity between the graph convolutional enhanced representations of the two endpoints of each edge as the structural credibility score of that edge; for each source interest point, sorting all outgoing edges according to their structural credibility scores, and retaining only the preset number of outgoing edges with the highest scores to obtain a pruned high-credibility topology; and performing graph representation learning on the pruned high-credibility topology to obtain the interest point topology embedding, thus completing the second-level denoising at the global topology level. The graph representation learning employs a gated graph neural network. During each update, neighborhood messages are aggregated first, and then gated state fusion is performed through update and reset gates. After multiple iterations, an interest point topological embedding is generated, completing secondary denoising at the global topological level. Specifically, the aggregation of neighborhood messages involves multiplying the previous hidden states of all outgoing neighbor nodes by their edge weights and summing the results during each update iteration to obtain a neighborhood message aggregation vector. The gated state fusion involves inputting the previous hidden state of the current node and the neighborhood message aggregation vector into a gated recurrent unit. The update gate controls the retention ratio of historical states, and the reset gate controls the introduction of new messages, resulting in a weighted fusion to generate a new hidden state for the node. The enhanced trajectory is fused with the interest point topology embedding to obtain recommended interest points.
2. The interest point recommendation method based on multi-level denoising as described in claim 1, characterized in that, The initial noise reduction specifically includes: The noisy check-in set is removed from the original check-in trajectory to obtain the denoised trajectory.
3. The interest point recommendation method based on multi-level denoising as described in claim 1, characterized in that, The semantic completion specifically includes: A pre-trained model is used to extract user embeddings, and the top preset number of users with the highest similarity are retrieved to obtain similar user trajectories.
4. The interest point recommendation method based on multi-level denoising as described in claim 3, characterized in that, The similarity is measured by calculating the cosine similarity between the target user and the embedded representations of each user in the global user set.
5. The interest point recommendation method based on multi-level denoising as described in claim 1, characterized in that, The topological embeddings of points of interest, user embeddings, time embeddings, and category embeddings of each check-in in the enhanced trajectory are concatenated and fused to construct a multimodal check-in representation sequence; The multimodal check-in representation sequence is input into the Transformer encoder for temporal modeling, and the next point of interest of the user is predicted by the multilayer perceptron to obtain the recommended point of interest.
6. An interest point recommendation system based on multi-level denoising, characterized in that, A method for recommending interest points based on multi-level denoising according to any one of claims 1-5, comprising: Data acquisition module: used to acquire users' original check-in history and the global user set; The first-level denoising module is used to perform semantic noise filtering on the original check-in trajectory to obtain a denoised trajectory, thus completing the initial denoising at the individual trajectory level. The trajectory enhancement module is used to extract similar user trajectories from the global user set and obtain enhanced trajectories through semantic completion. The second-level denoising module is used to construct a global interest point transfer graph based on the enhanced trajectory, extract node features, and then perform credibility evaluation and adaptive pruning on the edges in the graph; the pruned high-confidence topology structure is then subjected to graph representation learning to obtain the interest point topology embedding, thus completing the second-level denoising at the global topology level. Recommendation module: used to fuse the enhanced trajectory with the topological embedding of interest points to obtain recommended interest points.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Next interest point recommendation method based on global trajectory flow graph and graph contrast learning
CN119940489A