Hierarchical fine-grained human trajectory activity type inference method and related device

By combining a hierarchical approach with anchor rules, random forest models, and large language models for structured reasoning, this method solves the problems of low accuracy and poor stability in fine-grained activity type inference in existing technologies, and achieves reliable identification of non-rigid activities and improved adaptability to complex scenarios.

CN121765431APending Publication Date: 2026-03-31SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing activity type inference techniques suffer from low accuracy, poor stability, and limited adaptability in fine-grained activity type inference, making it difficult to achieve reliable fine-grained classification, especially in the identification of non-rigid activities.

Method used

A hierarchical approach is adopted. First, rigid activity types are identified through predefined anchor point rules. Then, a random forest model is used for motivation classification. Finally, structured inference is performed through transactional and leisure-oriented large language models. Explicit constraints and self-consistency checking mechanisms are introduced to ensure the controllability and stability of the inference.

Benefits of technology

It significantly improves the overall accuracy and robustness of fine-grained activity inference, especially optimizing performance in non-rigid activity recognition, enhancing adaptability to complex and multi-constraint scenarios, and supporting reliable fine-grained classification requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765431A_ABST
    Figure CN121765431A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of urban calculation and intelligent transportation, and discloses a hierarchical fine-grained human trajectory activity type inference method and related devices.The method comprises the steps that firstly, rigid activity types are recognized through predefined anchor point rules, and efficient labeling of regular segments is ensured; subsequently, the remaining staying segments mark transactional or leisure motivations through a motivation classification model to differentiate internal heterogeneity of non-rigid activities; and finally, performing structured inference on the marked fragment by a specific large language model of the motivator, and outputting a fine-grained non-rigid activity type. By the adoption of the method, the overall accuracy and robustness of fine-grained activity inference are remarkably improved, the performance is optimized especially on identification of non-rigid activities, adaptability to complex multi-constraint scenes is enhanced, and reliable fine-grained classification requirements are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of urban computing and intelligent transportation technology, and particularly relates to a hierarchical fine-grained method and related device for inferring human trajectory activity types. Background Technology

[0002] Semantic information about human movement trajectories (especially activity types such as staying at home, working, and dining) is core foundational data for transportation planning, urban governance, and public safety, playing a crucial role in supporting urban planning decisions, transportation system optimization, and risk assessment. Traditionally, these semantic attributes are primarily obtained through household travel surveys (HTS), but this method suffers from inherent drawbacks such as high collection costs, long cycles, and limited sample coverage, making it difficult to adapt to the current needs of refined and dynamic urban governance. With the development of information and communication technologies, large-scale passively collected trajectory data such as mobile phone signaling and GNSS trajectories have become an important supplement to HTS due to their advantages of strong continuity, high spatiotemporal resolution, and wide coverage; however, this type of data only contains spatiotemporal coordinate information and generally lacks high-level semantic labels such as activity types, making it difficult to directly serve urban analysis and decision-making. In recent years, Large Language Models (LLMs) have shown great potential in reasoning, decision-making, and understanding human behavior, providing a new direction for solving the semantic ambiguity problem of non-rigid activities. However, when used as an end-to-end black-box prediction model, it is prone to illusions and lacks controllable reasoning, making it difficult to ensure the reliability of fine-grained activity inference results. Meanwhile, urban planning and policy-making have continuously increased the requirements for semantic resolution of activity types. The activity type classification in travel survey forms has expanded to 15 categories or more, and coarse-grained classification can no longer meet the needs of urban complexity and governance precision.

[0003] Existing activity type inference technologies suffer from several significant drawbacks: First, they exhibit a singular technological paradigm, lacking effective coupling and collaboration mechanisms between different models or technologies, and lacking a unified inference framework. This results in various methods only functioning under specific conditions, making them unsuitable for complex and multi-constrained activity semantic inference tasks. Second, mainstream supervised learning methods are highly sensitive to data distribution, easily experiencing performance degradation in imbalanced categories and sparse sample scenarios, especially with a significant drop in accuracy in non-rigid activity recognition. Furthermore, the lack of explicit behavioral constraints and inference mechanisms can lead to inference results inconsistent with travel behavior logic in complex scenarios. Meanwhile, existing methods based on LL... Methods using M often employ a direct prompt-driven end-to-end reasoning model, lacking a structured reasoning process and effective inference constraints. This results in insufficient stability and controllability, and the semantic dependencies between activities are not adequately modeled, making it difficult to meet the inference needs of complex activities. Thirdly, the heterogeneity within non-rigid activities is not effectively modeled. Existing solutions typically treat non-rigid activities as a single category, ignoring the significant differences in sub-activity types such as indoor / outdoor and transactional / leisure activities. This prevents the model from characterizing the behavioral motivations, spatiotemporal features, and environmental dependencies of different sub-activities, further limiting the improvement of fine-grained activity type inference performance. In addition, existing methods generally suffer from "overstated accuracy." The overall accuracy mainly relies on rigid activities with strong regularity, such as "staying at home" and "going to work," while the recognition effect of non-rigid activities, which account for 60-80% of human daily behavior, is poor, making it difficult to support fine-grained classification needs.

[0004] It is evident that existing activity type inference techniques suffer from low accuracy, poor stability, and limited adaptability in fine-grained activity type inference, making it difficult to achieve reliable fine-grained activity type inference. Summary of the Invention

[0005] This invention provides a hierarchical fine-grained method and related apparatus for inferring human trajectory activity types. This method can effectively solve the problems of low accuracy, poor stability and limited adaptability of existing activity type inference techniques in fine-grained activity type inference. This method can achieve reliable fine-grained activity type inference.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A hierarchical, fine-grained method for inferring human trajectory activity types includes: Each dwell segment of the user's trajectory is judged one by one using predefined anchor point rules, and the dwell segments that match the anchor point rules are marked with the corresponding rigid activity type; wherein, the rigid activity type includes home activities, work activities and educational activities; The remaining stay segments after removing rigid activity types are input into a pre-built non-rigid activity motivation classification model, and the output stay segments are labeled with motivation tags; wherein, the motivation tags include transactional tags and leisure tags; The process involves inputting dwell segments labeled with motivational tags into a pre-built non-rigid activity type inference model, outputting the corresponding non-rigid activity type, and marking the dwell segments with motivational tags. The non-rigid activity type inference model is based on a large language model, including a transactional motivation large language model and a leisure motivation large language model. The transactional motivation large language model is used to filter corresponding dwell segments by transactional tags and infer the non-rigid activity type from these segments. The leisure motivation large language model is used to filter corresponding dwell segments by leisure tags and infer the non-rigid activity type from these segments.

[0007] Furthermore, the process employs predefined anchor point rules to evaluate each dwell segment of the user's trajectory, marking dwell segments that match the anchor point rules as corresponding rigid activity types. These anchor point rules include home activity identification rules, work activity identification rules, and educational activity identification rules, specifically defined as follows: Home activity identification rules: If a user stays at the same location for no less than a second preset hour within a first preset time period, and the location is surrounded by at least one residential interest point, and the user appears at the location at least a third preset number of days per week, then this location is identified as a residential anchor point, and the corresponding rigid activity type is the residential activity type. Work activity identification rules: If a user is marked as an employee and stays at the same location for no less than the fifth preset hour within the fourth preset time period, and appears at the location for at least the sixth preset number of days per week, then this location is identified as a work anchor point, and the corresponding rigid activity type is the work activity type. Educational activity identification rules: If a user is marked as a student and stays at the same location for no less than the eighth preset hour within the seventh preset time period, and appears at the location for at least the ninth preset number of days per week, and there is at least one school-related point of interest around the location, then this location is identified as an educational anchor point; the corresponding rigid activity type is the educational activity type.

[0008] Furthermore, the remaining dwell segments after removing rigid activity types are input into a pre-built non-rigid activity motivation classification model, which outputs dwell segments with motivational labels. The base model of the non-rigid activity motivation classification model is a random forest model, which includes a first random forest model and a second random forest model. The specific training process is as follows: Obtain historical dwell time segments and corresponding motivation tags for non-rigid activity types as historical data; Historical data is divided into weekday data and weekend data; The pre-constructed first random forest model is trained using weekday data, and the trained first random forest model is output. The pre-built second random forest model is trained using weekend data, and the trained second random forest model is output. The first random forest model is combined with the second random forest model to output a trained classification model for non-rigid activity motivation.

[0009] Further, the step of inputting the dwell segment with motivational tags into a pre-built non-rigid activity type inference model, outputting the non-rigid activity type corresponding to the dwell segment with motivational tags, and labeling the dwell segment with motivational tags includes: Input dwell segments with motivational labels into a pre-built non-rigid activity type inference model; By using motivation tags, the dwell time segments are automatically routed to the corresponding transaction motivation big language model or leisure motivation big language model; If the motivation label of the current segment is a transactional label, then in the set of candidate activities for transactional categories, the activity type with the highest probability is selected based on the conditional probability distribution of the transactional motivation big language model, and the activity type with the highest probability is output as a non-rigid activity type. If the motivation label of the current pause segment is a leisure label, then in the set of leisure candidate activities, the activity type with the highest probability is selected based on the conditional probability distribution of the leisure motivation big language model, and the activity type with the highest probability is output as a non-rigid activity type. Mark the pause segments with motivation tags as the corresponding non-rigid activity type.

[0010] Furthermore, the step of inputting dwell segments with motivational labels into a pre-built non-rigid activity type inference model, wherein the prompts during the inference process consist of a target constraint module, a feature contextualization module, and a few-shot prompt module; wherein: The target constraint module is used to explicitly limit the set of candidate activities for the current inference, allowing selection only from the corresponding motivation label group; The feature contextualization module is used to convert structured trajectory features into natural language, presenting travel time, spatial location, velocity status, and preceding and following activity chains in a contextualized manner. The few-sample prompting module is used to design a small number of prototype behavioral examples for each motivation label group to summarize typical travel patterns and guide the inference model of non-rigid activity types to calibrate the reasoning process.

[0011] Furthermore, the process of inputting dwell segments with motivational labels into a pre-built non-rigid activity type inference model also incorporates a self-consistency check mechanism, the specific process of which is as follows: By adjusting the temperature parameters of the inference model for non-rigid activity types, the inference was repeated three times for each dwell segment under different temperature settings, and the final activity type was determined by majority voting.

[0012] Furthermore, the remaining dwell segments after removing rigid activity types are input into a pre-constructed non-rigid activity motivation classification model, and the dwell segments with motivation labels are output. The non-rigid activity motivation classification model includes a multi-level motivation subclass classification structure, which can output dwell segments with motivation labels through a hierarchical classification method.

[0013] A hierarchical, fine-grained system for inferring human trajectory activity types, characterized by comprising: The first-level inference module is used to judge each stop segment of the user trajectory one by one using predefined anchor point rules, and to mark the stop segments that match the anchor point rules as the corresponding rigid activity type; wherein, the rigid activity type includes home activities, work activities and educational activities; The second-layer inference module is used to input the remaining stay segments after removing rigid activity types into a pre-built non-rigid activity motivation classification model, and output stay segments with motivation labels; wherein, the motivation labels include transaction type labels and leisure type labels; The third-layer inference module is used to input dwell segments with motivational tags into a pre-built non-rigid activity type inference model, output the non-rigid activity type corresponding to the dwell segments with motivational tags, and mark the dwell segments with motivational tags. The base model of the non-rigid activity type inference model adopts a large language model, including a transactional motivation large language model and a leisure motivation large language model. The transactional motivation large language model is used to filter corresponding dwell segments by transactional tags and infer the non-rigid activity type from these dwell segments. The leisure motivation large language model is used to filter corresponding dwell segments by leisure tags and infer the non-rigid activity type from these dwell segments.

[0014] A hierarchical, fine-grained human trajectory activity type inference device includes: Memory, used to store computer programs; A processor is used to implement the steps of the hierarchical fine-grained human trajectory activity type inference method described above when executing the computer program.

[0015] A computer-readable storage medium storing a computer program, which, when executed by a processor, is used to implement the steps of the above-described hierarchical fine-grained human trajectory activity type inference method.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a hierarchical, fine-grained method for inferring human trajectory activity types. First, predefined anchor point rules are used to identify rigid activity types, ensuring efficient labeling of regular segments. Then, remaining segments are labeled with transactional or leisure motivations using a motivation classification model to distinguish the internal heterogeneity of non-rigid activities. Finally, a motivation-specific large language model performs structured inference on the labeled segments, outputting fine-grained non-rigid activity types. This method integrates rules, machine learning, and LLM techniques through a hierarchical framework, achieving model collaboration and a unified inference process, overcoming the problem of a single technical paradigm. Motivation classification refines the sub-type differences of non-rigid activities, modeling behavioral motivations and feature dependencies, addressing the deficiency of ineffective handling of internal heterogeneity within activities. The LLM-based structured inference mechanism introduces explicit constraints, enhancing controllability and stability, and reducing errors inconsistent with travel logic. This method significantly improves the overall accuracy and robustness of fine-grained activity inference, especially optimizing performance in the identification of non-rigid activities, enhancing adaptability to complex, multi-constraint scenarios, and supporting reliable fine-grained classification requirements. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the overall structure of the three-layer activity type inference model provided in an embodiment of the present invention; Figure 2 A schematic diagram of rule-based rigid activity recognition provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the classification of non-rigid activity motivations based on machine learning, provided in an embodiment of the present invention. Figure 4 This is a schematic diagram illustrating the inference of non-rigid activity types based on LLM workflow provided in an embodiment of the present invention; Figure 5 The activity type transition matrix is ​​provided for the real data and inference results in the embodiments of the present invention; wherein, (a) is the real label; (b) is CRF; (c) is GPT-4; and (d) is V-LLM. Figure 6 The joint probability distribution of the activity start time and duration of the inference results provided in the embodiments of the present invention; Figure 7 A flowchart illustrating a hierarchical, fine-grained method for inferring human trajectory activity types, provided in an embodiment of the present invention; Figure 8This is a schematic diagram of a hierarchical fine-grained human trajectory activity type inference system provided in an embodiment of the present invention. Detailed Implementation

[0018] To further understand the content of this invention, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are merely illustrative and not limiting of the invention.

[0019] The technical terms involved in this invention are explained below: HTS stands for Household Travel Survey, which is a survey of family travel.

[0020] POI stands for Point of Interest.

[0021] MA: Mandatory Activity, also known as rigid activity.

[0022] NMA stands for Non-Mandatory Activity, which is a non-rigid activity.

[0023] LLM stands for Large Language Model.

[0024] GNSS stands for Global Navigation Satellite System.

[0025] NBC stands for Naive Bayes Classifier.

[0026] SVM stands for Support Vector Machine.

[0027] XGBoost: short for eXtreme Gradient Boosting.

[0028] ANN stands for Artificial Neural Network.

[0029] CRF stands for Conditional Random Field.

[0030] RF stands for Random Forest.

[0031] As mentioned in the background section, activity type information reflected in residents' travel behavior is a crucial foundation for understanding and characterizing human mobility patterns. With the widespread application of location-based trajectory data in urban research and the growing demand for refined urban decision-making, inferring fine-grained (or multi-category) activities from raw trajectory data lacking semantic annotation has become a key technical challenge in the field of human mobility analysis. Existing activity type inference techniques are mostly geared towards coarse-grained identification scenarios with a limited number of activity categories. They typically reduce the difficulty of modeling and inference by merging, simplifying, or classifying activity types. For example, multiple non-rigid activities are grouped into "other activities," or only a few representative or spatiotemporally regular activity types are used for inference (e.g., "pick-up and drop-off" activities typically occur in environments like train stations and airports). While these methods can achieve some success in overall accuracy evaluation, their applicability is mainly limited to coarse-grained activity analysis tasks. Existing activity type inference techniques can be mainly categorized as follows: The first type is rule-based activity type inference methods: Rule-based methods identify activity types by manually defining spatiotemporal discrimination logic, typically combining dwell time, time of occurrence, and spatial relationship with residence or workplace. For example, identifying long-staying points at night determines residential activities, while identifying repeated stays during daytime working hours determines work activities.

[0032] This type of method has a simple structure, clear logic, and a certain degree of interpretability, and can achieve high recognition accuracy for rigid activities such as staying at home and working when the rules are applicable. However, due to its heavy reliance on human experience and fixed threshold settings, it is difficult to adapt to complex and diverse travel behaviors, and its ability to recognize non-rigid activities is limited.

[0033] The second type is the activity type inference method based on probabilistic statistical models: Probabilistic statistical models estimate the probability of different activity types by explicitly modeling the uncertainty of activity occurrence. Typical methods include multinomial Logit models and nested Logit models. These methods typically treat activity type as a latent variable and combine it with contextual features such as time and spatial distance for probabilistic inference. While these models can characterize the randomness of activity occurrence to some extent, their feature representation capabilities are limited. They often rely on simplified statistical assumptions and struggle to fully utilize complex environmental semantics and behavioral contextual information, making them ill-suited for multi-category, fine-grained activity inference scenarios.

[0034] The third category is activity type inference methods based on machine learning: With the development of data-driven approaches, machine learning models have been widely introduced into activity type inference tasks. Unsupervised learning methods typically uncover latent behavioral patterns using Hidden Markov Models (HMMs) or representation learning methods in the absence of labeled data, and assign activity labels to trajectory segments based on spatiotemporal features. These methods can discover certain sequence structures, but their results are often difficult to interpret directly, and due to the lack of validation with real labels, their inference results can usually only be evaluated at the aggregation level.

[0035] Supervised and semi-supervised learning methods further utilize labeled activity data for training, with typical models including random forests, support vector machines, gradient boosting trees, and artificial neural networks. When labeled data is sufficient, these methods can achieve certain results in overall accuracy. However, these methods are highly sensitive to data distribution, and performance is prone to degradation under conditions of class imbalance and sparse samples, especially with a significant drop in accuracy for non-rigid activity identification. Furthermore, due to the lack of explicit behavioral constraints and inference mechanisms, these sampling-based models are prone to producing inferences inconsistent with basic travel behavior logic in complex scenarios, thus limiting their practical application value in fine-grained activity type inference tasks.

[0036] The fourth category involves attempts at activity-based semantic inference using large language models: With the development of large language models, some studies have attempted to leverage their semantic understanding and reasoning capabilities for activity type inference. For example, trajectory context can be converted into text and the large language model can directly output the activity type judgment. These methods show potential in commonsense understanding, but existing solutions often rely on direct prompt word invocation, lacking structured reasoning processes and inference constraints. This makes it difficult to guarantee the stability and controllability of the inference results, and they remain difficult to apply directly to complex inference scenarios. Furthermore, current methods based on large language models often employ end-to-end reasoning, neglecting the semantic dependencies between pairs of activities.

[0037] To achieve the above objectives, this embodiment provides a hierarchical fine-grained method for inferring human trajectory activity types. This method adopts the "divide and conquer" approach, introducing rule-based methods, machine learning classifiers, and large language model inference workflows at different levels. This breaks down the complex multi-category activity inference task into several semantically clear and less difficult sub-tasks, thereby achieving fine-grained inference of residents' activity types. It is particularly suitable for stable inference tasks involving non-rigid activities.

[0038] like Figure 7 As shown, this embodiment provides a hierarchical, fine-grained method for inferring human trajectory activity types, including: Each dwell segment of the user's trajectory is judged one by one using predefined anchor point rules, and the dwell segments that match the anchor point rules are marked with the corresponding rigid activity type; wherein, the rigid activity type includes home activities, work activities and educational activities; The remaining stay segments after removing rigid activity types are input into a pre-built non-rigid activity motivation classification model, and the output stay segments are labeled with motivation tags; wherein, the motivation tags include transactional tags and leisure tags; The process involves inputting dwell segments labeled with motivational tags into a pre-built non-rigid activity type inference model, outputting the corresponding non-rigid activity type, and marking the dwell segments with motivational tags. The non-rigid activity type inference model is based on a large language model, including a transactional motivation large language model and a leisure motivation large language model. The transactional motivation large language model is used to filter corresponding dwell segments by transactional tags and infer the non-rigid activity type from these segments. The leisure motivation large language model is used to filter corresponding dwell segments by leisure tags and infer the non-rigid activity type from these segments.

[0039] The inference method provided in this embodiment will be further explained below with reference to the accompanying drawings: In this inference method, to achieve effective multi-technology collaborative inference, the entire inference process is divided into three levels: The first layer: Based on rule-based methods, the activities corresponding to trajectory segments are distinguished as rigid or non-rigid activities, which is used to quickly identify activity types with stable spatiotemporal patterns. The second layer: For non-rigid activities, a weakly supervised machine learning model is introduced to classify the motivation of non-rigid activities based on trajectory features and environmental features, so as to reduce the space for subsequent semantic inference. The third layer: Based on the above results, a structured activity semantic space is constructed, and a motivation-guided LLM reasoning workflow is introduced to perform LLM-driven fine-grained activity type inference under explicit constraints and context conditions.

[0040] like Figure 1 As shown, to implement the above hierarchical inference steps, this embodiment proposes a three-layer activity type inference model that integrates a large language model, such as... Figure 1 As shown, Figure 1 This is the overall framework diagram of the three-layer activity type inference model. The model forms the core of the overall framework, used to decompose the original multi-category activity inference problem into three sub-tasks. The overall framework adopts a tree structure design, with data flowing sequentially from the root node to deeper levels (root node → first layer → second layer → third layer). Each layer corresponds to a clearly defined inference task. The overall process is as follows: The first step, the daily trajectory construction process, is the preprocessing stage: Before performing the three-layer inference task, the reconstructed trajectory data is first aggregated by user and further divided into daily trajectories organized by day. For each daily trajectory, the user's stop points are processed sequentially in chronological order, enabling activity type inference to be performed in a chain-like manner along the stop sequence within a single day. Specifically, the model first completes the activity type inference for each stop point (stop segment) for the day. After the trajectory for that user for that day is processed, it then iterates to infer the trajectory for the next day. Once all daily trajectory inferences for a user are completed, the model iterates to process the next user.

[0041] Step 2, Task 1: Determining Rigid and Non-rigid Activities (Root Node → First Layer): In the first-level inference, each dwell segment is judged according to predefined spatiotemporal rules. If the segment satisfies the criteria for rigid activity, it is directly marked as MA; otherwise, the segment is considered as non-rigid activity and is passed to the next level for further processing.

[0042] Step 3, Task 2: Classification of Non-Rigid Activity Motivations (First Level → Second Level): For non-rigid activity segments entering the second layer, a pre-trained classifier is used to further classify them into two motivation groups: "transactional" or "leisurely." At this stage, no specific activity type is determined; only motivation labels are generated, and activity segments with accompanying motivation information are passed to the third layer.

[0043] Step 4, Task 3: Fine-grained activity type inference (Second layer → Third layer): In the third-level inference, based on the motivation labels obtained in the previous level, the large language model inference workflow designed specifically for different motivation categories is invoked to perform fine-grained activity type inference in the corresponding candidate activity space, thereby determining the specific non-rigid activity category.

[0044] As can be seen, by adopting the above-mentioned hierarchical and progressive design approach, the model can gradually reduce the inference space and decrease the inference difficulty. The following is a further explanation of the three inference tasks mentioned above: Task 1: As Figure 2 As shown, rule-based rigid activity identification involves using predefined anchor point rules to evaluate each pause segment of the user's trajectory, and marking the pause segments that match the anchor point rules as corresponding rigid activity types. The specific steps are as follows: Rigid activities and non-rigid activities differ fundamentally in terms of human behavioral constraints and temporal flexibility. Existing research has shown that prioritizing the distinction between the two in the activity type inference process helps reduce the complexity of subsequent inferences and improve overall performance. Furthermore, in complex multi-intent inference scenarios, early elimination of activities with clear spatiotemporal patterns helps avoid instability in subsequent inference stages. Since rigid activities such as "staying at home," "going to work," and "education" typically have stable temporal patterns and spatial anchor features, they can be accurately identified using rule-based methods. Therefore, this study uses rule-based anchor identification as the first-layer screening mechanism. Specifically, potential spatial anchor locations are first identified in the user's complete trajectory; then, each stop point is judged sequentially according to time. Segments that match anchor rules are directly marked as the corresponding MA and proceed to the inference of the next segment; stop segments that do not match any anchor rules are considered NMA. After eliminating stop segments of rigid activity types, the remaining NMAs are passed to the next task for processing. In this way, effective separation of MA and NMA is achieved at the first layer.

[0045] The anchor point rules are as follows: (1) Rules for identifying home activities: If a user stays at the same location for no less than 4 hours continuously between 0:00 and 6:00, and there is at least one residential point of interest in the vicinity of the location, and the user appears at the location at least 3 days a week, then the location is identified as a residential anchor point. All activity segments that occur at this anchor point are marked as "home".

[0046] (2) Work activity identification rules: If a user is marked as an employee and stays at the same location for no less than 4 hours continuously between 7:00 and 17:00, and appears at that location at least 3 days a week, then that location is identified as a work anchor point. All activity segments that occur at this anchor point are marked as "work".

[0047] (3) Educational Activity Identification Rules: If a user is identified as a student and stays at the same location for no less than 4 hours continuously between 7:00 and 17:00, and appears at that location at least 3 days a week, and there is at least one school-related POI in the vicinity of that location, then that location is identified as an educational anchor point. All activity segments that occur at this anchor point are marked as "educational".

[0048] Task 2: Figure 3 As shown, the classification of non-rigid activity motivations based on machine learning involves using a pre-built non-rigid activity motivation classification model to infer the dwell segments corresponding to NMA and output dwell segments with motivation labels. The specific steps are as follows: In Task 1, Motive Actions (MAs) have been identified and stripped away, and the remaining endpoints (defaulting to Non-Motive Actions) enter the subsequent inference process, including this task. This task aims to reduce the semantic complexity within NMAs by dividing multi-class NMAs into subsets with higher semantic consistency, providing constraints for subsequent fine-grained inference. Specifically, this task categorizes all non-rigid activities into two main types based on differences in activity motivation: transactional activities (such as shopping, dining, business trips, medical visits, transportation, etc.) and leisure activities (such as entertainment, visiting, etc.). This motivational-level division effectively reduces the heterogeneity within NMAs, narrowing the subsequent inference space. In implementation, this task uses a Random Forest (RF) model to classify the motivations of non-rigid activities. As an ensemble learning method, Random Forest has good robustness to noise and changes in data distribution, making it suitable for the coarse-grained classification task in this stage. Considering the significant differences in activity behavior between weekdays and weekends, two independent RF models (i.e., the first Random Forest model and the second Random Forest model) are trained on weekday and weekend data respectively. During training, only NMA samples are used, and the original activity types are mapped to two motivational labels: "transactional" or "leisure," thus constructing a weakly supervised training set. After training, a dual random forest classifier is embedded in the overall inference process, outputting motivational labels for dwell segments identified as NMA in Task 1. This task, while explicitly modeling weekday-weekend differences, performs a structured classification of non-rigid activities at the motivational level, providing a more reliable inference space for subsequent fine-grained activity type inference based on a large language model. In this embodiment, the random forest model can be replaced by other models with similar classification capabilities (such as gradient boosting trees, artificial neural networks, etc.); the large language model used can also be replaced by other general or domain-specific large models with inference capabilities, as long as they participate in the inference process under structured constraints, all of which are optional implementation methods in this embodiment.

[0049] Task 3: (e.g.) Figure 4 As shown, non-rigid activity type inference based on LLM workflow is performed. Specifically, a pre-built non-rigid activity type inference model is used to infer the non-rigid activity type of dwell segments with motivational tags, outputting the corresponding non-rigid activity type and labeling the dwell segments with motivational tags. In Task 2, this embodiment divides the dwell points to be inferred into two semantically consistent inference spaces: transactional activities and leisure activities. Based on this, Task 3 further infers the specific activity type of the dwell segment within the corresponding semantic space. To this end, this embodiment proposes a "motivation-guided LLM inference workflow." Compared to common end-to-end LLM methods that rely solely on cue words, this workflow employs pre-structured inference steps, resulting in higher interpretability. This embodiment is the first to specifically utilize an LLM workflow for activity type inference research. The workflow adopts a parallel structure, consisting of two dedicated models matched with motivation labels: a transactional motivation LLM and a leisure motivation LLM. During this inference process, each dwell segment is automatically routed to the corresponding LLM based on its motivation label, ensuring the model always infers within a narrower, more semantically consistent candidate space. In each inference step, the model receives a set of spatiotemporal trajectory features of the dwell point and outputs the most probable activity type from the candidate activity set of the corresponding motivation group.

[0050] The specific inference rules are as follows: If the motivation label for this stop point is "transactional", then from the set of candidate activities for transactional activities, the activity type with the highest probability is selected based on the conditional probability distribution of the transaction motivation LLM. If the motivation label for the stop is "leisure", then from the set of leisure candidate activities, the activity type with the highest probability is selected based on the conditional probability distribution of the leisure motivation LLM.

[0051] The two LLMs perform inference within their respective candidate sets (in the pre-provided dataset, there are 6 transaction groups and 3 leisure groups). The clues used in the inference process consist of the following three core components: (1) Target constraint module: Target constraint is used to clearly limit the set of candidate activities for the current inference, and only allow selection within the corresponding motivation group to avoid misjudgment across motivation types.

[0052] (2) Feature Contextualization Module: The feature contextualization module transforms structured trajectory features into natural language descriptions, presenting information such as travel time, spatial location, velocity status, and preceding and following activity chains in a contextualized manner, rather than directly inputting numerical features, in order to enhance LLM's ability to understand behavioral semantics.

[0053] (3) Few sample prompting module: Design a small number of prototype behavior samples for each motivation group to summarize typical travel patterns and guide the model to calibrate the reasoning process, and solidify them into prompt words.

[0054] For example, in this embodiment, to further improve the stability of the inference results, a self-consistency check mechanism is introduced. Specifically, by adjusting the temperature parameter, the inference is repeated three times for each dwelling segment under different temperature settings, and the final activity type label is determined by majority voting. The finally determined activity label is assigned to the corresponding dwelling segment and used as a semantic context feature for the inference of subsequent event segments. The activity chain accumulates only within the single-day trajectory and does not cross days. It should be noted that the model processes all dwelling segments of a single user each day in chronological order until the semantic annotation of all the user's trajectories is completed. The entire inference workflow is implemented based on the n8n automation platform, and the large language model component is supported by ChatGPT-4.

[0055] Furthermore, this embodiment employs a three-layer hierarchical inference structure. However, in specific implementations, the number of layers can be adjusted according to the application scenario. For example, the motivation classification in the second layer can be further refined into multiple levels of motivation subclasses, or the rule discrimination in the first layer can be merged with the machine learning classification in the second layer. As long as the hierarchical inference idea of ​​"from coarse to fine, gradually narrowing the semantic space, and finally handing it over to a large model for reasoning" is still followed, it is an equivalent variation of this embodiment.

[0056] Therefore, this embodiment provides a hierarchical, fine-grained method for inferring human trajectory activity types, specifically a hierarchical activity type inference model for fine-grained activity type inference. Addressing the challenges of semantic ambiguity and high inference difficulty in multi-category, fine-grained activity type inference, this embodiment proposes a hierarchical, progressive activity type inference model. By decomposing the originally complex multi-category activity inference task into multiple semantically clear subtasks with progressively decreasing inference difficulty, and organizing the inference process in a tree structure, it achieves coarse-to-fine activity type recognition, significantly improving the stability and controllability of fine-grained activity inference. This method also proposes a multi-technology collaborative activity type inference framework, organically coupling rule-based methods, machine learning models, and large language models into an activity type inference framework. This allows different technologies to perform different inference functions according to their advantages within a unified process, thus overcoming the limitations of existing methods where a single technology paradigm works independently. It fully leverages the complementary advantages of multiple technologies in the activity type inference process, improving overall inference performance and applicability. Furthermore, this inference method is the first to introduce a large language model into the activity type inference task in the form of a structured reasoning workflow, rather than using it directly as an end-to-end prediction model. By constructing a constrained, branchable, and composable reasoning process, the large language model performs activity type inference within a reasonable semantic space, forming an inference paradigm that differs from the traditional prompt word invocation method. This reduces the illusion of large models and thus improves inference performance.

[0057] For example, a specific experiment was conducted on the inference method provided in this embodiment to compare and verify the performance of the method. Since the original trajectory data only contains spatiotemporal observation information and lacks activity semantic labels, this experiment uses HTS data as a reliable source of real labels to extract activity types. A certain city was selected as the study area. This city is a large city with a dense population and rich activity patterns, which can better verify the applicability of the proposed framework. The household travel survey data of this city used in this experiment was collected from November 2016 to January 2017, containing 3,850 activity records of 188 participants. Among them, 2,991 were rigid activities (including "at home", "at work", and "at school"), and 859 were non-rigid activities. The sample was divided into training set and test set at 70%:30%.

[0058] The survey identified 12 activity types, including home, work, education, dining, shopping, business trips, visiting friends, entertainment, picking up / dropping off others, medical visits, and other administrative and leisure activities, covering the main purposes of residents' daily travel. The study area was further divided into a 500m × 500m grid, and each stop point was mapped to a corresponding grid to obtain an area number. Examples of trajectories for two stop points (stop segments) are shown in Table 1. Table 1 shows examples of trajectories.

[0059] Geographical environmental characteristics are considered to have a significant impact on individual activity types; that is, residents' daily activities are often closely related to the land use structure of their area and the types of POIs in the surrounding area. Compared to land use data, POI data has a finer spatial resolution and classification granularity, and can more accurately represent location attributes and human activity characteristics. Currently, POI data can be obtained from various channels such as online maps, navigation applications, location-based service (LBS) platforms, and social networks, including Google Maps, Amap, and Foursquare. This experiment used POI data for this city obtained from the Amap open platform, with data collection time at the end of September 2018. The total number exceeded 1.7 million. Amap uses a three-level classification system for POIs, dividing them into 13 major categories, 153 intermediate categories, and 488 subcategories.

[0060] To systematically evaluate the effectiveness of this inference method, this experiment selected several representative baselines. First, in terms of traditional methods, seven models were selected: rule-based methods, Naive Bayes classifier (NBC), Support Vector Machine (SVM), Random Forest (RF), Extreme Gradient Boosting Tree (XGBoost), Artificial Neural Network (ANN), and Conditional Random Field (CRF). These models are widely used in existing activity inference research and have shown strong baseline performance on different datasets, representing the "statistical learning / traditional machine learning" paradigm. Second, to further examine the performance gain of this framework in large-scale model inference scenarios, two types of LLM inference baselines were constructed: (1) Vanilla-LLM (V-LLM): Directly inputs the daily trajectory and related features into the general large model to generate the complete activity chain of the user in one go, without designing additional inference strategies, to reflect the inference capability of the original LLM; (2) Vanilla-LLM+Chain: Based on Vanilla-LLM, a chain inference process is introduced. This involves taking an event segment as input, inferring the next activity segment by segment in chronological order, and updating it at each step using semantic information from previous activities. Additionally, to analyze the framework's scalability across different LLMs, three variants were implemented: Ours–Llama3-8B, Ours–GPT-3.5-turbo, and Ours–GPT-4-turbo.

[0061] In terms of evaluation metrics, this embodiment evaluates from the following three perspectives: (1) Classification metrics. Used to measure the overall accuracy of multi-class activity recognition and the ability to distinguish between different classes. Four metrics are used: accuracy, macro-precision, macro-recall, and macro-F1 score. Accuracy reflects the overall prediction accuracy. Macro metrics aggregate the precision, recall, and F1 scores of each class at the class level, which can more fairly evaluate the recognition performance of different activity types in the case of class imbalance. (2) Sequence metrics. To measure the semantic coherence and pattern similarity of the inferred activity sequence, Bilingual Evaluation Understudy (BLEU) is used to compare the degree of overlap between the inferred activity chain and the real activity chain in n-gram (BLEU-2, i.e., binary tuples, is used in this embodiment). This metric can characterize whether the "combination of two adjacent activities" is consistent with the real behavior pattern and is used to reflect the generation quality of the model at the activity sequence level. (3) Distribution metrics. At a more macro level, to assess whether the inferred results can reproduce the approximate proportions of various activities in real cities, the Activity Distribution (ActDis) similarity is used. Specifically, the overall proportions of 12 activity categories in the real data and the inferred results are statistically analyzed, forming two 12-dimensional probability distribution vectors. The difference between the two is measured using the Jensen-Shannon divergence. The smaller the ActDis value, the closer the inferred results are to the activity patterns of the real population in terms of "overall structure and intensity distribution."

[0062] As shown in Table 2, the model performance is first evaluated from an overall perspective. The three-layer inference framework (inference method) proposed in this embodiment achieves an accuracy of nearly 0.8 on the overall metrics, which is about 12% higher than the best-performing traditional method, CRF (0.71). Simultaneously, this framework achieves the best results on all metrics, including macro-precision, recall, F1, BLEU, and ActDis, indicating significant advantages in classification accuracy, activity sequence consistency, and overall distribution alignment. The overall accuracy of traditional machine learning models (such as SVM, RF, XGBoost, and ANN) is generally around 0.65, while CRF is slightly better due to its effective utilization of sequence structure. Interestingly, NBC performs even worse than rule-based methods due to its overly strong assumption of feature independence. Furthermore, the baseline performance of the two large-scale models is unsatisfactory: V-LLM is only 0.43, which improves to 0.49 after introducing chained inference, but is still lower than most statistical learning models, indicating that general large-scale models relying solely on cue word engineering are difficult to effectively perform this type of activity semantic inference task.

[0063] Further observation of the NMA metrics reveals a more significant improvement using our proposed inference method. Traditional baselines perform almost at a random level (approximately 0.10) on NMA, reflecting the difficulty of capturing the behavioral complexity of non-rigid activities in scenarios with limited samples. While the LLM baseline shows some improvement, V-LLM is only 0.18, and chained inference only reaches 0.26, indicating that although incremental inference can alleviate semantic drift, the current capabilities of LLM are still insufficient to support reliable end-to-end inference. In contrast, our framework achieves an accuracy of 0.5076 on NMA, significantly outperforming all baselines, more than four times the performance of the best traditional model (ANN), and nearly twice that of the best LLM baseline. This framework leads in precision, recall, F1, BLEU, and ActDis in NMA, fully demonstrating its advantages in handling complex, context-dependent NMA. Furthermore, among different large model variants, GPT-4 and GPT-4-turbo performed closest to each other and significantly outperformed GPT-3.5-turbo, validating the value of stronger language models in complex semantic inference. In contrast, the lightweight open-source model LLaMA3-8B performed significantly worse, indicating that large models with insufficient reasoning capabilities are ill-suited for fine-grained activity semantic judgment. The overall performance comparison between the models and non-rigid activity inference is shown in Table 2. In Table 2, the best results are marked in bold, and the second-best results are indicated by underline: Table 2 shows the comparison results of the overall model inference performance and non-rigid activity inference performance.

[0064] To test the model's ability to characterize behavioral structure at the sequence level, a comparative analysis was conducted on real activities, the model itself (based on GPT-4), the best traditional baseline (CRF), and the adjacent event activity transition patterns generated by V-LLM. Figure 5As shown, the transition matrix of real data exhibits a distinct high-density region on the left, primarily corresponding to transitions from various activities to "home" or "work," reflecting a typical "home-work-dominated" urban daily activity chain structure. Simultaneously, various NMAs (Non-Main Activity Modules) exhibit a clear diagonal structure in their own transition patterns, indicating a certain degree of persistence in individual NMA participation, meaning that the same type of activity tends to repeat itself. The CRF model significantly reinforces commuting transitions from "home" to "work," causing many low-frequency NMA transitions to be ignored, demonstrating its limited ability to characterize behavioral diversity. In contrast, the transition structure presented by this model is closer to the real data, accurately capturing both major commuting flows and preserving low-frequency non-commuting transitions. This indicates that the proposed framework can simultaneously represent routine behaviors and fine-grained activity dependencies. Furthermore, it was found that V-LLM significantly underestimates key commuting flows, indicating that relying solely on large models for end-to-end inference is insufficient to capture the temporal dependency structure between activities, further highlighting the limitations of LLM relying solely on simple cue word strategies in activity type inference tasks.

[0065] To further verify the temporal rationality of the model's inference behavior, a joint distribution analysis was conducted on the start time and duration of various activities in the inference results, specifically as follows: Figure 6 As shown in the figure, the model exhibits a pattern highly consistent with real-life daily behavior in terms of time dimension. For MA (Activity-Based Activity), the start time of "work" is mainly concentrated between 7:00 and 10:00, lasting an average of about 9 hours, which closely matches the typical commuting and office time structure. The start time of "staying at home" is more dispersed, mostly between 17:00 and 21:00, and its duration is usually 12–14 hours, corresponding to the lifestyle of returning home from get off work and staying until the next morning. For NMA (Non-Activity-Based Activity), "mealing" mostly occurs between 11:00 and 14:00, and its duration is usually less than 2 hours, consistent with the short duration of daily dining activities. The durations of "shopping," "picking up / dropping off," and "entertainment" are mostly between 2 and 3 hours, indicating that these activities are usually completed within a limited time window. It is noteworthy that the duration distribution of "other-leisure" shows a significant long-tail characteristic, sometimes reaching up to 20 hours, which may correspond to full-day outings on weekends or holidays. Furthermore, the start time and duration of both "official business" and "visiting friends" exhibit significant variability, reflecting their greater behavioral flexibility and heterogeneity. Overall, these distribution characteristics closely match real-life behavioral patterns, indicating that this model can not only identify activity types but also generate results consistent with real-world behavioral logic in the time dimension, thereby further enhancing the interpretability of the inference results.

[0066] like Figure 8As shown, this embodiment also provides a hierarchical fine-grained human trajectory activity type inference system, including: a first-layer inference module, used to judge each dwell segment of the user trajectory one by one using predefined anchor point rules, and to mark the dwell segments that match the anchor point rules as corresponding rigid activity types; wherein, the rigid activity types include home activities, work activities, and educational activities; a second-layer inference module, used to input the remaining dwell segments after removing rigid activity types into a pre-built non-rigid activity motivation classification model, and output dwell segments with motivation labels; wherein, the motivation labels include transactional labels and leisure labels; a third-layer inference module, used to classify dwell segments with motivation labels into non-rigid activity types. The dwell segments tagged with motivation are input into a pre-built non-rigid activity type inference model, which outputs the non-rigid activity type corresponding to the dwell segments with motivation tags and marks the dwell segments with motivation tags. The non-rigid activity type inference model is based on a large language model, including a transaction motivation large language model and a leisure motivation large language model. The transaction motivation large language model is used to filter corresponding dwell segments by transaction type tags and infer the non-rigid activity type from these dwell segments. The leisure motivation large language model is used to filter corresponding dwell segments by leisure type tags and infer the non-rigid activity type from these dwell segments.

[0067] The present invention also provides a hierarchical fine-grained human trajectory activity type inference device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of the hierarchical fine-grained human trajectory activity type inference method.

[0068] The present invention also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the hierarchical fine-grained human trajectory activity type inference method.

[0069] When the processor executes the computer program, it implements the above-mentioned hierarchical fine-grained human trajectory activity type inference steps, for example: using predefined anchor point rules to judge each dwell segment of the user trajectory one by one, and marking the dwell segments that match the anchor point rules as corresponding rigid activity types; wherein, the rigid activity types include home activities, work activities, and educational activities; inputting the remaining dwell segments after removing rigid activity types into a pre-built non-rigid activity motivation classification model, and outputting dwell segments with motivation labels; wherein, the motivation labels include transactional labels and leisure labels; inputting the dwell segments with motivation labels into a pre-built non-rigid activity type inference model, outputting the non-rigid activity types corresponding to the dwell segments with motivation labels, and marking the dwell segments with motivation labels; wherein, the basic model of the non-rigid activity type inference model adopts a large language model, including a transactional motivation large language model and a leisure motivation large language model; the transactional motivation large language model is used to filter the corresponding dwell segments by transactional labels and infer the dwell segments to output non-rigid activity types; the leisure motivation large language model is used to filter the corresponding dwell segments by leisure labels and infer the dwell segments to output non-rigid activity types.

[0070] Exemplarily, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing preset functions, the instruction segments describing the execution process of the computer program in the hierarchical fine-grained human trajectory activity type inference device. For example, the computer program can be divided into a first-layer inference module, a second-layer inference module, and a third-layer inference module, with the following specific functions: The first-layer inference module is used to judge each dwell segment of the user trajectory one by one using predefined anchor point rules, and to mark the dwell segments that match the anchor point rules as corresponding rigid activity types; wherein, the rigid activity types include home activities, work activities, and educational activities; The second-layer inference module is used to input the remaining dwell segments after removing rigid activity types into a pre-built non-rigid activity motivation classification model, and output dwell segments with motivation labels; wherein, the motivation labels include transactional labels and leisure labels; The third-layer inference module… This method inputs dwell segments labeled with motivational tags into a pre-built non-rigid activity type inference model, outputs the non-rigid activity type corresponding to the dwell segments labeled with motivational tags, and marks the dwell segments labeled with motivational tags. The non-rigid activity type inference model is based on a large language model, including a transactional motivation large language model and a leisure motivation large language model. The transactional motivation large language model is used to filter corresponding dwell segments by transactional tags and infer the non-rigid activity type from these dwell segments. The leisure motivation large language model is used to filter corresponding dwell segments by leisure tags and infer the non-rigid activity type from these dwell segments.

[0071] The hierarchical fine-grained human trajectory activity type inference device can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The hierarchical fine-grained human trajectory activity type inference device may include, but is not limited to, processors and memory. Those skilled in the art will understand that the above are examples of a hierarchical fine-grained human trajectory activity type inference device and do not constitute a limitation on the hierarchical fine-grained human trajectory activity type inference device. It may include more components than described above, or combine certain components, or different components. For example, the hierarchical fine-grained human trajectory activity type inference device may also include input / output devices, network access devices, buses, etc.

[0072] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or any conventional processor. This processor is the control center of the hierarchical fine-grained human trajectory activity type inference device, connecting various parts of the device via various interfaces and lines.

[0073] The memory can be used to store the computer program and / or modules. The processor implements various functions of the hierarchical fine-grained human trajectory activity type inference device by running or executing the computer program and / or modules stored in the memory and calling the data stored in the memory.

[0074] The memory may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function (such as sound playback, image playback, etc.). The data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices.

[0075] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the hierarchical fine-grained human trajectory activity type inference method.

[0076] If the modules / units integrated in the hierarchical fine-grained human trajectory activity type inference system are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0077] Based on this understanding, the present invention can implement all or part of the processes in the above-described hierarchical fine-grained human trajectory activity type inference method, or it can be accomplished by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above-described hierarchical fine-grained human trajectory activity type inference method. The computer program includes computer program code, which can be in the form of source code, object code, executable file, or a preset intermediate form, etc.

[0078] The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0079] It should be noted that the content contained in the computer-readable storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0080] In summary, this method has the following significant advantages compared to traditional inference methods: First, this method is applicable to more granular and multi-category activity type inference. Existing methods are mostly geared towards coarse-grained identification scenarios with a small number of activity categories, typically reducing the inference difficulty by merging or simplifying activity types, which is insufficient to effectively distinguish complex non-rigid activities. This invention, through a hierarchical and progressive inference design that coordinates multiple technical paradigms, can support more granular and multi-category activity type inference, and has better applicability in complex activity scenarios.

[0081] Secondly, it exhibits strong generalization ability and low dependence on data distribution. This invention relies on the reasoning capabilities of a large model, requiring only a small number of sample prompts to complete inference. Compared to mainstream supervised learning methods, this invention does not require a fully labeled dataset, has low dependence on dataset distribution, and thus improves generalization ability.

[0082] Third, the inference process is constrained, reducing the risk of hallucinations. Compared to existing simple prompt-based inference methods based on large language models, this invention introduces large language models into a structured reasoning workflow and completes the activity type determination under the constraints of preceding inference results. This reduces inference uncertainty and the risk of hallucinations caused by large models, thereby improving the controllability of inference results and the rationality of behavior.

[0083] The above embodiments are merely one of the implementation methods for achieving the technical solution of the present invention. The scope of protection claimed by the present invention is not limited to this embodiment, but also includes any variations, substitutions and other implementation methods that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the present invention.

Claims

1. A hierarchical fine-grained human trajectory activity type inference method, characterized in that, The application comprises the following steps: Each stay segment of the user trajectory is judged one by one by using a predefined anchor point rule, and the stay segment that hits the anchor point rule is marked with the corresponding rigid activity type; wherein the rigid activity type comprises home activity, work activity and education activity; The remaining stay segments after excluding the rigid activity type are input into a pre-constructed non-rigid activity motivation classification model to output stay segments with motivation labels; wherein the motivation labels comprise transactional labels and leisure labels; The stay segments with motivation labels are input into a pre-constructed non-rigid activity type inference model to output the non-rigid activity type corresponding to the stay segments with motivation labels, and the stay segments with motivation labels are marked; wherein the base model of the non-rigid activity type inference model adopts a large language model, comprising a transaction motivation large language model and a leisure motivation large language model; the transaction motivation large language model is used to filter the corresponding stay segment through the transactional label and infer the stay segment to output the non-rigid activity type; the leisure motivation large language model is used to filter the corresponding stay segment through the leisure label and infer the stay segment to output the non-rigid activity type.

2. The hierarchical fine-grained human trajectory activity type inference method of claim 1, wherein, In the step of judging each stay segment of the user trajectory one by one by using a predefined anchor point rule, and marking the stay segment that hits the anchor point rule with the corresponding rigid activity type, the anchor point rule comprises a home activity identification rule, a work activity identification rule and an education activity identification rule, which are specifically defined as follows: Home activity identification rule: if a user continuously stays in the same location for not less than a second preset hour within a first preset time period, and the surrounding area of the stay location includes at least one residential interest point, and the user appears at the stay location at least a third preset number of days per week, then the stay location is identified as a residential anchor point, and the corresponding rigid activity type is a residential activity type; Work activity identification rule: if a user is labeled as an employee, and continuously stays in the same location for not less than a fifth preset hour within a fourth preset time period, and appears at the stay location at least a sixth preset number of days per week, then the stay location is identified as a work anchor point, and the corresponding rigid activity type is a work activity type; Education activity identification rule: if a user is labeled as a student, and continuously stays in the same location for not less than an eighth preset hour within a seventh preset time period, and appears at the stay location at least a ninth preset number of days per week, and at least one school interest point exists in the surrounding area of the stay location, then the stay location is identified as an education anchor point; and the corresponding rigid activity type is an education activity type.

3. The hierarchical fine-grained human trajectory activity type inference method of claim 1, wherein, In the step of inputting the remaining stay segments after excluding the rigid activity type into a pre-constructed non-rigid activity motivation classification model to output stay segments with motivation labels, the base model of the non-rigid activity motivation classification model is a random forest model, and the non-rigid activity motivation classification model comprises a first random forest model and a second random forest model, and the specific training process is as follows: Obtain historical stay segments of non-rigid activity types and corresponding motivation labels as historical data; Divide the historical data into weekday data and weekend data; training the first random forest model pre-constructed by using the weekday data, and outputting the trained first random forest model; training the second random forest model pre-constructed by using the weekend data, and outputting the trained second random forest model; combining the first random forest model and the second random forest model, and outputting the trained non-rigid activity motivation classification model.

4. The hierarchical fine-grained human trajectory activity type inference method of claim 1, wherein, The method for inputting the stay segment with the motivation label into the pre-constructed non-rigid activity type inference model and outputting the non-rigid activity type corresponding to the stay segment with the motivation label comprises: inputting the stay segment with the motivation label into the pre-constructed non-rigid activity type inference model; routing the stay segment to the corresponding transaction motivation large language model or leisure motivation large language model through the motivation label; if the motivation label of the current stay segment is a transaction label, selecting the activity type with the maximum probability in the transaction candidate activity set based on the conditional probability distribution of the transaction motivation large language model, and outputting the activity type with the maximum probability as the non-rigid activity type; if the motivation label of the current stay segment is a leisure label, selecting the activity type with the maximum probability in the leisure candidate activity set based on the conditional probability distribution of the leisure motivation large language model, and outputting the activity type with the maximum probability as the non-rigid activity type; labeling the stay segment with the motivation label as the corresponding non-rigid activity type.

5. The hierarchical fine-grained human trajectory activity type inference method of claim 4, wherein, The prompt words in the inference process are composed of a target constraint module, a feature contextualization module and a few-shot prompt module; wherein: The target constraint module is used to explicitly limit the candidate activity set of the current inference, and only allows selection within the corresponding motivation label group; The feature contextualization module is used to convert structured trajectory features into natural language, and presents the travel time, spatial position, speed state and before-after activity chain in a situational way; The few-shot prompt module is used to design a small number of prototype behavior samples for each motivation label group to summarize typical travel patterns and guide the non-rigid activity type inference model to calibrate the reasoning process.

6. The layered fine-grained human trajectory activity type inference method of claim 4, wherein, The self-consistency checking mechanism is also introduced in the inference process, and the specific process of the self-consistency checking mechanism is as follows: By adjusting the temperature parameter of the non-rigid activity type inference model, the inference is repeated three times under different temperature settings for each stay segment, and the majority voting method is used to determine the final activity type.

7. The hierarchical fine-grained human trajectory activity type inference method of claim 1, wherein, The non-rigid activity motivation classification model includes a multi-level motivation sub-class classification structure, which can output the stay segment with the motivation label through layer-by-layer classification.

8. A hierarchical fine-grained human trajectory activity type inference system, characterized in that, Comprise: The first layer inference module is configured to determine each stay segment of the user trajectory one by one by using a predefined anchor point rule, and mark the stay segment that meets the anchor point rule as a rigid activity type; wherein the rigid activity type includes a home activity, a work activity and an education activity. The second layer inference module is configured to input the remaining stay segments after excluding the rigid activity type into a pre-constructed non-rigid activity motivation classification model, and output the stay segments with motivation labels; wherein the motivation labels include a transaction type label and a leisure type label. The third layer inference module is configured to input the stay segments with motivation labels into a pre-constructed non-rigid activity type inference model, output the non-rigid activity type corresponding to the stay segments with motivation labels, and mark the stay segments with motivation labels; wherein the non-rigid activity type inference model is based on a large language model, including a transaction motivation large language model and a leisure motivation large language model; the transaction motivation large language model is configured to filter the corresponding stay segment through the transaction type label and infer the stay segment to output the non-rigid activity type; and the leisure motivation large language model is configured to filter the corresponding stay segment through the leisure type label and infer the stay segment to output the non-rigid activity type.

9. A hierarchical fine-grained human trajectory activity type inference device, characterized by, The computer program is executed by the processor to implement the steps of the hierarchical fine-grained human trajectory activity type inference method according to any one of claims 1-7. The computer program is executed by the processor to implement the steps of the hierarchical fine-grained human trajectory activity type inference method according to any one of claims 1-7. ​ 10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. ​