Individual non-commuting activity identification method based on multi-source data fusion based on spatiotemporal big data
By integrating multi-source data and advanced algorithms, it can identify non-commuting activities of urban residents, solving the problems of low recognition accuracy and efficiency in existing technologies, achieving more comprehensive and efficient non-commuting activity analysis, and supporting urban planning and traffic management.
Patent Information
- Application Number
- CN202411236015.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-09-04
AI Technical Summary
When identifying non-commuting activities of urban residents, existing technologies have problems such as insufficient analysis of non-commuting activities, low data processing accuracy and efficiency, and lack of comprehensive consideration of individual attributes, activity attributes and built environment characteristics.
By integrating individual spatiotemporal trajectory data, resident travel survey data, urban point of interest data, and urban spatial raster data, and using advanced data mining and machine learning algorithms, we extract personal attribute characteristics, daily activity attribute characteristics, and built environment characteristics, train and test multiple models, and ultimately select the best-performing model for non-commuting activity identification.
It improves the accuracy and comprehensiveness of non-commuting activity identification, enhances the efficiency and intelligence of data processing, and provides deeper insights for urban planning and traffic management.
Smart Images

Figure CN119357765B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to, in particular to, a method for identifying non-commuting activities of individuals based on multi-source data fusion using spatiotemporal big data. Background Art
[0002] In today's rapidly developing urbanization, the growth of urban populations and the diversification of their activity patterns pose new challenges to urban planning and traffic management. The daily activities of urban residents, particularly non-commuting activities such as shopping and entertainment, are crucial for understanding the functional layout of cities and their quality of life. However, traditional resident travel survey methods, limited by sample size and data collection limitations, struggle to fully capture the complex activity patterns of urban residents.
[0003] With the rise of big data technologies, individual spatiotemporal data provides a new perspective for analyzing urban activity patterns. This data, with its extensive spatiotemporal coverage and representativeness of a diverse range of demographic attributes, can more realistically reflect the dynamics of urban residents' activities. However, accurately identifying non-commuting activity categories from this massive, unlabeled data remains a technical challenge.
[0004] The existing technology has the following deficiencies when dealing with such problems:
[0005] 1) Focusing on the identification of major activities such as commuting, while neglecting the analysis of non-commuting activities.
[0006] 2) In large-scale data processing, the accuracy and efficiency of existing methods need to be improved.
[0007] 3) Lack of comprehensive consideration of individual attributes, activity attributes and built environment characteristics. Summary of the Invention
[0008] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a method for identifying individual non-commuting activities based on multi-source data fusion based on individual spatiotemporal big data. By fusing multiple data sources, including but not limited to individual spatiotemporal trajectory data, resident travel survey data, urban point of interest (POI) data and urban spatial raster data, and using advanced data mining and machine learning algorithms, the method can effectively identify and classify individual non-commuting activities, thereby improving the efficiency and accuracy of identification.
[0009] The purpose of the present invention can be achieved by the following technical solutions:
[0010] The present invention provides a method for identifying non-commuting activities based on individual spatiotemporal big data and multi-source data fusion, comprising the following steps: S1, based on preset urban raster data, obtaining individual spatiotemporal big data containing raster spatial information, resident travel survey data and urban POI data; S2, based on the individual spatiotemporal big data, extracting a first personal attribute feature and a first single-day activity attribute feature; S3, based on the resident travel survey data, extracting a second personal attribute feature and a second single-day activity attribute feature; S4, based on the individual spatiotemporal big data and the urban POI data, extracting a first built environment feature, and based on the resident travel survey data and the urban POI data, extracting a second built environment feature; S5, merging the second personal attribute feature and the second single-day activity attribute feature into a single-day activity attribute feature. Preprocess the daily activity attribute characteristics and the second built environment characteristics; S6, based on the preprocessed second personal attribute characteristics, the second single-day activity attribute characteristics and the second built environment characteristics, train and test multiple non-commuting activity recognition models to be evaluated; S7, based on the test results, select the non-commuting activity recognition model to be evaluated with the best performance as the final individual spatiotemporal big data non-commuting activity recognition model; S8, preprocess the first personal attribute characteristics, the first single-day activity attribute characteristics and the first built environment characteristics; S9, based on the preprocessed first personal attribute characteristics, the first single-day activity attribute characteristics and the first built environment characteristics, use the individual spatiotemporal big data non-commuting activity recognition model to obtain the recognition result of non-commuting activities in the individual spatiotemporal big data.
[0011] As an optimal technical solution, S1 specifically includes: obtaining initial individual spatiotemporal big data, initial resident travel survey data and initial city POI data; spatially connecting the spatial information of user activity points in the initial individual spatiotemporal big data with the spatial information of the city raster data to obtain individual spatiotemporal big data containing raster spatial information; spatially connecting the spatial information of respondent activity points in the initial resident travel survey data with the spatial information of the city raster data to obtain resident travel survey data containing raster spatial information; spatially connecting the spatial information of city POI points in the initial city POI data with the spatial information of the city raster data to obtain city POI data containing raster spatial information.
[0012] As an optimal technical solution, the S2 specifically includes: based on individual spatiotemporal big data containing raster spatial information, directly extracting the first personal attribute feature, and simultaneously obtaining the user ID, activity start date, activity end date, activity start time and activity end time, the first personal attribute feature including the user's personal age, gender, and occupation; sorting according to the user ID, the activity start date, the activity end date, the activity start time and the activity end time to obtain the single-day activity sequence information of each user; based on the single-day activity sequence information, calculating the first single-day activity attribute feature, the first single-day activity attribute feature including the start time of this activity, the duration of this activity, whether the last activity was at home, whether the last activity was at work, the straight-line distance from home, the straight-line distance between home and workplace, whether the activity start date is a weekday, and the order of this activity in the resident's day.
[0013] As an optimal technical solution, the S3 specifically includes: based on individual spatiotemporal big data containing raster spatial information, directly extracting the second personal attribute feature, and simultaneously obtaining the respondent ID, activity start date, activity end date, activity start time and activity end time, the second personal attribute feature includes the respondent's personal age, gender, and occupation; sorting according to the respondent ID, the activity start date, the activity end date, the activity start time and the activity end time to obtain the single-day activity sequence information of each respondent; based on the single-day activity sequence information, calculating the second single-day activity attribute feature, the second single-day activity attribute feature includes the start time of this activity, the duration of this activity, whether the last activity was at home, whether the last activity was at work, the straight-line distance from home, the straight-line distance between home and workplace, whether the activity start date is a weekday, and the order of this activity in the resident's day.
[0014] As a preferred technical solution, S4 specifically includes: calculating the POI density of the grid where each POI type is located based on urban POI data containing raster spatial information; directly extracting the grids where the activity place, residence place and work place are located based on individual spatiotemporal big data containing raster spatial information and resident travel survey data containing raster spatial information; associating the POI density with the corresponding grids where the activity place, residence place and work place are located to obtain the first built environment feature and the second built environment feature, respectively.
[0015] As an optimal technical solution, the S8 specifically includes: merging the first personal attribute feature, the first single-day activity attribute feature and the first built environment feature; normalizing the first personal attribute feature, the first single-day activity attribute feature and the first built environment feature, and using the normalized data as the first feature set; obtaining the travel activity category field, and using the travel activity category field as the field to be inferred, merging it into the first feature set to obtain a prediction set.
[0016] As a preferred technical solution, S5 specifically includes: merging the second personal attribute characteristics, the second single-day activity attribute characteristics and the second built environment characteristics; normalizing the second personal attribute characteristics, the second single-day activity attribute characteristics and the second built environment characteristics, and using the normalized data as the second feature set; directly extracting non-commuting activities from the resident travel survey data as a second label set, and encoding them; merging the second feature set and the second label set as the data set of the multiple non-commuting activity recognition models to be evaluated, and dividing the data set into a training set and a test set.
[0017] As a preferred technical solution, the expression for the normalization process is:
[0018]
[0019] Where, Represents the first i Features, Indicates the i feature data, Indicates the i The maximum value of the feature data set, Indicates the i The minimum value of the feature dataset.
[0020] As a preferred technical solution, S6 specifically includes: inputting the training set into a plurality of pre-acquired non-commuting activity recognition models to be evaluated for training, wherein the plurality of non-commuting activity recognition models to be evaluated include an SVM model, a DT model, a RF model, a GBDT model, an ADAboost model, an XGBoost model, and a Transformer-based classification model.
[0021] As a preferred technical solution, the best performance judgment indicators include accuracy and Scores include:
[0022]
[0023]
[0024]
[0025]
[0026]
[0027] Where, is the predicted value of the label set, is the true value of the label set, Indicates whether the true value is predicted correctly, represents the test set, represents the number of test set samples, Indicates the accuracy, is the precision rate, For recall rate, The prediction result is the number of samples predicted correctly in one of the labels. The prediction result is the number of samples with incorrect predictions in one of the labels. The number of samples whose prediction result is not one of the labels is incorrectly predicted. express Fraction.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] 1. This method comprehensively utilizes multi-source data to extract multiple features, including personal attributes, daily activity attributes, and built environment characteristics. Combined with data fusion and preprocessing based on raster spatial information, the preprocessed features are incorporated into various mathematical models to identify non-commuting activities. This allows for more comprehensive and sophisticated data processing, thereby improving the accuracy and comprehensiveness of non-commuting activity identification and providing deeper insights for urban planning and traffic management.
[0030] 2. This invention adopts advanced mathematical models and algorithms, and utilizes preprocessed second personal attribute features, second daily activity attribute features, and second built environment features to train and test multiple non-commuting activity recognition models to be evaluated. The model with the best performance is then selected as the final non-commuting activity recognition model. This ensures optimal non-commuting activity recognition results in different scenarios while improving the efficiency and intelligence of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A schematic diagram of a basic flow chart of a method provided in an embodiment of the present invention;
[0032] Figure 2 This is a diagram showing the training results of various non-commuting activity recognition models to be evaluated provided in an embodiment of the present invention;
[0033] Figure 3A heat map showing the distribution of shopping and entertainment activity times of individual spatiotemporal big data users in 2023 (weekdays) provided in an embodiment of the present invention.
[0034] Figure 4 This is a heat map (weekdays) showing the distribution of other activities of individual spatiotemporal big data users in 2023, provided in an embodiment of the present invention.
[0035] Figure 5 A heat map showing the distribution of shopping and entertainment activity frequency of individual spatiotemporal big data users in 2023 (weekends) provided in an embodiment of the present invention;
[0036] Figure 6 This is a heat map (weekends) showing the distribution of other activities of individual spatiotemporal big data users in 2023, provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0038] Example
[0039] like Figure 1 As shown, this embodiment provides a method for identifying non-commuting activities of individuals based on multi-source data fusion using spatiotemporal big data. This method can more accurately identify and understand the non-commuting activities of urban residents, provide scientific decision-making support for urban development, and promote the optimization of urban transportation systems and the improvement of residents' quality of life. The specific steps include:
[0040] Step S1: Obtain the city's 500m×500m grid data, initial individual spatiotemporal big data with job and residence labels, initial resident travel survey data, and initial city POI data. Map the initial individual spatiotemporal big data, initial resident travel survey data, and initial city POI data to the city's 500m×500m grid data using spatial information, obtaining individual spatiotemporal big data containing grid spatial information, resident travel survey data, and city POI data. The city POI data includes city POI data from different years. Specifically:
[0041] Using ArcGIS software, we spatially connected the spatial information of user activity points in the initial individual spatiotemporal big data with the spatial information of the city's 500m×500m grid data. We also spatially connected the spatial information of the respondents' activity points in the initial resident travel survey data with the spatial information of the city's 500m×500m grid data, achieving a one-to-one correspondence between urban resident activity points and the grids where they are located.
[0042] Using ArcGIS software, the spatial information of urban POI points in the initial urban POI data is spatially connected with the spatial information of the urban 500m×500m grid data to achieve a one-to-one correspondence between urban POI points and the grids where they are located.
[0043] Step S2: Based on the individual spatiotemporal big data containing grid spatial information, extract the user's personal information and single-day activity sequence information, and then extract the first personal attribute feature and the first single-day activity attribute feature. Specifically:
[0044] In individual spatiotemporal big data containing raster spatial information, user personal information is directly given and can be directly extracted. The first personal attribute features extracted include the user's personal age, gender, and occupation. At the same time, the user ID, activity start date, activity end date, activity start time, and activity end time are obtained.
[0045] For the single-day activity attribute features, first sort according to the obtained user ID, activity start date, activity end date, activity start time (24-hour system), and activity end time (24-hour system) to obtain the single-day activity sequence information of each user. Then, based on the sequence, the single-day activity attribute features are calculated, including the start time of this activity, the duration of this activity, whether the last activity was at home, whether the last activity was at work, the straight-line distance from home, the straight-line distance between home and work, whether the activity start date is a weekday, and the order of this activity in the resident's day.
[0046] Step S3: Based on the resident travel survey data containing grid spatial information, extract the personal information and single-day activity sequence information of the respondents, and extract the second personal attribute feature and the second single-day activity attribute feature. Specifically:
[0047] In the resident travel survey data containing raster spatial information, the personal information of the respondents is directly given and can be directly extracted. The extracted second personal attribute features include the respondent's personal age, gender, and occupation. At the same time, the respondent's ID, activity start date, activity end date, activity start time, and activity end time are obtained.
[0048] For the single-day activity attribute characteristics, we first sort according to the obtained respondent ID, activity start date, activity end date, activity start time (24-hour system), and activity end time (24-hour system) to obtain the single-day activity sequence information of each user. Then, we calculate the single-day activity attribute characteristics based on the sequence, including the start time of this activity, the duration of this activity, whether the last activity was at home, whether the last activity was at work, the straight-line distance from home, the straight-line distance between home and work, whether the activity start date is a weekday, and the order of this activity in the resident's day.
[0049] Step S4: Extract the first built environment feature based on the individual spatiotemporal big data containing raster spatial information and the city POI data containing raster spatial information in the same year, and extract the second built environment feature based on the resident travel survey data containing raster spatial information and the city POI data containing raster spatial information in the same year. Specifically:
[0050] In the city POI data containing grid spatial information, for each POI type, the POI density of the grid where it is located is calculated using the following formula:
[0051]
[0052] in, Indicates the k In the grid i POI density of type, Indicates the k In the grid i Number of POIs of the type, represents the area of the kth grid ( km 2 );
[0053] In the individual spatiotemporal big data containing raster spatial information and the resident travel survey data containing raster spatial information, the grids of activity places, residence places and work places are directly obtained respectively.
[0054] The density information of various POIs in the grid is associated with the corresponding grids of activity, residence, and work to obtain the first and second built environment characteristics. The built environment characteristics include the density of regular bus stops within 500 meters and 800 meters of the center point of the grid where the activity, home, and work are located, as well as the density of POIs in the shopping, life service, catering, company, government, commercial housing, science, education and culture, health service, landmark, financial institution, and tourist attraction categories in the grid where the activity, home, and work are located.
[0055] Step S5: Preprocess the second personal attribute feature, the second daily activity attribute feature, and the second built environment feature. Specifically:
[0056] The second personal attribute feature, the second single-day activity attribute feature, and the second built environment feature are merged, that is, each activity of urban residents contains the above features;
[0057] All the above features are normalized, that is, the original data is converted to the range of [0,1]. The calculation formula is:
[0058]
[0059] in, Represents the first i Features, Indicates the i feature data, Indicates the i The maximum value of the feature data set, Indicates the i The minimum value of the feature data set;
[0060] The normalized data is used as the second feature set ;
[0061] Directly extract non-commuting activities (shopping and entertainment, other categories) from the resident travel survey data as the second label set , The secondary activities of shopping and entertainment were coded as 1, and the secondary activities of other categories were coded as 0;
[0062] Merge the second feature set With the second label set , as a dataset for various non-commuting activity recognition models to be evaluated, and divided into model training set and test set in a ratio of 8:2.
[0063] In step S6, based on the preprocessed second personal attribute feature, the second daily activity attribute feature, and the second built environment feature, a data set divided into a training set and a test set is obtained. The data set is used to train and test multiple non-commuting activity recognition models to be evaluated, so as to model the non-commuting activity recognition of individual spatiotemporal big data.
[0064] Specifically: the training set obtained by preprocessing is input into a variety of non-commuting activity recognition models to be evaluated for training. The various non-commuting activity recognition models to be evaluated include SVM model, DT model, RF model, GBDT model, ADAboost model, XGBoost model and Transformer-based classification model.
[0065] Step S7: Based on the test results, the performance of various mathematical models is output, and the non-commuting activity recognition model with the best performance is selected as the final individual spatiotemporal big data non-commuting activity recognition model. Specifically:
[0066] The overall accuracy of the indicators is used to describe the performance of multiple mathematical models, and the model with the best indicators is extracted as the subsequent individual spatiotemporal big data non-commuting activity identification model;
[0067] The indicators are: and ,include,
[0068]
[0069]
[0070] in, For the second label set The predicted value of For the second label set The true value of Indicates whether the true value is predicted correctly, represents the test set, represents the number of test set samples, Indicates accuracy;
[0071]
[0072]
[0073]
[0074] in, is the precision rate, For recall rate, The prediction result is the number of samples predicted correctly in a certain class label. The prediction result is the number of samples with incorrect predictions in a certain class of labels. The number of samples whose prediction results are incorrect for a certain type of label, express Fraction.
[0075] Step S8: Preprocess the first personal attribute feature, the first single-day activity attribute feature, and the first built environment feature. Specifically:
[0076] Merge the first personal attribute feature, the first single-day activity attribute feature, and the first built environment feature, that is, each non-commuting activity of each person in the individual spatiotemporal big data contains the above features;
[0077] Normalize all the above features using the same method as in step S5;
[0078] The normalized data is used as the first feature set ;
[0079] Get the travel activity category field and use it as the field to be inferred and merge it into the first feature set , forming a data prediction set.
[0080] In step S9, a data prediction set is formed based on the preprocessed first personal attribute feature, the first single-day activity attribute feature, and the first built environment feature. The prediction set data is input into the individual spatiotemporal big data non-commuting activity recognition model obtained in step S7 to obtain the recognition results of non-commuting activities in the individual spatiotemporal big data, including shopping and entertainment categories and other categories.
[0081] The individual spatiotemporal big data non-commuting activity identification method provided in this embodiment, based on multi-source data fusion, can more comprehensively capture personal attribute characteristics, single-day activity attribute characteristics, and built environment characteristics. At the same time, it combines more advanced mathematical models to improve the accuracy of secondary activity category inference, providing support for urban planning and traffic management.
[0082] Next, the effectiveness of the method proposed in this embodiment is verified:
[0083] The research data in this embodiment are 500m×500m raster data of a certain city, individual spatiotemporal big data with job and residence labels in the city in 2023 (data provided by the Mengxiang and Ke data platforms), resident travel survey data of the city in 2019, and urban POI data of the city in 2019 and 2023. The 2019 urban POI data and the 2019 resident travel survey data are used together for model training, and the 2023 urban POI data and the individual spatiotemporal big data with job and residence labels in 2023 are used together for non-commuting activity identification.
[0084] First, the resident travel survey data, urban POI data, and individual spatiotemporal big data are mapped onto a 500m×500m urban raster layer to unify the spatial information.
[0085] Based on resident travel survey data and city POI data, we extracted relevant influencing features required for modeling non-commuting activity identification using individual spatiotemporal big data. We also extracted relevant influencing features for identifying non-commuting activities using individual spatiotemporal big data using individual spatiotemporal big data and city POI data. It's important to note that the same feature types are used for modeling and inference, including personal attributes, daily activity attributes, and built environment characteristics. Detailed features are shown in Table 1.
[0086] Table 1 Feature set of the non-commuting activity recognition model for individual spatiotemporal big data
[0087]
[0088] After normalizing the acquired feature values, the feature data obtained from the resident survey data and the city POI data are used for model construction and need to be divided into a training set and a test set in an 8:2 ratio. The feature data obtained from the individual spatiotemporal big data and the city POI data are used as prediction data, which is the prediction set.
[0089] The training set is input into a variety of machine learning models, including SVM, DT, RF, GBDT, ADAboost, XGBoost and Transformer-based classification models, aiming to establish an inference model for the secondary activity categories of individual spatiotemporal big data. and To evaluate the performance of each model and select the best performing model.
[0090] The performance evaluation results are shown in Figure 2 It can be seen that the XGBoost model outperforms other models in all evaluation indicators, including SVM, DT, RF, GBDT, ADAboost and Transformer-based classification models. Specifically, the overall recognition accuracy of the XGBoost model reached 86%, with an F1 score of 0.89 in inferring the secondary activity categories of shopping and entertainment, and an F1 score of 0.83 in inferring the secondary activity categories of other categories. These results show that the XGBoost model can more accurately identify the secondary activity categories of shopping, entertainment and other types. Therefore, this embodiment ultimately selects the XGBoost model as the model with the best performance for further inference of secondary activity categories of individual spatiotemporal big data.
[0091] The individual attribute characteristics and daily activity attribute characteristics obtained from individual spatiotemporal big data containing raster spatial information and the built environment characteristics obtained from urban POI data containing raster spatial information were preprocessed and normalized and then imported into the optimal performance model to identify non-commuting activities in individual spatiotemporal big data. Then, based on the identified non-commuting activity information, a heat map of non-commuting activities in individual spatiotemporal big data was drawn (see Figures 3 to 6 The heat map results show that, whether on weekdays or weekends, shopping and entertainment activities are primarily concentrated in the central areas of a city's main urban area and five other new towns, while other activities are more widely distributed. This inference is consistent with real-world observations, confirming that the proposed non-commuting activity recognition technology can effectively serve the purpose of identifying non-commuting activities in individual spatiotemporal big data.
[0092] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A method for identifying non-commuting activities of individuals based on multi-source data fusion using spatiotemporal big data, characterized by: The following steps are involved: S1, based on the preset urban raster data, obtains individual spatiotemporal big data containing raster spatial information, resident travel survey data and urban POI data; S2, extracting the first personal attribute feature and the first single-day activity attribute feature based on the individual spatiotemporal big data; The S2 specifically includes: Based on the individual spatiotemporal big data containing raster spatial information, directly extract the first personal attribute feature, and simultaneously obtain the user ID, activity start date, activity end date, activity start time, and activity end time, wherein the first personal attribute feature includes the user's personal age, gender, and occupation; Sort by the user ID, the activity start date, the activity end date, the activity start time, and the activity end time to obtain the single-day activity sequence information of each user; Calculating first single-day activity attribute features based on the single-day activity sequence information, the first single-day activity attribute features including the start time of the current activity, the duration of the current activity, whether the last activity was at home, whether the last activity was at work, the straight-line distance from home, the straight-line distance between home and work, whether the activity start date is a weekday, and the order of the current activity in the resident's day; S3, extracting a second personal attribute feature and a second single-day activity attribute feature based on the resident travel survey data; The S3 specifically includes: Based on the individual spatiotemporal big data containing raster spatial information, the second personal attribute feature is directly extracted, and the respondent ID, activity start date, activity end date, activity start time and activity end time are obtained at the same time. The second personal attribute feature includes the respondent's personal age, gender, and occupation; Sort by the respondent ID, the activity start date, the activity end date, the activity start time, and the activity end time to obtain the single-day activity sequence information of each respondent; Calculating second single-day activity attribute features based on the single-day activity sequence information, the second single-day activity attribute features including the start time of the current activity, the duration of the current activity, whether the last activity was at home, whether the last activity was at work, the straight-line distance from home, the straight-line distance between home and work, whether the activity start date is a weekday, and the order of the current activity in the resident's day; S4, extracting a first built environment feature based on the individual spatiotemporal big data and the city POI data, and extracting a second built environment feature based on the resident travel survey data and the city POI data; The S4 specifically includes: Based on the urban POI data containing grid spatial information, the POI density of each POI type grid is calculated; Based on individual spatiotemporal big data containing raster spatial information and resident travel survey data containing raster spatial information, the grids of activity locations, residence locations, and workplace locations are directly extracted respectively; Associating the POI density with the grids of corresponding activity locations, residence locations, and workplace locations to obtain a first built environment feature and a second built environment feature, respectively; S5, preprocessing the second personal attribute feature, the second single-day activity attribute feature, and the second built environment feature; S6, based on the pre-processed second personal attribute features, second single-day activity attribute features and second built environment features, train and test multiple non-commuting activity recognition models to be evaluated; S7, based on the test results, select the non-commuting activity recognition model with the best performance as the final individual spatiotemporal big data non-commuting activity recognition model; S8, preprocessing the first personal attribute feature, the first single-day activity attribute feature, and the first built environment feature; S9, based on the preprocessed first personal attribute feature, the first single-day activity attribute feature and the first built environment feature, using the individual spatiotemporal big data non-commuting activity recognition model to obtain recognition results of non-commuting activities in the individual spatiotemporal big data.
2. The method for identifying non-commuting activities of individuals based on multi-source data fusion using spatiotemporal big data according to claim 1 is characterized in that: Said S1 specifically includes: Obtain initial individual spatiotemporal big data, initial resident travel survey data, and initial city POI data; Performing spatial connection between the user activity point spatial information in the initial individual spatiotemporal big data and the urban raster data spatial information to obtain individual spatiotemporal big data containing raster spatial information; Performing spatial connection between the spatial information of the activity points of the respondents in the initial resident travel survey data and the spatial information of the city raster data to obtain resident travel survey data containing raster spatial information; The city POI point spatial information in the initial city POI data is spatially connected with the city grid data spatial information to obtain city POI data containing grid spatial information.
3. The method for identifying non-commuting activities of individuals based on multi-source data fusion using spatiotemporal big data according to claim 1 is characterized in that: The S8 specifically includes: Merging the first personal attribute feature, the first single-day activity attribute feature, and the first built environment feature; Normalizing the first personal attribute feature, the first single-day activity attribute feature, and the first built environment feature, and using the normalized data as a first feature set; A travel activity category field is obtained, and the travel activity category field is used as a field to be inferred and merged into the first feature set to obtain a prediction set.
4. The method for identifying non-commuting activities of individuals based on multi-source data fusion using spatiotemporal big data according to claim 1 is characterized in that: The S5 specifically includes: Merging the second personal attribute feature, the second single-day activity attribute feature, and the second built environment feature; Normalizing the second personal attribute feature, the second single-day activity attribute feature, and the second built environment feature, and using the normalized data as a second feature set; Directly extracting non-commuting activities from the resident travel survey data as a second label set and encoding them; The second feature set and the second label set are combined as a data set for the multiple non-commuting activity recognition models to be evaluated, and the data set is divided into a training set and a test set.
5. The method for identifying non-commuting activities of individuals based on multi-source data fusion using spatiotemporal big data according to any one of claims 3-4, characterized in that: The expression of the normalization process is: Where, Represents the first i Features, Indicates the i feature data, Indicates the i The maximum value of the feature data set, Indicates the i The minimum value of the feature dataset.
6. The method for identifying non-commuting activities of individuals based on multi-source data fusion using spatiotemporal big data according to claim 4 is characterized in that: The S6 specifically includes: The training set is input into a plurality of pre-acquired non-commuting activity recognition models to be evaluated for training. The plurality of pre-acquired non-commuting activity recognition models to be evaluated include an SVM model, a DT model, a RF model, a GBDT model, an ADAboost model, an XGBoost model, and a Transformer-based classification model.
7. The method for identifying non-commuting activities of individuals based on multi-source data fusion using spatiotemporal big data according to claim 4 is characterized in that: The best performance indicators include accuracy and Scores include: Where, is the predicted value of the label set, is the true value of the label set, Indicates whether the true value is predicted correctly, represents the test set, represents the number of test set samples, Indicates the accuracy, is the precision rate, For recall rate, The prediction result is the number of samples predicted correctly in one of the labels. The prediction result is the number of samples with incorrect predictions in one of the labels. The number of samples whose prediction result is not one of the labels is incorrectly predicted. express Fraction.
Citation Information
Patent Citations
Urban functional area identification process based on space-time semantic mining
CN113627864A
Space division method and device, equipment and medium
CN114511125A