Individual activity position prediction data set construction method based on social media data
By building a multi-dimensional data framework and data processing technology, the sparse and privacy issues of social media data sets are solved, and high-quality, time-space-accurate individual activity position prediction data sets are achieved, and applications such as smart cities are supported.
Patent Information
- Application Number
- CN202510529417.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-07-25
AI Technical Summary
The existing social media data sets have problems such as sparse data, uneven spatial coverage, large coordinate deviations, and the risk of privacy leakage, resulting in insufficient accuracy and applicability of individual position prediction and behavioral pattern mining, and it is difficult to adapt to practical applications in different cities and cultural contexts.
By building a multi-dimensional data framework, data deduplication, missing value filling, abnormality detection and coordinate correction based on DBSCAN clustering and POI matching are carried out, and privacy protection measures such as hash encryption are used to perform data desensitization and spatiotemporal distribution analysis.
It significantly improves data integrity rate and spatial accuracy, ensures the coverage range, data integrity and spatial accuracy of the data set, provides high-quality, spatial and temporal and compliant data sets, and supports applications such as individual location prediction, smart city planning and public resource allocation.
Smart Images

Figure CN120372660A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of spatio-temporal data analysis, and more specifically, to a method for constructing a dataset for predicting individual activity locations based on social media data. Background Art
[0002] With the popularization of social media platforms, the check-in data of users on mobile terminals has become an important data source for reflecting individual activity trajectories. Traditional datasets often have problems such as data sparsity, uneven spatial coverage, large coordinate deviations, and privacy leakage risks, which seriously restrict the accuracy and applicability of data-driven individual location prediction and behavior pattern mining.
[0003] Although currently publicly available datasets such as GeoLife and Foursquare meet the research needs to a certain extent, due to the limitations of data collection methods and data processing methods, it is difficult to adapt to actual applications in different cities and different cultural backgrounds. The following are some relevant popular publicly available datasets: (1) GeoLife dataset: As an important basis for human behavior research, mobile pattern analysis, and traffic system optimization, the value of trajectory data has extended to multiple application fields such as anomaly detection and urban spatial planning. To promote related research, Microsoft Research released the landmark GeoLife trajectory dataset, which has now become the most widely used publicly available mobile trajectory benchmark data globally. This dataset has collected the mobile trajectories of 182 users, with more than 17,000 trajectory records, and the cumulative spatial span of trajectory points reaching more than 1.2 million kilometers. It covers continuous monitoring data of more than 50,000 hours in the time dimension. Mining the individual movement patterns in this dataset is of great significance for exploring human behavior characteristics, promoting the development of business intelligence, and academic research. However, due to restrictions such as location privacy protection, the acquisition of large-scale trajectory data always faces challenges, and existing publicly available datasets mostly rely on the collection of a limited group of volunteers. It is worth noting that the empirical research by the Hossein Amiri team based on Beijing regional data shows that: within a five-year observation period, only 45 users generated more than 100 valid stop points, and the peak number of daily active users did not exceed 25. Compared with the population base of 20 million in Beijing, this sample sparsity leads to a significant lack of data representativeness, especially when constructing a prediction model, it is difficult to effectively capture the complex human activity patterns at the urban scale. This significant gap between the data scale and the real scenario poses a fundamental challenge to the development of a general large-scale spatio-temporal prediction model based on the GeoLife dataset.
[0004] (2)Foursquare Dataset: Check-in data is a type of location-based service data that records users' visits to a certain area at a certain moment. Different from trajectory data that requires the collection of volunteers' behaviors, check-in data comes from mobile applications such as social media platforms, travel applications, and food service software. When users access a certain area through mobile devices, they perform operations similar to "check-in", such as sharing location information and evaluating a certain service. The relevant platforms can then obtain the corresponding check-in data. By analyzing check-in data, it can be found that many human activities in real life are predictable. For example, research shows that the next location a human will visit, the next time to return to a certain area, and the next speed state when driving a car can all be predicted by using the history of the behavior process, that is, there is predictability in the behavior pattern. The American technology company Foursquare released the Foursquare dataset, which contains check-in data from different cities. These data record users' check-in behaviors in various places, and its dataset includes information such as user information, venue information, check-in records, and social relationships. The Foursquare dataset is often used in business development and human behavior prediction, etc. Li et al. and Bao et al. helped merchants attract more customers by studying users' preferred POIs. Wen et al. found that users are more inclined to use keywords to express their preferences when planning a trip, so as to provide them with personalized travel routes. Gao et al. studied the temporal predictability of human online behavior by studying the Foursquare dataset. However, the Foursquare dataset is facing challenges in data processing, storage, and data quality issues due to its large data volume. Similarly, in order to handle such rich data, it has also improved the knowledge reserves of researchers in aspects such as spatial clustering and spatio-temporal association rule mining.
[0005] (3) Argoverse Motion Forecasting Dataset: Traffic data is usually collected by various traffic detection devices, sensors, GPS positioning systems, etc. By analyzing traffic data, information such as traffic flow, vehicle density, traffic accident situations, and traveler behavior data on arterial roads can be grasped, so as to be applied to traffic planning and tracking human behavior. The Argoverse Motion Forecasting Dataset is a large-scale autonomous driving dataset released by Argo AI, which is often used for predicting the future motion trajectories of vehicles and pedestrians. Chang et al. improved the accuracy of 3D object tracking and trajectory prediction based on this dataset through detailed map information such as lane directions, drivable areas, and ground heights, and predicted diverse trajectories according to R2P2. However, although the Argoverse Motion Forecasting Dataset provides rich category information in object detection, it does not give category labels in trajectory prediction, which limits the performance of the prediction model. At the same time, the Argoverse Motion Forecasting Dataset contains various sensor data and high-precision map data, and these data require very complex preprocessing and registration work.
[0006] Generally speaking, the problems of data sparsity and spatio-temporal discontinuity commonly existing in the above-mentioned publicly available datasets. How to solve the above problems and how to efficiently collect, strictly preprocess, and optimize the quality of social media check-in data while ensuring the data volume and data quality have become key technical problems to be solved urgently.
[0007] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention
[0008] The purpose of the present invention is to provide a method for constructing an individual activity location prediction dataset based on social media data, which can provide a dataset with high quality, spatio-temporal accuracy, and privacy compliance.
[0009] The present invention provides a method for constructing an individual activity location prediction dataset based on social media data, including the following steps: S01: According to the target research area, use the open API of the social media platform to collect data and construct a multi-dimensional data framework, and the multi-dimensional data framework includes original check-in data; S02: Preprocess the original check-in data to obtain preprocessed check-in data; S03: Evaluate and optimize the quality of the preprocessed check-in data to obtain an individual activity location prediction dataset.
[0010] Furthermore, step S01 specifically includes: S011: dividing the target research area into a plurality of equidistant grid units according to a preset accuracy; S012: collecting data using an open API of a social media platform based on the plurality of equidistant grid units to construct a multi-dimensional data framework.
[0011] Further, step S02 specifically includes: S021: performing structured conversion on the original sign-in data, parsing the information of each field in the sign-in record, and obtaining the sign-in data after structured conversion; S022: using a data deduplication algorithm to eliminate duplicate sign-in records in the sign-in data after structured conversion, ensuring that each record exists uniquely within the same time and space range, and obtaining deduplicated sign-in data; S023: using a missing value interpolation method to interpolate and complete the data faults in the sign-in trajectory of the deduplicated sign-in data, and obtaining the interpolated sign-in data; S024: using a clustering algorithm to detect outliers on the sign-in coordinates of the interpolated sign-in data, and using geocoding and POI matching methods to correct them, and obtaining corrected sign-in data; S025: using a desensitization algorithm to desensitize the user identification and sensitive information involved in the corrected sign-in data, and obtaining the pre-processed sign-in data.
[0012] Further, step S023 specifically includes: using a missing value interpolation method to interpolate and complete the data gaps in the sign-in track of the deduplicated sign-in data, and obtain the interpolated sign-in data, such as the formula: , in, is the interpolation moment, and is the value of the adjacent valid data point, , are the corresponding time points, is the time to be interpolated.
[0013] Furthermore, step S03 specifically includes: S031: calculating the coordinate completeness rate according to the preprocessed check-in data; S032: calculating the POI matching consistency according to the preprocessed check-in data; S033: performing spatiotemporal distribution analysis based on the coordinate completeness rate and the POI matching consistency using statistical analysis and visualization methods, performing weighted correction on the data of each region and each time period, and generating an individual activity location prediction data set.
[0014] Further, step S033 specifically includes: according to the coordinate integrity rate and the POI matching consistency, using statistical analysis and visualization methods to perform spatiotemporal distribution analysis, weighted correction of data in each region and each time period, and generating a data subset, such as formula: , Among them, is the weighted average data value, is the data quality weight of the region or time period, is the corresponding data value, is the quantity of the region or time period.
[0015] Furthermore, the method for constructing the individual activity location prediction data set based on social media data further includes: obtaining an individual activity location prediction database by using a database tool according to the individual activity location prediction data set.
[0016] Furthermore, the method for constructing the individual activity location prediction data set based on social media data further includes: updating the individual activity location prediction data set by using dynamic preprocessing parameters.
[0017] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned method for constructing an individual activity location prediction data set based on social media data are implemented.
[0018] Implementing the method for constructing an individual activity location prediction data set based on social media data provided by the present invention has the following beneficial effects: By constructing a multi-dimensional data framework, integrating the spatial grid sampling strategy, the Weibo open platform API, the dynamic proxy IP pool and the fingerprint obfuscation technology, the present invention realizes the efficient collection of check-in data in the target area; by adopting multiple data processing technologies such as deduplication, missing value filling, anomaly detection and coordinate correction based on DBSCAN clustering and POI matching for the original data, the present invention significantly improves the data integrity rate and spatial accuracy; the present invention uses privacy protection measures such as hash encryption to complete data desensitization; by spatio-temporal distribution analysis and data quality assessment, the present invention realizes the accurate control of the spatial coverage and feature consistency of the data set. The data set constructed by the present invention is superior to the traditional data set in terms of coverage, data integrity and spatio-temporal accuracy, and can provide more comprehensive and reliable data support for applications such as individual location prediction, smart city planning and public resource allocation; By constructing a multi-dimensional data framework and adopting a distributed data collection strategy, the present invention significantly improves the collection coverage rate and uniformity of check-in data in the target area; by using a structured data preprocessing process, including deduplication, missing value filling, and outlier correction, the present invention effectively improves the continuity and accuracy of the data; through data quality evaluation and optimization measures, the present invention ensures that the finally generated data set has a high coordinate integrity rate and a relatively high location matching consistency, providing more accurate input for the location prediction model; through the data storage and interface service module, the present invention realizes the efficient management, multi-format export, and online dynamic update of the data set, greatly improving the applicability and scalability of the data set in various application scenarios; while ensuring data quality, the present invention pays attention to privacy protection, and complies with relevant regulations through data desensitization processing, providing a solid data support and reliability guarantee for subsequent data-driven applications. The present invention overcomes the deficiencies existing in the data collection and preprocessing process of existing social media data sets, such as uneven data coverage, serious noise interference, many missing values, and imperfect privacy protection, etc., and can provide high-quality, spatio-temporally accurate, and privacy-compliant data support for applications such as location prediction, urban planning, and smart city management. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings: Figure 1 is a flowchart of a method for constructing an individual activity location prediction data set based on social media data provided by the present invention; Figure 2 is a general production flowchart of the data set provided by the present invention; Figure 3 is a process design diagram of the data collection method provided by the present invention; Figure 4 is a data cleaning flowchart provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] In order to have a clearer understanding of the technical features, objectives, and effects of the present invention, the specific embodiments of the present invention will now be described in detail with reference to the drawings.
[0021] Figure 1 shows a schematic diagram of a method for constructing an individual activity location prediction data set based on social media data in this embodiment. In this embodiment, the method for constructing an individual activity location prediction data set based on social media data includes the following steps: S01: According to the target research area, use the open API of the social media platform to collect data and construct a multi-dimensional data framework, and the multi-dimensional data framework includes original check-in data; In an exemplary embodiment, step S01 specifically includes: S011: Divide the target research area into multiple equally spaced grid cells according to a preset precision; In an exemplary embodiment, step S011 specifically includes: dividing the target research area into multiple equally spaced grid cells according to a preset precision, as shown in the formula: , where, is the number of grids, is the total area of the target area, is the area of a single grid cell; In an exemplary embodiment, the preset precision is 0.01°×0.01°; S012: According to the multiple equally spaced grid cells, use the open API of the social media platform to collect data and construct a multi-dimensional data framework; In an exemplary embodiment, the multi-dimensional data framework includes original check-in data, and the original check-in data includes geographical coordinates, timestamps, POI information, and related auxiliary fields; As an exemplary embodiment, in step S012, the OAuth 2.0 authentication mechanism is adopted, combined with the dynamic proxy IP and fingerprint obfuscation strategy, to batch obtain user check-in data (including geographical coordinates, timestamps, POI information, and related auxiliary fields) within a predetermined time range (for example, from August 2022 to September 2023); S02: Preprocess the original check-in data to obtain preprocessed check-in data; In an exemplary embodiment, step S02 specifically includes: S021: Perform structured conversion on the original check-in data, parse out the field information in the check-in record, and obtain structured-converted check-in data; S022: Use a data deduplication algorithm to eliminate duplicate check-in records in the structured-converted check-in data, ensuring that each record is uniquely present within the same time and space range, and obtain deduplicated check-in data; S023: Use a missing value imputation method to interpolate and complete the data gaps in the check-in trajectories of the deduplicated check-in data to obtain interpolated check-in data; In an exemplary embodiment, step S023 specifically includes: using a missing value imputation method to interpolate and complete the data gaps in the check-in trajectories of the deduplicated check-in data to obtain interpolated check-in data, as shown in the formula: , where, is the interpolation time, and are the values of adjacent valid data points, , are the corresponding time points respectively, is the interpolation time point; S024: Use the clustering algorithm to detect outliers in the signed-in coordinates of the interpolated signed-in data, and use the geocoding and POI matching method to correct them to obtain the corrected signed-in data; In an exemplary embodiment, the clustering algorithm is the DBSCAN clustering algorithm; As an exemplary embodiment, in step S024, based on the DBSCAN clustering algorithm, outliers in the signed-in coordinates are detected, and the distance metric uses the Euclidean distance formula: , And set the neighborhood radius ε and the minimum number of points MinPts to judge abnormal data; for the detected abnormal coordinates, correct them by combining geocoding and POI matching technology; S025: Use the desensitization algorithm to desensitize the user identification and sensitive information involved in the corrected signed-in data to obtain the preprocessed signed-in data; In an exemplary embodiment, the desensitization algorithm is the hash encryption algorithm; As an exemplary embodiment, in step S025, the user identification and sensitive information involved in the original data are processed using hash encryption or other desensitization algorithms to ensure data privacy and security; S03: Perform quality evaluation and optimization on the preprocessed signed-in data to obtain the individual activity location prediction data set; In an exemplary embodiment, step S03 specifically includes: S031: Calculate the coordinate completeness rate according to the preprocessed signed-in data; In an exemplary embodiment, the coordinate completeness rate is as follows: , where, is the coordinate completeness rate; is the number of valid signed-in records, is the total number of signed-in records; S032: Calculate the POI matching consistency according to the preprocessed signed-in data; In an exemplary embodiment, the POI matching consistency is as follows: , where, is the POI matching consistency, is the number of records matching the real POI; S033: Perform spatio-temporal distribution analysis using statistical analysis and visualization methods based on the coordinate completeness rate and the POI matching consistency, perform weighted correction on the data for each region and each time period, and generate an individual activity location prediction data set; In an exemplary embodiment, step S033 specifically includes: performing spatio-temporal distribution analysis using statistical analysis and visualization methods based on the coordinate completeness rate and the POI matching consistency, performing weighted correction on the data for each region and each time period, and generating a data subset, as shown in the formula: , where, is the weighted average data value, is the data quality weight for the region or time period, is the corresponding data value, is the number of regions or time periods.
[0022] In an exemplary embodiment, the above method for constructing an individual activity location prediction data set based on social media data further includes: obtaining an individual activity location prediction database using a database tool according to the individual activity location prediction data set; In an exemplary embodiment, the individual activity location prediction database includes a standard data table structure and a data interface; In an exemplary embodiment, the database tool is MySQL; In an exemplary embodiment, the above method for constructing an individual activity location prediction data set based on social media data further includes: updating the individual activity location prediction data set using dynamic preprocessing parameters; In an exemplary embodiment, the dynamic preprocessing parameters are as shown in the formula:
[0023] where, represents the dynamic preprocessing parameter, represents the old preprocessing parameter, is the difference between the new and old data quality indicators, is the adjustment coefficient.
[0024] In some embodiments, the above method for constructing an individual activity location prediction data set based on social media data can also be implemented in the following manner. In this embodiment, the method for constructing an individual activity location prediction data set based on social media data includes the following steps: S1: Construct a multi-dimensional data framework, integrate a spatial grid division strategy, a social media platform data interface, a dynamic proxy IP pool, and a fingerprint obfuscation technology to obtain a data set for constructing original check-in data; Step S1 specifically includes: S11: Divide the target research area into multiple equally spaced grid cells according to a preset precision (e.g., 0.01°×0.01°) to ensure the uniformity and representativeness of the check-in data sampling within the area. The number of grids is as shown in the formula:
[0025] where is the total area of the target area, is the area of a single grid cell; S12: Use the open API of the social media platform for data collection. Adopt the OAuth 2.0 authentication mechanism, combined with the dynamic proxy IP and fingerprint obfuscation strategy, to batch obtain user check-in data within a predetermined time range (e.g., from August 2022 to September 2023), including geographical coordinates, timestamps, POI information, and related auxiliary fields.
[0026] S2: Preprocess the original check-in data, specifically including data format conversion, duplicate data removal, missing value imputation, outlier detection, and coordinate correction. At the same time, perform privacy desensitization processing on user identifiers. Step S2 specifically includes: S21: Perform structured conversion on the collected original JSON-format data and parse out the field information in the check-in records. S22: Use the data deduplication algorithm to remove duplicate check-in records to ensure that each record exists uniquely within the same time and space range. S23: Adopt the missing value imputation method to interpolate and complete the data breaks in the check-in trajectory. The interpolation formula, for example, uses linear interpolation:
[0027] where and are the values of adjacent valid data points, , are the corresponding time points respectively, is the moment to be interpolated; S24: Based on the DBSCAN clustering algorithm, perform outlier detection on the check-in coordinates. The distance metric uses the Euclidean distance formula:
[0028] And set the neighborhood radius ε and the minimum number of points MinPts to judge abnormal data; for the detected abnormal coordinates, correct them by combining geocoding and POI matching technology. S25: Process the user identifiers and sensitive information involved in the original data using hash encryption or other desensitization algorithms to ensure data privacy and security.
[0029] S3: Conduct quality assessment and optimization on the preprocessed data. Based on indicators such as coordinate completeness rate, POI consistency, and spatio-temporal distribution balance, classify, screen, and weighted correct the data to form a high-quality data subset; Step S3 specifically includes: S31: Construct a data quality evaluation index system and calculate the coordinate completeness rate for the preprocessed data and POI matching consistency , where the coordinate completeness rate is defined as:
[0030] where is the number of valid check-in records, is the total number of check-in records; S32: Calculate the POI matching consistency, and the formula is:
[0031] where is the number of records matching the true POI; S33: Use statistical analysis and visualization methods to conduct spatio-temporal distribution analysis on the dataset, perform weighted correction on the data in each region and each time period, and generate a high-quality data subset. Its weighted average data value can be expressed as:
[0032] where is the data quality weight of the region or time period, is the corresponding data value.
[0033] S4: Store the dataset after data preprocessing and quality optimization in a structured manner, support the export of multiple data formats, and provide them for the training and application of subsequent individual activity location prediction models.
[0034] Step S4 specifically includes: S41: Store the dataset after preprocessing and quality assessment in a relational database (such as MySQL), establish a standard data table structure, including a check-in record table, a user information table, and a trajectory sequence table, and set necessary indexes to improve data query efficiency; S42: Support the export of the dataset in multiple formats such as CSV, Excel, GeoJSON, Parquet, etc., to facilitate downstream model training, data analysis, and visualization applications; S43: Provide a data interface or API service to realize the online call and update of data, and meet the real-time data requirements of the individual activity location prediction system.
[0035] Specifically, the above method for constructing an individual activity location prediction dataset based on social media data further includes performing periodic quality monitoring and updating on the constructed dataset, and dynamically adjusting the grid division strategy and preprocessing parameters according to newly collected data to ensure the timeliness and accuracy of the dataset in long-term applications. Its update mechanism can adjust the preprocessing parameters using the following formula:
[0036] Where represents the preprocessing parameter, is the difference between the quality indicators of the new and old data, is the adjustment coefficient.
[0037] In some embodiments, the above method for constructing an individual activity location prediction dataset based on social media data can also be implemented in the following manner. In this embodiment, the method for constructing an individual activity location prediction dataset based on social media data includes the following steps: Construct a multi-dimensional data framework. In this step, by integrating multiple information sources, a multi-dimensional data framework including space, time, and user behavior is constructed. The data framework covers geographical boundary information, grid division strategy, and check-in data of social media platforms within the target area to ensure that the collected data can comprehensively reflect the individual activity trajectories within the area. Specifically, when implementing, the target area is divided into several equidistant grid cells according to a preset precision, the open interface of the social media platform is used to collect user check-in records, and the continuity and effectiveness of data collection are ensured through dynamic proxy and fingerprint obfuscation technologies, while ensuring compliance with privacy protection requirements during the data collection process.
[0038] Data collection. Based on the constructed multi-dimensional data framework, in this step, the open API of the social media platform is used to batch collect user check-in data within the target area through an authentication mechanism and dynamic proxy method. The collected data includes not only the geographical coordinates, timestamps, and location names of check-ins, but may also include auxiliary information such as text and pictures posted by users. To improve the data collection efficiency, the system reasonably schedules the collection requests and uses a distributed collection strategy to break through the platform call limit to ensure obtaining as much effective data as possible within a predetermined time range.
[0039] Data preprocessing. To improve the quality of the dataset, in this step, a series of preprocessing operations are performed on the collected raw data, mainly including the following: 1. Perform structured conversion on the raw data to parse out each field information in the check-in record; 2. Use a deduplication algorithm to eliminate duplicate records to ensure the uniqueness of each check-in event; 3. Perform imputation processing on the data with missing values to fill in the data gaps through time series filling methods; 4. Identify and correct abnormal coordinate data using clustering algorithms and outlier detection techniques to ensure the accuracy of geographical location data; 5. For user privacy data, adopt hash encryption or other data desensitization means to ensure that the data meets privacy protection requirements during use.
[0040] Data quality evaluation and optimization. In this step, a data quality evaluation index system is constructed to quantitatively evaluate and optimize the preprocessed data. The system detects key indicators such as the spatial coverage rate, time continuity, coordinate integrity rate, and location matching consistency of the dataset, and uses statistical analysis and visualization means to deeply analyze the spatio-temporal distribution of the data. Based on the evaluation results, further eliminate low-quality data and perform weighted fusion processing on high-quality data to finally form a data subset with high accuracy and representativeness to meet the high requirements of the downstream location prediction model for data quality.
[0041] Data storage and interface services. To achieve efficient management and subsequent applications of data, this step stores the quality-optimized dataset in a relational database or a big data platform, and designs a standard data table structure, including check-in record tables, user information tables, and trajectory record tables, etc. Indexes are used for data storage to accelerate queries, and data export in multiple formats (such as CSV, Excel, GeoJSON, Parquet, etc.) is supported. At the same time, the system provides a standardized API interface to realize online data invocation and dynamic update to meet the needs of real-time data analysis and model training.
[0042] Dynamic update and maintenance of the dataset. To ensure the timeliness and accuracy of the dataset in long-term applications, this step establishes a periodic quality monitoring and update mechanism. By collecting the latest check-in data in real time, comparing historical data quality indicators, dynamically adjusting data collection parameters and preprocessing strategies, and continuously updating the dataset content. This mechanism ensures that the dataset can reflect the latest individual activity patterns and provides stable and reliable input data for the downstream prediction model.
[0043] In some embodiments, the above method for constructing an individual activity location prediction dataset based on social media data can also be implemented in the following way. In this embodiment, the method for constructing an individual activity location prediction dataset based on social media data includes the following steps: Please refer to Figure 2 , for the overall flowchart of the dataset. This figure mainly covers four core links: data collection, data processing, quality assessment, and data analysis. By constructing a multi-dimensional data framework, adopting a distributed data collection and anti-crawling strategy, through strict data preprocessing and quality inspection, a standardized dataset is finally output to meet the training needs of the downstream individual activity location prediction model. Specifically, it includes the following parts: For the data collection part, please refer to Figure 3 , Figure 3 which is the flowchart of the data collection method, showing the complete collection chain from data requirements to data processing and realizing the full-process management of data collection. The specific steps are as follows: Construction of a multi-dimensional data framework. According to the target application requirements, a multi-dimensional data framework is constructed, integrating the geographical boundaries, time information, and user behavior information of the target area. The target area is divided into multiple equidistant grid cells according to the preset spatial sampling accuracy (e.g., 0.01°×0.01°) to ensure the uniformity and representativeness of the check-in data sampling within the area. The number of grid cells can be calculated by the formula:
[0044] where is the total area of the target area, is the area of a single grid cell; Social media data collection. Using the open API of social media platforms (such as Weibo API), the OAuth 2.0 authentication mechanism is adopted for data invocation. At the same time, dynamic proxy IP and fingerprint obfuscation strategies are used to deal with anti-crawling measures to ensure the continuity of the collection process. User check-in data is batch-collected within a predetermined time range (e.g., from August 2022 to September 2023), and the collected data fields include geographical coordinates, timestamps, location names, and relevant auxiliary information. The anti-crawling countermeasure strategy is shown in Table 1.
[0045] Table 1: Anti-crawling countermeasure strategy
[0046] For the data processing part, please refer to Figure 4 , Figure 4 which is the data cleaning flowchart, showing the whole process of raw data parsing, data cleaning, data correction, data output, and data standardization output. The specific steps are as follows: Data format conversion. The collected raw data (such as JSON format) is parsed and converted into a structured data table, and each field (such as check-in time, coordinates, location name, etc.) is extracted to prepare the data for subsequent processing; Data deduplication. Using the deduplication algorithm based on time and space windows to eliminate duplicate check-in records and ensure the unique existence of each check-in event in the dataset; Missing value processing. The records with data breaks are complemented by interpolation methods to maintain data continuity. The linear interpolation formula is as follows:
[0047] where and are the values of adjacent valid data points, , are corresponding time points, is the time point to be interpolated; Outlier detection and coordinate correction: Use an algorithm based on density clustering (DBSCAN algorithm) to detect outliers in the check-in coordinates. Euclidean distance is used to calculate the distance between data points, and its formula is:
[0048] Then, judge the abnormal data by setting the neighborhood radius ε and the minimum number of points MinPts. Finally, correct the coordinates of the abnormal data by combining geocoding and POI matching technology to improve the data accuracy. Table 2 is the outlier processing strategy table.
[0049] Table 2: Outlier Processing Strategy Table
[0050] Privacy desensitization processing: Use hash encryption or other desensitization algorithms to process the data fields related to user identity and sensitive information to ensure that the data complies with privacy regulations during the collection, processing, and storage processes.
[0051] Quality assessment and verification part: Use data quality evaluation indicators (such as coordinate integrity rate, POI consistency, text availability rate, and time format unification rate) to evaluate the data quality and use data quality inspection strategies (such as manual sampling, automated testing, and delivery data specifications) to inspect the data quality, so as to form a high-quality data subset; The steps are as follows: Quality index calculation: Construct a data quality evaluation index system to quantitatively evaluate the preprocessed data. Table 3 is the data quality evaluation index result table.
[0052] Table 3: Data Quality Evaluation Index Result Table
[0053] The main indicators include: (1) Coordinate integrity rate, which reflects the proportion of valid geographical location information in the total number of records; (2) POI consistency, which evaluates the matching degree between the check-in location information and the real point of interest data; (3) Text availability rate, which measures the integrity and parsability of user-generated text information (such as check-in descriptions or comments); (4) Time format unification rate, which ensures that all timestamp data are stored and parsed in a unified format.
[0054] Data quality inspection: On the basis of quality index calculation, adopt a variety of data verification strategies for comprehensive inspection, including: (1) Manual sampling inspection: Randomly select a certain proportion of data samples and manually compare and check the accuracy and consistency of the data. (2) Automated testing: Use preset rules and scripts to perform automated verification on the data to promptly detect problems such as data anomalies or non-standard formats. (3) Data delivery specification verification: Based on the pre-established data delivery specifications, verify the format, units, field integrity, etc. of the data to ensure that the delivered data meets the standard requirements.
[0055] Dataset analysis section: Conduct in-depth screening, overall distribution, and statistical feature analysis on the data that has undergone standardization processing and quality assessment and inspection, providing a high-quality data foundation for subsequent urban behavior pattern mining, spatial modeling, and related applications. The specific steps are as follows: Data screening: According to the preset quality threshold and business requirements, eliminate low-quality data and retain the high-quality data subset. Overall distribution analysis: Use statistical analysis methods and visualization tools to comprehensively describe the distribution of the dataset in space and time, revealing the coverage range and potential patterns of the data. Statistical feature analysis: Perform statistical calculations on each indicator in the dataset to form a statistical report describing the basic characteristics of the dataset, providing a basis for subsequent model construction and data mining.
[0056] Data storage and interface service section: Store the dataset that has undergone data processing and quality optimization in a relational database (such as MySQL) or a big data platform in a structured manner, and design standard data table structures, including check-in record tables, user information tables, and trajectory sequence tables. To improve data query efficiency, set necessary indexes and support the export of the dataset in multiple formats such as CSV, Excel, GeoJSON, and Parquet for subsequent model training, data analysis, and visualization applications. At the same time, provide a standardized API interface to enable online data invocation and dynamic updates to meet the real-time data requirements of the individual activity location prediction system. The specific steps are as follows: Data storage design: According to the data content after preprocessing and quality optimization, design standard data table structures, mainly including: (1) Check-in record table: Store data such as user check-in time, geographical coordinates, and location names. (2) User information table: Store the anonymized user identifiers and related attributes. (3) Trajectory sequence table: Record the continuous check-in trajectories of users and the corresponding spatio-temporal information.
[0057] To improve data query efficiency, establish indexes (such as composite indexes based on time, geographical location, and user ID). Database selection and storage: Store the dataset in a relational database (such as MySQL) or a big data platform. Select a suitable storage solution based on the data volume and access requirements. When designing the database storage solution, consider data redundancy backup and partition storage to handle large-scale data access and subsequent query performance requirements; Data export support: Implement the multi-format data export function, supporting formats such as CSV, Excel, GeoJSON, and Parquet. Develop the export module to allow users to select data fields and formats according to their needs for downstream model training and data analysis; API interface development: Provide standardized API interfaces, including data query interfaces, data update interfaces, and data export interfaces. The interface design should support the RESTful style to ensure that the system can call data in real time and guarantee data transmission security through security mechanisms (such as Token authentication and HTTPS encryption). Implement interface call logs and error monitoring to promptly detect and resolve system exceptions.
[0058] Dataset dynamic update and maintenance part: To ensure the timeliness and accuracy of the dataset in long-term applications, this module establishes a periodic quality monitoring and update mechanism. By collecting the latest check-in data in real time and comparing it with historical data quality indicators, dynamically adjust data collection parameters and preprocessing strategies, and continuously update the dataset content. The update mechanism ensures that the dataset always reflects the latest individual activity patterns and provides stable and reliable input data for downstream prediction models. The specific steps are as follows: Periodic collection scheduling: Establish a timed task to regularly call the data collection module to obtain the latest social media check-in data. Adjust the collection parameters (such as grid division accuracy and collection frequency) to adapt to changes in data volume and ensure data real-time and coverage; Incremental data preprocessing: Perform the same preprocessing process on the newly collected data as the initial data (format conversion, deduplication, missing value imputation, anomaly detection, coordinate correction, and privacy desensitization), and integrate it with the existing database data. Adopt an incremental update strategy to ensure the continuity of the dataset and the consistency of historical data; Dynamic quality monitoring: Regularly calculate data quality evaluation indicators (such as coordinate integrity rate, POI consistency, text availability rate, and time format unification rate), and compare them with historical indicators. If a decrease in indicators is found, trigger an alarm, and the system automatically or manually intervenes to check the data collection and preprocessing processes and make necessary adjustments; Data update and version control: Manage the version of the updated dataset, record the time, data volume, and quality indicator changes of each update, and allow users to roll back to any historical version to ensure data consistency and audit traceability; Dynamic adjustment of preprocessing parameters: According to the difference between the new and old data quality indicators , using the formula
[0059] dynamically adjust the key parameters in data preprocessing (such as missing value filling strategy parameters, anomaly detection thresholds, etc.), where is the adjustment coefficient, regularly evaluate the effect of parameter adjustment, and record the adjustment log to provide a basis for system optimization.
[0060] This embodiment provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-mentioned method for constructing an individual activity location prediction data set based on social media data.
[0061] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.
Claims
1. A method for constructing an individual activity location prediction dataset based on social media data, characterized in that, It includes the following steps: S01: According to the target research area, use the open API of the social media platform to collect data and construct a multi-dimensional data framework, where the multi-dimensional data framework includes original check-in data; S02: Preprocess the original check-in data to obtain preprocessed check-in data; S03: Evaluate and optimize the quality of the preprocessed check-in data to obtain an individual activity location prediction data set.
2. The method for constructing an individual activity location prediction data set based on social media data according to claim 1, wherein, Step S01 specifically includes: S011: Divide the target research area into multiple equally spaced grid cells according to a preset precision; S012: According to the multiple equally spaced grid cells, use the open API of the social media platform to collect data and construct a multi-dimensional data framework.
3. The method for constructing an individual activity location prediction data set based on social media data according to claim 1, characterized in that, Step S02 specifically includes: S021: Perform a structured conversion on the original check-in data, parse out the field information in the check-in record, and obtain the structured converted check-in data; S022: Use a data deduplication algorithm to eliminate duplicate check-in records in the structured converted check-in data, ensuring that each record exists uniquely within the same time and space range, and obtain deduplicated check-in data; S023: Use a missing value imputation method to interpolate and complete the data gaps in the check-in trajectory of the deduplicated check-in data to obtain interpolated check-in data; S024: Use a clustering algorithm to detect outliers in the check-in coordinates of the interpolated check-in data, and use a geocoding and POI matching method for correction to obtain corrected check-in data; S025: Use a desensitization algorithm to desensitize the user identifiers and sensitive information involved in the corrected check-in data to obtain preprocessed check-in data.
4. The method for constructing an individual activity location prediction data set based on social media data according to claim 3, characterized in that, Step S023 specifically includes: Use a missing value imputation method to interpolate and complete the data gaps in the check-in trajectory of the deduplicated check-in data to obtain interpolated check-in data, as shown in the formula: , Wherein, is the moment of interpolation, and are the values of adjacent valid data points, 、 are the corresponding time points respectively, is the moment to be interpolated.
5. The method for constructing an individual activity location prediction dataset based on social media data according to claim 1, wherein Step S03 specifically includes: S031: Calculate the coordinate completeness rate according to the preprocessed check-in data; S032: Calculate the POI matching consistency according to the preprocessed check-in data; S033: According to the coordinate completeness rate and the POI matching consistency, use statistical analysis and visualization methods to perform spatio-temporal distribution analysis, perform weighted correction on the data in each region and each time period, and generate an individual activity location prediction data set.
6. The method for constructing an individual activity location prediction data set based on social media data according to claim 5, characterized in that Step S033 specifically includes: According to the coordinate completeness rate and the POI matching consistency, use statistical analysis and visualization methods to perform spatio-temporal distribution analysis, perform weighted correction on the data in each region and each time period, and generate a data subset, as shown in the formula: , Among them, is the weighted average data value, is the data quality weight of the region or time period, is the corresponding data value, is the quantity of the region or time period.
7. The method for constructing an individual activity location prediction data set based on social media data according to claim 1, characterized in that The method for constructing an individual activity location prediction data set based on social media data further includes: According to the individual activity location prediction data set, use database tools to obtain an individual activity location prediction database.
8. The method for constructing an individual activity location prediction data set based on social media data according to claim 1, wherein The method for constructing an individual activity location prediction data set based on social media data further includes: Use dynamic preprocessing parameters to update the individual activity location prediction data set.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for constructing an individual activity location prediction data set based on social media data according to any one of claims 1-8.