Personnel whole process activity mode identification and description method based on mobile phone signaling data
By combining activity chain segmentation of mobile phone signaling data and a combination of multiple machine learning methods, identifying and describing the full-process activity mode of urban residents, the problem of difficulty in mastering activity mode with high accuracy in traditional methods is solved, and the rapid and accurate portrayal of urban residents' activities is achieved, and important technical support is provided for the construction of smart cities.
Patent Information
- Application Number
- CN202510547321.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The prior art is difficult to master the activity patterns of urban residents with high accuracy and timeliness through traditional data collection and analysis methods, especially in extracting valuable information from massive and rough mobile phone signaling data.
By identifying the user's stay status, the user's activity chain is initially divided throughout the day, and the travel mode is identified by base station sequence matching and clustering, and combined with inference rules of speed and travel distance, the final version of the entire process activity chain is generated.
It realizes a rapid, comprehensive and accurate portrayal of urban residents' activity patterns, solves the problems of large amount of signaling data and low accuracy, provides a big data processing framework that supports real-time computing, and provides a scientific basis for decision-making such as urban transportation organization and optimization, emergency response, etc.
Smart Images

Figure CN120075745A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of data mining, and in particular relates to a method for identifying and describing a personnel full-process activity pattern based on mobile phone signaling data. Background Art
[0002] With the acceleration of urbanization and the rapid development of information technology, the concept of smart cities has become an important way to improve urban management efficiency and optimize the quality of life of residents. In the construction of smart cities, how to comprehensively and accurately grasp the activity patterns of urban residents has become a key issue that needs to be solved urgently. Traditional data collection and analysis methods, such as questionnaire surveys and traffic flow statistics, often have shortcomings such as long data collection cycles, limited coverage, and poor real-time performance, which are difficult to meet the needs of modern urban management for high precision and high timeliness.
[0003] In recent years, with the popularization of mobile communication technology and the widespread use of smart phones, mobile signaling data, as a new data source, has gradually attracted attention in the fields of urban planning and traffic management. Mobile signaling data refers to the records of location updates, calls, text messages, data connections, etc. generated by mobile phone users interacting with base stations during the use of mobile communication services. These data are not only massive, but also can reflect the movement trajectory and activity behavior of mobile phone users in real time, providing unprecedented possibilities for analyzing the activity patterns of urban residents.
[0004] However, mobile phone signaling data also faces challenges such as data redundancy, noise interference, and privacy protection. How to extract valuable information from these complex data and build a method system that can accurately describe the activity patterns of urban residents throughout the entire process is a hot topic and difficulty in current research. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides a method for identifying and describing the whole-process activity pattern of personnel based on mobile phone signaling data, comprising the following steps:
[0006] Step S1: Based on the user's signaling data throughout the day, by identifying the user's stay status, the user's activity chain throughout the day is preliminarily segmented into stay segments and movement segments of the activity chain;
[0007] Step S2: using base station sequence matching to identify the movement mode in the movement segment, including subway travel mode, high-speed rail travel mode, and bus travel mode, further segmenting the movement segment, filling the activity chain, and obtaining a basic version of the activity chain;
[0008] Step S3: On the basic version of the activity chain, identify travel modes of driving, walking, and cycling using clustering and inference rules based on speed and travel distance, and fill them into the basic version of the activity chain to obtain the synthesized version of the activity chain;
[0009] Step S4: Check and analyze the stay segments and add description tags; optimize and merge the synthesized version of the activity chain to generate the final version of the whole-process activity chain.
[0010] Beneficial effects:
[0011] The present invention provides a method for identifying and describing the whole-process activity patterns of personnel based on mobile phone signaling data. By extracting key spatio-temporal information from a large amount of rough signaling data and applying a big data processing framework that integrates multiple machine learning means, it realizes a fast, comprehensive, and accurate characterization of the activity patterns of all personnel within the target city scope. This method solves the difficulties of large data volume, low accuracy, and roughness of the signaling data itself, and achieves a high-precision travel mode recognition effect with a big data processing framework that supports real-time calculation, providing technical support for downstream tasks in human behavior understanding and pattern extraction, providing a scientific basis for decision-making in urban traffic organization and optimization, emergency response, commercial layout, etc., and having important significance for promoting the construction and development of smart cities. Description of the drawings
[0012] Figure 1 It is a schematic flow chart of a method for identifying and describing the whole-process activity patterns of personnel based on mobile phone signaling data according to the present invention;
[0013] Figure 2 It is a schematic diagram of the line GID grid sequence and the trajectory GID grid sequence;
[0014] Figure 3 It is a schematic diagram of the generation process of the basic version of the activity chain. Detailed implementation manners
[0015] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0016] Embodiment 1
[0017] As Figure 1 shown, a method for identifying and describing the whole-process activity patterns of personnel based on mobile phone signaling data provided by an embodiment of the present invention includes the following steps:
[0018] Step S1: Based on the user's signaling data throughout the day, by identifying the user's stay status, perform a preliminary activity semantic segmentation on the user's activity chain throughout the day to form the stay segments and movement segments of the activity chain;
[0019] Step S2: Adopt the method of base station sequence matching to identify the movement patterns in the movement segments, including subway travel patterns, high-speed rail travel patterns, and bus travel patterns. Further segment the part of the movement segments and fill them into the activity chain to obtain the basic version of the activity chain;
[0020] Step S3: On the basic version of the activity chain, adopt clustering and inference rules based on speed and travel distance to identify the travel patterns of driving, walking, and cycling, and fill them into the basic version of the activity chain to obtain the synthesized version of the activity chain;
[0021] Step S4: Check and analyze the stay segments and add description tags; for the synthesized version of the activity chain, perform optimization and merging to generate the final version of the full-process activity chain.
[0022] In one embodiment, the above Step S1: Based on the user's signaling data throughout the day, by identifying the user's stay status, perform a preliminary activity semantic segmentation on the user's activity chain throughout the day to form the stay segments and movement segments of the activity chain, specifically including:
[0023] Step S11: Put the signaling data into a sliding window, and determine whether the time period is a stay status according to the repetition degree of the longitude and latitude data in the sliding window, and calculate the stay points. The calculation formula is as follows:
[0024] (1)
[0025] (2)
[0026] Wherein, represents the repetition degree of the coordinate of the signaling data t in the window, represents the size of the window, is an indicator function, and are the coordinates of the signaling data numbered i and t respectively; represents the activity status of the signaling data t, including stay and movement;
[0027] Step S12: For each signaling data with a stay activity status, retain its longitude and latitude coordinates and retention time, weight them according to the retention time and perform clustering on the map, and use the center coordinate of each cluster after clustering as the stay point, specifically including:
[0028] Step S121: For each screened signaling data with a stay status, calculate its retention time in the sliding window ( ). The residence time can be obtained from the timestamps of the signaling data, i.e., the end timestamp ( ) minus the start timestamp ( ). This time difference represents the duration of the user's stay at that location. According to the length of the residence time, a weight is assigned to each stay signaling. The longer the residence time, the greater the weight. The purpose of weighting is to make the points with longer stay times have a greater impact on the clustering result in the subsequent clustering process. The formula for calculating the weight is as follows:
[0029] (3)
[0030] Where, represents the weight of the currently calculated signaling, represents the stay time of this signaling.
[0031] Step S122: Use the DBSCAN clustering algorithm for clustering. Input all the weighted stay signaling data (including: longitude and latitude and weight weight) into the selected clustering algorithm for clustering analysis.
[0032] Step S123: For each cluster obtained after clustering ( ), calculate its average value to obtain the center coordinates ( ). The cluster center coordinates will serve as the stay points, representing the typical stay locations of the user in that area. Output the center coordinates (i.e., the stay points) of each cluster obtained after clustering, and record the relevant information, including the ID of the cluster and the stay time period.
[0033] Step S13: Using the stay points as the segmentation points, divide the activity chain into multiple moving segments and stay segments to complete the preliminary semantic segmentation, providing a basis for subsequent activity chain segmentation and semantic annotation.
[0034] In one embodiment, in the above step S2: By using the method of base station sequence matching, identify the moving patterns in the moving segments, including subway travel patterns, high-speed rail travel patterns, and bus travel patterns, further divide the parts of the moving segments, and fill them into the activity chain to obtain the basic version of the activity chain, specifically including:
[0035] Step S21: Obtain subway, bus, and high-speed rail operation data, including: line description information data, station description information data, and line topology structure data, specifically including:
[0036] Step S211: Obtain subway line description information data, including: line ID, line name, line direction, whether the line is a loop line, starting and ending stations, and passing stations;
[0037] In the line description information data, each of the two running directions of a physical line corresponds to a line ID, that is, a line ID uniquely determines a line running in a fixed direction. For example, the line IDs corresponding to Beijing Subway Line 10 are "Line 10_Inner Loop" and "Line 10_Outer Loop";
[0038] Step S212: Obtain high-speed rail line description information data, including line ID, line name, line direction, whether it is a double-track line, starting and ending stations, and passing stations.
[0039] Step S213: Obtain high-speed rail line description information data, including line ID, line name, line direction, starting and ending stations, and passing stations.
[0040] Step S214: According to the poi information, obtain the longitude and latitude position coordinates of all stations.
[0041] Step S22: Generate running line data according to the line topologies of subways, buses, and high-speed rails. The form of each running line data is an ordered sequence of longitude and latitude coordinates; then map the longitude and latitude coordinates of the running line data into grids named in the form of GID to form a line GID grid sequence;
[0042] Step S23: Scan the longitude and latitude coordinates of all base stations, establish a mapping table between the base station IDs that may interact when traveling into a certain GID grid and the GID grid ID, and convert the signaling data into a trajectory GID grid sequence through this mapping table;
[0043] As Figure 2 shows a schematic diagram of the line GID grid sequence (left) and the trajectory GID grid sequence (right).
[0044] Step S24: Calculate the edit distance between the trajectory GID grid sequence and the line GID grid sequence to measure whether the trajectory of a person matches the line; if the edit distance does not exceed the matching threshold, it is determined that the travel segment is the corresponding subway, bus, or high-speed rail travel mode, and it is filled into the activity chain to obtain the basic version of the activity chain.
[0045] After multiple experiments, the embodiment of the present invention sets the matching threshold to 80%.
[0046] The definition and calculation pseudo-code of the edit distance are as follows:
[0047] The Edit Distance, also known as the Levenshtein distance, is a way to measure the difference between two strings. It is defined as the minimum number of edit operations required to convert one string into another, where the edit operations include inserting, deleting, and replacing characters. Here, the elements for the Edit Distance operation are sequences of GID grids, that is, it calculates the number of edits required to convert one grid sequence into another.
[0048] When calculating the Edit Distance, the dynamic programming algorithm is generally used, and its pseudocode is as follows:
[0049] FUNCTION EditDistance(gridSequence1, gridSequence2)
[0050] LET n = LENGTH(gridSequence1)
[0051] LET m = LENGTH(gridSequence2)
[0052] LET dp = ARRAY[0..n, 0..m] / / Create a two-dimensional array, where dp[i][j] represents the Edit Distance between the first i grid sequences of gridSequence1 and the first j grid sequences of gridSequence2
[0053] / / Initialize the first row and the first column
[0054] FOR i FROM 0 TO n
[0055] dp[i][0] = i / / It takes i deletions to change the first i grid sequences of gridSequence1 into an empty grid sequence string
[0056] END FOR
[0057] FOR j FROM 0 TO m
[0058] dp[0][j] = j / / It takes j insertions to change an empty grid sequence string into the first j grid sequences of gridSequence2
[0059] END FOR
[0060] / / Fill the dp array
[0061] FOR i FROM 1 TO n
[0062] FOR j FROM 1 TO m
[0063] If gridSequence1[i - 1] equals gridSequence2[j - 1]
[0064] cost = 0 / / The grid sequences are the same, no editing is required
[0065] Else
[0066] cost = 1 / / The grid sequences are different, 1 edit is required
[0067] End If
[0068] dp[i][j] = MIN(dp[i - 1][j] + 1, / / Delete operation
[0069] dp[i][j - 1] + 1, / / Insert operation
[0070] dp[i - 1][j - 1] + cost) / / Replace operation
[0071] End For
[0072] End For
[0073] / / dp[n][m] contains the edit distance between the entire gridSequence1 and gridSequence2
[0074] Return dp[n][m]
[0075] End Function
[0076] The pseudo - code for determining the travel mode of a travel segment is as follows:
[0077] Input: Sequence of line coordinates , sequence of user travel base stations .
[0078] Output: Corresponding travel mode
[0079] for :
[0080] = ( )
[0081] for :
[0082] if is in s or is absent from , where then
[0083] add to
[0084] for :
[0085] if is in then
[0086] add to
[0087] =
[0088] if then
[0089] belongs to a certain mode of transportation
[0090] Figure 3 Shows the generation process of the basic version of the activity chain.
[0091] In one embodiment, the above step S3: On the basic version of the activity chain, using clustering, inference rules based on speed and travel distance, identify the travel modes of driving, walking, and cycling, and fill them into the basic version of the activity chain to obtain the synthesized version of the activity chain, specifically including:
[0092] Step S31: Based on the basic version of the activity chain, select the start, end, and intermediate nodes of each travel segment, and calculate the average speed, travel time, and travel distance of each travel segment;
[0093] Step S32: Use the average speed, travel time, and travel distance as travel characteristics to cluster the travel segments;
[0094] Step S33: Select the clusters with low speed, short time, and short distance as walking segments, the clusters with medium speed, medium time, and medium distance as cycling segments, and the segments with high speed, long time, and long distance as driving segments to obtain the synthesized version of the activity chain.
[0095] In one embodiment, the above-mentioned step S4: Check and analyze the stay segments, and add description tags; optimize and merge the synthetic version of the activity chain to generate the final version of the whole-process activity chain, specifically including:
[0096] Step S41: Add tags for the stay mode to the stay segments according to the stay duration and stay moment, including: work, home, short stop outside, specifically including:
[0097] Step S411: Obtain the stay time of each point, and divide the stay duration into long-term stay and short-term stay according to a preset time threshold;
[0098] After multiple experiments, the value of this time threshold in the embodiment of the present invention is 1 hour.
[0099] Step S412: Analyze whether the start and end time periods of the stay are within the working hours of weekdays, and combine the time pattern (such as morning and evening rush hours) with whether it is a long stay to further infer whether the stay point is a work point;
[0100] Step S413: Add detailed mode tags to the stay points, such as "working on weekdays", "staying at home on weekends", "short stop outside on weekdays", etc.;
[0101] Step S42: Generate a travel sequence according to the synthetic version of the activity chain, remove and replace unreasonable travel mode sequences, specifically including:
[0102] Step S421: Detect public transportation modes according to the travel period, including subway, bus, high-speed rail. If it is in the non-operating time period (from zero o'clock to five o'clock every day), then replace it with driving, cycling and walking modes according to the speed;
[0103] Step S422: Mark driving, a non-public travel mode, and bus and subway as mutually exclusive options. If it is detected that the activity chain contains driving adjacent to bus or subway, then replace it with adjacent activity modes;
[0104] Step S43: Merge travel modes in the same state to generate the final version of the whole-process activity chain.
Claims
1. A method for identifying and describing the whole-process activity pattern of personnel based on mobile phone signaling data, characterized in that: include: Step S1: Based on the user's signaling data throughout the day, by identifying the user's stay status, the user's activity chain throughout the day is preliminarily segmented into stay segments and movement segments of the activity chain; Step S2: using base station sequence matching to identify the movement mode in the movement segment, including subway travel mode, high-speed rail travel mode, and bus travel mode, further segmenting the movement segment, filling the activity chain, and obtaining a basic version of the activity chain; Step S3: On the basic version activity chain, clustering, speed-based and travel distance-based inference rules are used to identify travel modes of driving, walking and cycling, and fill in the basic version activity chain to obtain a synthetic version activity chain; Step S4: Check and analyze the stay segment and add a description tag; optimize and merge the synthesized version activity chain to generate a final version of the full-process activity chain.
2. The method for identifying and describing the whole-process activity pattern of personnel based on mobile phone signaling data according to claim 1 is characterized in that: The step S1: based on the user's signaling data throughout the day, by identifying the user's stay status, the user's activity chain throughout the day is preliminarily segmented into stay segments and movement segments of the activity chain, specifically including: Step S11: Put the signaling data into a sliding window, determine whether the time period is a stay state according to the degree of repetition of the longitude and latitude of the data in the sliding window, and calculate the stay point. The calculation formula is as follows: (1) (2) in, Indicates the degree of repetition of the coordinates of the signaling data t within the window. Indicates the size of the window. is the indicative function, and are the coordinates of the signaling data numbered i and t respectively; Indicates the activity status of signaling data t, including stay and move; Step S12: for each signaling data whose activity state is stay, retain its latitude and longitude coordinates and retention time, weight them according to the retention time and cluster them on the map, and use the center coordinates of each cluster after clustering as the retention point; Step S13: Taking the dwell point as a segmentation point, the activity chain is segmented into a plurality of moving segments and staying segments, thereby completing preliminary semantic segmentation.
3. The method for identifying and describing the whole-process activity pattern of a person based on mobile phone signaling data according to claim 2 is characterized in that: The step S2: using base station sequence matching to identify the movement mode in the movement segment, including subway travel mode, high-speed rail travel mode, and bus travel mode, further segmenting the movement segment, filling the activity chain, and obtaining a basic version of the activity chain, specifically includes: Step S21: Obtain subway, bus, and high-speed rail operation data, including: line description information data, station description information data, and line topology data; Step S22: Generate running line data according to the line topology of subways, buses, and high-speed railways, each of which is in the form of a set of ordered longitude and latitude coordinate sequences; then map the longitude and latitude coordinates of the running line data to grids named in the form of GIDs to form a line GID grid sequence; Step S23: Scan the longitude and latitude coordinates of all base stations, establish a mapping table between the base station IDs that may interact when traveling into a certain GID grid and the GID grid ID, and convert the signaling data into a trajectory GID grid sequence through this mapping table; Step S24: Calculate the edit distance between the trajectory GID grid sequence and the line GID grid sequence to measure whether the trajectory of the person matches the line; if the edit distance does not exceed the matching threshold, determine that the travel segment is the corresponding subway, bus, or high-speed rail travel mode, fill it into the activity chain, and obtain a basic version of the activity chain.
4. The method for identifying and describing the whole-process activity pattern of personnel based on mobile phone signaling data according to claim 3 is characterized in that: The step S3: on the basic version activity chain, using clustering, inference rules based on speed and travel distance to identify the travel modes of driving, walking, and cycling, and filling in the basic version activity chain to obtain a synthetic version activity chain, specifically includes: Step S31: Based on the basic version activity chain, the start node, end node and intermediate node of each trip are selected to calculate the average speed, travel time and travel distance of each trip; Step S32: clustering the travel segments by taking the average speed, travel time and travel distance as travel features; Step S33: Select the clusters of low speed, short time and short distance as walking segments, the clusters of medium speed, medium time and medium distance as riding segments, and the clusters of high speed, long time and long distance as driving segments, to obtain a synthetic version of the activity chain.
5. The method for identifying and describing the whole-process activity pattern of personnel based on mobile phone signaling data according to claim 4 is characterized in that: The step S4: checking and analyzing the stay segment and adding a description tag; optimizing and merging the synthesized version activity chain to generate a final version of the full process activity chain, specifically including: Step S41: adding a label of the stay mode to the stay segment according to the stay duration and the stay time, including: work, home, and short stay outside; Step S42: generating a travel sequence according to the synthetic version activity chain, removing and replacing unreasonable travel mode sequences; Step S43: Merge travel modes in the same state to generate a final version of the full-process activity chain.
Citation Information
Patent Citations
Resident trip mode comprehensive judging method based on handset signaling data
CN105117789A
Mobile phone user trip mode identification method based on mobile phone signaling data and navigation route data
CN106197458A
Method for acquiring public transportation travel chain by using mobile phone GPS and electronic map data
CN111508228A
Method and system for identifying transportation mode based on mobile phone signaling
CN112542045A
Urban resident transportation mode identification method based on mobile phone signaling data
CN115206104A